Rna-guided nucleases and active fragments and variants thereof and methods of use

EP4750890A2Pending Publication Date: 2026-06-03LIFEEDIT THERAPEUTICS INC

Patent Information

Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
LIFEEDIT THERAPEUTICS INC
Filing Date
2024-07-27
Publication Date
2026-06-03

AI Technical Summary

Technical Problem

Existing genome editing technologies, such as meganucleases, zinc finger fusion proteins, and TALENs, require costly and inefficient generation of chimeric nucleases for each target sequence, whereas RNA-guided nucleases like CRISPR-Cas systems are less efficient in certain applications.

Method used

Development of engineered RNA-guided nuclease (RGN) variants, such as LPG10145, with increased editing efficiency, and the use of base editors and polymerase editors to modify target sequences in a sequence-specific manner.

Benefits of technology

The engineered RGN variants demonstrate enhanced gene editing efficiency, with some variants showing greater than 15% increase in activity, and enable precise modification of target sequences through various editing mechanisms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2024057272_30012025_PF_FP_ABST
    Figure IB2024057272_30012025_PF_FP_ABST
Patent Text Reader

Abstract

Compositions and methods for binding to a target sequence of interest are provided. The compositions find use in cleaving or modifying a target sequence of interest, visualization of a target sequence of interest, and modifying the expression of a sequence of interest. Compositions comprise RNA-guided nuclease (RGN) polypeptides, CRISPR RNAs, trans-activating CRISPR RNAs, guide RNAs, and nucleic acid molecules encoding the same. Vectors and host cells comprising the nucleic acid molecules are also provided. Further provided are RGN systems for binding a target sequence of interest, wherein the RGN system comprises an RNA-guided nuclease polypeptide and one or more guide RNAs.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] RNA-GUIDED NUCLEASES AND ACTIVE FRAGMENTS AND VARIANTS THEREOF AND

[0002] METHODS OF USE

[0003] CROSS-REFERENCE TO RELATED APPLICATIONS

[0004] This application claims priority to U.S. Provisional Application No. 63 / 516,127, filed July 27, 2023, which is incorporated by referenced herein in its entirety.

[0005] REFERENCE TO A SEQUENCE LISTING SUBMITTED ELECTRONICALLY AS AN XML FILE The instant application contains a Sequence Listing which has been submitted in xml format and is hereby incorporated by reference in its entirety. Said xml copy, created on July 24, 2024, is named L103438_1440WO_SL, and is 482,503 bytes in size.

[0006] FIELD OF THE INVENTION

[0007] The present invention relates to the field of molecular biology and gene editing.

[0008] BACKGROUND OF THE INVENTION

[0009] Targeted genome editing or modification is rapidly becoming an important tool for basic and applied research. Initial methods involved engineering nucleases such as meganucleases, zinc finger fusion proteins or TALENs, requiring the generation of chimeric nucleases with engineered, programmable, sequencespecific DNA-binding domains specific for each particular target sequence. RNA-guided nucleases, such as the Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-associated (Cas) proteins of the CRISPR-Cas bacterial system, allow for the targeting of specific sequences by complexing the nucleases with guide RNA that specifically hybridizes with a particular target sequence. Producing target-specific guide RNAs is less costly and more efficient than generating chimeric nucleases for each target sequence. Such RNA-guided nucleases can be used to edit genomes optionally through the introduction of a sequencespecific, double-stranded break that is repaired via error-prone non-homologous end-joining (NHEJ) to introduce a mutation at a specific genomic location. Alternatively, heterologous DNA may be introduced into the genomic site via homology-directed repair. RNA-guided nucleases (RGNs) can also be used for base editing when fused with a deaminase or prime editing when fused with reverse transcriptase.

[0010] Prime editing is a versatile and precise genome editing method that directly writes new genetic information into a specified DNA site using an RNA-guided DNA binding protein (e.g., RGN) working in association with a reverse transcriptase (described in, e.g., US 11,447,770BI; WO2021072328; WO2021226558; WO2020156575; W02021042047; US11193123; each incorporated by reference in its entirety herein). The prime editing system uses an RGN that is a nickase and a polymerase, and the system is programmed with a prime editing guide RNA that comprises a primer binding site (PBS) and a DNA synthesis template that serves as the template for the replacement strand comprising the edit. The prime editor nicks the non-target strand upstream of the sequence to be edited and upstream of the PAM, creating a 3' flap on the non-target strand. The PBS is complementary to the 3' flap of the non-target strand and hybridrization of the PBS and 3' flap of the non-target strand allows for the polymerization of the replacement strand containing the edit using the DNA synthesis template.

[0011] BRIEF SUMMARY OF THE INVENTION

[0012] Compositions and methods for binding a target sequence of interest are provided. The compositions find use in cleaving or modifying a target sequence of interest, detection of a target sequence of interest, and modifying the expression of a sequence of interest. Compositions comprise: RNA-guided nuclease (RGN) polypeptide variants that have increased editing efficiency as compared to their counterpart wild-type RGN polypeptide; base editors; polymerase editors (PEs); CRISPR RNAs (crRNAs); trans-activating CRISPR RNAs (tracrRNAs); guide RNAs (gRNAs); nucleic acid molecules encoding the same; vectors and host cells comprising the nucleic acid molecules; and pharmaceutical compositions comprising the same. Also provided are RGN systems and ribonucleoprotein complexes for binding a target sequence of interest, wherein the RGN system comprises an RNA-guided nuclease polypeptide and one or more guide RNAs. Polymerase editor (PE) systems comprising one or more polymerase editing guide RNAs (PEgRNAs), a polymerase, and an RGN polypeptide are also provided. Thus, methods disclosed herein are drawn to binding a target sequence of interest in a target polynucleotide, and in some embodiments, cleaving or modifying the target polynucleotide of interest. The target polynucleotide of interest can be modified, for example, as a result of non-homologous end joining, homology-directed repair with an introduced donor sequence, or base editing, or polymerase editing.

[0013] BRIEF DESCRIPTION OF THE FIGURES

[0014] FIG. 1 shows percent gene editing efficiency for 581 constructs in an initial screen. Each construct encoding a variant LPG10145 nuclease was delivered, along with a plasmid encoding a guide RNA, to HEK293T cells by plasmid lipofection. One guide RNA was tested in duplicate (n=2) in the initial screen. Tracking of Indels by DEcomposition (TIDE) analysis was performed to determine % gene editing efficiency 2 days post-transfection. The gene editing efficiency of wild-type LPG10145 nuclease was normalized to 1.

[0015] FIG. 2 shows that 123 variant LPG10145 nucleases having an increase in gene editing activity > 15% were obtained from the initial screen. The 123 variant LPG10145 nucleases were subjected to next generation sequencing (NGS) for confirmation of editing activity, which narrowed the 123 hits to 71 variants. The 71 variants were further tested with more guide RNAs. The gene editing efficiency of wildtype LPG10145 nuclease was normalized to 1. FIG. 3 shows gene editing activity (% Indel) for variant LPG10145 nucleases tested with 3 additional guide RNAs (for a total of 4 guide RNAs tested: guides A, B, C, D). Each construct encoding a variant LPG10145 nuclease was delivered, along with a plasmid encoding a guide RNA, to HEK293T cells by plasmid lipofection. Each guide RNA was tested in duplicate (n=2). Percent Indel (insertions / deletions) was determined by NGS 2 days post-transfection. The numbers on the x-axis for each guide from left to right are: 778, 969, 856, 822, 974, 55, 643, 653, 780, 954, 472, 968, 52, 86, 745, 973, 795, 541, 911, 533, 975, 774, and 872, and represent the amino acid position in LPG10145 nuclease of SEQ ID NO: 1 that is mutated.

[0016] FIG. 4 shows confirmation of gene editing activity by NGS of variant LPG10145 nucleases having an increase in gene editing activity from the initial screen. Twenty R variants increased gene editing > 20%. Each construct encoding a variant LPG10145 nuclease was delivered, along with a plasmid encoding a guide RNA, to HEK293T cells by plasmid lipofection. Three additional guide RNAs were tested (for a total of 4 guide RNAs tested). All guides were normalized and averaged. The gene editing efficiency of wild-type LPG10145 nuclease was normalized to 1.

[0017] FIG. 5 shows ranking of the top single variant LPG10145 nucleases with increased activity by statistical analysis. Left graph: Variant E778R at top to variant K872R, p < 0.001 ; variant K843R to variant K871R, p < 0.01. The right graph shows the statistical ranking of variants according to their position.

[0018] FIG. 6 shows a structural alignment of LPG10145 with a guide RNA. Many mutations, but not all, are at the interface with DNA / RNA. Set2 residues are labeled (unless disordered). The LPG10145 homology model is based on the structure of the RGN from .S', thermophilus (6M0W).

[0019] FIG. 7 shows that 6 variant LPG10145 nucleases were selected based on statistical analysis and structural modeling for a combinatorial library. The single variants (lx variant), double variant combinations (2x variants), triple variant combinations (3x variants), quadruple variant combinations (4x variants), quintuple variant combinations (5x variant), and the sextuple combination (6x variant) are shown. Sixty- three constructs encoding the 63 variant / combinatorial variant were generated and were tested with multiple different guide RNAs. Each construct encoding a variant / combinatorial variant LPG10145 nuclease was delivered, along with a plasmid encoding a guide RNA, to HEK293T cells by plasmid lipofection. Each guide RNA was tested in duplicate (n=2). Percent Indel (insertions / deletions) was determined by NGS 2 days post-transfection.

[0020] FIG. 8 shows increased gene editing by combinatorial variant LPG10145 nucleases with multiple different guide RNAs. Each construct encoding a combinatorial variant LPG10145 nuclease was delivered, along with a plasmid encoding a guide RNA, to HEK293T cells by plasmid lipofection. Each guide RNA was tested in duplicate (n=2). A total of seven guide RNAs (shown as A, Bl, B2, Cl, C2, DI, and D2) were tested. Percent Indel (insertions / deletions) was determined by NGS 2 days post-transfection. Six of the seven guide RNAs tested showed significant increase in editing with the variants. Higher editing guide RNAs see less of an effect. FIG. 9 shows combinatorial editing efficiency by combinatorial variant LPG10145 nucleases. The highest gene editing was obtained with 3x, 4x, and 5x variants. E778R and E969R were the common variants in the high editing populations. Each construct encoding a variant / combinatorial variant LPG10145 nuclease was delivered, along with a plasmid encoding a guide RNA, to HEK293T cells by plasmid lipofection. Each guide RNA was tested in duplicate (n=2). The gene editing efficiency of wild-type LPG10145 nuclease was normalized to 1.

[0021] FIG. 10 shows the top combinatorial variants with the highest significant editing. Fifteen combinatorial variant LPG10145 nucleases demonstrate the highest editing.

[0022] DETAILED DESCRIPTION

[0023] Many modifications and other embodiments of the inventions set forth herein will come to mind to one skilled in the art to which these inventions pertain having the benefit of the teachings presented in the foregoing descriptions and the associated drawings. Therefore, it is to be understood that the inventions are not to be limited to the specific embodiments disclosed and that modifications and other embodiments are intended to be included within the scope of the appended embodiments. Although specific terms are employed herein, they are used in a generic and descriptive sense only and not for purposes of limitation.

[0024] I. Overview

[0025] RNA-guided nucleases (RGNs) allow for the targeted manipulation of specific site(s) within a genome and are useful in the context of gene targeting for therapeutic and research applications. In a variety of organisms, including mammals, RNA-guided nucleases have been used for genome engineering by stimulating non-homologous end joining and homologous recombination, for example. The compositions and methods described herein are useful for creating single- or double-stranded breaks in polynucleotides, modifying polynucleotides, detecting a particular site within a polynucleotide, or modifying the expression of a particular gene.

[0026] The engineered variant LPG10145 RGNs, or active variants or fragments thereof, are directed to the target sequence by a guide RNA (gRNA) as part of a Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) RNA-guided nuclease system. The RGNs are considered “RNA-guided” because guide RNAs form a complex with the RNA-guided nucleases to direct the RNA-guided nuclease to bind to a target sequence and in some embodiments, introduce a single-stranded or double -stranded break at the target sequence. The engineered variant LPG10145 RGNs, or active variants or fragments thereof, disclosed herein can increase gene editing efficiency as compared to the corresponding wild-type LPG10145 RGN. In some embodiments, the increase in gene editing efficiency is at least 15%. After the target sequence has been cleaved, the break can be repaired such that the DNA sequence of the target sequence is modified during the repair process. Thus, provided herein are methods for using the RNA-guided nucleases to modify a target sequence in the DNA of host cells. For example, RNA-guided nucleases can be used to modify a target sequence at a genomic locus of eukaryotic cells or prokaryotic cells. In some embodiments, the variant LPG10145 RGNs can alter gene expression by modifying a target sequence.

[0027] Also provided herein are base editors, polymerase editors, base editing systems, polymerase editor systems, and methods of using the same for editing a target DNA molecule, wherein the editing systems comprise a DNA polymerase and an engineered variant LPG10145 RGN polypeptide as described herein, or an active variant or fragment thereof.

[0028] II. RNA-guided nucleases

[0029] Provided herein are RNA-guided nucleases. The term RNA-guided nuclease (RGN) refers to a polypeptide that binds to a particular target nucleotide sequence in a sequence-specific manner and is directed to the target nucleotide sequence by a guide RNA molecule that is complexed with the polypeptide and hybridizes with the target sequence. Although an RNA-guided nuclease can be capable of cleaving the target sequence upon binding, the term RNA-guided nuclease also encompasses nuclease-dead RNA-guided nucleases that are capable of binding to, but not cleaving, a target sequence. Cleavage of a target sequence by an RNA-guided nuclease can result in a single- or double-stranded break. RNA-guided nucleases only capable of cleaving a single strand of a double -stranded nucleic acid molecule are referred to herein as nickases.

[0030] The RNA-guided nucleases disclosed herein include engineered variants of the LPG10145 RNA- guided nuclease described in International Publ. No. WO 2023 / 139557, filed January 23, 2023, which is incorporated herein in its entirety. In some embodiments, the variant LPG10145 RGNs comprise an amino acid sequence set forth as SEQ ID NO: 1, wherein the amino acid residue at one or more of the following amino acid positions differs from the corresponding amino acid residue in SEQ ID NO: 1: 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975, and active fragments or variants thereof that retain the ability to bind to a target nucleotide sequence in an RNA-guided sequence-specific manner. When referring to a variant LPG10145 RGN that “comprises an amino acid sequence set forth as SEQ ID NO: 1, wherein the amino acid residue at one or more of the following amino acid positions differs”, it is intended to mean that the variant LPG10145 RGN comprises an amino acid sequence that is identical to the amino acid sequence set forth in SEQ ID NO: 1, except that amino acid residues at the recited positions differ from the corresponding amino acid residues in SEQ ID NO: 1. When referring to an amino acid position, the number of the amino acid position is counted from the amino-terminus of a given polypeptide. The first amino acid residue at the amino-terminus of a polypeptide is denoted as position 1. In some embodiments, the first amino acid residue is a Methionine. The position numbers then increase numerically with each amino acid residue from the amino terminus of the polypeptide to the carboxy terminus of the polypeptide, with the last amino acid position being the last amino acid residue at the carboxy terminus of the polypeptide. In some embodiments, the variant LPG10145 RGNs comprise an amino acid sequence that is identical to the amino acid sequence set forth in SEQ ID NO: 1, except that the amino acid residue at one or more of the following amino acid positions is a positively charged amino acid residue: 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975, and active fragments or variants thereof that retain the ability to bind to a target nucleotide sequence in an RNA-guided sequence-specific manner. In some embodiments, the variant LPG10145 RGNs comprise an amino acid sequence that is identical to the amino acid sequence set forth in SEQ ID NO: 1, except that the amino acid residue at one or more of the following amino acid positions is an R: 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975, and active fragments or variants thereof that retain the ability to bind to a target nucleotide sequence in an RNA-guided sequence -specific manner. When referring to a “variant LPG10145 RGN comprising an amino acid sequence set forth as SEQ ID NO: 1, wherein the amino acid residue at one or more of the following amino acid positions is a positively charged amino acid residue” or “a variant LPG10145 RGN comprising an amino acid sequence having at least x% (e.g., 85%) sequence identity to SEQ ID NO: 1, wherein the amino acid residue at one or more of the following amino acid positions is a positively charged amino acid residue”, it is intended that the positively charged amino acid residue in the variant LPG10145 RGN is a different positively charged amino acid residue if the amino acid residue at a particular position of SEQ ID NO: 1 already comprises a positively charged amino acid residue.

[0031] In some embodiments, the variant LPG10145 RGNs comprise an amino acid sequence that is identical to the amino acid sequence set forth in SEQ ID NO: 1, except that the amino acid residue at amino acid position 778 and / or 969 differs from the corresponding amino acid residue in SEQ ID NO: 1, and active fragments or variants thereof that retain the ability to bind to a target nucleotide sequence in an RNA-guided sequence -specific manner. In some embodiments, the variant LPG10145 RGNs comprise an amino acid sequence that is identical to the amino acid sequence set forth in SEQ ID NO: 1, except that the amino acid residue at amino acid position 778 and / or 969 is a positively charged amino acid residue, and active fragments or variants thereof that retain the ability to bind to a target nucleotide sequence in an RNA-guided sequence -specific manner. In some embodiments, the variant LPG10145 RGNs comprise an amino acid sequence that is identical to the amino acid sequence set forth in SEQ ID NO: 1, except that the amino acid residue at amino acid position 778 and / or 969 is an R, and active fragments or variants thereof that retain the ability to bind to a target nucleotide sequence in an RNA-guided sequence-specific manner.

[0032] In some embodiments, the variant LPG10145 RGNs comprise an amino acid sequence that is identical to the amino acid sequence set forth in SEQ ID NO: 1, except that amino acid residues at positions 778 and 856 differ from the corresponding amino acid residues in SEQ ID NO: 1, and active fragments or variants thereof that retain the ability to bind to a target nucleotide sequence in an RNA-guided sequencespecific manner. In some embodiments, the variant LPG10145 RGNs comprise an amino acid sequence that is identical to the amino acid sequence set forth in SEQ ID NO: 1, except that amino acid residues at positions 778 and 856 are positively charged amino acid residues, and active fragments or variants thereof that retain the ability to bind to a target nucleotide sequence in an RNA-guided sequence -specific manner.

[0033] In some embodiments, the variant LPG10145 RGNs comprise an amino acid sequence that is identical to the amino acid sequence set forth in SEQ ID NO: 1, except that amino acid residues at positions 55, 647, 778, and 969 differ from the corresponding amino acid residues in SEQ ID NO: 1, and active fragments or variants thereof that retain the ability to bind to a target nucleotide sequence in an RNA-guided sequence -specific manner. In some embodiments, the variant LPG10145 RGNs comprise an amino acid sequence that is identical to the amino acid sequence set forth in SEQ ID NO: 1, except that amino acid residues at positions 55, 647, 778, and 969 are positively charged amino acid residues, and active fragments or variants thereof that retain the ability to bind to a target nucleotide sequence in an RNA-guided sequencespecific manner.

[0034] In some embodiments, the variant LPG10145 RGNs comprise: (a) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 52 is an R; (b) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 55 is an R; (c) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 86 is an R; (d) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 472 is an R; (e) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 533 is an R; (f) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 541 is an R; (g) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 643 is substituted by an R; (h) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 647 is an R; (i) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 653 is an R; (j) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 745 is an R; (k) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position 774 is an R; (1) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 778 is an R; (m) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 780 is an R; (n) the amino acid sequence set forth as SEQ ID NO: 1, wherein A at amino acid position 795 is an R; (o) the amino acid sequence set forth as SEQ ID NO: 1, wherein Q at amino acid position 822 is an R; (p) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 843 is an R; (q) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 856 is an R; (r) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 871 is an R; (s) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 872 is an R; (t) the amino acid sequence set forth as SEQ ID NO: 1, wherein D at amino acid position 900 is an R; (u) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position 911 is an R; (v) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 913 is an R; (w) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position 954 is an R; (x) the amino acid sequence set forth as SEQ ID NO: 1, wherein V at amino acid position 958 is an R; (y) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 968 is an R; (z) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 969 is an R; (aa) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 973 is an R; (bb) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 974 is an R; (cc) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position 975 is an R; (dd) the amino acid sequence set forth as SEQ ID NO: 2; (ee) the amino acid sequence set forth as SEQ ID NO: 3; (ff) the amino acid sequence set forth as SEQ ID NO: 4; (gg) the amino acid sequence set forth as SEQ ID NO: 5; (hh) the amino acid sequence set forth as SEQ ID NO: 6; (ii) the amino acid sequence set forth as SEQ ID NO: 7; (jj) the amino acid sequence set forth as SEQ ID NO: 8; (kk) the amino acid sequence set forth as SEQ ID NO: 9; (11) the amino acid sequence set forth as SEQ ID NO: 10; (mm) the amino acid sequence set forth as SEQ ID NO: 11; (nn) the amino acid sequence set forth as SEQ ID NO: 12; (oo) the amino acid sequence set forth as SEQ ID NO: 13; (pp) the amino acid sequence set forth as SEQ ID NO: 14; (qq) the amino acid sequence set forth as SEQ ID NO: 15; and (rr) the amino acid sequence set forth as SEQ ID NO: 16, and active fragments or variants thereof that retain the ability to bind to a target nucleotide sequence in an RNA-guided sequence-specific manner.

[0035] In some embodiments, the variant LPG10145 RGNs comprise: (a) the amino acid sequence set forth as SEQ ID NO: 2; (b) the amino acid sequence set forth as SEQ ID NO: 3; (c) the amino acid sequence set forth as SEQ ID NO: 4; (d) the amino acid sequence set forth as SEQ ID NO: 5; (e) the amino acid sequence set forth as SEQ ID NO: 6; (f) the amino acid sequence set forth as SEQ ID NO: 7; (g) the amino acid sequence set forth as SEQ ID NO: 8; (h) the amino acid sequence set forth as SEQ ID NO: 9; (i) the amino acid sequence set forth as SEQ ID NO: 10; (j) the amino acid sequence set forth as SEQ ID NO: 11; (k) the amino acid sequence set forth as SEQ ID NO: 12; (1) the amino acid sequence set forth as SEQ ID NO: 13; (m) the amino acid sequence set forth as SEQ ID NO: 14; (n) the amino acid sequence set forth as SEQ ID NO: 15; and (o) the amino acid sequence set forth as SEQ ID NO: 16, and active fragments or variants thereof that retain the ability to bind to a target nucleotide sequence in an RNA-guided sequence-specific manner.

[0036] In some of these embodiments, the active fragment or variant of a variant LPG10145 RGN is capable of cleaving a single- or double -stranded target sequence. When referring to a variant LPGI0145 RGN that “comprises an amino acid sequence having at least x% (e.g., 85%) sequence identity to SEQ ID NO: 1, wherein the amino acid residue at one or more of the following amino acid positions differs”, it is intended to mean that the variant LPGI0145 RGN comprises an amino acid sequence that has at least x% (e.g., 85%) sequence identity to the amino acid sequence set forth in SEQ ID NO: 1, and the amino acid residues at the recited positions differ from the corresponding amino acid residues in SEQ ID NO: 1. In some embodiments, an active variant of a variant LPG10145 RGN comprises an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 1, wherein the amino acid residue at one or more of the following amino acid positions differs from the corresponding amino acid residue in SEQ ID NO: 1: 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975. In some embodiments, an active variant of a variant LPG10145 RGN comprises an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 1, wherein the amino acid residue at one or more of the following amino acid positions is a positively charged amino acid residue: 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975. In some embodiments, an active variant of a variant LPG10145 RGN comprises an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 1, wherein the amino acid residue at one or more of the following amino acid positions is an R: 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975.

[0037] In some of these embodiments, the active fragment or variant of a variant LPG10145 RGN is capable of cleaving a single- or double -stranded target sequence. The binding and / or cleaving of a target sequence by an active variant LPG10145 RGN can be dependent upon the active variant LPG10145 RGN recognizing a protospacer adjacent motif (PAM) adjacent and 3’ to the target sequence, and the PAM comprises a consensus sequence of NNGG. In some embodiments, an active variant of a variant LPG10145 RGN comprises an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 1, wherein the amino acid residue at amino acid position 778 and / or 969 differs from the corresponding amino acid residue in SEQ ID NO: 1. In some embodiments, an active variant of a variant LPG10145 RGN comprises an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 1, wherein the amino acid residue at amino acid position 778 and / or 969 is a positively charged amino acid residue. In some embodiments, an active variant of a variant LPG10145 RGN comprises an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 1, wherein the amino acid residue at amino acid position 778 and / or 969 is an R.

[0038] In some embodiments, an active variant of a variant LPG10145 RGN comprises an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 1, wherein amino acid residues at positions 778 and 856 differ from the corresponding amino acid residues in SEQ ID NO: 1. In some embodiments, an active variant of a variant LPG10145 RGN comprises an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 1, wherein the amino acid residues at positions

[0039] 778 and 856 are positively charged amino acid residues. In some embodiments, an active variant of a variant LPG10145 RGN comprises an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 1, wherein amino acid residues at positions 55, 647, 778, and 969 differ from the corresponding amino acid residues in SEQ ID NO: 1. In some embodiments, an active variant of a variant LPG10145 RGN comprises an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 1, wherein amino acid residues at positions 55, 647, 778, and 969 are positively charged amino acid residues.

[0040] In some embodiments, an active variant of a variant LPG10145 RGN comprises an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to: (a) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 52 is an R; (b) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 55 is an R; (c) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 86 is an R; (d) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 472 is an R; (e) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 533 is an R; (f) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 541 is an R; (g) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 643 is substituted by an R; (h) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 647 is an R; (i) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 653 is an R; (j) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 745 is an R; (k) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position 774 is an R; (1) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 778 is an R; (m) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 780 is an R; (n) the amino acid sequence set forth as SEQ ID NO: 1, wherein A at amino acid position 795 is an R; (o) the amino acid sequence set forth as SEQ ID NO: 1, wherein Q at amino acid position 822 is an R; (p) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 843 is an R; (q) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 856 is an R; (r) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 871 is an R; (s) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 872 is an R; (t) the amino acid sequence set forth as SEQ ID NO: 1, wherein D at amino acid position 900 is an R; (u) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position 911 is an R; (v) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 913 is an R; (w) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position 954 is an R; (x) the amino acid sequence set forth as SEQ ID NO: 1, wherein V at amino acid position 958 is an R; (y) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 968 is an R; (z) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 969 is an R; (aa) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 973 is an R; (bb) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 974 is an R; (cc) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position 975 is an R; (dd) the amino acid sequence set forth as SEQ ID NO: 2, wherein Q at amino acid position 822 is an R, G at amino acid position 856 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 778 is an R; (ee) the amino acid sequence set forth as SEQ ID NO: 3, wherein Q at amino acid position 822 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 778 is substituted by an R; (ff) the amino acid sequence set forth as SEQ ID NO: 4, wherein Q at amino acid position 822 is an R, S at amino acid position 55 is an R, and E at amino acid position 778 is an R; (gg) the amino acid sequence set forth as SEQ ID NO: 5, wherein Q at amino acid position 822 is an R, S at amino acid position 647 is an R, and E at amino acid position 778 is an R; (hh) the amino acid sequence set forth as SEQ ID NO: 6, wherein Q at amino acid position 822 is an R, G at amino acid position 856 is an R, and E at amino acid position 778 is an R; (ii) the amino acid sequence set forth as SEQ ID NO: 7, wherein G at amino acid position 856 is an R, and E at amino acid position 778 is an R; (jj) the amino acid sequence set forth as SEQ ID NO: 8, wherein E at amino acid position 778 is an R, Q at amino acid position 822 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R; (kk) the amino acid sequence set forth as SEQ ID NO: 9, wherein E at amino acid position 778 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R; (11) the amino acid sequence set forth as SEQ ID NO: 10, wherein E at amino acid position 778 is an R, G at amino acid position 856 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R; (mm) the amino acid sequence set forth as SEQ ID NO: 11, wherein E at amino acid position 778 is an R, G at amino acid position 856 is an R, S at amino acid position 647 is an R, and E at amino acid position 969 is an R; (nn) the amino acid sequence set forth as SEQ ID NO: 12, wherein E at amino acid position 778 is an R, Q at amino acid position 822 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R; (oo) the amino acid sequence set forth as SEQ ID NO: 13, wherein E at amino acid position 778 is an R, Q at amino acid position 822 is an R, S at amino acid position 647 is an R, and E at amino acid position 969 is an R; (pp) the amino acid sequence set forth as SEQ ID NO: 14, wherein E at amino acid position 778 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R; (qq) the amino acid sequence set forth as SEQ ID NO: 15, wherein E at amino acid position 778 is an R, G at amino acid position 856 is an R, and E at amino acid position 969 is an R; and (rr) the amino acid sequence set forth as SEQ ID NO: 16, wherein E at amino acid position 778 is an R, and E at amino acid position 969 is an R.

[0041] In some embodiments, an active variant of a variant LPG10145 RGN comprises an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%,

[0042] 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to: (a) the amino acid sequence set forth as SEQ ID NO: 2, wherein Q at amino acid position 822 is an R, G at amino acid position 856 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 778 is an R; (b) the amino acid sequence set forth as SEQ ID NO: 3, wherein Q at amino acid position 822 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 778 is substituted by an R; (c) the amino acid sequence set forth as SEQ ID NO: 4, wherein Q at amino acid position 822 is an R, S at amino acid position 55 is an R, and E at amino acid position 778 is an R; (d) the amino acid sequence set forth as SEQ ID NO: 5, wherein Q at amino acid position 822 is an R, S at amino acid position 647 is an R, and E at amino acid position 778 is an R; (e) the amino acid sequence set forth as SEQ ID NO: 6, wherein Q at amino acid position 822 is an R, G at amino acid position 856 is an R, and E at amino acid position 778 is an R; (f) the amino acid sequence set forth as SEQ ID NO: 7, wherein G at amino acid position 856 is an R, and E at amino acid position 778 is an R; (g) the amino acid sequence set forth as SEQ ID NO: 8, wherein E at amino acid position 778 is an R, Q at amino acid position 822 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R; (h) the amino acid sequence set forth as SEQ ID NO: 9, wherein E at amino acid position 778 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R; (i) the amino acid sequence set forth as SEQ ID NO: 10, wherein E at amino acid position 778 is an R, G at amino acid position 856 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R; (j) the amino acid sequence set forth as SEQ ID NO: 11, wherein E at amino acid position 778 is an R, G at amino acid position 856 is an R, S at amino acid position 647 is an R, and E at amino acid position 969 is an R; (k) the amino acid sequence set forth as SEQ ID NO: 12, wherein E at amino acid position 778 is an R, Q at amino acid position 822 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R; (1) the amino acid sequence set forth as SEQ ID NO: 13, wherein E at amino acid position 778 is an R, Q at amino acid position 822 is an R, S at amino acid position 647 is an R, and E at amino acid position 969 is an R; (m) the amino acid sequence set forth as SEQ ID NO: 14, wherein E at amino acid position 778 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R; (n) the amino acid sequence set forth as SEQ ID NO: 15, wherein E at amino acid position 778 is an R, G at amino acid position 856 is an R, and E at amino acid position 969 is an R; and (o) the amino acid sequence set forth as SEQ ID NO: 16, wherein E at amino acid position 778 is an R, and E at amino acid position 969 is an R.

[0043] In some embodiments, an active fragment of a variant LPG10145 RGN comprises at least 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050 or more contiguous amino acid residues of an amino acid sequence set forth as SEQ ID NO: 1, wherein the amino acid residue at one or more of the following amino acid positions differs from the corresponding amino acid residue in SEQ ID NO: 1: 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975. In some embodiments, an active fragment of a variant LPG10145 RGN comprises at least 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050 or more contiguous amino acid residues of an amino acid sequence set forth as SEQ ID NO: 1, wherein the amino acid residue at one or more of the following amino acid positions is a positively charged amino acid residue: 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975. In some embodiments, an active fragment of a variant LPG10145 RGN comprises at least 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050 or more contiguous amino acid residues of an amino acid sequence set forth as SEQ ID NO: 1, wherein the amino acid residue at one or more of the following amino acid positions is an R: 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975.

[0044] In some embodiments, an active fragment of a variant LPG10145 RGN comprises at least 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050 or more contiguous amino acid residues of an amino acid sequence selected from: (a) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 52 is an R; (b) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 55 is an R; (c) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 86 is an R; (d) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 472 is an R; (e) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 533 is an R; (f) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 541 is an R; (g) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 643 is substituted by an R; (h) the amino acid sequence set forth as SEQ

[0045] ID NO: 1, wherein S at amino acid position 647 is an R; (i) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 653 is an R; (j) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 745 is an R; (k) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position 774 is an R; (1) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 778 is an R; (m) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 780 is an R; (n) the amino acid sequence set forth as SEQ ID NO: 1, wherein A at amino acid position 795 is an R; (o) the amino acid sequence set forth as SEQ ID NO: 1, wherein Q at amino acid position 822 is an R; (p) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 843 is an R; (q) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 856 is an R; (r) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 871 is an R; (s) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 872 is an R; (t) the amino acid sequence set forth as SEQ ID NO: 1, wherein D at amino acid position 900 is an R; (u) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position 911 is an R; (v) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 913 is an R; (w) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position 954 is an R; (x) the amino acid sequence set forth as SEQ ID NO: 1, wherein V at amino acid position 958 is an R; (y) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 968 is an R; (z) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 969 is an R; (aa) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 973 is an R; (bb) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 974 is an R; (cc) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position 975 is an R; (dd) the amino acid sequence set forth as SEQ ID NO: 2; (ee) the amino acid sequence set forth as SEQ ID NO: 3; (ff) the amino acid sequence set forth as SEQ ID NO: 4; (gg) the amino acid sequence set forth as SEQ ID NO: 5; (hh) the amino acid sequence set forth as SEQ ID NO: 6; (ii) the amino acid sequence set forth as SEQ ID NO: 7; (jj) the amino acid sequence set forth as SEQ ID NO: 8; (kk) the amino acid sequence set forth as SEQ ID NO: 9; (11) the amino acid sequence set forth as SEQ ID NO: 10; (mm) the amino acid sequence set forth as SEQ ID NO: 11; (nn) the amino acid sequence set forth as SEQ ID NO: 12; (oo) the amino acid sequence set forth as SEQ ID NO: 13; (pp) the amino acid sequence set forth as SEQ ID NO: 14; (qq) the amino acid sequence set forth as SEQ ID NO: 15; and (rr) the amino acid sequence set forth as SEQ ID NO: 16.

[0046] In some embodiments, an active fragment of a variant LPG10145 RGN comprises at least 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050 or more contiguous amino acid residues of an amino acid sequence set forth as SEQ ID NO: 1, wherein the amino acid residue at amino acid position 778 and / or 969 differs from the corresponding amino acid residue in SEQ ID NO: 1. In some embodiments, an active fragment of a variant LPG10145 RGN comprises at least 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050 or more contiguous amino acid residues of an amino acid sequence set forth as SEQ ID NO: 1, wherein the amino acid residue at amino acid position 778 and / or 969 is a positively charged amino acid residue. In some embodiments, an active fragment of a variant LPG10145 RGN comprises at least 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050 or more contiguous amino acid residues of an amino acid sequence set forth as SEQ ID NO: 1, wherein the amino acid residue at amino acid position 778 and / or 969 is an R.

[0047] In some embodiments, an active fragment of a variant LPG10145 RGN comprises at least 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050 or more contiguous amino acid residues of an amino acid sequence set forth as SEQ ID NO: 1, wherein amino acid residues at positions 778 and 856 differ from the corresponding amino acid residues in SEQ ID NO: 1. In some embodiments, an active fragment of a variant LPG10145 RGN comprises at least 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050 or more contiguous amino acid residues of an amino acid sequence set forth as SEQ ID NO: 1, wherein the amino acid residues at positions 778 and 856 are positively charged amino acid residues.

[0048] In some embodiments, an active fragment of a variant LPG10145 RGN comprises at least 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050 or more contiguous amino acid residues of an amino acid sequence set forth as SEQ ID NO: 1, wherein amino acid residues at positions 55, 647, 778, and 969 differ from the corresponding amino acid residues in SEQ ID NO: 1. In some embodiments, an active fragment of a variant LPG10145 RGN comprises at least 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050 or more contiguous amino acid residues of an amino acid sequence set forth as SEQ ID NO: 1, wherein the amino acid residues at positions 55, 647, 778, and 969 are positively charged amino acid residues.

[0049] In some embodiments, an active fragment of a variant LPG10145 RGN comprises at least 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050 or more contiguous amino acid residues of an amino acid sequence selected from: (a) the amino acid sequence set forth as SEQ ID NO: 2; (b) the amino acid sequence set forth as SEQ ID NO: 3; (c) the amino acid sequence set forth as SEQ ID NO: 4; (d) the amino acid sequence set forth as SEQ ID NO: 5; (e) the amino acid sequence set forth as SEQ ID NO: 6; (f) the amino acid sequence set forth as SEQ ID NO: 7; (g) the amino acid sequence set forth as SEQ ID NO: 8; (h) the amino acid sequence set forth as SEQ ID NO: 9; (i) the amino acid sequence set forth as SEQ ID NO: 10; (j) the amino acid sequence set forth as SEQ ID NO: 11; (k) the amino acid sequence set forth as SEQ ID NO: 12; (1) the amino acid sequence set forth as SEQ ID NO: 13; (m) the amino acid sequence set forth as SEQ ID NO: 14; (n) the amino acid sequence set forth as SEQ ID NO: 15; and (o) the amino acid sequence set forth as SEQ ID NO: 16.

[0050] An RGN polypeptide of the disclosure can comprise one or more of the following domains: a linker domain (linker domain 1, linker domain 2), a wedge (WED) domain, a RuvC nuclease domain, an HNH nuclease domain, a Rec domain, or a PAM-interacting (PI) domain. The RuvC domain can include a RuvCI, a RuvCII, a RuvCIII domain, or a combination thereof. A Rec or recognition lobe mediates nucleic acid binding through multiple Rec domains (e.g., Recl-3) by sensing nucleic acids, regulating the HNH conformational transition, and locking the catalytic HNH domain at the cleavage site. A wedge domain is responsible for the recognition of guide RNA scaffolds. A PAM-interacting domain is the domain of an RGN polypeptide that binds to a PAM site. Non-limiting examples of domains within LPG10145 RGN (SEQ ID NO: 1) or a variant LPG10145 RGN (SEQ ID NOs: 2-16, 182-196, and 271-285) include: RuvC-I from amino acid residues 1 to 42; BH from amino acid residues 43 to 79; RECI from amino acid residues 80 to 236; REC2 from amino acid residues 237 to 476; RuvC-II from amino acid residues 477 to 524; LI from amino acid residues 525 to 560; HNH from amino acid residues 561 to 676; L2 from amino acid residues 677 to 690; RuvC-III from amino acid residues 691 to 828; WED from amino acid residues 829 to 976; and PI from amino acid residues 977 to 1130. The general domains of RGN polypeptides can be determined via structural comparison to RGN polypeptides with defined domains.

[0051] An active variant or fragment of a variant LPG10145 RGN disclosed herein can comprise a RuvCI domain comprising amino acid residues 1 to 42 or that differs from amino acid residues 1 to 42 by 1, 2, or 3 amino acid residues, wherein the amino acid positions are in reference to SEQ ID NO: 1 or to a variant LPG10145 RGN, or an active variant or fragment thereof, disclosed herein (e.g., any one of SEQ ID NOs: 2- 16, 182-196, and 271-285). For example, an active variant or fragment of a variant LPG10145 RGN disclosed herein can comprise an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to any one of SEQ ID NOs: 2-16, 182-196, and 271-285 and can comprise a RuvCI domain comprising amino acid residues 1 to 42 or that differs from amino acid residues 1 to 42 of the RGN sequence (i.e. SEQ ID NOs: 2-16, 182-196, and 271-285) by 1, 2, or 3 amino acid residues.

[0052] An active variant or fragment of a variant LPG10145 RGN disclosed herein can comprise a BH domain comprising amino acid residues 43 to 79 or that differs from amino acid residues 43 to 79 by 1, 2, or 3 amino acid residues, wherein the amino acid positions are in reference to SEQ ID NO: 1 or to a variant LPG10145 RGN, or an active variant or fragment thereof, disclosed herein (e.g., any one of SEQ ID NOs: 2- 16, 182-196, and 271-285). For example, an active variant or fragment of a variant LPG10145 RGN disclosed herein can comprise an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to any one of SEQ ID NOs: 2-16, 182-196, and 271-285 and can comprise a BH domain comprising amino acid residues 43 to 79 or that differs from amino acid residues 43 to 79 of the RGN sequence (i.e. SEQ ID NOs: 2-16, 182-196, and 271-285) by 1, 2, or 3 amino acid residues.

[0053] An active variant or fragment of a variant LPG10145 RGN disclosed herein can comprise a RECI domain comprising amino acid residues 80 to 236 or that differs from amino acid residues 80 to 236 by 1, 2, or 3 amino acid residues, wherein the amino acid positions are in reference to SEQ ID NO: 1 or to a variant LPG10145 RGN, or an active variant or fragment thereof, disclosed herein (e.g., any one of SEQ ID NOs: 2- 16, 182-196, and 271-285). For example, an active variant or fragment of a variant LPG10145 RGN disclosed herein can comprise an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to any one of SEQ ID NOs: 2-16, 182-196, and 271-285 and can comprise a RECI domain comprising amino acid residues 80 to 236 or that differs from amino acid residues 80 to 236 of the RGN sequence (i.e. SEQ ID NOs: 2-16, 182-196, and 271-285) by 1, 2, or 3 amino acid residues.

[0054] An active variant or fragment of a variant LPG10145 RGN disclosed herein can comprise a REC2 domain comprising amino acid residues 237 to 476 or that differs from amino acid residues 237 to 476 by 1, 2, or 3 amino acid residues, wherein the amino acid positions are in reference to SEQ ID NO: 1 or to a variant LPG10145 RGN, or an active variant or fragment thereof, disclosed herein (e.g., any one of SEQ ID NOs: 2-16, 182-196, and 271-285). For example, an active variant or fragment of a variant LPG10145 RGN disclosed herein can comprise an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to any one of SEQ ID NOs: 2-16, 182-196, and 271-285 and can comprise a REC2 domain comprising amino acid residues 237 to 476 or that differs from amino acid residues 237 to 476 of the RGN sequence (i.e. SEQ ID NOs: 2-16, 182-196, and 271-285) by 1, 2, or 3 amino acid residues.

[0055] An active variant or fragment of a variant LPG10145 RGN disclosed herein can comprise a RuvCII domain comprising amino acid residues 477 to 524 or that differs from amino acid residues 477 to 524 by 1, 2, or 3 amino acid residues, wherein the amino acid positions are in reference to SEQ ID NO: 1 or to a variant LPG10145 RGN, or an active variant or fragment thereof, disclosed herein (e.g., any one of SEQ ID NOs: 2-16, 182-196, and 271-285). For example, an active variant or fragment of a variant LPG10145 RGN disclosed herein can comprise an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to any one of SEQ ID NOs: 2-16, 182-196, and 271-285 and can comprise a RuvCII domain comprising amino acid residues 477 to 524 or that differs from amino acid residues 477 to 524 of the RGN sequence (i.e. SEQ ID NOs: 2-16, 182-196, and 271-285) by 1, 2, or 3 amino acid residues.

[0056] An active variant or fragment of a variant LPG10145 RGN disclosed herein can comprise an LI domain comprising amino acid residues 525 to 560 or that differs from amino acid residues 525 to 560 by 1, 2, or 3 amino acid residues, wherein the amino acid positions are in reference to SEQ ID NO: 1 or to a variant LPG10145 RGN, or an active variant or fragment thereof, disclosed herein (e.g., any one of SEQ ID NOs: 2-16, 182-196, and 271-285). For example, an active variant or fragment of a variant LPG10145 RGN disclosed herein can comprise an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to any one of SEQ ID NOs: 2-16, 182-196, and 271-285 and can comprise an LI domain comprising amino acid residues 525 to 560 or that differs from amino acid residues 525 to 560 of the RGN sequence (i.e. SEQ ID NOs: 2-16, 182-196, and 271-285) by 1, 2, or 3 amino acid residues.

[0057] An active variant or fragment of a variant LPG10145 RGN disclosed herein can comprise an HNH domain comprising amino acid residues 561 to 676 or that differs from amino acid residues 561 to 676 by 1, 2, or 3 amino acid residues, wherein the amino acid positions are in reference to SEQ ID NO: 1 or to a variant LPG10145 RGN, or an active variant or fragment thereof, disclosed herein (e.g., any one of SEQ ID NOs: 2-16, 182-196, and 271-285). For example, an active variant or fragment of a variant LPG10145 RGN disclosed herein can comprise an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to any one of SEQ ID NOs: 2-16, 182-196, and 271-285 and can comprise an HNH domain comprising amino acid residues 561 to 676 or that differs from amino acid residues 561 to 676 of the RGN sequence (i.e. SEQ ID NOs: 2-16, 182-196, and 271-285) by 1, 2, or 3 amino acid residues.

[0058] An active variant or fragment of a variant LPG10145 RGN disclosed herein can comprise an L2 domain comprising amino acid residues 677 to 690 or that differs from amino acid residues 677 to 690 by 1, 2, or 3 amino acid residues, wherein the amino acid positions are in reference to SEQ ID NO: 1 or to a variant LPG10145 RGN, or an active variant or fragment thereof, disclosed herein (e.g., any one of SEQ ID NOs: 2-16, 182-196, and 271-285). For example, an active variant or fragment of a variant LPG10145 RGN disclosed herein can comprise an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to any one of SEQ ID NOs: 2-16, 182-196, and 271-285 and can comprise an L2 domain comprising amino acid residues 677 to 690 or that differs from amino acid residues 677 to 690 of the RGN sequence (i.e. SEQ ID NOs: 2-16, 182-196, and 271-285) by 1, 2, or 3 amino acid residues.

[0059] An active variant or fragment of a variant LPG10145 RGN disclosed herein can comprise a RuvCIII domain comprising amino acid residues 691 to 828 or that differs from amino acid residues 691 to 828 by 1, 2, or 3 amino acid residues, wherein the amino acid positions are in reference to SEQ ID NO: 1 or to a variant LPG10145 RGN, or an active variant or fragment thereof, disclosed herein (e.g., any one of SEQ ID NOs: 2-16, 182-196, and 271-285). For example, an active variant or fragment of a variant LPG10145 RGN disclosed herein can comprise an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to any one of SEQ ID NOs: 2-16, 182-196, and 271-285 and can comprise a RuvCIII domain comprising amino acid residues 691 to 828 or that differs from amino acid residues 691 to 828 of the RGN sequence (i.e. SEQ ID NOs: 2-16, 182-196, and 271-285) by 1, 2, or 3 amino acid residues.

[0060] An active variant or fragment of a variant LPG10145 RGN disclosed herein can comprise a WED domain comprising amino acid residues 829 to 976 or that differs from amino acid residues 829 to 976 by 1, 2, or 3 amino acid residues, wherein the amino acid positions are in reference to SEQ ID NO: 1 or to a variant LPG10145 RGN, or an active variant or fragment thereof, disclosed herein (e.g., any one of SEQ ID NOs: 2-16, 182-196, and 271-285). For example, an active variant or fragment of a variant LPG10145 RGN disclosed herein can comprise an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to any one of SEQ ID NOs: 2-16, 182-196, and 271-285 and can comprise a WED domain comprising amino acid residues 829 to 976 or that differs from amino acid residues 829 to 976 of the RGN sequence (i.e. SEQ ID NOs: 2-16, 182-196, and 271-285) by 1, 2, or 3 amino acid residues.

[0061] An active variant or fragment of a variant LPG10145 RGN disclosed herein can comprise a PI domain comprising amino acid residues 977 to 1130 or that differs from amino acid residues 977 to 1130 by

[0062] 1, 2, or 3 amino acid residues, wherein the amino acid positions are in reference to SEQ ID NO: 1 or to a variant LPG10145 RGN, or an active variant or fragment thereof, disclosed herein (e.g., any one of SEQ ID NOs: 2-16, 182-196, and 271-285). For example, an active variant or fragment of a variant LPG10145 RGN disclosed herein can comprise an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to any one of SEQ ID NOs: 2-16, 182-196, and 271-285 and can comprise a PI domain comprising amino acid residues 977 to 1130 or that differs from amino acid residues 977 to 1130 of the RGN sequence (i.e. SEQ ID NOs: 2-16, 182-196, and 271-285) by 1,

[0063] 2, or 3 amino acid residues.

[0064] In some embodiments, the present disclosure provides an active variant or fragment of a variant LPG10145 RGN that has increased nuclease activity as compared to the RGN of SEQ ID NO: 1. In some embodiments, an active variant or fragment of a variant LPG10145 RGN disclosed herein comprising an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to any one of SEQ ID NOs: 2-16, 182-196, and 271-285 has increased nuclease activity as compared to the RGN of SEQ ID NO: 1. In some embodiments, an active variant or fragment of a variant LPG10145 RGN disclosed herein comprising the amino acid sequence of any one of SEQ ID NOs: 2-16, 182-196, and 271-285 has increased nuclease activity as compared to the RGN of SEQ ID NO: 1. In some embodiments, the present disclosure provides an active variant or fragment of a variant LPG10145 RGN that has nuclease activity that is from about 80% to about 500%, from about 80% to about 200%, or from about 90% to about 150%, or from about 95% to about 120%, or is at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 100%, at least 101%, at least 102%, at least 103%, at least 104%, at least 105%, at least 106%, at least 107%, at least 108%, at least 109%, at least 110%, at least 111%, at least 112%, at least 113%, at least 114%, at least 115%, at least 116%, at least 117%, at least 118%, at least 119%, at least 120%, at least 125%, at least 130%, at least 135%, at least 140%, at least 145%, at least 150%, at least 160%, at least 170%, at least 180%, at least 190%, at least 200%, at least 250%, at least 300%, at least 350%, at least 400%, at least 450%, at least 500%, or more, of the nuclease activity of a reference LPG10145 RGN. In some embodiments, a reference LPG10145 RGN has the amino acid sequence set forth as SEQ ID NO: 1. In some embodiments, a reference LPG10145 RGN is a variant of SEQ ID NO: 1 that lacks the corresponding mutations of the active variant or fragment thereof that has nuclease activity, e.g., as described above. In some embodiments, a reference LPG10145 RGN is a non-identical variant LPG10145 RGN. In some embodiments of this aspect, the active variant of a variant LPG10145 RGN can comprise an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 1, wherein the amino acid residue at one or more of the following amino acid positions differs from the corresponding amino acid residue in SEQ ID NO: 1: 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975. In some embodiments, an active variant of a variant LPG10145 RGN can comprise an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 1, wherein the amino acid residue at one or more of the following amino acid positions is a positively charged amino acid residue: 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975. In some embodiments, an active variant of a variant LPG10145 RGN can comprise an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 1, wherein the amino acid residue at one or more of the following amino acid positions is an R: 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975. In some embodiments, the present disclosure provides an active variant or fragment of a variant LPG10145 RGN that has increased nuclease activity as compared to the RGN of SEQ ID NO: 1. In some embodiments, an active variant or fragment of a variant LPG10145 RGN disclosed herein comprising an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to any one of SEQ ID NOs: 2-16, 182-196, and 271-285 has increased nuclease activity as compared to the RGN of SEQ ID NO: 1. In some embodiments, an active variant or fragment of a variant LPG10145 RGN disclosed herein comprising the amino acid sequence of any one of SEQ ID NOs: 2-16, 182-196, and 271-285 has increased nuclease activity as compared to the RGN of SEQ ID NO: 1. In some embodiments of the aspect wherein an active variant or fragment of a variant LPG10145 RGN disclosed herein can have nuclease activity that is from about 80% to about 500%, from about 80% to about 200%, or from about 90% to about 150%, or from about 95% to about 120%, or is at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 100%, at least 101%, at least 102%, at least 103%, at least 104%, at least 105%, at least 106%, at least 107%, at least 108%, at least 109%, at least 110%, at least 111%, at least

[0065] 112%, at least 113%, at least 114%, at least 115%, at least 116%, at least 117%, at least 118%, at least

[0066] 119%, at least 120%, at least 125%, at least 130%, at least 135%, at least 140%, at least 145%, at least

[0067] 150%, at least 160%, at least 170%, at least 180%, at least 190%, at least 200%, at least 250%, at least

[0068] 300%, at least 350%, at least 400%, at least 450%, at least 500%, or more, of the nuclease activity of a reference LPG10145 RGN, the active variant of a variant LPG10145 RGN can comprise an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 1, wherein the amino acid residue at amino acid position 778 and / or 969 differs from the corresponding amino acid residue in SEQ ID NO: 1. In some embodiments, an active variant of a variant LPG10145 RGN can comprise an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 1, wherein the amino acid residue at amino acid position 778 and / or 969 is a positively charged amino acid residue. In some embodiments, an active variant of a variant LPG10145 RGN can comprise an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 1, wherein the amino acid residue at amino acid position 778 and / or 969 is an R. In some embodiments, an active variant of a variant LPG10145 RGN comprises an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 1, wherein amino acid residues at positions 778 and 856 differ from the corresponding amino acid residues in SEQ ID NO: 1. In some embodiments, an active variant of a variant LPG10145 RGN comprises an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 1, wherein the amino acid residues at positions 778 and 856 are positively charged amino acid residues. In some embodiments, an active variant of a variant LPG10145 RGN comprises an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 1, wherein amino acid residues at positions 55, 647, 778, and 969 differ from the corresponding amino acid residues in SEQ ID NO: 1. In some embodiments, an active variant of a variant LPG10145 RGN comprises an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 1, wherein amino acid residues at positions 55, 647, 778, and 969 are positively charged amino acid residues.

[0069] In some embodiments, the present disclosure provides an active variant or fragment of a variant LPG10145 RGN that has increased nuclease activity as compared to the RGN of SEQ ID NO: 1. In some embodiments, an active variant or fragment of a variant LPG10145 RGN disclosed herein comprising an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to any one of SEQ ID NOs: 2-16, 182-196, and 271-285 has increased nuclease activity as compared to the RGN of SEQ ID NO: 1. In some embodiments, an active variant or fragment of a variant LPG10145 RGN disclosed herein comprising the amino acid sequence of any one of SEQ ID NOs: 2-16, 182-196, and 271-285 has increased nuclease activity as compared to the RGN of SEQ ID NO: 1. In some embodiments of the aspect wherein an active variant or fragment of a variant LPG10145 RGN disclosed herein can have nuclease activity that is from about 80% to about 500%, from about 80% to about 200%, or from about 90% to about 150%, or from about 95% to about 120%, or is at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 100%, at least 101%, at least 102%, at least 103%, at least 104%, at least 105%, at least 106%, at least 107%, at least 108%, at least 109%, at least 110%, at least 111%, at least

[0070] 112%, at least 113%, at least 114%, at least 115%, at least 116%, at least 117%, at least 118%, at least

[0071] 119%, at least 120%, at least 125%, at least 130%, at least 135%, at least 140%, at least 145%, at least

[0072] 150%, at least 160%, at least 170%, at least 180%, at least 190%, at least 200%, at least 250%, at least

[0073] 300%, at least 350%, at least 400%, at least 450%, at least 500%, or more, of the nuclease activity of a reference LPG10145 RGN, the active variant of a variant LPG10145 RGN can comprise an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to: (a) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 52 is an R; (b) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 55 is an R; (c) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 86 is an R; (d) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 472 is an R; (e) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 533 is an R; (f) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 541 is an R; (g) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 643 is substituted by an R; (h) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 647 is an R; (i) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 653 is an R; (j) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 745 is an R; (k) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position 774 is an R; (1) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 778 is an R; (m) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 780 is an R; (n) the amino acid sequence set forth as SEQ ID NO: 1, wherein A at amino acid position 795 is an R; (o) the amino acid sequence set forth as SEQ ID NO: 1, wherein Q at amino acid position 822 is an R; (p) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 843 is an R; (q) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 856 is an R; (r) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 871 is an R; (s) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 872 is an R; (t) the amino acid sequence set forth as SEQ ID NO: 1, wherein D at amino acid position 900 is an R; (u) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position 911 is an R; (v) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 913 is an R; (w) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position 954 is an R; (x) the amino acid sequence set forth as SEQ ID NO: 1, wherein V at amino acid position 958 is an R; (y) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 968 is an R; (z) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 969 is an R; (aa) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 973 is an R; (bb) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 974 is an R; (cc) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position 975 is an R; (dd) the amino acid sequence set forth as SEQ ID NO: 2; (ee) the amino acid sequence set forth as SEQ ID NO: 3; (ff) the amino acid sequence set forth as SEQ ID NO: 4; (gg) the amino acid sequence set forth as SEQ ID NO: 5; (hh) the amino acid sequence set forth as SEQ ID NO: 6; (ii) the amino acid sequence set forth as SEQ ID NO: 7; (jj) the amino acid sequence set forth as SEQ ID NO: 8; (kk) the amino acid sequence set forth as SEQ ID NO: 9; (11) the amino acid sequence set forth as SEQ ID NO: 10; (mm) the amino acid sequence set forth as SEQ ID NO: 11; (nn) the amino acid sequence set forth as SEQ ID NO: 12; (oo) the amino acid sequence set forth as SEQ ID NO: 13; (pp) the amino acid sequence set forth as SEQ ID NO: 14; (qq) the amino acid sequence set forth as SEQ ID NO: 15; and (rr) the amino acid sequence set forth as SEQ ID NO: 16.

[0074] In some embodiments, the present disclosure provides an active variant or fragment of a variant LPG10145 RGN that has increased nuclease activity as compared to the RGN of SEQ ID NO: 1. In some embodiments, an active variant or fragment of a variant LPG10145 RGN disclosed herein comprising an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to any one of SEQ ID NOs: 2-16, 182-196, and 271-285 has increased nuclease activity as compared to the RGN of SEQ ID NO: 1. In some embodiments, an active variant or fragment of a variant LPG10145 RGN disclosed herein comprising the amino acid sequence of any one of SEQ ID NOs: 2-16, 182-196, and 271-285 has increased nuclease activity as compared to the RGN of SEQ ID NO: 1. In some embodiments of the aspect wherein an active variant or fragment of a variant LPG10145 RGN disclosed herein can have nuclease activity that is from about 80% to about 500%, from about 80% to about 200%, or from about 90% to about 150%, or from about 95% to about 120%, or is at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 100%, at least 101%, at least 102%, at least 103%, at least 104%, at least 105%, at least 106%, at least 107%, at least 108%, at least 109%, at least 110%, at least 111%, at least

[0075] 112%, at least 113%, at least 114%, at least 115%, at least 116%, at least 117%, at least 118%, at least

[0076] 119%, at least 120%, at least 125%, at least 130%, at least 135%, at least 140%, at least 145%, at least

[0077] 150%, at least 160%, at least 170%, at least 180%, at least 190%, at least 200%, at least 250%, at least

[0078] 300%, at least 350%, at least 400%, at least 450%, at least 500%, or more, of the nuclease activity of a reference LPG10145 RGN, the active variant of a variant LPG10145 RGN can comprise an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to: (a) the amino acid sequence set forth as SEQ ID NO: 2; (b) the amino acid sequence set forth as SEQ ID NO: 3; (c) the amino acid sequence set forth as SEQ ID NO: 4; (d) the amino acid sequence set forth as SEQ ID NO: 5; (e) the amino acid sequence set forth as SEQ ID NO: 6; (f) the amino acid sequence set forth as SEQ ID NO: 7; (g) the amino acid sequence set forth as SEQ ID NO: 8SEQ ID NO: 8; (h) the amino acid sequence set forth as SEQ ID NO: 9; (i) the amino acid sequence set forth as SEQ ID NO: 10; (j) the amino acid sequence set forth as SEQ ID NO: 11; (k) the amino acid sequence set forth as SEQ ID NO: 12; (1) the amino acid sequence set forth as SEQ ID NO: 13; (m) the amino acid sequence set forth as SEQ ID NO: 14; (n) the amino acid sequence set forth as SEQ ID NO: 15; and (o) the amino acid sequence set forth as SEQ ID NO: 16.

[0079] RNA-guided nucleases provided herein can in embodiments comprise at least one nuclease domain (e.g., DNase, RNase domain) and at least one RNA recognition and / or RNA binding domain to interact with guide RNAs. In some embodiments, the RGN comprises only one active nuclease domain and thus functions as a nickase. The RGN nuclease domain that is active in an RGN nickase can be a RuvC domain or an HNH domain. An RGN nickase can comprise an inactivated HNH nuclease domain or can lack an HNH domain. Further domains that can be found in RNA-guided nucleases provided herein include, but are not limited to: DNA binding domains, helicase domains, protein-protein interaction domains, and dimerization domains. In specific embodiments, the RNA-guided nucleases provided herein can comprise at least 70%, 75%, 80%,

[0080] 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to one or more of a DNA binding domain, helicase domain, protein-protein interaction domain, and dimerization domain. In some embodiments, variant LPG10145 RGNs of the disclosure do not comprise amino acid residues in the PAM recognition domain, the REC 1 domain, the REC2 domain, and / or the Wing domain that differ from the corresponding residues in the wild-type LPG10145 RGN (e.g., SEQ ID NO: 1).

[0081] A target nucleotide sequence is bound by an RNA-guided nuclease provided herein and hybridizes with the guide RNA associated with the RNA-guided nuclease. The target sequence can then be subsequently cleaved by the RNA-guided nuclease if the polypeptide possesses nuclease activity. The terms “cleave” or “cleavage” refer to the hydrolysis of at least one phosphodiester bond within the backbone of a target nucleotide sequence that can result in either single-stranded or double-stranded breaks within the target sequence. The presently disclosed RGNs can cleave nucleotides within a polynucleotide, functioning as an endonuclease or can be an exonuclease, removing successive nucleotides from the end (the 5' and / or the 3' end) of a polynucleotide. In some embodiments, the disclosed RGNs can cleave nucleotides of a target sequence within any position of a polynucleotide and thus function as both an endonuclease and exonuclease. The cleavage of a target polynucleotide by the presently disclosed RGNs can result in staggered breaks or blunt ends. A staggered cut in a polynucleotide leads to two sticky ends or overhanging ends, and is formed when the nuclease cuts each strand of a polynucleotide such that the cuts are not directly opposite each other. For each sticky end of the cut polynucleotide, one strand (i.e. the overhanging strand) is longer than the other (typically by at least a few nucleotides), such that the longer strand has bases which are left unpaired. The longer strand of an overhanging end of a cleaved polynucleotide can have one unpaired nucleotide, two unpaired nucleotides, 3 unpaired nucleotides, 4 unpaired nucleotides, 5 unpaired nucleotides, or more unpaired nucleotides. In some embodiments, the longer strand of an overhanging end of a cleaved polynucleotide can have one unpaired nucleotide. The overhanging end of a cleaved polynucleotide can be a 3' overhang or a 5' overhang. In some embodiments, the overhanging end of a cleaved polynucleotide is a 3' overhang. In some embodiments, the overhanging end of a cleaved polynucleotide is a 5' overhang. In some embodiments, a variant LPG10145 RGN, or an active variant or fragment thereof, of the disclosure cleaves a target polynucleotide to form a staggered cut, wherein the staggered cut creates a 3' overhang with one unpaired nucleotide. By contrast, a blunt cut generates two blunt ends, such that each blunt end of the cut polynucleotide has both strands that are of equal length - i.e. there are no unpaired bases on either strand of a blunt end.

[0082] The presently disclosed RNA-guided nucleases are variants of wild-type polypeptides. The wildtype RGN can be modified to alter nuclease activity or alter PAM specificity, for example. In some embodiments, the RNA-guided nuclease is not naturally-occurring.

[0083] In certain embodiments, the RNA-guided nuclease functions as a nickase, only cleaving a single strand of the target nucleotide sequence. Such RNA-guided nucleases have a single functioning nuclease domain. In particular embodiments, the nickase is capable of cleaving the positive strand or negative strand. In some of these embodiments, additional nuclease domains have been mutated such that the nuclease activity is reduced or eliminated.

[0084] In other embodiments, the RNA-guided nuclease lacks nuclease activity altogether and is referred to herein as nuclease-dead or nuclease inactive. Any method known in the art for introducing mutations into an amino acid sequence, such as PCR-mediated mutagenesis and site-directed mutagenesis, can be used for generating nickases or nuclease-dead RGNs. See, e.g., U.S. Publ. No. 2014 / 0068797 and U.S. Pat. No. 9,790,490; each of which is incorporated by reference in its entirety. A non-limiting example of a nickase is the UPG10145 D16A nickase, which is set forth herein as SEQ ID NO: 124. Nickases which comprise a mutation in the RuvC domain and have a functional UNH domain are useful in base editing wherein the nickase is fused to a base editing polypeptide such as a deaminase. Another non-limiting example of a nickase is the EPG10145 H611A nickase, which is set forth herein as SEQ ID NO: 125. Nickases which comprise a mutation in the HNH domain and have a functional RuvC domain are useful in polymerase (i.e. prime) editing wherein the nickase is fused to a polymerase (i.e. prime) editing polypeptide such as a reverse transcriptase. A non-limiting example of a nuclease dead RGN is the LPG10145 D16A H611A sequence set forth as SEQ ID NO: 126. Variant LPG10145 RGNs of the disclosure can thus additionally comprise mutations in one or more nuclease domains that confer nickase function to the variant LPG10145 RGNs. In some embodiments, a variant LPG10145 RGN of the disclosure comprises an alanine (or another nonconserved amino acid residue) at a position corresponding to 634 of SEQ ID NO: 1.

[0085] In some embodiments, variant LPG10145 RGNs of the disclosure that comprise the H611A mutation in the HNH domain comprises an amino acid sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity to any one of SEQ ID NOs: 182- 196. In some embodiments, variant LPG10145 RGNs of the disclosure that comprise the D16A mutation comprises the amino acid sequence set forth as any one of SEQ ID NOs: 271-285. In some embodiments, variant LPG10145 RGNs of the disclosure that comprise the D16A mutation comprises an amino acid sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity to any one of SEQ ID NOs: 271-285. In some embodiments, variant LPG10145 RGNs of the disclosure that comprise the D16A mutation comprises the amino acid sequence set forth as any one of SEQ ID NOs: 271-285.

[0086] RNA-guided nucleases that lack nuclease activity can be used to deliver a fused polypeptide, polynucleotide, or small molecule payload to a particular genomic location. In some of these embodiments, the RGN polypeptide or guide RNA can be fused to a detectable label to allow for detection of a particular sequence. As a non-limiting example, a nuclease-dead RGN can be fused to a detectable label (e.g., fluorescent protein) and targeted to a particular sequence associated with a disease to allow for detection of the disease-associated sequence.

[0087] Alternatively, nuclease-dead RGNs can be targeted to particular genomic locations to alter the expression of a desired sequence. In some embodiments, the binding of a nuclease-dead RNA-guided nuclease to a target sequence results in the reduction in expression of the target sequence or a gene under transcriptional control by the target sequence by interfering with the binding of RNA polymerase or transcription factors within the targeted genomic region. In some embodiments, the RGN (e.g., a nuclease- dead RGN) or its complexed guide RNA further comprises an expression modulator that, upon binding to a target sequence, serves to either repress or activate the expression of the target sequence or a gene under transcriptional control by the target sequence. In some of these embodiments, the expression modulator modulates the expression of the target sequence or regulated gene through epigenetic mechanisms.

[0088] In other embodiments, the nuclease-dead RGNs or an RGN with nickase activity can be targeted to particular genomic locations to modify the sequence of a target polynucleotide through fusion to a baseediting polypeptide, for example a deaminase polypeptide or active variant or fragment thereof, that directly chemically modifies (e.g., deaminates) a nucleobase, resulting in conversion from one nucleobase to another. The base-editing polypeptide can be fused to the RGN at its N-terminal or C-terminal end. Additionally, the base-editing polypeptide may be fused to the RGN via a peptide linker. A non-limiting example of a deaminase polypeptide that is useful for such compositions and methods includes a cytosine deaminase or an adenine deaminase (such as the adenine deaminase base editor described in Gaudelli et al. (2017) Nature 551:464-471, U.S. Publ. Nos. 2017 / 0121693 and 2018 / 0073012, and International Publ. No. WO 2018 / 027078, or any of the deaminases disclosed in International Publ. No. WO 2020 / 139783, International Publ. No. WO 2022 / 056254, International Publ. No. WO 2022 / 204093, and International Publ. No. WO 2024 / 095245, each of which is herein incorporated by reference in its entirety). In some embodiments, the deaminase polypeptide that is useful for such compositions and methods is a cytosine deaminase or an adenine deaminase comprising an amino acid sequence selected from any one of SEQ ID NOs: 42-113, and 257. In one embodiment, the deaminase polypeptide that is useful for such compositions and methods is a cytosine deaminase or an adenine deaminase having a sequence that is at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater identity to any one of the amino acid sequences set forth as SEQ ID NOs: 42-113, and 257. In some embodiments, the deaminase polypeptide that is useful for such the presently disclosed compositions and methods is a deaminase disclosed in Table 17 of International Publ. No. WO 2020 / 139783, which is incorporated herein by reference in its entirety. Further, it is known in the art that certain fusion proteins between an RGN and a base-editing enzyme (e.g., cytosine deaminase) may also comprise at least one uracil stabilizing polypeptide that increases the mutation rate of a cytidine, deoxycytidine, or cytosine to a thymidine, deoxythymidine, or thymine in a nucleic acid molecule by a deaminase. Non-limiting examples of uracil stabilizing polypeptides include those disclosed in International Publ. No. WO 2022 / 015969, which is herein incorporated by reference in its entirety, including USP2 (SEQ ID NO: 40), and a uracil glycosylase inhibitor (UGI) domain (SEQ ID NO: 35), which may increase base editing efficiency. Therefore, a fusion protein may comprise an RGN described herein or variant thereof, a deaminase, and optionally at least one uracil stabilizing polypeptide, such as UGI or USP2. In certain embodiments, the RGN that is fused to the base-editing polypeptide is a nickase that cleaves the DNA strand that is not acted upon by the base-editing polypeptide (e.g., deaminase).

[0089] A variant LPG10145 RGN, or active variant or fragment thereof, of the disclosure can comprise a protospacer adjacent motif (PAM)-interacting (PI) domain that contributes to recognition of a PAM site in a target polynucleotide. In some embodiments, a variant LPG10145 RGN, or active variant or fragment thereof, recognizes and binds a consensus nucleotide sequence set forth as NNGG adjacent to the target polynucleotide. The PI domain can comprise 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150 or more amino acid residues. In some embodiments, the PI domain of an RGN, or an active variant or fragment thereof, of the disclosure is located within the carboxy (C)-terminal region of the RGN. The C-terminal region comprising the PI domain of an RGN, or an active variant or fragment thereof, of the disclosure can include the C-terminal 151 amino acid residues, the C-terminal 150 amino acid residues, the C-terminal 140 amino acid residues, the C-terminal 135 amino acid residues, the C-terminal 132 amino acid residues, the C-terminal 130 amino acid residues, the C-terminal 125 amino acid residues, the C-terminal 120 amino acid residues, the C-terminal 110 amino acid residues, the C-terminal 100 amino acid residues, the C-terminal 90 amino acid residues, the C-terminal 80 amino acid residues, the C-terminal 70 amino acid residues, the C-terminal 60 amino acid residues, the C-terminal 50 amino acid residues, the C-terminal 40 amino acid residues, the C-terminal 30 amino acid residues, the C- terminal 20 amino acid residues, or the C-terminal 10 amino acid residues of the RGN. In some embodiments, the PI domain of an RGN, or an active variant or fragment thereof, of the disclosure is within or includes amino acid residues 977-1130 of the RGN. In some embodiments, the PI domain of a variant LPG10145 RGN polypeptide, or an active variant or fragment thereof, of the disclosure has the amino acid sequence set forth as SEQ ID NO: 253.

[0090] III. Fusion Proteins Comprising variant LPG10145 RGN polypeptides

[0091] The present disclosure provides fusion proteins comprising the presently disclosed engineered variant LPG10145 RGN polypeptides, or active variants or fragments thereof, operably fused to at least one heterologous polypeptide, as well as polynucleotides encoding the fusion proteins. When used to refer to the joining of two protein coding regions (either by fusion or insertion), by “operably linked” or “operably fused” is intended that the coding regions are in the same reading frame, even if one is inserted into another. In some embodiments, polypeptides that are “operably fused” or “operably linked” means that the structure and / or biological activity of each individual peptide is also present in the fusion.

[0092] As used herein, “heterologous”, in reference to a polypeptide that is heterologous to another polypeptide (e.g., a variant LPG10145 RGN), is a polypeptide that is not operably fused to the presently described variant LPG10145 RGN in nature. The heterologous polypeptide can originate from a foreign species or from the same species. The heterologous polypeptide can be in its native form or is substantially modified from its native form in composition and / or genomic locus by deliberate human intervention. The heterologous polypeptide can be any polypeptide, including but not limited to a localization signal, a cellpenetrating domain, a polymerase (e.g., reverse transcriptase), a base editing polypeptide, an effector domain (e.g., a cleavage domain, a deaminase domain, or an expression modulator domain), a detectable label (e.g., fluorescent protein), or a purification tag, or active variant or fragment thereof. The heterologous polypeptide can be operably fused to the amino (N)-terminus, to the carboxy (C)-terminus, or to an internal location of a variant LPG10145 RGN polypeptide, or active variant or fragment thereof, described herein. In some embodiments, the heterologous polypeptide is operably fused to a variant LPG10145 RGN polypeptide by a peptide linker as described herein.

[0093] In some embodiments, a variant LPG10145 RGN polypeptide, or active variant or fragment thereof, of the disclosure may be fused to a polymerase (i.e. prime) editing polypeptide (e.g., DNA polymerase or reverse transcriptase) to generate a polymerase editor (PE). The polymerase editing polypeptide (e.g., DNA polymerase or reverse transcriptase) can be operably fused to the amino (N)-terminus, to the carboxy (C)- terminus, or to an internal location of a variant LPG10145 RGN polypeptide, or active variant or fragment thereof, described herein. In some embodiments, the polymerase editing polypeptide (e.g., DNA polymerase or reverse transcriptase) is operably fused to a variant LPG10145 RGN polypeptide by a peptide linker as described herein. Polymerase (i.e.p) editing is a versatile and precise genome editing method that directly writes new genetic information into a specified DNA site using a nucleic acid programmable DNA binding protein working in association with a polymerase. The polymerase (i.e. prime) editing system uses an RGN that is a nickase, and the system is programmed with a polymerase (i.e. prime) editing (PE) guide RNA (“PEgRNA”). The PEgRNA is a guide RNA that both specifies the target sequence and provides the template for polymerization of the replacement strand containing the edit by way of an extension engineered onto the guide RNA (e.g., at the 5' or 3' end, or at an internal portion of the guide RNA). The RGN nickase / polymerase (i.e. prime) editing polypeptide fusion is guided to the target sequence by the PEgRNA and nicks the target strand upstream of sequence to be edited and upstream of the PAM, creating a 3' flap on the target strand. The PEgRNA includes a primer binding site (PBS) that is complementary to the 3' flap of the target strand. In some embodiments, a PBS is at least about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides in length. In certain embodiments, the pegRNA comprises a PBS that is at least 5 (e.g., at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 28, 19, or 20) nucleotides in length. In some embodiments, the pegRNA may comprise a PBS that is at least 8 nucleotides in length. Hybridrization of the PBS and 3' flap of the target strand allows polymerization of the replacement strand containing the edit using the extension of the PEgRNA as template. The extension of the PEgRNA can be formed from RNA or DNA. In the case of an RNA extension, the polymerase of the polymerase (i.e. prime) editor can be an RNA-dependent DNA polymerase (such as a reverse transcriptase). In the case of a DNA extension, the polymerase of the polymerase (i.e. prime) editor may be a DNA-dependent DNA polymerase.

[0094] The replacement strand containing the desired edit (e.g., a single nucleobase substitution) shares the same sequence as the target strand of the target sequence to be edited (with the exception that it includes the desired edit). Through DNA repair and / or replication machinery, the target strand of the target sequence is replaced by the newly synthesized replacement strand containing the desired edit. In some cases, polymerase (i.e. prime) editing may be thought of as a “search-and-replace” genome editing technology since the polymerase (i.e. prime) editors not only search and locate the desired target sequence to be edited, but at the same time, encode a replacement strand containing a desired edit which is installed in place of the corresponding target strand of the target sequence. Thus, in some embodiments, a guide RNA of the disclosure comprises an extension comprising an edit template for polymerase (i.e. prime) editing. In some embodiments, a polymerase (i.e. prime) editing polypeptide that can be fused to an RGN includes a DNA polymerase. In certain embodiments, the DNA polymerase is a reverse transcriptase. In certain embodiments, the RGN is a nickase. A fusion protein comprising a variant LPG10145 RGN polypeptide, or active variant or fragment thereof, disclosed herein and a reverse transcriptase can comprise an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or more sequence identity to any one of SEQ ID NOs: 135, and 162-181. In some embodiments, a fusion protein comprising a variant LPG10145 RGN polypeptide, or active variant or fragment thereof, disclosed herein and a reverse transcriptase comprises the amino acid sequence set forth as any one of SEQ ID NOs: 135, and 162-181.

[0095] A variant LPG10145 RGN polypeptide, or active variant or fragment thereof, of the disclosure may be operably fused to a base-editing polypeptide as described herein. The base-editing polypeptide (e.g., deaminase) can be operably fused to the amino (N)-terminus, to the carboxy (C)-terminus, or to an internal location of a variant LPG10145 RGN polypeptide, or active variant or fragment thereof, described herein. The base-editing polypeptide (e.g., deaminase) may be operably fused to the RGN via a peptide linker. In some embodiments, the base editing polypeptide in a fusion protein of the disclosure is a cytosine deaminase or an adenine deaminase comprising an amino acid sequence set forth as any one of SEQ ID NOs: 42-113, and 257. In some embodiments, the base editing polypeptide in a fusion protein of the disclosure is a cytosine deaminase or an adenine deaminase having an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater sequence identity to the amino acid sequence set forth as any one of SEQ ID NOs: 42-113, and 257. In some embodiments, the base editing polypeptide in a fusion protein of the disclosure is a deaminase described in Gaudelli et al. (2017) Nature 551:464-471, U.S. Publ. Nos. 2017 / 0121693 and 2018 / 0073012, and International Publ. No. WO 2018 / 027078, or any of the deaminases disclosed in International Publ. No. WO 2020 / 139783, International Publ. No. WO 2022 / 056254, International Publ. No. WO 2022 / 204093, and International Publ. No. WO 2024 / 095245, each of which is herein incorporated by reference in its entirety. The fusion protein comprising an engineered variant LPG10145 RGN polypeptide and a base-editing polypeptide (e.g., cytosine or adenine deaminase) may further comprise at least one uracil stabilizing polypeptide that increases the mutation rate of a nucleobase (e.g., cytidine, deoxycytidine, or cytosine to a thymidine, deoxythymidine, or thymine) in a nucleic acid molecule by a deaminase. In some embodiments, a USP comprises USP2 (SEQ ID NO: 40) or a uracil glycosylase inhibitor (UGI) domain (SEQ ID NO: 35), which may increase base editing efficiency. In some embodiments, the variant LPG10145 RGN polypeptide that is fused to the base-editing polypeptide is a nickase that cleaves the DNA strand that is not acted upon by the base-editing polypeptide (e.g., deaminase). In some embodiments, the variant LPG10145 RGN polypeptide that is fused to the base-editing polypeptide is a nuclease-dead RGN polypeptide. A fusion protein comprising a variant LPG10145 RGN polypeptide, or active variant or fragment thereof, disclosed herein and a deaminase can comprise an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or more sequence identity to any one of SEQ ID NOs: 258-270. In some embodiments, a fusion protein comprising a variant LPG10145 RGN polypeptide, or active variant or fragment thereof, disclosed herein and a deaminase comprises the amino acid sequence set forth as any one of SEQ ID NOs: 258-270.

[0096] The presently disclosed variant LPG10145 RNA-guided nucleases, or active variants or fragments thereof, can comprise at least one nuclear localization signal (NLS) to enhance transport of the RGN to the nucleus of a cell. Nuclear localization signals are known in the art and generally comprise a stretch of basic amino acids (see, e.g., Lange et al., J. Biol. Chem. (2007) 282:5101-5105). In some embodiments, the RGN comprises 2, 3, 4, 5, 6 or more nuclear localization signals. The nuclear localization signal(s) can be a heterologous NLS. Non-limiting examples of nuclear localization signals useful for the presently disclosed RGNs are the nuclear localization signals of SV40 Large T-antigen, nucleoplasmin, and c-Myc (see, e.g., Ray et al. (2015) Bioconjug Chem 26(6): 1004-7). In some embodiments, the NLS has an amino acid sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or more sequence identity to any one of SEQ ID NOs: 36, 37, 234, or 235. In particular embodiments, the RGN comprises the NLS sequence set forth as SEQ ID NO: 36, 37, 234, or 235. The RGN can comprise one or more NLS sequences operably fused at its N-terminus, C- terminus, or both the N-terminus and C-terminus. For example, the RGN can comprise two NLS sequences at the N-terminal region and four NLS sequences at the C-terminal region. Other localization signal sequences known in the art that localize polypeptides to particular subcellular location(s) can also be used to target the presently disclosed variant LPG10145 RGNs, or active variants or fragments thereof, including, but not limited to, plastid localization sequences, mitochondrial localization sequences, and dual-targeting signal sequences that target to both the plastid and mitochondria (see, e.g., Nassoury and Morse (2005) Biochim Biophys Acta 1743:5-19; Kunze and Berger (2015) Front Physiol dx.doi.org / 10.3389 / fphys.2015.00259; Herrmann and Neupert (2003) IUBMB Life 55:219-225; Soil (2002) Curr Opin Plant Biol 5:529-535; Carrie and Small (2013) Biochim Biophys Acta 1833:253-259; Carrie et al. (2009) FEBSJN16 1187-1195; Silva-Filho (2003) Curr Opin Plant Biol 6:589-595; Peeters and Small (2001) Biochim Biophys Acta 1541:54-63; Murcha t a / . (2014) J Exp Bot 65:6301-6335; Mackenzie (2005) Trends Cell Biol 15:548-554; Glaser et al. (1998) Plant Mol Biol 38:311-338).

[0097] In certain embodiments, the presently disclosed variant LPG10145 RNA-guided nucleases, or active variants or fragments thereof, comprise at least one cell-penetrating domain that facilitates cellular uptake of the RGN. Cell-penetrating domains are known in the art and generally comprise stretches of positively charged amino acid residues (i.e., polycationic cell -penetrating domains), alternating polar amino acid residues and non-polar amino acid residues (i.e., amphipathic cell-penetrating domains), or hydrophobic amino acid residues (i.e., hydrophobic cell-penetrating domains) (see, e.g., Milletti F. (2012) Drug Discov Today 17:850-860). A non-limiting example of a cell -penetrating domain is the trans-activating transcriptional activator (TAT) from the human immunodeficiency virus 1.

[0098] The nuclear localization signal, plastid localization signal, mitochondrial localization signal, dualtargeting localization signal, and / or cell-penetrating domain can be located at the amino-terminus (N- terminus), the carboxyl -terminus (C-terminus), or in an internal location of the RNA-guided nuclease.

[0099] The presently disclosed variant LPG10145 RGNs, or active variants or fragments thereof, can be fused to an effector domain, such as a cleavage domain, a deaminase domain, or an expression modulator domain, either directly or indirectly via a linker peptide. Such a domain can be located at the N-terminus, the C-terminus, or an internal location of the RNA-guided nuclease. In some of these embodiments, the RGN component of the fusion protein is a nuclease-dead RGN or a nickase.

[0100] In some embodiments, the RGN fusion protein comprises a cleavage domain, which is any domain that is capable of cleaving a polynucleotide (i.e., RNA, DNA, or RNA / DNA hybrid) and includes, but is not limited to, restriction endonucleases and homing endonucleases, such as Type IIS endonucleases (e.g., Fokl) (see, e.g., Belfort et al. (1997) Nucleic Acids Res. 25:3379-3388; Linn et al. (eds.) Nucleases, Cold Spring Harbor Laboratory Press, 1993).

[0101] In other embodiments, the RGN fusion protein comprises a deaminase domain that deaminates a nucleobase, resulting in conversion from one nucleobase to another, and includes, but is not limited to, a cytosine deaminase or an adenine deaminase (see, e.g., Gaudelli et al. (2017) Nature 551:464-471, U.S. Publ. Nos. 2017 / 0121693 and 2018 / 0073012, U.S. Patent No. 9,840,699, and International Publ. No. WO / 2018 / 027078, or any of the deaminases disclosed in International Publ. No. WO 2020 / 139783, International Publ. No. WO 2022 / 056254, and International Publ. No. WO 2022 / 204093, each of which is herein incorporated by reference in its entirety).

[0102] In some embodiments, the effector domain of the RGN fusion protein can be an expression modulator domain, which is a domain that either serves to upregulate or downregulate transcription. The expression modulator domain can be an epigenetic modification domain, a transcriptional repressor domain or a transcriptional activation domain.

[0103] In some of these embodiments, the expression modulator of the RGN fusion protein comprises an epigenetic modification domain that covalently modifies DNA or histone proteins to alter histone structure and / or chromosomal structure without altering the DNA sequence, leading to changes in gene expression (z.e., upregulation or downregulation). Non-limiting examples of epigenetic modifications include acetylation or methylation of lysine residues, arginine methylation, serine and threonine phosphorylation, and lysine ubiquitination and sumoylation of histone proteins, and methylation and hydroxymethylation of cytosine residues in DNA. Non-limiting examples of epigenetic modification domains include histone acetyltransferase domains, histone deacetylase domains, histone methyltransferase domains, histone demethylase domains, DNA methyltransferase domains, and DNA demethylase domains.

[0104] In other embodiments, the expression modulator of the fusion protein comprises a transcriptional repressor domain, which interacts with transcriptional control elements and / or transcriptional regulatory proteins, such as RNA polymerases and transcription factors, to reduce or terminate transcription of at least one gene. Transcriptional repressor domains are known in the art and include, but are not limited to, Spl- like repressors, IKB, and Kriippel associated box (KRAB) domains.

[0105] In yet other embodiments, the expression modulator of the fusion protein comprises a transcriptional activation domain, which interacts with transcriptional control elements and / or transcriptional regulatory proteins, such as RNA polymerases and transcription factors, to increase or activate transcription of at least one gene. Transcriptional activation domains are known in the art and include, but are not limited to, a herpes simplex virus VP 16 activation domain and an NF AT activation domain.

[0106] The presently disclosed variant LPG10145 RGN polypeptides, or active variants or fragments thereof, can comprise a detectable label or a purification tag. The detectable label or purification tag can be located at the N-terminus, the C-terminus, or an internal location of the RNA-guided nuclease, either directly or indirectly via a linker peptide. In some of these embodiments, the RGN component of the fusion protein is a nuclease-dead RGN. In other embodiments, the RGN component of the fusion protein is an RGN with nickase activity.

[0107] A detectable label is a molecule that can be visualized or otherwise observed. The detectable label may be fused to the RGN as a fusion protein (e.g., fluorescent protein) or may be a small molecule conjugated to the RGN polypeptide that can be detected visually or by other means. Detectable labels that can be fused to the presently disclosed RGNs as a fusion protein include any detectable protein domain, including but not limited to, a fluorescent protein or a protein domain that can be detected with a specific antibody. Non-limiting examples of fluorescent proteins include green fluorescent proteins (e.g., GFP, EGFP, ZsGreenl) and yellow fluorescent proteins (e.g., YFP, EYFP, ZsYellowl). Non-limiting examples of small molecule detectable labels include radioactive labels, such as3H and35S.

[0108] Variant LPG10145 RGN polypeptides, or active variants or fragments thereof, of the disclosure can also comprise a purification tag, which is any molecule that can be utilized to isolate a protein or fused protein from a mixture (e.g., biological sample, culture medium). Non-limiting examples of purification tags include biotin, myc, maltose binding protein (MBP), glutathione-S-transferase (GST), and 3X FLAG tag.

[0109] Variant LPG10145 RNA-guided nucleases, or active variants or fragments thereof, of the disclosure that are fused to a heterologous polypeptide or domain can be separated or joined by a linker. The term "linker," as used herein, refers to a chemical group or a molecule linking two molecules or moieties, e.g. , a binding domain and a cleavage domain of a nuclease. In some embodiments, a linker joins a gRNA binding domain of an RNA guided nuclease and a base-editing polypeptide, such as a deaminase. In some embodiments, a linker joins a nuclease-dead RGN and a deaminase. Typically, the linker is positioned between, or flanked by, two groups, molecules, or other moieties and connected to each one via a covalent bond, thus connecting the two. In some embodiments, the linker is an organic molecule, group, polymer, or chemical moiety.

[0110] In some embodiments, the linker is an amino acid or a plurality of amino acids (e.g., a peptide or protein). A peptide linker of the disclosure can connect one polypeptide to another in a fusion protein. For example, a peptide linker can connect a variant LPGI0145 RGN polypeptide and a heterologous polypeptide (e.g., a polymerase, a base editing polypeptide, an effector domain). A fusion protein comprising 3 polypeptides (e.g., a variant LPGI0145 RGN polypeptide, a polymerase, and a detectable label) can comprise at least one peptide linker. In some embodiments, a peptide linker comprises at least one NLS. In some embodiments, a peptide linker comprises 2 NLSs. The peptide linker can be operably fused at the N- terminus, the C-terminus, or both the N-terminus and C-terminus of the variant LPGI0145 RGN polypeptide or the heterologous polypeptide (e.g., a polymerase, a base editing polypeptide, an effector domain). In some embodiments, a peptide linker has a formula of -(SGGS)x-NLSm-(SGGS)y-NLSn-(SGGS)z-, wherein each of x, y, or z is 0, 1, 2, 3, or 4; and wherein each of m or n is 0 or 1. In certain embodiments, a peptide linker has a formula of -(SGGS)x-NLSm-(SGGS)y-NLSn-(SGGS)z-, wherein each of x or y is 0, 1, 2, 3, or 4, and y is 0; and wherein one of m or n is 0, and the other is 1. In other embodiments, the peptide linker has a formula of -(SGGS)x-NLSm-(SGGS)y-NLSn-(SGGS)z-, wherein each of x, y, or z is 0, 1, 2, 3, or 4; and wherein each of m or n is 1. In some embodiments, a peptide linker comprises one or more copies of amino acid sequence SGGS (SEQ ID NO: 241). In some embodiments, the peptide linker is 4-100 amino acids in length, for example, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30- 35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids in length. Longer or shorter linkers are also contemplated. In some embodiments, a peptide linker has a length of at least 13 amino acids, including but not limited to about 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, or more amino acids. In some embodiments, the peptide linker has an amino acid sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or more sequence identity to any one of SEQ ID NOs: 236-241. In some embodiments, the peptide linker has the amino acid sequence of any one of SEQ ID NOs: 236-241.

[0111] In some embodiments, the heterologous polypeptide comprises a polymerase (e.g., reverse transcriptase), and the peptide linker between the variant LPG10145 RGN polypeptide and the polymerase (e.g., reverse transcriptase) comprises at least 2 amino acids, at least 3 amino acids, at least 4 amino acids, at least 5 amino acids, at least 6 amino acids, at least 7 amino acids, at least 8 amino acids, at least 9 amino acids, at least 10 amino acids, at least 11 amino acids, at least 12 amino acids, or at least 13 amino acids. In some embodiments, the peptide linker between the variant LPG10145 RGN polypeptide and the polymerase (e.g., reverse transcriptase) comprises at least 13 amino acids, including but not limited to about 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, or more amino acids. In some embodiments, the fusion protein comprises a variant LPG10145 RGN polypeptide connected to a polymerase (e.g., reverse transcriptase) by a peptide linker having an amino acid sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or more sequence identity to any one of SEQ ID NOs: 236-241. In some embodiments, the fusion protein comprises a variant LPG10145 RGN polypeptide connected to a polymerase (e.g., reverse transcriptase) by a peptide linker having the amino acid sequence set forth as any one of SEQ ID NOs: 236-241.

[0112] A fusion protein of the disclosure in some embodiments comprises a presently disclosed variant LPG10145 RGN polypeptide and a presently disclosed heterologous polypeptide (e.g., a polymerase, a base editing polypeptide, an effector domain). In some embodiments, a fusion protein of the disclosure comprises from amino terminus to carboxy terminus: a variant LPG10145 RGN polypeptide and a heterologous polypeptide (e.g., a polymerase, a base editing polypeptide, an effector domain). In some embodiments, a fusion protein of the disclosure comprises from amino terminus to carboxy terminus: a heterologous polypeptide (e.g., a polymerase, a base editing polypeptide, an effector domain) and a variant LPG10145 RGN polypeptide.

[0113] The fusion protein can in some embodiments comprises a heterologous polypeptide (e.g., a polymerase, a base editing polypeptide, an effector domain) inserted within a variant LPG10145 RGN polypeptide. In some embodiments, the heterologous polypeptide (e.g., a polymerase, a base editing polypeptide, an effector domain) is inserted between surface amino acid residues. The heterologous polypeptide (e.g., a polymerase, a base editing polypeptide, an effector domain) can be inserted within or between a linker domain 2, a wedge (WED) domain, a RuvC domain, an HNH domain, a Rec-2 domain, or a PAM-interacting (PI) domain of the RGN as described in more detail elsewhere herein.

[0114] The heterologous polypeptide (e.g., a polymerase, a base editing polypeptide, an effector domain) in some embodiments is inserted within a variant LPG10145 RGN polypeptide immediately after the amino acid position selected from the group consisting of: i) amino acid position corresponding to position 347 of SEQ ID NO: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, or 285; ii) amino acid position corresponding to position 524 of SEQ ID NO: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, or 285; iii) amino acid position corresponding to position 640 of SEQ ID NO: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, or 285; iv) amino acid position corresponding to position 666 of SEQ ID NO: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, or 285; v) amino acid position corresponding to position 680 of SEQ ID NO: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, or 285; vi) amino acid position corresponding to position 740 of SEQ ID NO: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, or 285; vii) amino acid position corresponding to position 785 of SEQ ID NO: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, or 285; viii) amino acid position corresponding to position 910 of SEQ ID NO: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, or 285; and ix) amino acid position corresponding to position 1077 of SEQ ID NO: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, or 285.

[0115] IV. Guide RNA

[0116] The present disclosure provides guide RNAs and polynucleotides encoding the same. The term “guide RNA” refers to a nucleotide sequence having sufficient complementarity with a target nucleotide sequence to hybridize with the target sequence and direct sequence-specific binding of an associated RNA- guided nuclease to the target nucleotide sequence. More specifically, when the target nucleotide sequence is double-stranded as is the case with DNA, the target nucleotide sequence comprises a target strand and a nontarget strand (which comprises the PAM sequence). In these embodiments, the guide RNA has sufficient complementarity with the target strand of a double -stranded target sequence (e.g., target DNA sequence) such that the guide RNA hybridizes with the target strand and directs sequence-specific binding of an associated RGN to the target sequence (e.g., target DNA sequence). Therefore, in some embodiments, a guide RNA includes a spacer that is identical to the sequence of the non-target strand except that uracil (U) replaces thymidine (T) in the guide RNA. In embodiments where multiplex gene editing is used and there are multiple guide RNAs, each of the one or more guide RNA has sufficient complementarity with the target strand of a particular target sequence and is capable of hybridizing to the target strand of that target sequence. Thus, “a corresponding target sequence” for a guide RNA refers to the target sequence that the guide RNA has sufficient complementarity with and is capable of hybridizing to.

[0117] Thus, an RGN’s respective guide RNA is one or more RNA molecules (generally, one or two), that can bind to the RGN and guide the RGN to bind to a particular target nucleotide sequence, and in those embodiments wherein the RGN has nickase or nuclease activity, also cleave the target strand and / or the nontarget strand. In general, a guide RNA comprises a CRISPR RNA (crRNA) and a trans-activating CRISPR RNA (tracrRNA). Native guide RNAs that comprise both a crRNA and a tracrRNA generally comprise two separate RNA molecules that hybridize to each other through the repeat sequence of the crRNA and the antirepeat sequence of the tracrRNA. A guide RNA can encompass a polymerase editing (PE) guide RNA.

[0118] Native direct repeat sequences within a CRISPR array generally range in length from 28 to 37 base pairs, although the length can vary between about 23 bp to about 55 bp. Spacer sequences within a CRISPR array generally range from about 32 to about 38 bp in length, although the length can be between about 21 bp to about 72 bp. Each CRISPR array generally comprises less than 50 units of the CRISPR repeat-spacer sequence. The CRISPRs are transcribed as part of a long transcript termed the primary CRISPR transcript, which comprises much of the CRISPR array. The primary CRISPR transcript is cleaved by Cas proteins to produce crRNAs or in some cases, to produce pre-crRNAs that are further processed by additional Cas proteins into mature crRNAs. Mature crRNAs comprise a spacer sequence and a CRISPR repeat sequence. In some embodiments in which pre-crRNAs are processed into mature (or processed) crRNAs, maturation involves the removal of about one to about six or more 5', 3', or 5' and 3' nucleotides. For the purposes of genome editing or targeting a particular target nucleotide sequence of interest, these nucleotides that are removed during maturation of the pre-crRNA molecule are not necessary for generating or designing a guide RNA.

[0119] A guide RNA of the disclosure can comprise at least one chemical modification. The at least one chemical modification includes: a bridged nucleic acid (BNA) modification; 2'-O-methyl (2'-O-Me) modification; 2'-O-methoxy-ethyl (2'MOE) modification; 2'-fluoro (2'-F) modification; 2'F-4'Ca-OMe modification; 2',4'-di-Ca-OMe modification; 2'-O-methyl 3'phosphorothioate (MS) modification; 2'-O- methyl 3'thiophosphonoacetate (MSP) modification; 2'-O-methyl 3'phosphonoacetate (MP) modification; and phosphorothioate (PS) modification; or a combination thereof. In some embodiments, the BNA comprises a 2', 4' BNA modification. In some embodiments, the 2', 4' BNA modification is selected from the group consisting of: locked nucleic acid (LNA) modification, BNANC[N-Me] modification, 2'-O,4'-C- ethylene bridged nucleic acid (2',4'-ENA) modification, and S-constrained ethyl (cEt) modification. In some embodiments, the 2', 4' BNA is a LNA modification. In some embodiments, the 2', 4' BNA is a cEt modification. In some embodiments, the at least one chemical modification comprises a BNA modification, 2'-0-Me modification, or PS modification. Chemical modifications of spacers, crRNA repeats, crRNAs, tracrRNAs, and guide RNAs are described in International Application Publication No. WO 2024 / 042489, which is hereby incorporated by reference in its entirety herein.

[0120] The present disclosure provides compositions comprising guide RNAs comprising CRISPR RNAs (crRNAs). A CRISPR RNA (crRNA) comprises a spacer and a CRISPR repeat. The “spacer” is the nucleotide sequence that directly hybridizes with the target nucleotide sequence of interest. The spacer is engineered to be fully or partially complementary with the target sequence of interest. In various embodiments, the spacer can comprise from about 8 nucleotides to about 30 nucleotides, or more. For example, the spacer sequence can be about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, or more nucleotides in length. In some embodiments, the spacer is 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or more nucleotides in length. In some embodiments, the spacer is about 10 to about 26 nucleotides in length, or about 12 to about 30 nucleotides in length. In some embodiments, the spacer is 19-27 nucleotides in length. In particular embodiments, the spacer is about 30 nucleotides in length. In some embodiments, the spacer is 30 nucleotides in length. In some embodiments, the spacer is about 20-25 nucleotides in length. In some embodiments, the degree of complementarity between a spacer and its corresponding target sequence, when optimally aligned using a suitable alignment algorithm, is between 50% and 99% or more, including but not limited to about or more than about 50%, about 60%, about 70%, about 75%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or more. In particular embodiments, the degree of complementarity between a spacer and its corresponding target sequence, when optimally aligned using a suitable alignment algorithm, is 50%, 60%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more. In some embodiments, the spacer can be identical in sequence to the non-target strand of a target sequence. In some of those embodiments wherein the target sequence is a target DNA sequence, the spacer can be identical in sequence to the non-target strand of the target DNA sequence, with the exception of the thymidines (Ts) in the non-target strand being replaced by uracils (Us) in the spacer. In particular embodiments, the spacer is free of secondary structure, which can be predicted using any suitable polynucleotide folding algorithm known in the art, including but not limited to mFold (see, e.g., Zuker and Stiegler (1981) Nucleic Acids Res. 9: 133-148) and RNAfold (see, e.g., Gruber et al. (2008) Cell 106(l):23- 24). Along with a spacer, a crRNA further comprises a CRISPR RNA (crRNA) repeat. Generally, a CRISPR RNA repeat comprises a nucleotide sequence that forms a structure, either on its own or in concert with a hybridized tracrRNA, that is recognized by the RGN polypeptide. For LPG10145 guide crRNAs, the spacer is 5' of the crRNA repeat. In various embodiments, the CRISPR RNA repeat can comprise from about 8 nucleotides to about 30 nucleotides, or more. For example, the CRISPR repeat can be about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, or more nucleotides in length. In particular embodiments, the CRISPR repeat is 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or more nucleotides in length. In some embodiments, a crRNA repeat comprises a total length of 19 to 40 nucleotides (nt). In some embodiments, a crRNA repeat comprises a total length of at most 19 nt, 20 nt, 21 nt, 22 nt, 23 nt, 24 nt, 25 nt, 26 nt, 27 nt, 28 nt, 29 nt, or 30 nt. In some embodiments, the crRNA repeat is about 19 nt or 21 nt. In some embodiments, the degree of complementarity between a CRISPR repeat and its corresponding tracrRNA, when optimally aligned using a suitable alignment algorithm, is about or more than about 50%, about 60%, about 70%, about 75%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or more. In particular embodiments, the degree of complementarity between a CRISPR repeat and its corresponding tracrRNA, when optimally aligned using a suitable alignment algorithm, is 50%, 60%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more.

[0121] In particular embodiments, the CRISPR repeat comprises the nucleotide sequence of SEQ ID NO: 33, 244, or 245, or an active variant or fragment thereof that when comprised within a guide RNA, is capable of directing the sequence-specific binding of an associated variant LPG10145 RGN provided herein to a target sequence of interest. In certain embodiments, an active CRISPR repeat sequence variant of a wild-type sequence comprises a nucleotide sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the nucleotide sequence set forth as SEQ ID NO: 33, 244, or 245. In certain embodiments, an active CRISPR repeat fragment of a wild-type sequence comprises at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, or 22 contiguous nucleotides of the nucleotide sequence set forth as SEQ ID NO: 33, 244, or 245. The CRISPR repeat can comprise a nucleotide sequence set forth as SEQ ID NO: 33, 244, or 245, or that differs from SEQ ID NO: 33, 244, or 245 by 1 to 5 nucleotides. In some embodiments, the CRISPR repeat comprises a nucleotide sequence that differs from SEQ ID NO: 33, 244, or 245 by 5 nucleotides. In some embodiments, the CRISPR repeat comprises a nucleotide sequence that differs from SEQ ID NO: 33, 244, or 245 by 4 nucleotides. In some embodiments, the CRISPR repeat comprises a nucleotide sequence that differs from SEQ ID NO: 33, 244, or 245 by 3 nucleotides. In some embodiments, the CRISPR repeat comprises a nucleotide sequence that differs from SEQ ID NO: 33, 244, or 245 by 2 nucleotides. In some embodiments, the CRISPR repeat comprises a nucleotide sequence that differs from SEQ ID NO: 33, 244, or 245 by 1 nucleotide. In some embodiments, the CRISPR repeat comprises the nucleotide sequence set forth as SEQ ID NO: 33, 244, or 245.

[0122] In those embodiments wherein the RGN has the amino acid sequence set forth as SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, or 285, the crRNA repeat of the associated gRNA can have the nucleotide sequence set forth as SEQ ID NO: 33, 244, or 245, or an active variant or fragment thereof.

[0123] In certain embodiments, the crRNA is not naturally-occurring. In some of these embodiments, the specific CRISPR repeat is not linked to the engineered spacer in nature and the CRISPR repeat is considered heterologous to the spacer. In certain embodiments, the spacer is an engineered sequence that is not naturally occurring.

[0124] The guide RNA can further comprise a trans-activating CRISPR RNA (tracrRNA). A tracrRNA molecule comprises a nucleotide sequence comprising a region that has sufficient complementarity to hybridize to a CRISPR repeat of a crRNA, which is referred to herein as the anti-repeat region. In some embodiments, the tracrRNA molecule further comprises a region with secondary structure (e.g., stem-loop) or forms secondary structure upon hybridizing with its corresponding crRNA. In particular embodiments, the region of the tracrRNA that is fully or partially complementary to a CRISPR repeat sequence is at the 5' end of the molecule and the 3' end of the tracrRNA comprises secondary structure. This region of secondary structure generally comprises several hairpin structures, including the nexus hairpin, which is found adjacent to the anti -repeat sequence. The nexus forms the core of the interactions between the guide RNA and the RGN, and is at the intersection between the guide RNA, the RGN, and the target DNA. The nexus hairpin often has a conserved nucleotide sequence in the base of the hairpin stem, with the motif UNANNC found in many nexus hairpins in tracrRNAs. In some embodiments, a tracrRNA comprises a non-canonical sequence in the base of the hairpin stem of its nexus hairpin, including UNANNA, UNANNG, and CNANNC. There are often terminal hairpins at the 3' end of the tracrRNA that can vary in structure and number, but often comprise a GC-rich Rho-independent transcriptional terminator hairpin followed by a string of U’s at the 3' end. See, for example, Briner et al. (2014) Molecular Cell 56:333-339, Briner and Barrangou (2016) Cold Spring Harb Pro ct doi: 10. 1101 / pdb.top090902, and U.S. Publication No. 2017 / 0275648, each of which is herein incorporated by reference in its entirety.

[0125] In various embodiments, the anti-repeat region of the tracrRNA that is fully or partially complementary to the CRISPR repeat comprises from about 8 nucleotides to about 30 nucleotides, or more. For example, the region of base pairing between the tracrRNA anti-repeat and the CRISPR repeat can be about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, or more nucleotides in length. In particular embodiments, the region of base pairing between the tracrRNA anti-repeat and the CRISPR repeat is 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or more nucleotides in length. In some embodiments, the degree of complementarity between a CRISPR repeat and its corresponding tracrRNA anti-repeat, when optimally aligned using a suitable alignment algorithm, is about or more than about 50%, about 60%, about 70%, about 75%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or more. In particular embodiments, the degree of complementarity between a CRISPR repeat and its corresponding tracrRNA anti-repeat, when optimally aligned using a suitable alignment algorithm, is 50%, 60%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more.

[0126] In various embodiments, the entire tracrRNA can comprise from about 60 nucleotides to more than about 210 nucleotides. For example, the tracrRNA can be about 60, about 65, about 70, about 75, about 80, about 85, about 90, about 95, about 100, about 105, about 110, about 115, about 120, about 125, about 130, about 135, about 140, about 150, about 160, about 170, about 180, about 190, about 200, about 210, or more nucleotides in length. In particular embodiments, the tracrRNA is 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 150, 160, 170, 180, 190, 200, 210 or more nucleotides in length. In particular embodiments, the tracrRNA is about 70 to about 105 nucleotides in length, including about 70, about 71, about 72, about 73, about 74, about 75, about 76, about 77, about 78, about 79, about 80, about 81, about 82, about 83, about 84, about 85, about 86, about 87, about 88, about 89, about 90, about 91, about 92, about 93, about 94, about 95, about 96, about 97, about 98, about 99, about 100, about 101, about 102, about

[0127] 103, about 104, and about 105 nucleotides in length. In particular embodiments, the tracrRNA is 70 to 105 nucleotides in length, including 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, and 105 nucleotides in length.

[0128] In particular embodiments, the tracrRNA comprises the nucleotide sequence of SEQ ID NO: 34, 246, 247, or 248, or an active variant or fragment thereof that when comprised within a guide RNA is capable of directing the sequence -specific binding of an associated RNA-guided nuclease provided herein to a target sequence of interest. In certain embodiments, an active tracrRNA sequence variant of a wild-type sequence comprises a nucleotide sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the nucleotide sequences set forth as SEQ ID NO: 34, 246, 247, or 248. In certain embodiments, an active tracrRNA sequence fragment of a wild-type sequence comprises at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, or more contiguous nucleotides of the nucleotide sequence set forth as SEQ ID NO: 34, 246, 247, or 248

[0129] In those embodiments wherein the RGN has the amino acid sequence set forth as SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, or 285, the tracrRNA of the associated gRNA can have the nucleotide sequence set forth as SEQ ID NO: 34, 246, 247, or 248, or an active variant or fragment thereof. An active variant or fragment of a variant LPG10145 RGN disclosed herein can bind a tracrRNA comprising the nucleotide sequence set forth as SEQ ID NO: 34, 246, 247, or 248, or an active variant or fragment thereof.

[0130] Two polynucleotide sequences can be considered to be substantially complementary when the two sequences hybridize to each other under stringent conditions. Likewise, an RGN is considered to bind to a particular target sequence within a sequence-specific manner if the guide RNA bound to the RGN binds to the target sequence under stringent conditions. By "stringent conditions" or "stringent hybridization conditions" is intended conditions under which the two polynucleotide sequences will hybridize to each other to a detectably greater degree than to other sequences (e.g. , at least 2-fold over background). Stringent conditions are sequence-dependent and will be different in different circumstances. Typically, stringent conditions will be those in which the salt concentration is less than about 1.5 M Na ion, typically about 0.01 to 1.0 M Na ion concentration (or other salts) at pH 7.0 to 8.3, and the temperature is at least about 30°C for short sequences (e.g., 10 to 50 nucleotides) and at least about 60°C for long sequences (e.g., greater than 50 nucleotides). Stringent conditions may also be achieved with the addition of destabilizing agents such as formamide. Exemplary low stringency conditions include hybridization with a buffer solution of 30 to 35% formamide, 1 M NaCl, 1% SDS (sodium dodecyl sulfate) at 37°C, and a wash in IX to 2X SSC (20X SSC = 3.0 M NaCl / 0.3 M trisodium citrate) at 50 to 55°C. Exemplary moderate stringency conditions include hybridization in 40 to 45% formamide, 1.0 M NaCl, 1% SDS at 37°C, and a wash in 0.5X to IX SSC at 55 to 60°C. Exemplary high stringency conditions include hybridization in 50% formamide, 1 M NaCl, 1% SDS at 37°C, and a wash in 0. IX SSC at 60 to 65°C. Optionally, wash buffers may comprise about 0.1% to about 1% SDS. Duration of hybridization is generally less than about 24 hours, usually about 4 to about 12 hours. The duration of the wash time will be at least a length of time sufficient to reach equilibrium.

[0131] The Tm is the temperature (under defined ionic strength and pH) at which 50% of a complementary target sequence hybridizes to a perfectly matched sequence. For DNA-DNA hybrids, the Tm can be approximated from the equation of Meinkoth and Wahl (1984) Anal. Biochem. 138:267-284: Tm = 81.5°C + 16.6 (log M) + 0.41 (%GC) - 0.61 (% form) - 500 / L; where M is the molarity of monovalent cations, %GC is the percentage of guanosine and cytosine nucleotides in the DNA, % form is the percentage of formamide in the hybridization solution, and L is the length of the hybrid in base pairs. Generally, stringent conditions are selected to be about 5 °C lower than the thermal melting point (Tm) for the specific sequence and its complement at a defined ionic strength and pH. However, severely stringent conditions can utilize a hybridization and / or wash at 1, 2, 3, or 4°C lower than the thermal melting point (Tm); moderately stringent conditions can utilize a hybridization and / or wash at 6, 7, 8, 9, or 10°C lower than the thermal melting point (Tm); low stringency conditions can utilize a hybridization and / or wash at 11, 12, 13, 14, 15, or 20°C lower than the thermal melting point (Tm). Using the equation, hybridization and wash compositions, and desired Tm, those of ordinary skill will understand that variations in the stringency of hybridization and / or wash solutions are inherently described. An extensive guide to the hybridization of nucleic acids is found in Tijssen (1993) Laboratory Techniques in Biochemistry and Molecular Biology — Hybridization with Nucleic Acid Probes, Part I, Chapter 2 (Elsevier, New York); and Ausubel et al., eds. (1995) Current Protocols in Molecular Biology, Chapter 2 (Greene Publishing and Wiley-Interscience, New York). See Sambrook et al. (1989) Molecular Cloning: A Laboratory Manual (2d ed., Cold Spring Harbor Laboratory Press, Plainview, New York).

[0132] The term “sequence specific” can also refer to the binding of a RGN polypeptide or RGN polypeptide fusion to a target sequence at a greater affinity than binding to a randomized background sequence.

[0133] The guide RNA can be a single guide RNA (sgRNA) or a dual -guide RNA (dgRNA). A single guide RNA comprises the crRNA and tracrRNA on a single molecule of RNA, whereas a dual -guide RNA system comprises a crRNA and a tracrRNA present on two distinct RNA molecules, hybridized to one another through at least a portion of the CRISPR repeat sequence of the crRNA and at least a portion of the tracrRNA, which may be fully or partially complementary to the CRISPR repeat sequence of the crRNA. In some of those embodiments wherein the guide RNA is a single guide RNA, the crRNA and tracrRNA are separated by a linker nucleotide sequence. In general, the linker nucleotide sequence is one that does not include complementary bases in order to avoid the formation of secondary structure within or comprising nucleotides of the linker nucleotide sequence. In some embodiments, the linker nucleotide sequence between the crRNA and tracrRNA is at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, or more nucleotides in length. In particular embodiments, the linker nucleotide sequence of a single guide RNA is at least 4 nucleotides in length. In certain embodiments, the linker nucleotide sequence is the nucleotide sequence AAAG.

[0134] In some embodiments, the guide RNA is a single guide RNA (sgRNA) having the backbone sequence (comprising a crRNA repeat, an optional linker nucleotide sequence, and a tracrRNA) of any one of SEQ ID NOs: 249-252, or an active variant or fragment thereof. In certain embodiments, an active sgRNA backbone sequence variant comprises a nucleotide sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the nucleotide sequence set forth as any one of SEQ ID NOs: 249-252. In certain embodiments, an active sgRNA backbone sequence fragment comprises at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, or more contiguous nucleotides of the nucleotide sequence set forth as any one of SEQ ID NOs: 249-252. In those embodiments wherein the RGN has the amino acid sequence set forth as SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, or 285, the sgRNA backbone can have the nucleotide sequence set forth as any one of SEQ ID NOs: 249-252, or an active variant or fragment thereof. The single guide RNA or dual-guide RNA can be synthesized chemically or via in vitro transcription. Assays for determining sequence-specific binding between an RGN and a guide RNA are known in the art and include, but are not limited to, in vitro binding assays between an expressed RGN and the guide RNA, which can be tagged with a detectable label (e.g., biotin) and used in a pull-down detection assay in which the guide RNA:RGN complex is captured via the detectable label (e.g., with streptavidin beads). A control guide RNA with an unrelated sequence or structure to the guide RNA can be used as a negative control for non-specific binding of the RGN to RNA.

[0135] In certain embodiments, the guide RNA can be introduced into a target cell, organelle, or embryo as an RNA molecule. The guide RNA can be transcribed in vitro or chemically synthesized. In other embodiments, a nucleotide sequence encoding the guide RNA is introduced into the cell, organelle, or embryo. In some of these embodiments, the nucleotide sequence encoding the guide RNA is operably linked to a promoter (e.g., an RNA polymerase III promoter). The promoter can be a native promoter or heterologous to the guide RNA-encoding nucleotide sequence.

[0136] In various embodiments, the guide RNA can be introduced into a target cell, organelle, or embryo as a ribonucleoprotein complex, as described herein, wherein the guide RNA is bound to an RNA-guided nuclease polypeptide.

[0137] The guide RNA directs an associated RNA-guided nuclease to a particular target nucleotide sequence of interest through hybridization of the guide RNA to the target nucleotide sequence. A target nucleotide sequence can comprise DNA, RNA, or a combination of both and can be single-stranded or double -stranded. A target nucleotide sequence can be genomic DNA (z.e., chromosomal DNA), plasmid DNA, or an RNA molecule (e.g., messenger RNA, ribosomal RNA, transfer RNA, micro RNA, small interfering RNA). The target nucleotide sequence can be bound (and in some embodiments, cleaved) by an RNA-guided nuclease in vitro or in a cell. The chromosomal sequence targeted by the RGN can be a nuclear, plastid or mitochondrial chromosomal sequence. In some embodiments, the target nucleotide sequence is unique in the target genome. In some embodiments, the target sequence is double-stranded and comprises a target strand and a non-target strand.

[0138] The target sequence is adjacent to a protospacer adjacent motif (PAM) and the target strand of the target sequence is the strand that comprises the PAM. The PAM is immediately adjacent to the target sequence and often comprise Ns, which represent any nucleotide. In some embodiments, the PAM comprises about 1 to about 10 Ns, including about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, or about 10 Ns. In some embodiments, a PAM comprises 1 to 10 Ns, including 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 Ns. In general, the PAM can be 5' or 3' of the target sequence on its non-target strand. In some embodiments, the PAM is 3' of the target sequence on its non-target strand for the presently disclosed guide RNAs and RGN systems. Generally, the PAM is a consensus sequence of about 3-4 nucleotides, but in some embodiments, it can be 2, 3, 4, 5, 6, 7, 8, 9, or more nucleotides in length. In some embodiments, the PAM sequence recognized by the presently disclosed variant LPG10145 RGNs comprises the consensus sequence set forth as NNGG.

[0139] In some embodiments, the variant LPG10145 RGNs comprise an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein the amino acid residue at one or more of the following amino acid positions differs from the corresponding amino acid residue in SEQ ID NO: 1: 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975. In some embodiments, the variant LPG10145 RGNs comprise an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein the amino acid residue at one or more of the following amino acid positions is a positively charged amino acid residue: 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975. In some embodiments, the variant LPG10145 RGNs comprise an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein the amino acid residue at one or more of the following amino acid positions is an R: 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975.

[0140] In some embodiments, the variant LPG10145 RGNs comprise an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein the amino acid residue at amino acid position 778 and / or 969 differs from the corresponding amino acid residue in SEQ ID NO: 1. In some embodiments, the variant LPG10145 RGNs comprise an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein the amino acid residue at amino acid position 778 and / or 969 is a positively charged amino acid residue. In some embodiments, the variant LPG10145 RGNs comprise an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein the amino acid residue at amino acid position 778 and / or 969 is an R.

[0141] In some embodiments, the variant LPG10145 RGNs comprise an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein amino acid residues at positions 778 and 856 differ from the corresponding amino acid residues in SEQ ID NO: 1. In some embodiments, the variant LPG10145 RGNs comprise an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein the amino acid residues at positions 778 and 856 are positively charged amino acid residues.

[0142] In some embodiments, the variant LPG10145 RGNs comprise an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein amino acid residues at positions 55, 647, 778, and 969 differ from the corresponding amino acid residues in SEQ ID NO: 1. In some embodiments, the variant LPG10145 RGNs comprise an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein amino acid residues at positions 55, 647, 778, and 969 are positively charged amino acid residues.

[0143] In particular embodiments, a variant LPG10145 RGN having an amino acid sequence selected from: (a) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 52 is an R; (b) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 55 is an R; (c) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 86 is an R; (d) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 472 is an R; (e) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 533 is an R; (f) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 541 is an R; (g) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 643 is substituted by an R; (h) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 647 is an R; (i) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 653 is an R; (j) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 745 is an R; (k) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position 774 is an R; (1) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 778 is an R; (m) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 780 is an R; (n) the amino acid sequence set forth as SEQ ID NO: 1, wherein A at amino acid position 795 is an R; (o) the amino acid sequence set forth as SEQ ID NO: 1, wherein Q at amino acid position 822 is an R; (p) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 843 is an R; (q) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 856 is an R; (r) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 871 is an R; (s) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 872 is an R; (t) the amino acid sequence set forth as SEQ ID NO: 1, wherein D at amino acid position 900 is an R; (u) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position 911 is an R; (v) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 913 is an R; (w) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position 954 is an R; (x) the amino acid sequence set forth as SEQ ID NO: 1, wherein V at amino acid position 958 is an R; (y) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 968 is an R; (z) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 969 is an R; (aa) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 973 is an R; (bb) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 974 is an R; (cc) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position 975 is an R; (dd) the amino acid sequence set forth as SEQ ID NO: 2; (ee) the amino acid sequence set forth as SEQ ID NO: 3; (ff) the amino acid sequence set forth as SEQ ID NO: 4; (gg) the amino acid sequence set forth as SEQ ID NO: 5; (hh) the amino acid sequence set forth as SEQ ID NO: 6; (ii) the amino acid sequence set forth as SEQ ID NO: 7; (jj) the amino acid sequence set forth as SEQ ID NO: 8; (kk) the amino acid sequence set forth as SEQ ID NO: 9; (11) the amino acid sequence set forth as SEQ ID NO: 10; (mm) the amino acid sequence set forth as SEQ ID NO: 11; (nn) the amino acid sequence set forth as SEQ ID NO: 12; (oo) the amino acid sequence set forth as SEQ ID NO: 13; (pp) the amino acid sequence set forth as SEQ ID NO: 14; (qq) the amino acid sequence set forth as SEQ ID NO: 15; and (rr) the amino acid sequence set forth as SEQ ID NO: 16, or an active variant or fragment thereof, binds a target nucleotide sequence adjacent to a PAM sequence set forth as NNGG. In particular embodiments, a variant LPG10145 RGN having an amino acid sequence selected from: (a) the amino acid sequence set forth as SEQ ID NO: 2; (b) the amino acid sequence set forth as SEQ ID NO: 3; (c) the amino acid sequence set forth as SEQ ID NO: 4; (d) the amino acid sequence set forth as SEQ ID NO: 5; (e) the amino acid sequence set forth as SEQ ID NO: 6; (f) the amino acid sequence set forth as SEQ ID NO: 7; (g) the amino acid sequence set forth as SEQ ID NO: 8SEQ ID NO: 8; (h) the amino acid sequence set forth as SEQ ID NO: 9; (i) the amino acid sequence set forth as SEQ ID NO: 10; (j) the amino acid sequence set forth as SEQ ID NO: 11; (k) the amino acid sequence set forth as SEQ ID NO: 12; (1) the amino acid sequence set forth as SEQ ID NO: 13; (m) the amino acid sequence set forth as SEQ ID NO: 14; (n) the amino acid sequence set forth as SEQ ID NO: 15; and (o) the amino acid sequence set forth as SEQ ID NO: 16, or an active variant or fragment thereof, binds a target nucleotide sequence adjacent to a PAM sequence set forth as NNGG.

[0144] In some embodiments, the variant LPG10145 RGN binds to a guide sequence comprising a CRISPR repeat set forth as SEQ ID NO: 33, 244, or 245, or an active variant or fragment thereof, and a tracrRNA sequence set forth as SEQ ID NO: 34, 246, 247, or 248, or an active variant or fragment thereof. The RGN systems are described further in Examples 1-6 of the present specification.

[0145] It is well-known in the art that PAM sequence specificity for a given nuclease enzyme is affected by enzyme concentration (see, e.g., Karvelis et al. (2015) Genome Biol 16:253), which may be modified by altering the promoter used to express the RGN, or the amount of ribonucleoprotein complex delivered to the cell, organelle, or embryo.

[0146] Upon recognizing its corresponding PAM sequence, the RGN, if active, can cleave the target nucleotide sequence at a specific cleavage site. As used herein, a cleavage site is made up of the two particular nucleotides within a target nucleotide sequence at which the target strand, non-target strand, or both strands of the target nucleotide sequence is cleaved by an RGN. The cleavage site can comprise the 1stand 2nd, 2ndand 3rd. 3rdand 4th, 4thand 5th, 5thand 6th, 7thand 8th, or 8thand 9thnucleotides from the PAM in either the 5' or 3' direction. In some embodiments, the cleavage site may be over 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides from the PAM in either the 5' or 3' direction. As RGNs can cleave a target nucleotide sequence resulting in staggered ends, in some embodiments, the cleavage site is defined based on the distance of the two nucleotides from the PAM on the non-target strand of the target polynucleotide and for the target strand, the distance of the two nucleotides from the complement of the PAM.

[0147] V. Polymerases

[0148] Compositions of the disclosure, including fusion proteins, polymerase editors (PEs), and PE systems comprising engineered variant LPG10145 RGN polypeptides, can comprise polymerases (e.g., DNA polymerases, reverse transcriptases). As used herein, a “polymerase” is an enzyme that catalyzes the formation of a nucleic acid polymer. A polymerase can be an RNA polymerase (catalyzing an RNA polymer) or a DNA polymerase (catalyzing a DNA polymer). In some embodiments, the polymerase of the polymerase editor or system is a DNA polymerase. The PE or PE system can comprise a DNA-dependent DNA polymerase (uses DNA as a template) or an RNA-dependent DNA polymerase (uses RNA as a template). In some embodiments, the DNA polymerase of the presently disclosed PEs and PE systems is an RNA-dependent DNA polymerase (i.e., reverse transcriptase).

[0149] Reverse transcriptases (RTs) are a class of enzymes that catalyze the transcription of RNA into DNA, a process known as reverse transcription. This enzymatic activity is critical in the life cycles of retroviruses, such as Human Immunodeficiency Virus (HIV), and in the replication of various mobile genetic elements, including retrotransposons. First, the RT uses its RNA-dependent DNA polymerase activity to convert single-stranded RNA (ssRNA) templates into complementary DNA (cDNA). RTs can also possess RNase H activity, which degrades the RNA strand of an RNA-DNA hybrid, providing a template for the synthesis of the second DNA strand. The RT then synthesizes the second DNA strand through its DNA-dependent DNA polymerase activity, resulting in a double -stranded DNA (dsDNA) molecule that can integrate into the host genome. As used herein, a “reverse transcriptase” or “RT” is an enzyme that has polymerase activity to catalyze the formation of a nucleic acid polymer. In some embodiments, the RT synthesizes a nucleic acid polymer using a template nucleic acid molecule. In some embodiments, the polymerase activity is a DNA polymerase activity. In some embodiments, an RT catalyzes the addition of nucleotides to a nicked polynucleotide strand, using a template.

[0150] RTs include retroviral RTs such as HIV-1 RT, hepatitis B RT, and Murine Leukemia Virus (MLV)- RT (also known as Moloney Murine Leukemia Virus (MMLV)-RT). MLV-RT serves as a model for understanding the basic mechanisms of reverse transcription. Reverse transcriptases typically exhibit a "right hand" structure with three main domains: a ‘finger’ domain involved in binding the template -primer and dNTPs; the ‘palm’ domain containing the active site with highly conserved motifs responsible for catalysis; and the ‘thumb’, which maintains the enzyme's interaction with the nucleic acid substrate. RTs also have distinct polymerase and RNase H domains, where the polymerase domain is responsible for nucleic acid molecule (e.g., DNA) synthesis, and the RNase H domain degrades the RNA strand in RNA-DNA hybrids to allow second-strand DNA synthesis. RTs are important tools in molecular biology and biotechnology, with uses in RT-PCR, cDNA synthesis from RNA, amplification and quantification of RNA, RNA Sequencing (RNA-Seq), preparing cDNA libraries from RNA samples, and gene cloning and expression studies.

[0151] A polymerase (e.g., reverse transcriptase) of the disclosure includes but is not limited to the polymerases (e.g., reverse transcriptases) described herein and variants or fragments thereof, including but not limited to a reverse transcriptase comprising an amino acid sequence having at least 40%, at least 41%, at least 42%, at least 43%, at least 44%, at least 45%, at least 46%, at least 47%, at least 48%, at least 49%, at least 50%, at least 51%, at least 52%, at least 53%, at least 54%, at least 55%, at least 56%, at least 57%, at least 58%, at least 59%, at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity to the amino acid sequence set forth as SEQ ID NO: 254 or 255.

[0152] An RT of the presently disclosed compositions and methods can lack an RNase H domain.

[0153] VI. Polymerase Editors (PEs)

[0154] As used herein, a “polymerase editor” or “PE” refers to a protein or a plurality of proteins comprising an RGN polypeptide and a polymerase editing polypeptide (e.g., DNA polymerase or reverse transcriptase) that, along with a polymerase editing guide RNA (PEgRNA) that comprises an extension arm comprising a primer binding site (PBS) and a DNA synthesis template comprising a desired edit, is capable of editing a double-stranded polynucleotide through the replacement of a target sequence using the DNA synthesis template as a template for the polymerase. In certain embodiments, the RGN polypeptide and the polymerase editing polypeptide (e.g., DNA polymerase or reverse transcriptase) are operably linked (by fusion or insertion). In other embodiments, the RGN polypeptide and the polymerase editing polypeptide (e.g., DNA polymerase or reverse transcriptase) are not operably linked. In one particular embodiment, the RGN polypeptide and the polymerase editing polypeptide (e.g., DNA polymerase or reverse transcriptase) are two separate polypeptides. In some embodiments, the PE does not require the introduction of a doublestranded break, but rather utilizes an RGN nickase that nicks the non-target strand upstream of the sequence to be edited and upstream of the PAM, creating a 3' flap on the non-target strand. The PBS of the PEgRNA is complementary to the 3' flap of the non-target strand and hybridrization of the PBS and 3' flap of the non- target strand allows for the polymerization of the replacement strand containing the edit using the DNA synthesis template and polymerase. Those polymerase editors that utilize a reverse transcriptase as the polymerase are referred to herein as “RT editors” or “RTEs”.

[0155] The presently disclosed polymerase editors (PEs) comprise a polymerase (e.g., RT) and an engineered variant LPG10145 RGN polypeptide. The RT includes RTs, or active variants or fragments thereof, as described herein, including but not limited to an RT comprising an amino acid sequence having at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or more sequence identity to the amino acid sequence set forth as SEQ ID NO: 254 or 255.

[0156] The engineered variant LPG10145 RGN polypeptides, or active variants or fragments thereof, include those described herein. The binding and / or cleaving activity of an active variant or fragment of a variant LPG10145 RGN disclosed herein can be dependent upon recognizing a protospacer adjacent motif (PAM) adjacent and 3’ to the target sequence. In some embodiments, the PAM comprises a consensus sequence of NNGG. In some embodiments, the variant LPG10145 RGNs comprise an amino acid sequence having at least 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 1, wherein the amino acid residue at one or more of the following amino acid positions differs from the corresponding amino acid residue in SEQ ID NO: 1: 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975, and active fragments or variants thereof that retain the ability to bind to a target nucleotide sequence in an RNA- guided sequence -specific manner.

[0157] In some embodiments, the variant LPG10145 RGNs comprise an amino acid sequence having at least 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 1, wherein the amino acid residue at one or more of the following amino acid positions is a positively charged amino acid residue: 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975, and active fragments or variants thereof that retain the ability to bind to a target nucleotide sequence in an RNA-guided sequence -specific manner. In some embodiments, the variant LPG10145 RGNs comprise an amino acid sequence having at least 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 1, wherein the amino acid residue at one or more of the following amino acid positions is an R: 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975, and active fragments or variants thereof that retain the ability to bind to a target nucleotide sequence in an RNA-guided sequence-specific manner.

[0158] In some embodiments, the variant LPG10145 RGNs comprise an amino acid sequence having at least 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 1, wherein the amino acid residue at amino acid position 778 and / or 969 differs from the corresponding amino acid residue in SEQ ID NO: 1, and active fragments or variants thereof that retain the ability to bind to a target nucleotide sequence in an RNA-guided sequence-specific manner. In some embodiments, the variant LPG10145 RGNs comprise an amino acid sequence having at least 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 1, wherein the amino acid residue at amino acid position 778 and / or 969 is a positively charged amino acid residue, and active fragments or variants thereof that retain the ability to bind to a target nucleotide sequence in an RNA-guided sequence -specific manner. In some embodiments, the variant LPG10145 RGNs comprise an amino acid sequence having at least 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 1, wherein the amino acid residue at amino acid position 778 and / or 969 is an R, and active fragments or variants thereof that retain the ability to bind to a target nucleotide sequence in an RNA- guided sequence -specific manner.

[0159] In some embodiments, the variant LPG10145 RGNs comprise an amino acid sequence having at least 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 1, wherein amino acid residues at positions 778 and 856 differ from the corresponding amino acid residues in SEQ ID NO: 1. In some embodiments, the variant LPG10145 RGNs comprise an amino acid sequence having at least 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 1, wherein amino acid residues at positions 778 and 856 are positively charged amino acid residues.

[0160] In some embodiments, the variant LPG10145 RGNs comprise an amino acid sequence having at least 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 1, wherein amino acid residues at positions 55, 647, 778, and 969 differ from the corresponding amino acid residues in SEQ ID NO: 1. In some embodiments, the variant LPG10145 RGNs comprise an amino acid sequence having at least 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to SEQ ID NO: 1, wherein amino acid residues at positions 55, 647, 778, and 969 are positively charged amino acid residues.

[0161] In some embodiments, the variant LPG10145 RGNs comprise an amino acid sequence selected from: (a) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 52 is an R; (b) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 55 is an R; (c) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 86 is an R; (d) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 472 is an R; (e) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 533 is an R; (f) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 541 is an R; (g) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 643 is substituted by an R; (h) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 647 is an R; (i) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 653 is an R; (j) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 745 is an R; (k) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position 774 is an R; (1) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 778 is an R; (m) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 780 is an R; (n) the amino acid sequence set forth as SEQ ID NO: 1, wherein A at amino acid position 795 is an R; (o) the amino acid sequence set forth as SEQ ID NO: 1, wherein Q at amino acid position 822 is an R; (p) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 843 is an R; (q) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 856 is an R; (r) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 871 is an R; (s) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 872 is an R; (t) the amino acid sequence set forth as SEQ ID NO: 1, wherein D at amino acid position 900 is an R; (u) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position 911 is an R; (v) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 913 is an R; (w) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position 954 is an R; (x) the amino acid sequence set forth as SEQ ID NO: 1, wherein V at amino acid position 958 is an R; (y) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 968 is an R; (z) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 969 is an R; (aa) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 973 is an R; (bb) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 974 is an R; (cc) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position 975 is an R; (dd) the amino acid sequence set forth as SEQ ID NO: 2; (ee) the amino acid sequence set forth as SEQ ID NO: 3; (ff) the amino acid sequence set forth as SEQ ID NO: 4; (gg) the amino acid sequence set forth as SEQ ID NO: 5; (hh) the amino acid sequence set forth as SEQ ID NO: 6; (ii) the amino acid sequence set forth as SEQ ID NO: 7; (jj) the amino acid sequence set forth as SEQ ID NO: 8; (kk) the amino acid sequence set forth as SEQ ID NO: 9; (11) the amino acid sequence set forth as SEQ ID NO: 10; (mm) the amino acid sequence set forth as SEQ ID NO: 11; (nn) the amino acid sequence set forth as SEQ ID NO: 12; (oo) the amino acid sequence set forth as SEQ ID NO: 13; (pp) the amino acid sequence set forth as SEQ ID NO: 14; (qq) the amino acid sequence set forth as SEQ ID NO: 15; and (rr) the amino acid sequence set forth as SEQ ID NO: 16.

[0162] In some embodiments, the variant LPG10145 RGNs comprise an amino acid sequence selected from: (a) the amino acid sequence set forth as SEQ ID NO: 2; (b) the amino acid sequence set forth as SEQ ID NO: 3; (c) the amino acid sequence set forth as SEQ ID NO: 4; (d) the amino acid sequence set forth as SEQ ID NO: 5; (e) the amino acid sequence set forth as SEQ ID NO: 6; (f) the amino acid sequence set forth as SEQ ID NO: 7; (g) the amino acid sequence set forth as SEQ ID NO: 8; (h) the amino acid sequence set forth as SEQ ID NO: 9; (i) the amino acid sequence set forth as SEQ ID NO: 10; (j) the amino acid sequence set forth as SEQ ID NO: 11; (k) the amino acid sequence set forth as SEQ ID NO: 12; (1) the amino acid sequence set forth as SEQ ID NO: 13; (m) the amino acid sequence set forth as SEQ ID NO: 14; (n) the amino acid sequence set forth as SEQ ID NO: 15; and (o) the amino acid sequence set forth as SEQ ID NO: 16, and active fragments or variants thereof that retain the ability to bind to a target nucleotide sequence in an RNA-guided sequence-specific manner.

[0163] 1. Various Formats of Polymerase Editors

[0164] The presently disclosed PEs can be provided in trans, wherein the polymerase (e.g., RT) and variant LPGI0145 RGN polypeptide are separate polypeptides. In some embodiments, the polymerase (e.g., RT) and variant LPGI0145 RGN polypeptide are transcribed together and have a sequence encoding a selfcleaving peptide (e.g., 2A peptide such as P2A) in between, such that translation results in two separate polypeptides. Any self-cleaving peptide known in the art can be used in such embodiments, including but not limited to, 2A peptides, which is a class of 18-22 amino acid long peptides that may function through ribosomal skipping during translation. Non-limiting examples of 2A peptides are T2A, P2A, E2A, and F2A.

[0165] In some embodiments, the presently disclosed PEs can comprise a polymerase (e.g., RT) operably fused to a variant LPGI0145 RGN polypeptide, wherein the RT and variant LPGI0145 RGN polypeptide are fused to each other end-to-end or wherein the polymerase (e.g., RT) is inserted into the variant LPG10145 RGN polypeptide, such as those inlaid base editors described in International Appl. Publ. No. WO 2024 / 095245, which is herein incorporated by reference in its entirety. In an end-to-end fusion, the polymerase (e.g., RT) can be fused to the amino terminus of the variant LPG10145 RGN polypeptide or the carboxy terminus of the variant LPG10145 RGN polypeptide. The presently disclosed PEs can comprise from amino terminus to carboxy terminus: the polymerase (e.g., RT) and the variant LPG10145 RGN polypeptide; or the variant LPG10145 RGN polypeptide and the polymerase (e.g., RT).

[0166] In those embodiments wherein the polymerase (e.g., RT) is inserted within a variant LPG10145 RGN polypeptide, the polymerase (e.g., RT) is inserted between surface amino acid residues of the variant LPG10145 RGN polypeptide. The polymerase (e.g., RT) can be inserted within or between a linker domain 2, a wedge (WED) domain, a RuvC domain, an HNH domain, a Rec-2 domain, or a PAM-interacting (PI) domain. In some embodiments, the RuvC domain is the RuvCIII domain. A Rec or recognition lobe mediates nucleic acid binding through multiple Rec domains (e.g., Recl-3) by sensing nucleic acids, regulates the HNH conformational transition, and locks the catalytic HNH domain at the cleavage site. A wedge domain is responsible for the recognition of guide RNA scaffolds. A PAM-interacting domain is the domain of an RGN polypeptide that binds to a PAM site. Non-limiting examples of domains within LPG10145 RGN (SEQ ID NO: 1) or a variant LPG10145 RGN (SEQ ID NOs: 2-16, 124-126, 182-196, and 271-285) include: RuvC -I from amino acid residues 1 to 42; BH from amino acid residues 43 to 79; RECI from amino acid residues 80 to 236; REC2 from amino acid residues 237 to 476; RuvC-II from amino acid residues 477 to 524; LI from amino acid residues 525 to 560; HNH from amino acid residues 561 to 676; L2 from amino acid residues 677 to 690; RuvC -III from amino acid residues 691 to 828; WED from amino acid residues 829 to 976; and PI from amino acid residues 977 to 1130. The general domains of RGN polypeptides can be determined via structural comparison to RGN polypeptides with defined domains.

[0167] In those embodiments wherein the PE comprises a variant LPG10145 RGN polypeptide comprising the amino acid sequence set forth as SEQ ID NO: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 182, 183,

[0168] 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 271, 272, 273, 274, 275, 276, 277, 278,

[0169] 279, 280, 281, 282, 283, 284, or 285, or an active variant or fragment thereof, the polymerase (e.g., RT) can be inserted within the variant LPG10145 RGN polypeptide, or an active variant or fragment thereof, immediately after the amino acid position selected from the group consisting of: i) amino acid position corresponding to position 347 of SEQ ID NO: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 182, 183, 184,

[0170] 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 271, 272, 273, 274, 275, 276, 277, 278, 279,

[0171] 280, 281, 282, 283, 284, or 285; ii) amino acid position corresponding to position 524 of SEQ ID NO: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, or 285; iii) amino acid position corresponding to position 640 of SEQ ID NO: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, or 285; iv) amino acid position corresponding to position 666 of SEQ ID NO: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, or 285; v) amino acid position corresponding to position 680 of SEQ ID NO: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16,

[0172] 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 271, 272, 273, 274, 275, 276,

[0173] 277, 278, 279, 280, 281, 282, 283, 284, or 285; vi) amino acid position corresponding to position 740 of SEQ ID NO: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 182, 183, 184, 185, 186, 187, 188, 189, 190,

[0174] 191, 192, 193, 194, 195, 196, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, or 285; vii) amino acid position corresponding to position 785 of SEQ ID NO: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, or 285; viii) amino acid position corresponding to position 910 of SEQ ID NO: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, or 285; and ix) amino acid position corresponding to position 1077 of SEQ ID NO: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, or 285.

[0175] The polymerase (e.g., RT) may be fused directly to the variant LPG10145 RGN polypeptide or a peptide linker can connect the polymerase (e.g., RT) and the variant LPG10145 RGN polypeptide. In those embodiments wherein the polymerase (e.g., RT) is inserted into the variant LPG10145 RGN polypeptide, there can be peptide linkers on one or both ends of the polymerase (e.g., RT). Any suitable peptide linker can be used to connect the polymerase (e.g., RT) and variant LPG10145 RGN polypeptide, but one suitable peptide linker comprises one or more copies of SGGS (SEQ ID NO: 241). In some embodiments, the peptide linker comprises 1 SGGS (SEQ ID NO: 241) sequence, 2 SGGS (SEQ ID NO: 241) sequences, 3 SGGS (SEQ ID NO: 241) sequences, 4 SGGS (SEQ ID NO: 241) sequences, or more, such that the linker sequence can be 4, 8, 12, or 16 amino acids long. The linker between the polymerase (e.g., RT) and variant LPG10145 RGN polypeptide can be 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, or more amino acids in length. The peptide linker separating the polymerase (e.g., RT) and variant LPG10145 RGN polypeptide can also comprise an NLS, such as but not limited to those disclosed elsewhere herein, including SEQ ID NO: 36, 37, 234, or 235, wherein the NLSs can be connected by peptide linkers (such as SGGS, SEQ ID NO: 241). In some embodiments, the peptide linker separating the polymerase (e.g., RT) and variant LPG10145 RGN polypeptide comprises more than one localization sequence, such as 2, 3, or more localization sequences. In some embodiments, a peptide linker has a formula of -(SGGS)x-NLSm-(SGGS)y-NLSn-(SGGS)z-, wherein each of x, y, or z is 0, 1, 2, 3, or 4; and wherein each of m or n is 0 or 1. In certain embodiments, the peptide linker has a formula of -(SGGS)x-NLSm-(SGGS)y-NLSn-(SGGS)z-, wherein each of x or y is 0, 1, 2, 3, or 4, and y is 0; and wherein one of m or n is 0, and the other is 1. In other embodiments, the peptide linker has a formula of -(SGGS)x-NLSm-(SGGS)y-NLSn-(SGGS)z-, wherein each of x, y, or z is 0, 1, 2, 3, or 4; and wherein each of m or n is 1.

[0176] The presently disclosed polymerases (e.g., RTs), variant LPG10145 RGN polypeptides, fusion proteins, or PEs can comprise at least one nuclear localization signal (NLS) to enhance transport of the protein to the nucleus of a cell. Nuclear localization signals are known in the art and generally comprise a stretch of basic amino acids (see, e.g., Lange et al., J. Biol. Chem. (2007) 282:5101-5105). In some embodiments, the polymerase (e.g., RT), variant LPG10145 RGN polypeptide, fusion protein, or PE comprises 2, 3, 4, 5, 6 or more nuclear localization signals. The nuclear localization signal(s) can be a heterologous NLS. Non-limiting examples of nuclear localization signals useful for the presently disclosed polymerases (e.g., RTs), variant LPG10145 RGN polypeptides, fusion proteins, or PEs are the nuclear localization signals of SV40 Large T-antigen, nucleoplasmin, and c-Myc (see, e.g., Ray et al. (2015) Bioconjug Chem 26(6): 1004-7). In some embodiments, the polymerase (e.g., RT), variant LPG10145 RGN polypeptide, fusion protein, or PE comprises the NLS sequence set forth as SEQ ID NO: 36, 37, 234, and / or 235. The polymerase (e.g., RT), variant LPG10145 RGN polypeptide, fusion protein, or PE can comprise one or more NLS sequences at its N-terminus, C- terminus, or both the N-terminus and C-terminus. For example, the polymerase (e.g., RT), variant LPG10145 RGN polypeptide, fusion protein, or PE can comprise two NLS sequences at the N-terminal region and four NLS sequences at the C-terminal region. In some embodiments, a peptide linker can connect the NLS to the polymerase (e.g., RT), variant LPG10145 RGN polypeptide, or fusion protein.

[0177] Other localization signal sequences known in the art that localize polypeptides to particular subcellular location(s) can also be used to target the polymerases (e.g., RTs), variant LPG10145 RGN polypeptides, fusion proteins, or PEs, including, but not limited to, plastid localization sequences, mitochondrial localization sequences, and dual-targeting signal sequences that target to both the plastid and mitochondria (see, e.g., Nassoury and Morse (2005) Biochim Biophys Acta 1743:5-19; Kunze and Berger (2015) Front Physiol dx.doi.org / 10.3389 / fphys.2015.00259; Herrmann and Neupert (2003) IUBMB Life 55:219-225; Soil (2002) Curr Opin Plant Biol 5:529-535; Carrie and Small (2013) Biochim Biophys Acta 1833:253-259; Carrie et al. (2009) FEBSJT16'. 1187-1195; Silva-Filho (2003) Curr Opin Plant Biol 6:589- 595; Peeters and Small (2001) Biochim Biophys Acta 1541:54-63; Murcha et aZ. (2014) J Exp Bot 65:6301- 6335; Mackenzie (2005) Trends Cell Biol 15:548-554; Glaser et al. (1998) Plant Mol Biol 38:311-338).

[0178] Polymerases (e.g., RTs), variant LPG10145 RGN polypeptides, fusion proteins, or PEs can comprise at least one cell-penetrating domain that facilitates cellular uptake of the polymerases (e.g., RTs), variant LPG10145 RGN polypeptides, fusion proteins, or PEs. Cell-penetrating domains are known in the art and generally comprise stretches of positively charged amino acid residues (i.e., polycationic cell-penetrating domains), alternating polar amino acid residues and non-polar amino acid residues (i.e., amphipathic cellpenetrating domains), or hydrophobic amino acid residues (i.e., hydrophobic cell-penetrating domains) (see, e.g., Milletti F. (2012) Drug Discov Today 17:850-860). A non-limiting example of a cell-penetrating domain is the trans-activating transcriptional activator (TAT) from the human immunodeficiency virus 1.

[0179] The nuclear localization signal, plastid localization signal, mitochondrial localization signal, dualtargeting localization signal, and / or cell-penetrating domain can be located at the amino-terminus (N- terminus), the carboxyl -terminus (C-terminus), and / or in an internal location of the polymerase (e.g., RT), variant LPG10145 RGN polypeptide, fusion protein, or PE.

[0180] Polymerases (e.g., RTs), variant LPG10145 RGN polypeptides, fusion proteins, or PEs can also comprise a purification tag, which is any molecule that can be utilized to isolate a protein or fused protein from a mixture (e.g., biological sample, culture medium). Non-limiting examples of purification tags include biotin, myc, maltose binding protein (MBP), glutathione-S-transferase (GST), and 3X FLAG tag.

[0181] A PE of the present disclosure can comprise an amino acid sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater sequence identity to the amino acid sequence set forth as any one of SEQ ID NOs: 135, and 162-181. In some embodiments, a PE of the disclosure comprises the amino acid sequence set forth as any one of SEQ ID NOs: 135, and 162-181.

[0182] 2. PEgRNA

[0183] A PE system utilizes a polymerase editing guide RNA (“PEgRNA”). The PEgRNA is a guide RNA that both specifies the target sequence and provides the template for polymerization of the replacement strand containing a desired edit by way of an extension engineered onto the RGN guide RNA or a part thereof, referred to herein as an extension arm. The PEgRNA can be a single guide RNA, wherein the extension arm can be at the 5' or 3' end, or at an internal portion of the guide RNA, or multiple polynucleotides (e.g., a dual guide RNA). In embodiments wherein the PEgRNA is a dual guide RNA, the extension arm can be at the 5' or 3' end, or at an internal portion of the crRNA or tracrRNA molecule. The template for polymerization within an extension arm is referred to herein as the DNA synthesis template. In those embodiments wherein the polymerase of the PE is an RT, the DNA synthesis template can be referred to as the reverse transcriptase template (RTT). The RGN is guided to the target sequence by the PEgRNA and in those embodiments wherein the RGN is a nickase with an inactivated HNH domain and active RuvC domain, the RGN nickase nicks the non-target strand upstream of the sequence to be edited and upstream of the PAM, creating a 3' flap on the non-target strand. The PEgRNA includes a primer binding site (PBS) that is complementary to the 3' flap of the non-target strand. The PBS can be at least about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides in length. In certain embodiments, the PEgRNA comprises a PBS that is at least 5 (e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20) nucleotides in length. In some embodiments, the PEgRNA may comprise a PBS that is 9, 11, 12, 13, 15, or 18 nucleotides in length. Hybridization of the PBS and 3' flap of the non-target strand allows polymerization of the replacement strand containing the edit using the DNA synthesis template in the extension of the PEgRNA. The DNA synthesis template can be at least about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more nucleotides in length. In certain embodiments, the PEgRNA comprises a DNA synthesis template that is at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, or at least 50 nucleotides in length. In some embodiments, the PEgRNA comprises a DNA synthesis template that is 19, 22, 23, 24, 25, 26, 29, 30, 32, 34, 37, 38, 39, 40, 42, or 46 nucleotides in length. The DNA synthesis template comprises the desired edit, which can be a substitution of one or more nucleotides, a deletion of one or more nucleotides, or an addition of one or more nucleotides.

[0184] The extension arm of the PEgRNA can be formed from RNA or DNA. In the case of an RNA extension, the polymerase of the polymerase editor can be an RNA-dependent DNA polymerase (such as a reverse transcriptase). In the case of a DNA extension, the polymerase of the polymerase editor may be a DNA-dependent DNA polymerase.

[0185] The replacement strand containing the desired edit (e.g., substitution, deletion, or addition) shares the same sequence as the non-target strand of the target sequence to be edited (with the exception that it includes the desired edit). Through DNA repair and / or replication machinery, the non-target strand of the target sequence is replaced by the newly synthesized replacement strand containing the desired edit. In some cases, polymerase editing may be thought of as a “search-and-replace” genome editing technology since the polymerase editors not only search and locate the desired target sequence to be edited, but at the same time, encode a replacement strand containing a desired edit which is installed in place of the corresponding non-target strand of the target sequence. Thus, in some embodiments, a guide RNA of the disclosure comprises an extension comprising an edit template for polymerase editing.

[0186] An RTT of the present disclosure can comprise a nucleotide sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater sequence identity to the nucleotide sequence set forth as any one of SEQ ID NOs: 226- 233. In some embodiments, an RTT of the present disclosure comprises the nucleotide sequence set forth as any one of SEQ ID NOs: 226-233. A PBS of the present disclosure can comprise a nucleotide sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater sequence identity to the nucleotide sequence set forth as any one of SEQ ID NOs: 221-223. In some embodiments, a PBS of the present disclosure comprises the nucleotide sequence set forth as any one of SEQ ID NOs: 221-223. A PEgRNA of the present disclosure can comprise a nucleotide sequence having at least 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater sequence identity to the nucleotide sequence set forth as any one of SEQ ID NOs: 129-132, and 197-220. In some embodiments, a PEgRNA of the present disclosure comprises the nucleotide sequence set forth as any one of SEQ ID NOs: 129-132, and 197-220. j. Nicking guide RNA

[0187] In order to reduce the possibility that the edit introduced by a polymerase editor is removed due to mismatch repair of the edited strand, a nicking guide RNA can be used. A “nicking guide RNA” is a guide RNA that targets a sequence within the unedited strand at a site nearby and opposite to the original nick and guides the RGN nickase of the PE system to this unedited strand to introduce a single -stranded nick. The nicking guide RNA can be designed to match the edited sequence introduced by the PEgRNA, but not the original unedited sequence, to ensure that the nicking occurs after the editing event on the non-target strand takes place.

[0188] VII. Nucleotides Encoding RNA-guided nucleases, base editing polypeptides, polymerases, fusion proteins, base editors, polymerase editors, CRISPR RNA, tracrRNA, and / or guide RNA

[0189] The present disclosure provides polynucleotides comprising the presently disclosed CRISPR RNAs, tracrRNAs, and / or sgRNAs and polynucleotides comprising a nucleotide sequence encoding the presently disclosed variant LPG10145 RGNs, CRISPR RNAs, tracrRNAs, sgRNAs, and / or PEgRNAs. Systems of the disclosure can comprise polynucleotides comprising or encoding guide RNAs or PEgRNAs and polynucleotides comprising a nucleotide sequence encoding variant LPG10145 RGN polypeptides, fusion proteins comprising the same, polymerases (e.g., RTs), base editing polypeptides, base editors, and / or PEs. Presently disclosed polynucleotides include those comprising or encoding a CRISPR repeat comprising the nucleotide sequence set forth as SEQ ID NO: 33, 244, or 245, or an active variant or fragment thereof that when comprised within a guide RNA is capable of directing the sequence -specific binding of an associated RNA-guided nuclease to a target sequence of interest. Also disclosed are polynucleotides comprising or encoding a tracrRNA comprising the nucleotide sequence set forth as SEQ ID NO: 34, 246, 247, or 248, or an active variant or fragment thereof that when comprised within a guide RNA is capable of directing the sequence -specific binding of an associated RNA-guided nuclease to a target sequence of interest.

[0190] Polynucleotides are also provided that encode an RNA-guided nuclease comprising the amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein the amino acid residue at one or more of the following amino acid positions differs from the corresponding amino acid residue in SEQ ID NO: 1: 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975. Polynucleotides are provided that encode an RNA-guided nuclease comprising the amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein the amino acid residue at one or more of the following amino acid positions is a positively charged amino acid residue: 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975. Polynucleotides are provided that encode an RNA-guided nuclease comprising the amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein the amino acid residue at one or more of the following amino acid positions is an R: 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975.

[0191] Polynucleotides are also provided that encode an RNA-guided nuclease comprising the amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein the amino acid residue at amino acid position 778 and / or 969 differs from the corresponding amino acid residue in SEQ ID NO: 1. Polynucleotides are provided that encode an RNA-guided nuclease comprising the amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein the amino acid residue at amino acid position 778 and / or 969 is a positively charged amino acid residue. Polynucleotides are provided that encode an RNA-guided nuclease comprising the amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein the amino acid residue at amino acid position 778 and / or 969 is an R.

[0192] Polynucleotides are also provided that encode an RNA-guided nuclease comprising the amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein amino acid residues at positions 778 and 856 differ from the corresponding amino acid residues in SEQ ID NO: 1. In some embodiments, polynucleotides are provided that encode an RNA-guided nuclease comprising the amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein the amino acid residues at positions 778 and 856 are positively charged amino acid residues.

[0193] Polynucleotides are also provided that encode an RNA-guided nuclease comprising the amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein amino acid residues at positions 55, 647, 778, and 969 differ from the corresponding amino acid residues in SEQ ID NO: 1. In some embodiments, polynucleotides are also provided that encode an RNA-guided nuclease comprising the amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein amino acid residues at positions 55, 647, 778, and 969 are positively charged amino acid residues.

[0194] In some embodiments, a polynucleotide is provided that encodes an RNA-guided nuclease comprising the amino acid sequence selected from: (a) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 52 is an R; (b) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 55 is an R; (c) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 86 is an R; (d) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 472 is an R; (e) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 533 is an R; (f) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 541 is an R; (g) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 643 is substituted by an R; (h) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 647 is an R; (i) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 653 is an R; (j) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 745 is an R; (k) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position 774 is an R; (1) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 778 is an R; (m) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 780 is an R; (n) the amino acid sequence set forth as SEQ ID NO: 1, wherein A at amino acid position 795 is an R; (o) the amino acid sequence set forth as SEQ ID NO: 1, wherein Q at amino acid position 822 is an R; (p) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 843 is an R; (q) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 856 is an R; (r) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 871 is an R; (s) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 872 is an R; (t) the amino acid sequence set forth as SEQ ID NO: 1, wherein D at amino acid position 900 is an R; (u) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position 911 is an R; (v) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 913 is an R; (w) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position 954 is an R; (x) the amino acid sequence set forth as SEQ ID NO: 1, wherein V at amino acid position 958 is an R; (y) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 968 is an R; (z) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 969 is an R; (aa) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 973 is an R; (bb) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 974 is an R; (cc) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position 975 is an R; (dd) the amino acid sequence set forth as SEQ ID NO: 2; (ee) the amino acid sequence set forth as SEQ ID NO: 3; (ff) the amino acid sequence set forth as SEQ ID NO: 4; (gg) the amino acid sequence set forth as SEQ ID NO: 5; (hh) the amino acid sequence set forth as SEQ ID NO: 6; (ii) the amino acid sequence set forth as SEQ ID NO: 7; (jj) the amino acid sequence set forth as SEQ ID NO: 8; (kk) the amino acid sequence set forth as SEQ ID NO: 9; (11) the amino acid sequence set forth as SEQ ID NO: 10; (mm) the amino acid sequence set forth as SEQ ID NO: 11; (nn) the amino acid sequence set forth as SEQ ID NO: 12; (oo) the amino acid sequence set forth as SEQ ID NO: 13; (pp) the amino acid sequence set forth as SEQ ID NO: 14; (qq) the amino acid sequence set forth as SEQ ID NO: 15; and (rr) the amino acid sequence set forth as SEQ ID NO: 16, and active fragments or variants thereof that retain the ability to bind to a target nucleotide sequence in an RNA-guided sequence -specific manner.

[0195] In some embodiments, a polynucleotide is provided that encodes an RNA-guided nuclease comprising the amino acid sequence selected from: (a) the amino acid sequence set forth as SEQ ID NO: 2; (b) the amino acid sequence set forth as SEQ ID NO: 3; (c) the amino acid sequence set forth as SEQ ID NO: 4; (d) the amino acid sequence set forth as SEQ ID NO: 5; (e) the amino acid sequence set forth as SEQ ID NO: 6; (f) the amino acid sequence set forth as SEQ ID NO: 7; (g) the amino acid sequence set forth as SEQ ID NO: 8; (h) the amino acid sequence set forth as SEQ ID NO: 9; (i) the amino acid sequence set forth as SEQ ID NO: 10; (j) the amino acid sequence set forth as SEQ ID NO: 11; (k) the amino acid sequence set forth as SEQ ID NO: 12; (1) the amino acid sequence set forth as SEQ ID NO: 13; (m) the amino acid sequence set forth as SEQ ID NO: 14; (n) the amino acid sequence set forth as SEQ ID NO: 15; and (o) the amino acid sequence set forth as SEQ ID NO: 16, and active fragments or variants thereof that retain the ability to bind to a target nucleotide sequence in an RNA-guided sequence-specific manner.

[0196] The use of the term "polynucleotide" or “nucleic acid molecule” is not intended to limit the present disclosure to polynucleotides comprising DNA. Those of ordinary skill in the art will recognize that polynucleotides can comprise ribonucleotides (RNA) and combinations of ribonucleotides and deoxyribonucleotides. Such deoxyribonucleotides and ribonucleotides include both naturally occurring molecules and synthetic analogues. These include peptide nucleic acids (PNAs), PNA-DNA chimers, locked nucleic acids (LNAs), and phosphothiorate linked sequences. The polynucleotides disclosed herein also encompass all forms of sequences including, but not limited to, single -stranded forms, double -stranded forms, DNA-RNA hybrids, triplex structures, stem-and-loop structures, and the like.

[0197] In some embodiments, the polynucleotide encoding a presently disclosed variant LPG10145 RGN polypeptide, fusion protein comprising the same, polymerase (e.g., RT), base editing polypeptides, base editors, and / or PE is an mRNA (messenger RNA) molecule. An mRNA refers to any polynucleotide which encodes a polypeptide of interest and which is capable of being translated to produce the encoded polypeptide of interest in vitro, in vivo, in situ, or ex vivo. In embodiments, the basic components of an mRNA molecule include at least a coding region, a 5'UTR, a 3 'UTR, a 5' cap and a poly-A tail. A 5' UTR, situated 5' of a coding sequence and transcribed as part of an mRNA, may comprise various regulatory elements, including, e.g., 5' cap structure, G-quadruplex structure (G4), stem-loop structure, and internal ribosome entry sites (IRES), which can control translation initiation of the mRNA. A 3' UTR, situated 3' of a coding sequence and transcribed as part of an mRNA, can be involved in numerous regulatory processes including transcript cleavage, stability and polyadenylation, translation, and mRNA localization. The 3' UTR can serve as a binding site for numerous regulatory proteins and small non-coding RNAs, e.g., microRNAs. A 5' UTR and / or a 3' UTR heterologous to an mRNA originates from an organism or species that is different from that of the mRNA, or if from the same organism or species as the mRNA, is substantially modified from its native form in composition and / or genomic locus by deliberate human intervention.

[0198] In some embodiments, inclusion of a 5' UTR and / or a 3' UTR heterologous to an mRNA encoding a variant LPG10145 RGN polypeptide, fusion protein comprising same, polymerase (e.g., RT), and / or PE of the disclosure improves polypeptide synthesis from the mRNA in a tissue (e.g., liver, or cells in vitro, such as stem cells, hepatocytes or lymphocytes). Heterologous 5' UTRs and / or 3' UTRs may, for example, increase protein synthesis by increasing the time that the mRNA remains in translating polysomes (message stability) and / or the rate at which ribosomes initiate translation on the mRNA (message translation efficiency). Thus, inclusion of a 5' UTR and / or a 3' UTR heterologous to an mRNA encoding variant LPG10145 RGN polypeptide, fusion protein comprising the same, polymerase (e.g., RT), base editing polypeptides, base editors, and / or PE of the disclosure can lead to prolonged and / or increased polypeptide synthesis, enabling improved editing of a target polynucleotide by the PE or PE system. In some embodiments, the enhanced polypeptide synthesis from an mRNA occurs in a tissue-specific manner. Heterologous UTR sequences are described, for example, in US 2023 / 0050143 and US 2017 / 0252461.

[0199] In some embodiments, an mRNA encoding a variant LPG10145 RGN, fusion protein comprising the same, polymerase (e.g., RT), base editing polypeptides, base editors, and / or PE useful in the presently disclosed methods and compositions can include one or more structural and / or chemical modifications or alterations which impart useful properties to the polynucleotide. For instance, a useful property of an mRNA includes the lack of a substantial induction of the innate immune response of a cell into which the mRNA is introduced. A “structural” feature or modification is one in which two or more linked nucleotides are inserted, deleted, duplicated, inverted or randomized in an mRNA without significant chemical modification to the nucleotides themselves. Because chemical bonds will necessarily be broken and reformed to effect a structural modification, structural modifications are of a chemical nature and hence are chemical modifications. However, structural modifications will result in a different sequence of nucleotides. Chemical modifications to mRNA can involve inclusion of 5 -methylcytosine, Nl-methyl-pseudouridine, pseudouridine, 2-thiouridine, 4-thiouridine, 5-methoxyuridine, 2'Fluoroguanosine, 2'Fluorouridine, 5- bromouridine, 5-(2-carbomethoxyvinyl) uridine, 5-[3(l-E-propenylamino)] uridine, a-thiocytidine, N6- methyladenosine, 5 -methylcytidine, N4-acetylcytidine, 5 -formylcytidine, or combinations thereof, in an mRNA.

[0200] The nucleic acid molecules encoding variant LPG10145 RGNs, fusion proteins comprising the same, polymerases (e.g., RTs), base editing polypeptides, base editors, and / or PEs can be codon optimized for expression in an organism of interest. A "codon-optimized” coding sequence is a polynucleotide coding sequence having its frequency of codon usage designed to mimic the frequency of preferred codon usage or transcription conditions of a particular host cell. Expression in the particular host cell or organism is enhanced as a result of the alteration of one or more codons at the nucleic acid level such that the translated amino acid sequence is not changed. Nucleic acid molecules can be codon optimized, either wholly or in part. Codon tables and other references providing preference information for a wide range of organisms are available in the art (see, e.g., Campbell and Gowri (1990) Plant Physiol. 92: 1-11 for a discussion of plantpreferred codon usage). Methods are available in the art for synthesizing plant-preferred genes or mammalian (for example human) codon-optimized coding sequences. See, for example, U.S. Patent Nos. 5,380,831, and 5,436,391, and Murray et al. (1989) Nucleic Acids Res. 17:477-498, herein incorporated by reference.

[0201] Polynucleotides encoding the variant LPG10145 RGNs, fusion proteins comprising the same, polymerases (e.g., RTs), base editing polypeptides, base editors, PEs, crRNAs, tracrRNAs, sgRNAs, and / or PEgRNAs provided herein can be provided in expression cassettes for in vitro expression or expression in a cell, organelle, embryo, or organism of interest. The cassette will include 5' and 3' regulatory sequences operably linked to a polynucleotide encoding a variant LPG10145 RGN, fusion protein comprising the same, polymerase (e.g., RT), base editing polypeptides, base editors, PE, crRNA, tracrRNAs, sgRNAs, and / or PEgRNAs provided herein that allows for expression of the polynucleotide. The cassette may additionally contain at least one additional gene or genetic element to be cotransformed into the organism. Where additional genes or elements are included, the components are operably linked. The term “operably linked” is intended to mean a functional linkage between two or more elements. For example, an operable linkage between a promoter and a coding region of interest (e.g., region coding for an RGN, crRNA, tracrRNAs, sgRNAs, and / or PEgRNAs) is a functional link that allows for expression of the coding region of interest. Operably linked elements may be contiguous or non-contiguous. When used to refer to the joining of two protein coding regions, by “operably linked” or “operably fused” is intended that the coding regions are in the same reading frame, even if one is inserted into another. In some embodiments, polypeptides that are “operably fused” or “operably linked” means that the structure and / or biological activity of each individual peptide is also present in the fusion. Alternatively, the additional gene(s) or element(s) can be provided on multiple expression cassettes. For example, the nucleotide sequence encoding a presently disclosed variant LPG10145 RGN, fusion protein comprising the same, polymerase (e.g., RT), base editing polypeptide, base editor, or PE can be present on one expression cassette, whereas the nucleotide sequence encoding a crRNA, tracrRNA, or complete guide RNA (or PEgRNA) can be on a separate expression cassette. Such an expression cassette is provided with a plurality of restriction sites and / or recombination sites for insertion of the polynucleotides to be under the transcriptional regulation of the regulatory regions. The expression cassette may additionally contain a selectable marker gene.

[0202] The expression cassette will include in the 5'-3' direction of transcription, a transcriptional (and, in some embodiments, translational) initiation region (z.e., a promoter), a variant LPG10145 RGN-, a fusion protein-, a polymerase- (e.g., RT-), base editing polypeptide-, base editor-, PE-, crRNA-, tracrRNA-, sgRNA-, and / or PEgRNA- encoding polynucleotide of the invention, and a transcriptional (and in some embodiments, translational) termination region (i. e. , termination region) functional in the organism of interest. The promoters of the invention are capable of directing or driving expression of a coding sequence in a host cell. The regulatory regions (e.g., promoters, transcriptional regulatory regions, and translational termination regions) may be endogenous or heterologous to the host cell or to each other. As used herein, “heterologous” in reference to a regulatory region that is heterologous to another regulatory region or to the host cell, is a regulatory region that is not found with another regulatory region or in the host cell in nature. The heterologous regulatory region can originate from a foreign species or from the same species. The heterologous regulatory region can be in its native form or is substantially modified from its native form in composition and / or genomic locus by deliberate human intervention. For example, a chimeric gene comprises a coding sequence operably linked to a transcription initiation region that is heterologous to the coding sequence. Similarly, a nucleic acid molecule that is heterologous to another nucleic acid molecule is a nucleic acid molecule that is not found with another nucleic acid molecule in nature. The heterologous nucleic acid molecule can originate from a foreign species or from the same species. The heterologous nucleic acid molecule can be in its native form or is substantially modified from its native form in composition and / or genomic locus by deliberate human intervention. For example, an mRNA encoding a variant LPG10145 RGN polypeptide can comprise a heterologous 5’ UTR, wherein the 5’ UTR is not present as part of the mRNA encoding the variant LPG10145 RGN polypeptide in nature.

[0203] Convenient termination regions include ones from simian virus (SV40), human growth hormone (hGH), bovine growth hormone (BGH), and rabbit beta-globin (rbGlob). See also Proudfoot (1991) Cell 64:671-674; Munroe et al. (1990) Gene 91: 151-158; Schek et al. (1992) Molecular and Cellular Biology 12( 12): 5386-5393; Gil and Proudfoot (1987) Cell 49(3)399-406; Goodwin and Rottman (1992) The Journal of Biological Chemistry 267(23): 16330-16334; and Lanoix and Acheson (1988) EMBO J. 7(8): 2515-2522. Additional termination regions are available from the Ti-plasmid of A. iumefaciens. such as the octopine synthase and nopaline synthase termination regions. See also Guerineau et al. (1991) Mol. Gen. Genet. 262: 141-144; Proudfoot (1991) Cell 64:671-674; Sanfacon et al. (1991) Genes Dev. 5: 141-149; Mogen et a / . (1990) Plant Cell 2: 1261-1272; Munroe et al. (1990) Gene 91: 151-158; Ballas et al. (1989) Nucleic Acids Res. 17:7891-7903; and Joshi et al. (1987) Nucleic Acids Res. 15:9627-9639.

[0204] Additional regulatory signals include, but are not limited to, transcriptional initiation start sites, operators, activators, enhancers, other regulatory elements, ribosomal binding sites, an initiation codon, termination signals, and the like. See, for example, U.S. Pat. Nos. 5,039,523 and 4,853,331; EPO 0480762A2; Sambrook et al. (1992) Molecular Cloning: A Laboratory Manual, ed. Maniatis et al. (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y.), hereinafter "Sambrook 11"; Davis et al., eds. (1980) Advanced Bacterial Genetics (Cold Spring Harbor Laboratory Press), Cold Spring Harbor, N.Y., and the references cited therein.

[0205] In preparing the expression cassette, the various DNA fragments may be manipulated, so as to provide for the DNA sequences in the proper orientation and, as appropriate, in the proper reading frame. Toward this end, adapters or linkers may be employed to join the DNA fragments or other manipulations may be involved to provide for convenient restriction sites, removal of superfluous DNA, removal of restriction sites, or the like. For this purpose, in vitro mutagenesis, primer repair, restriction, annealing, resubstitutions, e.g., transitions and transversions, may be involved.

[0206] A number of promoters can be used in the practice of the invention. The promoters can be selected based on the desired outcome. The nucleic acids can be combined with constitutive, inducible, growth stage-specific, cell type-specific, tissue-preferred, tissue-specific, or other promoters for expression in the organism of interest. See, for example, promoters set forth in WO 99 / 43838 and in US Patent Nos: 8,575,425; 7,790,846; 8,147,856; 8,586832; 7,772,369; 7,534,939; 6,072,050; 5,659,026; 5,608,149; 5,608,144; 5,604,121; 5,569,597; 5,466,785; 5,399,680; 5,268,463; 5,608,142; and 6,177,611; herein incorporated by reference. Exemplary constitutive promoters for expression in cells of the present disclosure include: an SV40 early promoter; a mouse mammary tumor virus long terminal repeat (LTR) promoter; adenovirus major late promoter (Ad MLP); a herpes simplex virus (HSV) promoter; a cytomegalovirus (CMV) promoter such as the CMV immediate early promoter region (CMVIE); a rous sarcoma virus (RSV) promoter; a human ubiquitin C promoter (UBC); a human U6 small nuclear promoter (U6); an enhanced U6 promoter; a human Hl promoter from RNA polymerase III (Hl); a human elongation factor la promoter (EF1A); a human betaactin promoter (ACTB); a human or mouse phosphoglycerate kinase 1 promoter (PGK); a chicken P-Actin promoter coupled with CMV early enhancer (CAGG); a yeast transcription elongation factor promoter (TEF1); and the like. See, for example, Miyagishi et al. (2002) Nature Biotechnology 20:497-500; Xia et al. (2003) Nucleic Acids Res. 31(17):el00-el00; Pasleau et al. (1985) Gene 38:227-232; Martin-Gallardo et al. (1988) Gene 70: 51-56; Oellig and Seliger (1990) J Neurosci Res 26: 390-396; Manthorpe et al. (1993) Hum Gene Ther 4: 419-431; Yew et al. (1997) Hum Gene Ther 8: 575-584; Xu et al. (2001) Gene 272: 149-156; Nguyen et al. (2008) J Surg Res 148: 60-66; Costa et al. (2005) Nat Meth. 2:259-260; Lam and Truong (2020) ACS Synth. Biol. 9( 10) :2625-2631.

[0207] For expression in plants, constitutive promoters also include CaMV 35S promoter (Odell et al. (1985) Nature 313:810-812); rice actin (McElroy et al. (1990) Plant Cell 2: 163-171); ubiquitin (Christensen et al. (1989) Plant Mol. Biol. 12:619-632 and Christensen et a / . (1992) Plant Mol. Biol. 18:675-689); pEMU (Last et al. (1991) Theor. Appl. Genet. 81:581-588); and MAS (Velten t a / . (1984) EMBO J. 3:2723-2730).

[0208] Examples of inducible promoters are the Adhl promoter which is inducible by hypoxia or cold stress, the Hsp70 promoter which is inducible by heat stress, the PPDK promoter and the pepcarboxylase promoter which are both inducible by light. Also useful are promoters which are chemically inducible, such as the In2-2 promoter which is safener induced (U.S. Pat. No. 5,364,780), the Axigl promoter which is auxin induced and tapetum specific but also active in callus (PCT US01 / 22169), the steroid-responsive promoters (see, for example, the ERE promoter which is estrogen induced, and the glucocorticoid-inducible promoter in Schena et a / . (1991) Proc. Natl. Acad. Sci. USA 88: 10421-10425 and McNellis et al. (1998) Plant J. 14(2): 247-257) and tetracycline-inducible and tetracycline-repressible promoters (see, for example, Gatz et al. (1991) Mol. Gen. Genet. 227:229-237, and U.S. Pat. Nos. 5,814,618 and 5,789,156), herein incorporated by reference.

[0209] Tissue-specific or tissue-preferred promoters can be utilized to target expression of an expression construct within a particular tissue. In some embodiments, the tissue-specific or tissue-preferred promoters are active in mammalian tissue. Examples of tissue-specific or tissue-preferred promoters include promoters that initiate transcription preferentially in certain tissues, such as white blood cells (e.g., CD4 T cell), heart, kidney, liver, CNS, eye, pancreas, skeletal muscle, and testis. In certain embodiments, the tissue-specific or tissue-preferred promoters are active in plant tissue. Examples of promoters under developmental control in plants include promoters that initiate transcription preferentially in certain tissues, such as leaves, roots, fruit, seeds, or flowers. A "tissue specific" promoter is a promoter that initiates transcription only in certain tissues. Unlike constitutive expression of genes, tissue-specific expression is the result of several interacting levels of gene regulation. As such, promoters from homologous or closely related plant species can be preferable to use to achieve efficient and reliable expression of transgenes in particular tissues. In some embodiments, the expression comprises a tissue-preferred promoter. A "tissue preferred" promoter is a promoter that initiates transcription preferentially, but not necessarily entirely or solely in certain tissues.

[0210] In some embodiments, the nucleic acid molecules encoding a variant LPGI0145 RGN, fusion protein comprising the same, polymerase (e.g., RT), base editing polypeptide, base editor, PE, crRNA, tracrRNA, sgRNA, and / or PEgRNA comprise a cell type-specific promoter. A "cell type specific" promoter is a promoter that primarily drives expression in certain cell types in one or more organs. Some examples of cells in which cell type specific promoters may be primarily active include, for example, a primary cell, a neuronal cell, a glial cell, an adipocyte, a cardiomyocyte, a smooth muscle cell, a photoreceptor cell, and a retinal ganglia cell. Some examples of plant cells in which cell type specific promoters functional in plants may be primarily active include, for example, BETL cells, vascular cells in roots, leaves, stalk cells, and stem cells. The nucleic acid molecules can also include cell type preferred promoters. A "cell type preferred" promoter is a promoter that primarily drives expression mostly, but not necessarily entirely or solely in certain cell types in one or more organs. Some examples of cells in which cell type preferred promoters may be preferentially active include, for example, a primary cell, a neuron, an adipocyte, a cardiomyocyte, a smooth muscle cell, and a photoreceptor cell. Some examples of plant cells in which cell type preferred promoters functional in plants may be preferentially active include, for example, BETL cells, vascular cells in roots, leaves, stalk cells, and stem cells.

[0211] The nucleic acid sequences encoding the variant LPGI0145 RGNs, fusion proteins comprising the same, polymerases (e.g., RTs), base editing polypeptides, base editors, PEs, crRNAs, tracrRNAs, sgRNAs, and / or PEgRNAs can be operably linked to a promoter sequence that is recognized by a phage RNA polymerase for example, for in vitro mRNA synthesis. In such embodiments, the in vv / ro-transcribcd RNA can be purified for use in the methods described herein. For example, the promoter sequence can be a T7, T3, or SP6 promoter sequence or a variation of a T7, T3, or SP6 promoter sequence. In such embodiments, the expressed protein and / or RNAs can be purified for use in the methods of genome modification described herein.

[0212] In certain embodiments, the polynucleotide encoding the variant LPGI0145 RGN, fusion protein comprising the same, polymerase (e.g., RT), base editing polypeptide, base editor, PE, crRNA, tracrRNA, sgRNA, and / or PEgRNA also can be linked to a polyadenylation signal (e.g., SV40 polyA signal and other signals functional in plants) and / or at least one transcriptional termination sequence. Additionally, the sequence encoding the RGN also can be linked to sequence(s) encoding at least one nuclear localization signal, at least one cell-penetrating domain, and / or at least one signal peptide capable of trafficking proteins to particular subcellular locations, as described elsewhere herein. The polynucleotide encoding the variant LPG10145 RGN, fusion protein comprising the same, polymerase (e.g., RT), base editing polypeptide, base editor, PE, crRNA, tracrRNA, sgRNA, and / or PEgRNA can be present in a vector or multiple vectors. A “vector” refers to a polynucleotide composition for transferring, delivering, or introducing a nucleic acid into a host cell. Suitable vectors include plasmid vectors, phagemids, cosmids, artificial / mini-chromosomes, transposons, and viral vectors (e.g., lentiviral vectors, adeno-associated viral vectors, baculoviral vector). The vector can comprise additional expression control sequences (e.g., enhancer sequences, Kozak sequences, polyadenylation sequences, transcriptional termination sequences), selectable marker sequences (e.g., antibiotic resistance genes), origins of replication, and the like. Additional information can be found in "Current Protocols in Molecular Biology" Ausubel et al., John Wiley & Sons, New York, 2003 or "Molecular Cloning: A Laboratory Manual" Sambrook & Russell, Cold Spring Harbor Press, Cold Spring Harbor, N.Y., 3rd edition, 2001.

[0213] The vector can also comprise a selectable marker gene for the selection of transformed cells. Selectable marker genes are utilized for the selection of transformed cells or tissues. Marker genes include genes encoding antibiotic resistance, such as those encoding neomycin phosphotransferase II (NEO) and hygromycin phosphotransferase (HPT), as well as genes conferring resistance to herbicidal compounds, such as glufosinate ammonium, bromoxynil, imidazolinones, and 2,4-dichlorophenoxyacetate (2,4-D). Marker genes can include genes that allow selection for growth on a particular nutrient or substance, such as dihydrofolate reductase (DHFR; Simonsen and Levinson (1983) Proc. Natl. Acad. Sci. U.S.A. 80:2495-2499), histidinol dehydrogenase (hisD; Hartman and Mulligan (1988) Proc. Natl. Acad. Sci. U.S.A. 85:8047-8051), puromycin-N-acetyl transferase (PAC or puro; de la Luna et al. (1988) Gene 62: 121- 126), thymidine kinase (TK; Littlefield (1964) Science 145:709-710), and xanthine-guanine phosphoribosyltransferase (XGPRT or gpt; Mulligan and Berg (1981) Proc. Natl. Acad. Sci. U.S.A. 78:2072- 2076).

[0214] In some embodiments, the expression cassette or vector comprising the polynucleotide encoding the variant LPG10145 RGN polypeptide, fusion protein comprising the same, polymerase (e.g., RT), base editing polypeptide, base editor, and / or PE, can further comprise a polynucleotide encoding a crRNA and / or a tracrRNA, or the crRNA and tracrRNA combined to create a sgRNA or combined to be a part of a PEgRNA. The polynucleotide sequence(s) encoding the crRNA, tracrRNA, gRNA, and / or PEgRNA can be operably linked to at least one transcriptional control sequence for expression of the crRNA, tracrRNA, gRNA, and / or PEgRNA in the organism or host cell of interest. For example, the polynucleotide encoding the crRNA, tracrRNA, gRNA, and / or PEgRNA can be operably linked to a promoter sequence that is recognized by RNA polymerase III (Pol III). Examples of suitable Pol III promoters include, but are not limited to, mammalian U6, U3, Hl, and 7SL RNA promoters and rice U6 and U3 promoters, such as the human U6 promoter set forth as SEQ ID NO: 41, as well as the promoters disclosed in International Publication No. WO 2022 / 261394, which is herein incorporated by reference in its entirety, including those set forth herein as SEQ ID NOs: 114-123.

[0215] As indicated, expression constructs comprising nucleotide sequences encoding the variant LPG10145 RGNs, fusion proteins comprising the same, polymerases (e.g., RTs), base editing polypeptides, base editors, and / or PEs, crRNAs, tracrRNAs, sgRNAs, and / or PEgRNAs can be used to transform organisms of interest. Methods for transformation involve introducing a nucleotide construct into an organism of interest. By "introducing" is intended to introduce the nucleotide construct to the host cell in such a manner that the construct gains access to the interior of the host cell. The methods of the invention do not require a particular method for introducing a nucleotide construct to a host organism, only that the nucleotide construct gains access to the interior of at least one cell of the host organism. The host cell can be a eukaryotic or prokaryotic cell. In particular embodiments, the eukaryotic host cell is a plant cell, a mammalian cell, an avian cell, or an insect cell. In some embodiments, the eukaryotic cell that comprises or expresses a presently disclosed variant LPG10145 RGN, fusion protein comprising the same, polymerase (e.g., RT), base editing polypeptide, base editor, or PE, or that has been modified by a presently disclosed variant LPG10145 RGN, fusion protein comprising the same, polymerase (e.g., RT), base editing polypeptide, base editor, or PE, is a human cell. In some embodiments, the eukaryotic cell that comprises or expresses a presently disclosed variant LPG10145 RGN, fusion protein comprising the same, polymerase (e.g., RT), base editing polypeptide, base editor, or PE, or that has been modified by a presently disclosed variant LPG10145 RGN, fusion protein comprising the same, polymerase (e.g., RT), base editing polypeptide, base editor, or PE, is a primary cell. The term "primary cell" refers to a cell isolated directly from a multicellular organism. Primary cells typically have undergone very few population doublings and are therefore more representative of the main functional component of the tissue from which they are derived in comparison to continuous (tumor or artificially immortalized) cell lines. In some cases, primary cells are cells that have been isolated and then used immediately. In other cases, primary cells cannot divide indefinitely and thus cannot be cultured for long periods of time in vitro. In some embodiments, a primary cell is a primary T cell. In some embodiments, the eukaryotic cell that comprises or expresses a presently disclosed variant LPG10145 RGN, fusion protein comprising the same, polymerase (e.g., RT), base editing polypeptide, base editor, or PE, or that has been modified by a presently disclosed variant LPG10145 RGN, fusion protein comprising the same, polymerase (e.g., RT), base editing polypeptide, base editor, or PE, is a cell of hematopoietic origin, such as an immune cell (i.e., a cell of the innate or adaptive immune system) including but not limited to a B cell, a T cell, a natural killer (NK) cell, a pluripotent stem cell, an induced pluripotent stem cell, a chimeric antigen receptor T (CAR-T) cell, a monocyte, a macrophage, and a dendritic cell. In some embodiments, the eukaryotic cell that comprises or expresses a presently disclosed variant LPG10145 RGN, fusion protein comprising the same, polymerase (e.g., RT), base editing polypeptide, base editor, or PE, or that has been modified by a presently disclosed variant LPG10145 RGN, fusion protein comprising the same, polymerase (e.g., RT), base editing polypeptide, base editor, or PE, is an ocular cell, muscle cell (e.g., skeletal muscle cell), epithelial cell (e.g., lung epithelial cell), diseased cell (e.g., tumor cell). Methods for introducing nucleotide constructs into plants and other host cells are known in the art including, but not limited to, stable transformation methods, transient transformation methods, and virus- mediated methods.

[0216] The methods result in a transformed organism, such as a plant, including whole plants, as well as plant organs (e.g., leaves, stems, roots, etc.), seeds, plant cells, propagules, embryos and progeny of the same. Plant cells can be differentiated or undifferentiated (e.g. callus, suspension culture cells, protoplasts, leaf cells, root cells, phloem cells, pollen).

[0217] "Transgenic organisms" or "transformed organisms" or "stably transformed" organisms or cells or tissues refers to organisms that have incorporated or integrated a polynucleotide encoding a variant LPG10145 RGN, fusion protein comprising the same, polymerase (e.g., RT), base editing polypeptide, base editor, PE, crRNA, tracrRNA, gRNA, and / or PEgRNA of the invention. It is recognized that other exogenous or endogenous nucleic acid sequences or DNA fragments may also be incorporated into the host cell. Agrobacterium-and biolistic-mediated transformation remain the two predominantly employed approaches for transformation of plant cells. However, transformation of a host cell may be performed by infection, transfection, microinjection, electroporation, microprojection, biolistics or particle bombardment, electroporation, silica / carbon fibers, ultrasound mediated, PEG mediated, calcium phosphate coprecipitation, polycation DMSO technique, DEAE dextran procedure, and viral mediated, liposome mediated and the like. Viral-mediated introduction of a polynucleotide encoding a variant LPG10145 RGN, fusion protein comprising the same, polymerase (e.g., RT), base editing polypeptide, base editor, PE, crRNA, tracrRNA, gRNA, and / or PEgRNA includes retroviral, lentiviral, adenoviral, and adeno-associated viral mediated introduction and expression, as well as the use of Caulimoviruses, Geminiviruses, and RNA plant viruses.

[0218] Transformation protocols as well as protocols for introducing polypeptides or polynucleotide sequences into plants may vary depending on the type of host cell (e.g, monocot or dicot plant cell) targeted for transformation. Methods for transformation are known in the art and include those set forth in US Patent Nos: 8,575,425; 7,692,068; 8,802,934; 7,541,517; each of which is herein incorporated by reference. See, also, Rakoczy-Trojanowska, M. (2002) Cell Mol Biol Lett. 7:849-858; Jones et al. (2005) Plant Methods 1:5; Rivera et al. (2012) Physics of Life Reviews 9:308-345; Bartlett et al. (2008) Plant Methods 4: 1-12; Bates, G.W. (1999) Methods in Molecular Biology 111:359-366; Binns and Thomashow (1988) Annual Reviews in Microbiology 42:575-606; Christou, P. (1992) The Plant Journal 2:275-281; Christou, P. (1995) Euphytica 85: 13-27; Tzfira et al. (2004) TRENDS in Genetics 20:375-383; Yao et al. (2006) Journal of Zupan and Zambryski (1995) Plant Physiology 107: 1041-1047; Jones et al. (2005) Plant Methods 1:5.

[0219] Transformation may result in stable or transient incorporation of the nucleic acid into the cell. "Stable transformation" is intended to mean that the nucleotide construct introduced into a host cell integrates into the genome of the host cell and is capable of being inherited by the progeny thereof. "Transient transformation" is intended to mean that a polynucleotide is introduced into the host cell and does not integrate into the genome of the host cell.

[0220] Methods for transformation of chloroplasts are known in the art. See, for example, Svab et al. (1990) Proc. Nail. Acad. Sci. USA 87:8526-8530; Svab and Maliga (1993) Proc. Natl. Acad. Sci. USA 90:913-917; Svab and Maliga (1993) EMBO J. 12:601-606. The method relies on particle gun delivery of DNA containing a selectable marker and targeting of the DNA to the plastid genome through homologous recombination. Additionally, plastid transformation can be accomplished by transactivation of a silent plastid-borne transgene by tissue-preferred expression of a nuclear-encoded and plastid-directed RNA polymerase. Such a system has been reported in McBride et al. (1994) Proc. Natl. Acad. Sci. USA 91:7301- 7305.

[0221] The cells that have been transformed may be grown into a transgenic organism, such as a plant, in accordance with conventional ways. See, for example, McCormick et al. (1986) Plant Cell Reports 5:81-84. These plants may then be grown, and either pollinated with the same transformed strain or different strains, and the resulting hybrid having constitutive expression of the desired phenotypic characteristic identified. Two or more generations may be grown to ensure that the polynucleotide encoding a variant LPG10145 RGN, fusion protein comprising the same, polymerase (e.g., RT), base editing polypeptide, base editor, PE, crRNA, tracrRNA, gRNA, and / or PEgRNA is stably maintained and inherited and then seeds harvested to ensure the presence of the polynucleotide encoding a variant LPG10145 RGN, fusion protein comprising the same, polymerase (e.g., RT), base editing polypeptide, base editor, PE, crRNA, tracrRNA, gRNA, and / or PEgRNAs. In this manner, the present invention provides a transformed plant or plant part having a nucleotide construct of the invention, for example, an expression cassette of the invention, stably incorporated into their genome. Seed having an expression cassette of the disclosure stably incorporated into their genome can be referred to as "transgenic seed" .

[0222] Alternatively, cells that have been transformed may be introduced into an organism. These cells could have originated from the organism, wherein the cells are transformed in an ex vivo approach. These cells can be autologous (originated and returned to the same subject), allogeneic (the donor and recipient subjects are of the same species).

[0223] The sequences provided herein may be used for transformation of any plant species, including, but not limited to, monocots and dicots. Examples of plants of interest include, but are not limited to, com (maize), sorghum, wheat, sunflower, tomato, crucifers, peppers, potato, cotton, rice, soybean, sugarbeet, sugarcane, tobacco, barley, and oilseed rape, Brassica sp., alfalfa, rye, millet, safflower, peanuts, sweet potato, cassava, coffee, coconut, pineapple, citrus trees, cocoa, tea, banana, avocado, fig, guava, mango, olive, papaya, cashew, macadamia, almond, oats, vegetables, ornamentals, and conifers.

[0224] Vegetables include, but are not limited to, tomatoes, lettuce, green beans, lima beans, peas, and members of the genus Curcumis such as cucumber, cantaloupe, and musk melon. Ornamentals include, but are not limited to, azalea, hydrangea, hibiscus, roses, tulips, daffodils, petunias, carnation, poinsettia, and chrysanthemum. Preferably, plants of the present invention are crop plants (for example, maize, sorghum, wheat, sunflower, tomato, crucifers, peppers, potato, cotton, rice, soybean, sugarbeet, sugarcane, tobacco, barley, oilseed rape, etc.).

[0225] As used herein, the term plant includes plant cells, plant protoplasts, plant cell tissue cultures from which plants can be regenerated, plant calli, plant clumps, and plant cells that are intact in plants or parts of plants such as embryos, pollen, ovules, seeds, leaves, flowers, branches, fruit, kernels, ears, cobs, husks, stalks, roots, root tips, anthers, and the like. Grain is intended to mean the mature seed produced by commercial growers for purposes other than growing or reproducing the species. Progeny, variants, and mutants of the regenerated plants are also included within the scope of the invention, provided that these parts comprise the introduced polynucleotides. Further provided is a processed plant product or byproduct that retains the sequences disclosed herein, including for example, soymeal.

[0226] The polynucleotides encoding the variant LPG10145 RGNs, fusion proteins comprising the same, polymerases (e.g., RTs), base editing polypeptides, base editors, PEs, crRNAs, tracrRNAs, gRNAs, and / or PEgRNAs or comprising the crRNAs, tracrRNAs, gRNAs, and / or PEgRNAs can also be used to transform any prokaryotic species, including but not limited to, archaea and bacteria (e.g., Bacillus sp., Klebsiella sp. Streptomyces sp., Rhizobium sp., Escherichia sp., Pseudomonas sp., Salmonella sp., Shigella sp., Vibrio sp., Yersinia sp., Mycoplasma sp., Agrobacterium, Lactobacillus sp.).

[0227] The polynucleotides encoding the variant LPG10145 RGNs, fusion proteins comprising the same, polymerases (e.g., RTs), base editing polypeptides, base editors, PEs, crRNAs, tracrRNAs, gRNAs, and / or PEgRNAs or comprising the crRNAs, tracrRNAs, gRNAs, and / or PEgRNAs can be used to transform any eukaryotic species, including but not limited to animals (e.g., mammals, insects, fish, birds, and reptiles), fungi, amoeba, algae, and yeast.

[0228] Conventional viral and non-viral based gene transfer methods can be used to introduce nucleic acids in mammalian, insect, or avian cells or target tissues. Such methods can be used to administer nucleic acids encoding components of an RGN, base editing, or PE system to cells in culture, or in a host organism. Non- viral vector delivery systems include DNA plasmids, RNA (e.g. a transcript of a vector described herein), naked nucleic acid, and nucleic acid complexed with a delivery vehicle, such as a liposome. Viral vector delivery systems include DNA and RNA viruses, which have either episomal or integrated genomes after delivery to the cell. For a review of gene therapy procedures, see Anderson, Science 256: 808- 813 (1992); Nabel & Feigner, TIBTECH 11:211-217 (1993); Mitani & Caskey, TIBTECH 11: 162-166 (1993); Dillon, TIBTECH 11: 167-175 (1993); Miller, Nature 357:455-460 (1992); Van Brunt, Biotechnology 6(10): 1149- 1154 (1988); Vigne, Restorative Neurology and Neuroscience 8:35-36 (1995); Kremer & Perricaudet, British Medical Bulletin 51 ( 1): 31-44 (1995); Haddada et al., in Current Topics in Microbiology and Immunology, Doerfler and Bohm (eds) (1995); and Yu et al., Gene Therapy 1: 13-26 (1994).

[0229] Methods of non-viral delivery of nucleic acids include lipofection, nucleofection, microinjection, biolistics, virosomes, liposomes, immunoliposomes, polycation or lipid: nucleic acid conjugates, naked DNA, artificial virions, and agent-enhanced uptake of DNA. Lipofection is described in e.g., U.S. Pat. Nos. 5,049,386, 4,946,787; and 4,897,355) and lipofection reagents are sold commercially (e.g., Transfectam ™ and Lipofectin™). Cationic and neutral lipids that are suitable for efficient receptorrecognition lipofection of polynucleotides include those of Feigner, WO 91 / 17424; WO 91 / 16024. Delivery can be to cells (e.g. in vitro or ex vivo administration) or target tissues (e.g. in vivo administration). The preparation of lipidmucleic acid complexes, including targeted liposomes such as immunolipid complexes, is well known to one of skill in the art (see, e.g., Crystal, Science 270:404-410 (1995); Blaese et al., Cancer Gene Ther. 2:291- 297 (1995); Behr et al., Bioconjugate Chem. 5:382-389 (1994); Remy et al., Bioconjugate Chem. 5:647-654 (1994); Gao et al., Gene Therapy 2:710-722 (1995); Ahmad et al., Cancer Res. 52:4817-4820 (1992); U.S. Pat. Nos. 4,186,183, 4,217,344, 4,235,871, 4,261,975, 4,485,054, 4,501,728, 4,774,085, 4,837,028, and 4,946,787).

[0230] The use of RNA or DNA viral based systems for the delivery of nucleic acids takes advantage of highly evolved processes for targeting a virus to specific cells in the body and trafficking the viral payload to the nucleus. Viral vectors can be administered directly to patients (in vivo) or they can be used to treat cells in vitro, and the modified cells may optionally be administered to patients (ex vivo). Conventional viral based systems could include retroviral, lentivirus, adenoviral, adeno-associated and herpes simplex virus vectors for gene transfer. Integration in the host genome is possible with the retrovirus, lentivirus, and adeno-associated virus gene transfer methods, often resulting in long term expression of the inserted transgene. Additionally, high transduction efficiencies have been observed in many different cell types and target tissues.

[0231] The tropism of a retrovirus can be altered by incorporating foreign envelope proteins, expanding the potential target population of target cells. Uentiviral vectors are retroviral vectors that are able to transduce or infect non-dividing cells and typically produce high viral titers. Selection of a retroviral gene transfer system would therefore depend on the target tissue. Retroviral vectors are comprised of cis-acting long terminal repeats with packaging capacity for up to 6-10 kb of foreign sequence. The minimum cis-acting UTRs are sufficient for replication and packaging of the vectors, which are then used to integrate the therapeutic gene into the target cell to provide permanent transgene expression. Widely used retroviral vectors include those based upon murine leukemia virus (MuUV), gibbon ape leukemia virus (GaUV), Simian Immuno deficiency virus (SIV), human immuno deficiency virus (HIV), and combinations thereof (see, e.g., Buchscher et al., J. Viral. 66:2731-2739 (1992); Johann et al., J. Viral. 66: 1635-1640 (1992); Sommnerfelt et al., Viral. 176:58-59 (1990); Wilson et al., J. Viral. 63:2374-2378 (1989); Miller et al., 1. Viral. 65:2220-2224 (1991); WO 1994 / 026877).

[0232] In applications where transient expression is preferred, adenoviral based systems may be used. Adenoviral based vectors are capable of very high transduction efficiency in many cell types and do not require cell division. With such vectors, high titer and levels of expression have been obtained. This vector can be produced in large quantities in a relatively simple system. Adeno-associated virus ("AAV") vectors may also be used to transduce cells with target nucleic acids, e.g., in the in vitro production of nucleic acids and peptides, and for in vivo and ex vivo gene therapy procedures (see, e.g., West et al., Virology 160:38-47 (1987); U.S. Pat. No. 4,797,368; WO 93 / 24641; Katin, Human Gene Therapy 5:793-801 (1994);

[0233] Muzyczka, 1. Clin. Invest. 94: 1351 (1994).

[0234] The term “adeno-associated virus” or “AAV” as used herein refers to a member of the class of viruses associated with this name and belonging to the genus dependoparvovirus, family Parvoviridae. Multiple serotypes of this virus are known to be suitable for gene delivery; all known serotypes can infect cells from various tissue types. At least 11, sequentially numbered, have been described. Non-limiting exemplary serotypes useful in the compositions and methods disclosed herein include any of the 11 serotypes (e.g., AAV2, AAV5, AAV6, AAV8, AAV9), or variant serotypes, e.g., AAV-DJ. AAV is advantageous over other viral vectors for in vivo delivery of genes (e.g., encoding gene editing components) due to their low toxicity and low probability of causing insertional mutagenesis because it typically does not integrate into the host genome. AAV has a packaging limit of about 4.5 to 4.75 Kb.

[0235] Construction of recombinant AAV vectors are described in a number of publications, including U.S. Pat. No. 5, 173,414; Tratschin et al., Mol. Cell. Biol. 5:3251-3260 (1985); Tratschin, et al., Mol. Cell. Biol. 4:2072-2081 (1984); Hermonat & Muzyczka, PNAS 81:6466-6470 (1984); and Samulski et al., 1. Viral. 63:03822-3828 (1989). Packaging cells are typically used to form virus particles that are capable of infecting a host cell. Such cells include 293 cells, which package adenovirus, and q / J2 cells or PA317 cells, which package retrovirus.

[0236] Viral vectors used in gene therapy are usually generated by producing a cell line that packages a nucleic acid vector into a viral particle. The vectors typically contain the minimal viral sequences required for packaging and subsequent integration into a host, other viral sequences being replaced by an expression cassette for the polynucleotide(s) to be expressed. The missing viral functions are typically supplied in trans by the packaging cell line. For example, AAV vectors used in gene therapy typically only possess ITR sequences from the AAV genome which are required for packaging and integration into the host genome. Viral DNA is packaged in a cell line, which contains a helper plasmid encoding the other AAV genes, namely rep and cap, but lacking ITR sequences.

[0237] The cell line may also be infected with adenovirus as a helper. The helper virus promotes replication of the AAV vector and expression of AAV genes from the helper plasmid. The helper plasmid is not packaged in significant amounts due to a lack of ITR sequences. Contamination with adenovirus can be reduced by, e.g., heat treatment to which adenovirus is more sensitive than AAV. Additional methods for the delivery of nucleic acids to cells are known to those skilled in the art. See, for example, US20030087817, incorporated herein by reference.

[0238] In some embodiments, a host cell is transiently or non-transiently transfected with one or more vectors described herein. In some embodiments, a cell is transfected as it naturally occurs in a subject. In some embodiments, a cell that is transfected is taken from a subject. In some embodiments, a cell that is transfected is a eukaryotic cell. In some embodiments, the eukaryotic cell is an animal cell (e.g., mammals, humans, insects, fish, birds, and reptiles). In some embodiments, a cell that is transfected is a human cell. In some embodiments, a cell that is transfected is a cell of hematopoietic origin, such as an immune cell (i.e., a cell of the innate or adaptive immune system) including but not limited to a B cell, a T cell, a natural killer (NK) cell, a pluripotent stem cell, an induced pluripotent stem cell, a chimeric antigen receptor T (CAR-T) cell, a monocyte, a macrophage, and a dendritic cell.

[0239] In some embodiments, the cell is derived from cells taken from a subject, such as a cell line. In some embodiments, the cell or cell line is prokaryotic. In some embodiments, the cell or cell line is eukaryotic. In some embodiments, the cell or cell line may be mammalian, such as for example human, monkey, mouse, cow, swine, goat, hamster, rat, cat, or dog. In further embodiments, the cell or cell line is derived from insect, avian, plant, or fungal species. A wide variety of cell lines for tissue culture are known in the art. Examples of cell lines include, but are not limited to, C8161, CCRF-CEM, MOLT, mIMCD-3, NHDF, HeLaS3, Huhl, Huh4, Huh7, HUVEC, HASMC, HEKn, HEKa, MiaPaCell, Panel, PC-3, TF1, CTLL-2, CIR, Rat6, CVI, RPTE, A1O, T24, 182, A375, ARH-77, Calul, SW480, SW620, SKOV3, SK-UT, CaCo2, P388D1, SEM-K2, WEHI- 231, HB56, TIB55, lurkat, 145.01, LRMB, Bcl-1, BC-3, IC21, DLD2, Raw264.7, NRK, NRK-52E, MRC5, MEF, Hep G2, HeLa B, HeLa T4. COS, COS-1, COS-6, C0S-M6A, BS-C-1 monkey kidney epithelial, BALB / 3T3 mouse embryo fibroblast, 3T3 Swiss, 3T3-L1, 132-d5 human fetal fibroblasts; 10.1 mouse fibroblasts, 293-T, 3T3, 721, 9L, A2780, A2780ADR, A2780cis, A172, A20, A253, A431, A-549, ALC, B16, B35, BCP-I cells, BEAS-2B, bEnd.3, BHK-21, BR 293, BxPC3, C3H-10T1 / 2, C6 / 36, Cal-27, CHO, CHO-7, CHO-IR, CHO-K1, CHO-K2, CHO-T, CHO Dhfir- / -, COR-L23, COR- L23 / CPR, COR-L235010, CORL23 / R23, COS-7, COV-434, CML Tl, CMT, CT26, D17, DH82, DU145, DuCaP, EL4, EM2, EM3, EMT6 / AR1, EMT6 / AR10.0, FM3, H1299, H69, HB54, HB55, HCA2, HEK-293, HeLa, Hepalclc7, HL-60, HMEC, HT-29, lurkat, IY cells, K562 cells, Ku812, KCL22, KG1, KYO1, LNCap, Ma-Mel 1-48, MC-38, MCF-7, MCF-10A, MDA-MB-231, MDA-MB-468, MDA-MB-435, MDCKII, MDCKII, MOR / 0.2R, MONO-MAC 6, MTD-1A, MyEnd, NCI-H69 / CPR, NCI-H69 / LX10, NCI- H69 / LX20, NCI-H69 / LX4, NIH-3T3, NALM-1, NW- 145, OPCN / OPCT cell lines, Peer, PNT-1A / PNT 2, RenCa, RIN-5F, RMA / RMAS, Saos-2 cells, Sf-9, SkBr3, T2, T-47D, T84, THP1 cell line, U373, U87, U937, VCaP, Vero cells, WM39, WT-49, X63, YAC-1, YAR, and transgenic varieties thereof. Cell lines are available from a variety of sources known to those with skill in the art (see, e.g., the American Type Culture Collection (ATCC) (Manassas, Va.)).

[0240] In some embodiments, a cell transfected with one or more polynucleotides or vectors described herein is used to establish a new cell line comprising one or more vector-derived sequences. In some embodiments, a cell transiently transfected with the components of an RGN, base editing, or PE system as described herein (such as by transient transfection of one or more vectors, or transfection with RNA), and modified through the activity of an RGN, base editing, or PE system, is used to establish a new cell line comprising cells containing the modification but lacking any other exogenous sequence. In some embodiments, cells transiently or non-transiently transfected with one or more vectors described herein, or cell lines derived from such cells are used in assessing one or more test compounds.

[0241] In some embodiments, one or more vectors described herein are used to produce a non-human transgenic animal or transgenic plant. In some embodiments, the transgenic animal is a mammal, such as a mouse, rat, hamster, rabbit, cow, or pig. In some embodiments, the transgenic animal is a bird, such as a chicken or a duck. In some embodiments, the transgenic animal is an insect, such as a mosquito or a tick.

[0242] VIII. Variants and Fragments of Polypeptides and Polynucleotides

[0243] The present disclosure provides active engineered variants of a naturally-occurring (i. e. , wild-type) LPG10145 RNA-guided nuclease. In some embodiments, the wild-type LPG10145 RGN comprises the amino acid sequence set forth as SEQ ID NO: 1. In some embodiments, a variant LPG10145 RGN of the disclosure comprises an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein the amino acid residue at one or more of the following amino acid positions differs from the corresponding amino acid residue in SEQ ID NO: 1 at one or more corresponding amino acid residues.

[0244] When referring to a variant LPG10145 RGN that “comprises an amino acid sequence having at least x% (e.g., 85%) sequence identity to SEQ ID NO: 1, wherein the amino acid residue at one or more of the following amino acid positions differs”, it is intended to mean that the variant LPG10145 RGN comprises an amino acid sequence that has at least x% (e.g., 85%) sequence identity to the amino acid sequence set forth in SEQ ID NO: 1, and the amino acid residues at the recited positions differ from the corresponding amino acid residues in SEQ ID NO: 1. When referring to a variant LPG10145 RGN that “comprises an amino acid sequence set forth as SEQ ID NO: 1, wherein the amino acid residue at one or more of the following amino acid positions differs”, it is intended to mean that the variant LPG10145 RGN comprises an amino acid sequence that is identical to the amino acid sequence set forth in SEQ ID NO: 1, except that amino acid residues at the recited positions differ from the corresponding amino acid residues in SEQ ID NO: 1.

[0245] In some embodiments, a variant LPG10145 RGN of the disclosure comprises an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein the amino acid residue at one or more of the following amino acid positions differs from the corresponding amino acid residue in SEQ ID NO: 1: 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975. In some embodiments, a variant LPG10145 RGN of the disclosure comprises an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein the amino acid residue at one or more of the following amino acid positions is a positively charged amino acid residue: 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975. In some embodiments, a variant LPG10145 RGN of the disclosure comprises an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein the amino acid residue at one or more of the following amino acid positions is an R: 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975.

[0246] In some embodiments, a variant LPG10145 RGN of the disclosure comprises an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein the amino acid residue at amino acid position 778 and / or 969 differs from the corresponding amino acid residue in SEQ ID NO: 1. In some embodiments, a variant LPG10145 RGN of the disclosure comprises an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein the amino acid residue at amino acid position 778 and / or 969 is a positively charged amino acid residue. In some embodiments, a variant LPG10145 RGN of the disclosure comprises an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein the amino acid residue at amino acid position 778 and / or 969 is an R.

[0247] In some embodiments, a variant LPG10145 RGN of the disclosure comprises an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein amino acid residues at positions 778 and 856 differ from the corresponding amino acid residues in SEQ ID NO: 1. In some embodiments, a variant LPG10145 RGN of the disclosure comprises an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein the amino acid residues at positions 778 and 856 are positively charged amino acid residues.

[0248] In some embodiments, a variant LPG10145 RGN of the disclosure comprises an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein amino acid residues at positions 55, 647, 778, and 969 differ from the corresponding amino acid residues in SEQ ID NO: 1. In some embodiments, a variant LPG10145 RGN of the disclosure comprises an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein amino acid residues at positions 55, 647, 778, and 969 are positively charged amino acid residues.

[0249] In some embodiments, a variant LPG10145 RGN of the disclosure comprises an amino acid sequence selected from: (a) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 52 is an R; (b) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 55 is an R; (c) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 86 is an R; (d) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 472 is an R; (e) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 533 is an R; (f) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 541 is an R; (g) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 643 is substituted by an R; (h) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 647 is an R; (i) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 653 is an R; (j) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 745 is an R; (k) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position 774 is an R; (1) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 778 is an R; (m) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 780 is an R; (n) the amino acid sequence set forth as SEQ ID NO: 1, wherein A at amino acid position 795 is an R; (o) the amino acid sequence set forth as SEQ ID NO: 1, wherein Q at amino acid position 822 is an R; (p) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 843 is an R; (q) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 856 is an R; (r) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 871 is an R; (s) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 872 is an R; (t) the amino acid sequence set forth as SEQ ID NO: 1, wherein D at amino acid position 900 is an R; (u) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position 911 is an R; (v) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 913 is an R; (w) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position 954 is an R; (x) the amino acid sequence set forth as SEQ ID NO: 1, wherein V at amino acid position 958 is an R; (y) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 968 is an R; (z) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 969 is an R; (aa) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 973 is an R; (bb) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 974 is an R; (cc) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position 975 is an R; (dd) the amino acid sequence set forth as SEQ ID NO: 2; (ee) the amino acid sequence set forth as SEQ ID NO: 3; (ff) the amino acid sequence set forth as SEQ ID NO: 4; (gg) the amino acid sequence set forth as SEQ ID NO: 5; (hh) the amino acid sequence set forth as SEQ ID NO: 6; (ii) the amino acid sequence set forth as SEQ ID NO: 7; (jj) the amino acid sequence set forth as SEQ ID NO: 8; (kk) the amino acid sequence set forth as SEQ ID NO: 9; (11) the amino acid sequence set forth as SEQ ID NO: 10; (mm) the amino acid sequence set forth as SEQ ID NO: 11; (nn) the amino acid sequence set forth as SEQ ID NO: 12; (oo) the amino acid sequence set forth as SEQ ID NO: 13; (pp) the amino acid sequence set forth as SEQ ID NO: 14; (qq) the amino acid sequence set forth as SEQ ID NO: 15; and (rr) the amino acid sequence set forth as SEQ ID NO: 16, or an active variant or fragment thereof.

[0250] In some embodiments, a variant LPG10145 RGN of the disclosure comprises an amino acid sequence selected from: (a) the amino acid sequence set forth as SEQ ID NO: 2; (b) the amino acid sequence set forth as SEQ ID NO: 3; (c) the amino acid sequence set forth as SEQ ID NO: 4; (d) the amino acid sequence set forth as SEQ ID NO: 5; (e) the amino acid sequence set forth as SEQ ID NO: 6; (f) the amino acid sequence set forth as SEQ ID NO: 7; (g) the amino acid sequence set forth as SEQ ID NO: 8; (h) the amino acid sequence set forth as SEQ ID NO: 9; (i) the amino acid sequence set forth as SEQ ID NO: 10; (j) the amino acid sequence set forth as SEQ ID NO: 11; (k) the amino acid sequence set forth as SEQ ID NO: 12; (1) the amino acid sequence set forth as SEQ ID NO: 13; (m) the amino acid sequence set forth as SEQ ID NO: 14; (n) the amino acid sequence set forth as SEQ ID NO: 15; and (o) the amino acid sequence set forth as SEQ ID NO: 16, or an active variant or fragment thereof. In some embodiments, the present disclosure provides active variants and fragments of naturally- occurring CRISPR repeats, such as the sequence set forth as SEQ ID NO: 33, 244, or 245, and active variants and fragments of naturally-occurring tracrRNAs, such as the sequence set forth as SEQ ID NO: 34, 246, 247, or 248, and polynucleotides encoding the same.

[0251] While the activity of a variant or fragment may be altered compared to the polynucleotide or polypeptide of interest, the variant and fragment should retain the functionality of the polynucleotide or polypeptide of interest. For example, a variant or fragment may have increased activity, decreased activity, different spectrum of activity or any other alteration in activity when compared to the polynucleotide or polypeptide of interest.

[0252] Variant LPG10145 RGN polypeptides, such as those disclosed herein, will retain sequence-specific, RNA-guided DNA-binding activity. In particular embodiments, fragments and variants of variant LPG10145 RGN polypeptides, such as those disclosed herein, will retain nuclease activity (single-stranded or double-stranded).

[0253] The binding and / or cleaving activity of an active variant or fragment of a variant LPG10145 RGN disclosed herein can be dependent upon recognizing a protospacer adjacent motif (PAM) adjacent and 3’ to the target sequence. In some embodiments, the PAM comprises a consensus sequence of NNGG.

[0254] In some embodiments, the present disclosure provides an active variant or fragment of a variant LPG10145 RGN that has increased nuclease activity as compared to the RGN of SEQ ID NO: 1. In some embodiments, an active variant or fragment of a variant LPG10145 RGN disclosed herein comprising an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to any one of SEQ ID NOs: 2-16, 182-196, and 271-285 has increased nuclease activity as compared to the RGN of SEQ ID NO: 1. In some embodiments, an active variant or fragment of a variant LPG10145 RGN disclosed herein comprising the amino acid sequence of any one of SEQ ID NOs: 2-16, 182-196, and 271-285 has increased nuclease activity as compared to the RGN of SEQ ID NO: 1. In some embodiments, an active variant or fragment of a variant LPG10145 RGN disclosed herein has nuclease activity that is from about 80% to about 500%, from about 80% to about 200%, or from about 90% to about 150%, or from about 95% to about 120%, or is at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 100%, at least 101%, at least 102%, at least 103%, at least 104%, at least 105%, at least 106%, at least

[0255] 107%, at least 108%, at least 109%, at least 110%, at least 111%, at least 112%, at least 113%, at least

[0256] 114%, at least 115%, at least 116%, at least 117%, at least 118%, at least 119%, at least 120%, at least

[0257] 125%, at least 130%, at least 135%, at least 140%, at least 145%, at least 150%, at least 160%, at least

[0258] 170%, at least 180%, at least 190%, at least 200%, at least 250%, at least 300%, at least 350%, at least

[0259] 400%, at least 450%, at least 500%, or more, of the nuclease activity of a reference LPG10145 RGN. In some embodiments, a reference LPG10145 RGN has the amino acid sequence set forth as SEQ ID NO: 1. In some embodiments, a reference LPG10145 RGN is a variant of SEQ ID NO: 1 that lacks the corresponding mutations of the active variant or fragment thereof that has nuclease activity, e.g., as described above. In some embodiments, a reference LPG10145 RGN is a non-identical variant LPG10145 RGN. In some embodiments, a reference LPG10145 RGN is a non-identical variant LPG10145 RGN. Nuclease activity can be measured by assays known to one of skill in the art, including but not limited to, Tracking of Indels by DEcomposition (TIDE) analysis, flow cytometry, in vitro or in vivo cleavage assays wherein cleavage is confirmed using PCR, sequencing, or gel electrophoresis, with or without the attachment of an appropriate label (e.g., radioisotope, fluorescent substance) to the target sequence to facilitate detection of degradation products. In some embodiments, the nicking triggered exponential amplification reaction (NTEXPAR) assay can be used (see, e.g., Zhang et al. (2016) Chem. Set. 7:4951-4957). In vivo cleavage can be evaluated using the Surveyor assay (Guschin et al. (2010) Methods Mol Biol 649:247-256). In some embodiments, the efficiency of cleaving a target sequence is assessed by measuring the percentage of a target sequence or cells comprising the target sequence that comprise altered expression of the target sequence or of a polypeptide encoded by the target sequence. In some embodiments, the expression is measured by quantitative PCR, microarray, RNA-seq, flow cytometry, immunoblot, enzyme-linked immunosorbent assay (ELISA), protein immunoprecipitation, immunostaining, high performance liquid chromatography (HPLC), liquid chromatography-mass spectrometry (LC / MS), mass spectrometry, or a combination thereof. In some embodiments, the target sequence encodes a cell surface expressed protein, and the efficiency of cleaving the target sequence is assessed by measuring the percentage of cells comprising a reduction of the cell surface expressed protein as measured by flow cytometry.

[0260] In some embodiments, an active variant or fragment of a variant LPG10145 RGN disclosed herein, when operably fused to a deaminase, has increased base editing activity as compared to the RGN of SEQ ID NO: 1. In some embodiments, an active variant or fragment of a variant LPG10145 RGN disclosed herein comprising an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to any one of SEQ ID NOs: 2-16, 182-196, and 271-285, when operably fused to a deaminase, has increased base editing activity as compared to the RGN of SEQ ID NO: 1. In some embodiments, an active variant or fragment of a variant LPG10145 RGN disclosed herein comprising the amino acid sequence of any one of SEQ ID NOs: 2-16, 182-196, and 271-285, when operably fused to a deaminase, has increased base editing activity as compared to the RGN of SEQ ID NO: 1. In some embodiments, an active variant or fragment of a variant LPG10145 RGN disclosed herein, when operably fused to a deaminase, has base editing activity that is from about 80% to about 500%, from about 80% to about 200%, or from about 90% to about 150%, or from about 95% to about 120%, or is at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 100%, at least 101%, at least 102%, at least 103%, at least 104%, at least 105%, at least 106%, at least 107%, at least 108%, at least 109%, at least 110%, at least 111%, at least 112%, at least 113%, at least 114%, at least 115%, at least 116%, at least 117%, at least

[0261] 118%, at least 119%, at least 120%, at least 125%, at least 130%, at least 135%, at least 140%, at least

[0262] 145%, at least 150%, at least 160%, at least 170%, at least 180%, at least 190%, at least 200%, at least

[0263] 250%, at least 300%, at least 350%, at least 400%, at least 450%, at least 500%, or more, of the base editing activity of a reference LPG10145 RGN, when operably fused to the same deaminase. In some embodiments, a reference LPG10145 RGN has the amino acid sequence set forth as SEQ ID NO: 1. In some embodiments, a reference LPG10145 RGN is a variant of SEQ ID NO: 1 that lacks the corresponding mutations of the active variant or fragment thereof that has base editing activity, e.g., as described above. In some embodiments, a reference LPG10145 RGN is a non-identical variant LPG10145 RGN. Base editing activity can be measured by assays known to one of skill in the art, including but not limited to, transfection of mammalian cells with a base editor (e.g., comprising a variant LPG10145 RGN and a deaminase) and a guide RNA and detecting base editing by PCR amplification of the target sequence and next generation sequencing (NGS), as described in Example 4 of the present disclosure. Base editing assays are also described in International Publication No. WO 2022 / 056254, which is herein incorporated by reference in its entirety. In some embodiments, the efficiency of base editing a target sequence is assessed by measuring the percentage of a target sequence or cells comprising the target sequence that comprise altered expression of the target sequence or of a polypeptide encoded by the target sequence. In some embodiments, the expression is measured by quantitative PCR, microarray, RNA-seq, flow cytometry, immunoblot, enzyme- linked immunosorbent assay (ELISA), protein immunoprecipitation, immunostaining, high performance liquid chromatography (HPLC), liquid chromatography-mass spectrometry (LC / MS), mass spectrometry, or a combination thereof.

[0264] In some embodiments, an active variant or fragment of a variant LPG10145 RGN disclosed herein, when operably fused to a polymerase (e.g., RT), has increased editing activity as compared to SEQ ID NO: 1. In some embodiments, an active variant or fragment of a variant LPG10145 RGN disclosed herein comprising an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to any one of SEQ ID NOs: 2-16, 182-196, and 271-285, when operably fused to a polymerase (e.g., RT), has increased editing activity as compared to the RGN of SEQ ID NO: 1. In some embodiments, an active variant or fragment of a variant LPG10145 RGN disclosed herein comprising the amino acid sequence of any one of SEQ ID NOs: 2-16, 182-196, and 271-285, when operably fused to a polymerase (e.g., RT), has increased editing activity as compared to the RGN of SEQ ID NO: 1. In some embodiments, an active variant or fragment of a variant LPG10145 RGN disclosed herein, when operably fused to a polymerase (e.g., RT), has editing activity that is from about 80% to about 500%, from about 80% to about 200%, or from about 90% to about 150%, or from about 95% to about 120%, or is at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 100%, at least 101%, at least 102%, at least 103%, at least 104%, at least 105%, at least 106%, at least 107%, at least 108%, at least 109%, at least

[0265] 110%, at least 111%, at least 112%, at least 113%, at least 114%, at least 115%, at least 116%, at least

[0266] 117%, at least 118%, at least 119%, at least 120%, at least 125%, at least 130%, at least 135%, at least

[0267] 140%, at least 145%, at least 150%, at least 160%, at least 170%, at least 180%, at least 190%, at least

[0268] 200%, at least 250%, at least 300%, at least 350%, at least 400%, at least 450%, at least 500%, or more, of the editing activity of a reference LPG10145 RGN, when operably fused to the same polymerase. In some embodiments, a reference LPG10145 RGN has the amino acid sequence set forth as SEQ ID NO: 1. In some embodiments, a reference LPG10145 RGN is a variant of SEQ ID NO: 1 that lacks the corresponding mutations of the active variant or fragment thereof that has editing activity, e.g., as described above. In some embodiments, a reference LPG10145 RGN is a non-identical variant LPG10145 RGN. Editing activity can be measured by assays known to one of skill in the art, including but not limited to, transfection of mammalian cells with a polymerase editor (e.g., comprising a variant LPG10145 RGN and a reverse transcriptase) and a PEgRNA and detecting editing by PCR amplification of the target sequence and next generation sequencing (NGS), as described in Examples 5 and 6 of the present disclosure. The editing rate (RT Edit %) can be determined by calculating the percent of read counts assigned to alleles containing the desired edit and, optionally, no additional edits. In some embodiments, the efficiency of editing a target sequence is assessed by measuring the percentage of a target sequence or cells comprising the target sequence that comprise altered expression of the target sequence or of a polypeptide encoded by the target sequence. In some embodiments, the expression is measured by quantitative PCR, microarray, RNA-seq, flow cytometry, immunoblot, enzyme-linked immunosorbent assay (ELISA), protein immunoprecipitation, immunostaining, high performance liquid chromatography (HPLC), liquid chromatography-mass spectrometry (LC / MS), mass spectrometry, or a combination thereof.

[0269] Active variants and fragments of CRISPR repeats, such as those disclosed herein, will retain the ability, when part of a guide RNA (comprising a tracrRNA), to bind to and guide an RNA-guided nuclease or base editor or PE comprising the same (complexed with the guide RNA) to a target nucleotide sequence in a sequence-specific manner.

[0270] Active variants and fragments of tracrRNAs, such as those disclosed herein, will retain the ability, when part of a guide RNA (comprising a CRISPR RNA), to guide an RNA-guided nuclease or base editor or PE comprising the same (complexed with the guide RNA) to a target nucleotide sequence in a sequencespecific manner.

[0271] Active variants and fragments of PEgRNAs disclosed herein, will retain the ability to bind and guide a PE to a target nucleotide sequence in a sequence-specific manner.

[0272] Active variants and fragments of base editors disclosed herein, will retain the ability to, when associated with a guide RNA, chemically modify (e.g., deaminate) a nucleobase, resulting in conversion from one nucleobase to another. Active variants and fragments of PEs disclosed herein, will retain the ability to, when associated with a PEgRNA, edit a double-stranded polynucleotide through the replacement of a target sequence using the template sequence of the PEgRNA.

[0273] The term “fragment” refers to a portion of a polynucleotide or polypeptide sequence of the invention. "Fragments", “active fragments”, or "biologically active portions" include polynucleotides comprising a sufficient number of contiguous nucleotides to retain the biological activity (z. e. , binding to and directing an RGN in a sequence-specific manner to a target nucleotide sequence when comprised within a guide RNA). "Fragments" or "biologically active portions" include polypeptides comprising a sufficient number of contiguous amino acid residues to retain the biological activity (i.e. , binding to a target nucleotide sequence in a sequence -specific manner when complexed with a guide RNA). Fragments of the RGN proteins include those that are shorter than the full-length sequences due to the use of an alternate downstream start site. Such biologically active portions can be prepared by recombinant techniques and evaluated for activity.

[0274] A biologically active portion of an RGN protein can be a polypeptide that comprises, for example, 10, 25, 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, or more contiguous amino acid residues of an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein the amino acid residue at one or more of the following amino acid positions differs from the corresponding amino acid residue in SEQ ID NO: 1: 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975. A biologically active portion of an RGN protein can be a polypeptide that comprises, for example, 10, 25, 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, or more contiguous amino acid residues of an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein the amino acid residue at one or more of the following amino acid positions is a positively charged amino acid residue: 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975. A biologically active portion of an RGN protein can be a polypeptide that comprises, for example, 10, 25, 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, or more contiguous amino acid residues of an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein the amino acid residue at one or more of the following amino acid positions is an R: 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975.

[0275] A biologically active portion of an RGN protein can be a polypeptide that comprises, for example, 10, 25, 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, or more contiguous amino acid residues of an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein the amino acid residue at amino acid position 778 and / or 969 differs from the corresponding amino acid residue in SEQ ID NO: 1. A biologically active portion of an RGN protein can be a polypeptide that comprises, for example, 10, 25, 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, or more contiguous amino acid residues of an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein the amino acid residue at amino acid position 778 and / or 969 is a positively charged amino acid residue. A biologically active portion of an RGN protein can be a polypeptide that comprises, for example, 10, 25, 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, or more contiguous amino acid residues of an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein the amino acid residue at amino acid position 778 and / or 969 is an R.

[0276] A biologically active portion of an RGN protein can be a polypeptide that comprises, for example, 10, 25, 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, or more contiguous amino acid residues of an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein amino acid residues at positions 778 and 856 differ from the corresponding amino acid residues in SEQ ID NO: 1. A biologically active portion of an RGN protein can be a polypeptide that comprises, for example, 10, 25, 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, or more contiguous amino acid residues of an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein amino acid residues at positions 778 and 856 are positively charged amino acid residues.

[0277] A biologically active portion of an RGN protein can be a polypeptide that comprises, for example, 10, 25, 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, or more contiguous amino acid residues of an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein amino acid residues at positions 55, 647, 778, and 969 differ from the corresponding amino acid residues in SEQ ID NO: 1. A biologically active portion of an RGN protein can be a polypeptide that comprises, for example, 10, 25, 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, or more contiguous amino acid residues of an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein amino acid residues at positions 55, 647, 778, and 969 are positively charged amino acid residues.

[0278] In some embodiments, a biologically active portion of an RGN protein can be a polypeptide that comprises, for example, 10, 25, 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, or more contiguous amino acid residues of an amino acid sequence selected from: (a) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 52 is an R; (b) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 55 is an R; (c) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 86 is an R; (d) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 472 is an R; (e) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 533 is an R; (f) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 541 is an R; (g) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 643 is substituted by an R; (h) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 647 is an R; (i) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 653 is an R; (j) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 745 is an R; (k) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position 774 is an R; (1) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 778 is an R; (m) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 780 is an R; (n) the amino acid sequence set forth as SEQ ID NO: 1, wherein A at amino acid position 795 is an R; (o) the amino acid sequence set forth as SEQ ID NO: 1, wherein Q at amino acid position 822 is an R; (p) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 843 is an R; (q) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 856 is an R; (r) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 871 is an R; (s) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 872 is an R; (t) the amino acid sequence set forth as SEQ ID NO: 1, wherein D at amino acid position 900 is an R; (u) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position 911 is an R; (v) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 913 is an R; (w) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position 954 is an R; (x) the amino acid sequence set forth as SEQ ID NO: 1, wherein V at amino acid position 958 is an R; (y) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 968 is an R; (z) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 969 is an R; (aa) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 973 is an R; (bb) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 974 is an R; (cc) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position 975 is an R; (dd) the amino acid sequence set forth as SEQ ID NO: 2; (ee) the amino acid sequence set forth as SEQ ID NO: 3; (ff) the amino acid sequence set forth as SEQ ID NO: 4; (gg) the amino acid sequence set forth as SEQ ID NO: 5; (hh) the amino acid sequence set forth as SEQ ID NO: 6; (ii) the amino acid sequence set forth as SEQ ID NO: 7; (jj) the amino acid sequence set forth as SEQ ID NO: 8; (kk) the amino acid sequence set forth as SEQ ID NO: 9; (11) the amino acid sequence set forth as SEQ ID NO: 10; (mm) the amino acid sequence set forth as SEQ ID NO: 11; (nn) the amino acid sequence set forth as SEQ ID NO: 12; (oo) the amino acid sequence set forth as SEQ ID NO: 13; (pp) the amino acid sequence set forth as SEQ ID NO: 14; (qq) the amino acid sequence set forth as SEQ ID NO: 15; and (rr) the amino acid sequence set forth as SEQ ID NO: 16.

[0279] In some embodiments, a biologically active portion of an RGN protein can be a polypeptide that comprises, for example, 10, 25, 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, or more contiguous amino acid residues of an amino acid sequence selected from: (a) the amino acid sequence set forth as SEQ ID NO: 2; (b) the amino acid sequence set forth as SEQ ID NO: 3; (c) the amino acid sequence set forth as SEQ ID NO: 4; (d) the amino acid sequence set forth as SEQ ID NO: 5; (e) the amino acid sequence set forth as SEQ ID NO: 6; (f) the amino acid sequence set forth as SEQ ID NO: 7; (g) the amino acid sequence set forth as SEQ ID NO: 8; (h) the amino acid sequence set forth as SEQ ID NO: 9; (i) the amino acid sequence set forth as SEQ ID NO: 10; (j) the amino acid sequence set forth as SEQ ID NO: 11; (k) the amino acid sequence set forth as SEQ ID NO: 12; (1) the amino acid sequence set forth as SEQ ID NO: 13; (m) the amino acid sequence set forth as SEQ ID NO: 14; (n) the amino acid sequence set forth as SEQ ID NO: 15; and (o) the amino acid sequence set forth as SEQ ID NO: 16. Such biologically active portions can be prepared by recombinant techniques and evaluated for sequence -specific, RNA-guided DNA-binding activity.

[0280] A biologically active fragment of a CRISPR repeat can comprise at least 8 contiguous amino acids of SEQ ID NO: 33, 244, or 245. A biologically active portion of a CRISPR repeat can be a polynucleotide that comprises, for example, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 contiguous nucleotides of SEQ ID NO: 33, 244, or 245. A biologically active portion of a tracrRNA can be a polynucleotide that comprises, for example, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, or more contiguous nucleotides of SEQ ID NO: 34, 246, 247, or 248. A biologically active fragment of a guide RNA backbone (comprising the CRISPR repeat and tracrRNA, and optionally a nucleotide linker) can be a polynucleotide that comprises, for example, 40, 45, 50, 55, 60, 65, 70, 75, 80, or more contiguous nucleotides of any one of SEQ ID NOs: 249-252. A biologically active fragment of a PEgRNA can be a polynucleotide that comprises, for example, 100, 105, 110, 115, 120, 125, 130, or more contiguous nucleotides of any one of SEQ ID NOs: 129-132, and 197-220.

[0281] A biologically active portion of a polymerase (e.g., RT) can be a polypeptide that comprises, for example, 10, 25, 50, 100, 150, 200, 250, 300, 350, 400, 450, or more contiguous amino acid residues of SEQ ID NO: 254 or 255.

[0282] A biologically active portion of a base editing polypeptide (e.g., deaminase) can be a polypeptide that comprises, for example, 10, 25, 50, 75, 100, 125, or more contiguous amino acid residues of any one of SEQ ID NOs: 42-113, and 257.

[0283] A biologically active portion of a base editor can be a polypeptide that comprises, for example, 10, 25, 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, 1300, 1350, 1400, 1450, 1500, 1550, 1600, 1650, or more contiguous amino acid residues of any one of SEQ ID NOs: 258-270.

[0284] A biologically active portion of a PE can be a polypeptide that comprises, for example, 10, 25, 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, 1300, 1350, 1400, 1450, 1500, 1550, 1600, 1650, or more contiguous amino acid residues of any one of SEQ ID NOs: 135, and 162-181.

[0285] In general, "variants" or “active variant” is intended to mean substantially similar sequences. For polynucleotides, a variant comprises a deletion and / or addition of one or more nucleotides at one or more internal sites within the native polynucleotide and / or a substitution of one or more nucleotides at one or more sites in the native polynucleotide. As used herein, a "native" or “wild type” polynucleotide or polypeptide comprises a naturally occurring nucleotide sequence or amino acid sequence, respectively. For polynucleotides, conservative variants include those sequences that, because of the degeneracy of the genetic code, encode the native amino acid sequence of the gene of interest. Naturally occurring allelic variants such as these can be identified with the use of well-known molecular biology techniques, as, for example, with polymerase chain reaction (PCR) and hybridization techniques as outlined below. Variant polynucleotides also include synthetically derived polynucleotides, such as those generated, for example, by using site-directed mutagenesis but which still encode the polypeptide or the polynucleotide of interest. Generally, variants of a particular polynucleotide disclosed herein will have at least about 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to that particular polynucleotide as determined by sequence alignment programs and parameters described elsewhere herein.

[0286] Variants of a particular polynucleotide disclosed herein (z.e., the reference polynucleotide) can also be evaluated by comparison of the percent sequence identity between the polypeptide encoded by a variant polynucleotide and the polypeptide encoded by the reference polynucleotide. Percent sequence identity between any two polypeptides can be calculated using sequence alignment programs and parameters described elsewhere herein. Where any given pair of polynucleotides disclosed herein is evaluated by comparison of the percent sequence identity shared by the two polypeptides they encode, the percent sequence identity between the two encoded polypeptides is at least about 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity.

[0287] In some embodiments, the presently disclosed polynucleotides encode an RNA-guided nuclease polypeptide comprising an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater identity to the amino acid sequence set forth as SEQ ID NO: 1, wherein the amino acid residue at one or more of the following amino acid positions differs from the corresponding amino acid residue in SEQ ID NO: 1: 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975. In some embodiments, the presently disclosed polynucleotides encode an RNA-guided nuclease polypeptide comprising an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater identity to the amino acid sequence set forth as SEQ ID NO: 1, wherein the amino acid residue at one or more of the following amino acid positions is a positively charged amino acid residue: 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975. In some embodiments, the presently disclosed polynucleotides encode an RNA-guided nuclease polypeptide comprising an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater identity to the amino acid sequence set forth as SEQ ID NO: 1, wherein the amino acid residue at one or more of the following amino acid positions is an R: 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975.

[0288] In some embodiments, the presently disclosed polynucleotides encode an RNA-guided nuclease polypeptide comprising an amino acid sequence havi...

Claims

THAT WHICH IS CLAIMED:

1. A nucleic acid molecule comprising a polynucleotide encoding an RNA-guided nuclease (RGN) polypeptide, wherein said RGN polypeptide comprises an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, and wherein the amino acid residue at one or more of amino acid positions 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975 of said amino acid sequence differs from the corresponding amino acid residue in SEQ ID NO: 1.

2. The nucleic acid molecule of claim 1, wherein said RGN polypeptide comprises an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, and wherein the amino acid residue at one or more of amino acid positions 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975of said amino acid sequence is a positively charged amino acid residue.

3. The nucleic acid molecule of claim 1 or 2, wherein said RGN polypeptide comprises an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, and wherein the amino acid residue at one or more of amino acid positions 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975of said amino acid sequence is an R.

4. A nucleic acid molecule comprising a polynucleotide encoding an RNA-guided nuclease (RGN) polypeptide, wherein said RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence selected from:(a) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 52 is an R;(b) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 55 is an R;(c) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 86 is an R;(d) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 472 is an R;(e) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 533 is an R;(f) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 541 is an R;(g) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 643 is substituted by an R;(h) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 647(i) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 653(j) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 745(k) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position 774(l) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 778(m) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 780(n) the amino acid sequence set forth as SEQ ID NO: 1, wherein A at amino acid position 795(o) the amino acid sequence set forth as SEQ ID NO: 1, wherein Q at amino acid position 822(p) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 843(q) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 856(r) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 871(s) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 872(t) the amino acid sequence set forth as SEQ ID NO: 1, wherein D at amino acid position 900(u) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position 911(v) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 913(w) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position R;(x) the amino acid sequence set forth as SEQ ID NO: 1, wherein V at amino acid position 958(y) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 968(z) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 969 is an R;(aa) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position973 is an R;(bb) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position974 is an R;(cc) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position975 is an R;(dd) the amino acid sequence set forth as SEQ ID NO: 2, wherein Q at amino acid position 822 is an R, G at amino acid position 856 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 778 is an R;(ee) the amino acid sequence set forth as SEQ ID NO: 3, wherein Q at amino acid position 822 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 778 is substituted by an R;(ff) the amino acid sequence set forth as SEQ ID NO: 4, wherein Q at amino acid position 822 is an R, S at amino acid position 55 is an R, and E at amino acid position 778 is an R;(gg) the amino acid sequence set forth as SEQ ID NO: 5, wherein Q at amino acid position 822 is an R, S at amino acid position 647 is an R, and E at amino acid position 778 is an R;(hh) the amino acid sequence set forth as SEQ ID NO: 6, wherein Q at amino acid position 822 is an R, G at amino acid position 856 is an R, and E at amino acid position 778 is an R;(ii) the amino acid sequence set forth as SEQ ID NO: 7, wherein G at amino acid position 856 is an R, and E at amino acid position 778 is an R;(jj) the amino acid sequence set forth as SEQ ID NO: 8, wherein E at amino acid position 778 is an R, Q at amino acid position 822 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(kk) the amino acid sequence set forth as SEQ ID NO: 9, wherein E at amino acid position 778 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(11) the amino acid sequence set forth as SEQ ID NO: 10, wherein E at amino acid position 778 is an R, G at amino acid position 856 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(mm) the amino acid sequence set forth as SEQ ID NO: 11, wherein E at amino acid position 778 is an R, G at amino acid position 856 is an R, S at amino acid position 647 is an R, and E at amino acid position 969 is an R;(nn) the amino acid sequence set forth as SEQ ID NO: 12, wherein E at amino acid position 778 is an R, Q at amino acid position 822 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(oo) the amino acid sequence set forth as SEQ ID NO: 13, wherein E at amino acid position 778 is an R, Q at amino acid position 822 is an R, S at amino acid position 647 is an R, and E at amino acid position 969 is an R;(pp) the amino acid sequence set forth as SEQ ID NO: 14, wherein E at amino acid position 778 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(qq) the amino acid sequence set forth as SEQ ID NO: 15, wherein E at amino acid position 778 is an R, G at amino acid position 856 is an R, and E at amino acid position 969 is an R; and(rr) the amino acid sequence set forth as SEQ ID NO: 16, wherein E at amino acid position 778 is an R, and E at amino acid position 969 is an R.

5. The nucleic acid molecule of claim 4, wherein said RGN polypeptide comprises an amino acid sequence having at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to the amino acid sequence selected from:(a) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 52 is an R;(b) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 55 is an R;(c) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 86 is an R;(d) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 472 is an R;(e) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 533 is an R;(f) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 541 is an R;(g) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 643 is substituted by an R;(h) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 647 is an R;(i) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 653 is an R;(j) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 745 is an R;(k) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position 774(l) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 778(m) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 780(n) the amino acid sequence set forth as SEQ ID NO: 1, wherein A at amino acid position 795(o) the amino acid sequence set forth as SEQ ID NO: 1, wherein Q at amino acid position 822(p) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 843(q) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 856(r) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 871(s) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 872(t) the amino acid sequence set forth as SEQ ID NO: 1, wherein D at amino acid position 900(u) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position 911(v) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 913(w) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position R;(x) the amino acid sequence set forth as SEQ ID NO: 1, wherein V at amino acid position 958(y) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 968(z) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 969(aa) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position R;(bb) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position R;(cc) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position 975 is an R;(dd) the amino acid sequence set forth as SEQ ID NO: 2, wherein Q at amino acid position 822 is an R, G at amino acid position 856 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 778 is an R;(ee) the amino acid sequence set forth as SEQ ID NO: 3, wherein Q at amino acid position 822 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 778 is substituted by an R;(ff) the amino acid sequence set forth as SEQ ID NO: 4, wherein Q at amino acid position 822 is an R, S at amino acid position 55 is an R, and E at amino acid position 778 is an R;(gg) the amino acid sequence set forth as SEQ ID NO: 5, wherein Q at amino acid position 822 is an R, S at amino acid position 647 is an R, and E at amino acid position 778 is an R;(hh) the amino acid sequence set forth as SEQ ID NO: 6, wherein Q at amino acid position 822 is an R, G at amino acid position 856 is an R, and E at amino acid position 778 is an R;(ii) the amino acid sequence set forth as SEQ ID NO: 7, wherein G at amino acid position 856 is an R, and E at amino acid position 778 is an R;(jj) the amino acid sequence set forth as SEQ ID NO: 8, wherein E at amino acid position 778 is an R, Q at amino acid position 822 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(kk) the amino acid sequence set forth as SEQ ID NO: 9, wherein E at amino acid position 778 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(11) the amino acid sequence set forth as SEQ ID NO: 10, wherein E at amino acid position 778 is an R, G at amino acid position 856 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(mm) the amino acid sequence set forth as SEQ ID NO: 11, wherein E at amino acid position 778 is an R, G at amino acid position 856 is an R, S at amino acid position 647 is an R, and E at amino acid position 969 is an R;(nn) the amino acid sequence set forth as SEQ ID NO: 12, wherein E at amino acid position 778 is an R, Q at amino acid position 822 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(oo) the amino acid sequence set forth as SEQ ID NO: 13, wherein E at amino acid position 778 is an R, Q at amino acid position 822 is an R, S at amino acid position 647 is an R, and E at amino acid position 969 is an R;(pp) the amino acid sequence set forth as SEQ ID NO: 14, wherein E at amino acid position 778 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(qq) the amino acid sequence set forth as SEQ ID NO: 15, wherein E at amino acid position 778 is an R, G at amino acid position 856 is an R, and E at amino acid position 969 is an R; and(rr) the amino acid sequence set forth as SEQ ID NO: 16, wherein E at amino acid position 778 is an R, and E at amino acid position 969 is an R.

6. The nucleic acid molecule of claim 5, wherein said RGN polypeptide comprises an amino acid sequence having 100% sequence identity to the amino acid sequence selected from:(a) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 52 is an R;(b) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 55 is an R;(c) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 86 is an R;(d) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 472 is an R;(e) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 533 is an R;(f) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 541 is an R;(g) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 643 is substituted by an R;(h) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 647 is an R;(i) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 653 is an R;(j) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 745 is an R;(k) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position 774 is an R;(l) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 778 is an R;(m) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 780 is an R;(n) the amino acid sequence set forth as SEQ ID NO: 1, wherein A at amino acid position 795 is an R;(o) the amino acid sequence set forth as SEQ ID NO: 1, wherein Q at amino acid position 822 is an R;(p) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 843(q) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 856(r) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 871(s) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 872(t) the amino acid sequence set forth as SEQ ID NO: 1, wherein D at amino acid position 900(u) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position 911(v) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 913(w) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position R;(x) the amino acid sequence set forth as SEQ ID NO: 1, wherein V at amino acid position 958(y) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 968(z) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 969(aa) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position R;(bb) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position R;(cc) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position R;(dd) the amino acid sequence set forth as SEQ ID NO: 2;(ee) the amino acid sequence set forth as SEQ ID NO: 3;(ff) the amino acid sequence set forth as SEQ ID NO: 4;(gg) the amino acid sequence set forth as SEQ ID NO: 5;(hh) the amino acid sequence set forth as SEQ ID NO: 6;(ii) the amino acid sequence set forth as SEQ ID NO: 7;(jj) the amino acid sequence set forth as SEQ ID NO: 8;(kk) the amino acid sequence set forth as SEQ ID NO: 9;(11) the amino acid sequence set forth as SEQ ID NO: 10;(mm) the amino acid sequence set forth as SEQ ID NO: 11;(nn) the amino acid sequence set forth as SEQ ID NO: 12;(oo) the amino acid sequence set forth as SEQ ID NO: 13;(pp) the amino acid sequence set forth as SEQ ID NO: 14;(qq) the amino acid sequence set forth as SEQ ID NO: 15; and(rr) the amino acid sequence set forth as SEQ ID NO: 16.

7. The nucleic acid molecule of any one of claims 1-6, wherein said polynucleotide encoding an RGN polypeptide is operably linked to a promoter heterologous to said polynucleotide.

8. The nucleic acid molecule of any one of claims 1-7, wherein said RGN polypeptide is nuclease inactive or is a nickase.

9. The nucleic acid molecule of claim 8, wherein said RGN polypeptide comprises a D16A and / or a H611A mutation(s).

10. The nucleic acid molecule of claim 8 or 9, wherein said RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 182-196, and 271-285.

11. The nucleic acid molecule of any one of claims 1-10, wherein said RGN polypeptide is operably linked to a heterologous polypeptide.

12. The nucleic acid molecule of claim 11, wherein said heterologous polypeptide is operably linked to the N-terminus, to the C-terminus, or to an internal location of said RGN polypeptide.

13. The nucleic acid molecule of claim 11 or 12, wherein said heterologous polypeptide is a polymerase editing polypeptide.

14. The nucleic acid molecule of claim 13, wherein said polymerase editing polypeptide comprises a DNA polymerase.

15. The nucleic acid molecule of claim 13, wherein said polymerase editing polypeptide comprises a reverse transcriptase.

16. The nucleic acid molecule of claim 15, wherein said reverse transcriptase has at least 90% sequence identity to the amino acid sequence of SEQ ID NO: 254 or 255.

17. The nucleic acid molecule of claim 11 or 12, wherein the heterologous polypeptide is a base-editing polypeptide.

18. The nucleic acid molecule of claim 17, wherein the base-editing polypeptide is a deaminase.

19. The nucleic acid molecule of claim 18, wherein the deaminase has at least 90% sequence identity to an amino acid sequence of any one of SEQ ID NOs: 42-113, and 257.

20. The nucleic acid molecule of claim 11 or 12, wherein the heterologous polypeptide is an effector domain, a detectable label, or a purification tag.

21. The nucleic acid molecule of claim 20, wherein the effector domain is a cleavage domain, a deaminase domain, or an expression modulator domain.

22. The nucleic acid molecule of any one of claims 1-21, wherein the polynucleotide encoding the RGN polypeptide is codon optimized for expression in a eukaryotic cell.

23. A vector comprising the nucleic acid molecule of any one of claims 1-22.

24. The vector of claim 23, further comprising at least one nucleotide sequence encoding a guide RNA.

25. The vector of claim 24, wherein the guide RNA comprises a CRISPR RNA (crRNA) comprising a CRISPR repeat comprising a nucleotide sequence set forth as SEQ ID NO: 33, 244, or 245, or that differs from SEQ ID NO: 33, 244, or 245 by 1 to 5 nucleotides.

26. The vector of claim 24 or 25, wherein the guide RNA comprises a tracrRNA comprising a nucleotide sequence having at least 90% sequence identity to SEQ ID NO: 34, 246, 247, or 248.

27. The vector of any one of claims 24-26, where said guide RNA is a single guide RNA.

28. The vector of claim 27, wherein said single guide RNA further comprises an extension comprising an edit template for polymerase editing.

29. The vector of any one of claims 24-26, wherein said guide RNA is a dual-guide RNA.

30. A cell comprising the nucleic acid molecule of any one of claims 1-22 or the vector of any one of claims 23-29.

31. A method for making an RGN polypeptide comprising culturing the cell of claim 30 under conditions in which the RGN polypeptide is expressed.

32. An RNA-guided nuclease (RGN) polypeptide, wherein said RGN polypeptide comprises an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, and wherein the amino acid residue at one or more of amino acid positions 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975 of said amino acid sequence differs from the corresponding amino acid residue in SEQ ID NO: 1.

33. The RGN polypeptide of claim 32, wherein said RGN polypeptide comprises an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, and wherein the amino acid residue at one or more of amino acid positions 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975 of said amino acid sequence is a positively charged amino acid residue.

34. The RGN polypeptide of claim 32 or 33, wherein said RGN polypeptide comprises an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, and wherein the aminoacid residue at one or more of amino acid positions 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975of said amino acid sequence is an R.

35. An RNA-guided nuclease (RGN) polypeptide, wherein said RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence selected from:(a) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 52 is an R;(b) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 55 is an R;(c) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 86 is an R;(d) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 472 is an R;(e) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 533 is an R;(f) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 541 is an R;(g) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 643 is substituted by an R;(h) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 647 is an R;(i) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 653 is an R;(j) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 745 is an R;(k) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position 774 is an R;(l) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 778 is an R;(m) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 780 is an R;(n) the amino acid sequence set forth as SEQ ID NO: 1, wherein A at amino acid position 795 is an R;(o) the amino acid sequence set forth as SEQ ID NO: 1, wherein Q at amino acid position 822 is an R;(p) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 843 is an R;(q) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 856 is an R;(r) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 871 is an R;(s) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 872 is an R;(t) the amino acid sequence set forth as SEQ ID NO: 1, wherein D at amino acid position 900 is an R;(u) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position 911 is an R;(v) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 913 is an R;(w) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position 954 is an R;(x) the amino acid sequence set forth as SEQ ID NO: 1, wherein V at amino acid position 958 is an R;(y) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 968 is an R;(z) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 969 is an R;(aa) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position973 is an R;(bb) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position974 is an R;(cc) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position975 is an R;(dd) the amino acid sequence set forth as SEQ ID NO: 2, wherein Q at amino acid position 822 is an R, G at amino acid position 856 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 778 is an R;(ee) the amino acid sequence set forth as SEQ ID NO: 3, wherein Q at amino acid position 822 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 778 is substituted by an R;(ff) the amino acid sequence set forth as SEQ ID NO: 4, wherein Q at amino acid position 822 is an R, S at amino acid position 55 is an R, and E at amino acid position 778 is an R;(gg) the amino acid sequence set forth as SEQ ID NO: 5, wherein Q at amino acid position 822 is an R, S at amino acid position 647 is an R, and E at amino acid position 778 is an R;(hh) the amino acid sequence set forth as SEQ ID NO: 6, wherein Q at amino acid position 822 is an R, G at amino acid position 856 is an R, and E at amino acid position 778 is an R;(ii) the amino acid sequence set forth as SEQ ID NO: 7, wherein G at amino acid position 856 is an R, and E at amino acid position 778 is an R;(jj) the amino acid sequence set forth as SEQ ID NO: 8, wherein E at amino acid position 778 is an R, Q at amino acid position 822 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(kk) the amino acid sequence set forth as SEQ ID NO: 9, wherein E at amino acid position 778 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(11) the amino acid sequence set forth as SEQ ID NO: 10, wherein E at amino acid position 778 is an R, G at amino acid position 856 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(mm) the amino acid sequence set forth as SEQ ID NO: 11, wherein E at amino acid position 778 is an R, G at amino acid position 856 is an R, S at amino acid position 647 is an R, and E at amino acid position 969 is an R;(nn) the amino acid sequence set forth as SEQ ID NO: 12, wherein E at amino acid position 778 is an R, Q at amino acid position 822 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(oo) the amino acid sequence set forth as SEQ ID NO: 13, wherein E at amino acid position 778 is an R, Q at amino acid position 822 is an R, S at amino acid position 647 is an R, and E at amino acid position 969 is an R;(pp) the amino acid sequence set forth as SEQ ID NO: 14, wherein E at amino acid position 778 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(qq) the amino acid sequence set forth as SEQ ID NO: 15, wherein E at amino acid position 778 is an R, G at amino acid position 856 is an R, and E at amino acid position 969 is an R; and(rr) the amino acid sequence set forth as SEQ ID NO: 16, wherein E at amino acid position 778 is an R, and E at amino acid position 969 is an R.

36. The RGN polypeptide of claim 35, wherein said RGN polypeptide comprises an amino acid sequence having at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to the amino acid sequence selected from:(a) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 52 is an R;(b) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 55(c) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 86(d) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 472(e) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 533(f) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 541(g) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 643tuted by an R;(h) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 647(i) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 653(j) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 745(k) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position 774(l) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 778(m) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 780(n) the amino acid sequence set forth as SEQ ID NO: 1, wherein A at amino acid position 795(o) the amino acid sequence set forth as SEQ ID NO: 1, wherein Q at amino acid position 822(p) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 843(q) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 856(r) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 871(s) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 872(t) the amino acid sequence set forth as SEQ ID NO: 1, wherein D at amino acid position 900 is an R;(u) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position 911 is an R;(v) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 913 is an R;(w) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position 954 is an R;(x) the amino acid sequence set forth as SEQ ID NO: 1, wherein V at amino acid position 958 is an R;(y) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 968 is an R;(z) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 969 is an R;(aa) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position973 is an R;(bb) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position974 is an R;(cc) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position975 is an R;(dd) the amino acid sequence set forth as SEQ ID NO: 2, wherein Q at amino acid position 822 is an R, G at amino acid position 856 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 778 is an R;(ee) the amino acid sequence set forth as SEQ ID NO: 3, wherein Q at amino acid position 822 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 778 is substituted by an R;(ff) the amino acid sequence set forth as SEQ ID NO: 4, wherein Q at amino acid position 822 is an R, S at amino acid position 55 is an R, and E at amino acid position 778 is an R;(gg) the amino acid sequence set forth as SEQ ID NO: 5, wherein Q at amino acid position 822 is an R, S at amino acid position 647 is an R, and E at amino acid position 778 is an R;(hh) the amino acid sequence set forth as SEQ ID NO: 6, wherein Q at amino acid position 822 is an R, G at amino acid position 856 is an R, and E at amino acid position 778 is an R;(ii) the amino acid sequence set forth as SEQ ID NO: 7, wherein G at amino acid position 856 is an R, and E at amino acid position 778 is an R;(jj) the amino acid sequence set forth as SEQ ID NO: 8, wherein E at amino acid position 778 is an R, Q at amino acid position 822 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(kk) the amino acid sequence set forth as SEQ ID NO: 9, wherein E at amino acid position 778 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(11) the amino acid sequence set forth as SEQ ID NO: 10, wherein E at amino acid position 778 is an R, G at amino acid position 856 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(mm) the amino acid sequence set forth as SEQ ID NO: 11, wherein E at amino acid position 778 is an R, G at amino acid position 856 is an R, S at amino acid position 647 is an R, and E at amino acid position 969 is an R;(nn) the amino acid sequence set forth as SEQ ID NO: 12, wherein E at amino acid position 778 is an R, Q at amino acid position 822 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(oo) the amino acid sequence set forth as SEQ ID NO: 13, wherein E at amino acid position 778 is an R, Q at amino acid position 822 is an R, S at amino acid position 647 is an R, and E at amino acid position 969 is an R;(pp) the amino acid sequence set forth as SEQ ID NO: 14, wherein E at amino acid position 778 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(qq) the amino acid sequence set forth as SEQ ID NO: 15, wherein E at amino acid position 778 is an R, G at amino acid position 856 is an R, and E at amino acid position 969 is an R; and(rr) the amino acid sequence set forth as SEQ ID NO: 16, wherein E at amino acid position 778 is an R, and E at amino acid position 969 is an R.

37. The RGN polypeptide of claim 36, wherein said RGN polypeptide comprises an amino acid sequence having 100% sequence identity to the amino acid sequence selected from:(a) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 52 is an R;(b) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 55 is an R;(c) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 86 is an R;(d) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 472 is an R;(e) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 533 is an R;(f) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 541(g) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 643tuted by an R;(h) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 647(i) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 653(j) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 745(k) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position 774(l) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 778(m) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 780(n) the amino acid sequence set forth as SEQ ID NO: 1, wherein A at amino acid position 795(o) the amino acid sequence set forth as SEQ ID NO: 1, wherein Q at amino acid position 822(p) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 843(q) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 856(r) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 871(s) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 872(t) the amino acid sequence set forth as SEQ ID NO: 1, wherein D at amino acid position 900(u) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position 911(v) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 913(w) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position R;(x) the amino acid sequence set forth as SEQ ID NO: 1, wherein V at amino acid position 958 is an R;(y) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 968 is an R;(z) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 969 is an R;(aa) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position973 is an R;(bb) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position974 is an R;(cc) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position975 is an R;(dd) the amino acid sequence set forth as SEQ ID NO: 2;(ee) the amino acid sequence set forth as SEQ ID NO: 3;(ff) the amino acid sequence set forth as SEQ ID NO: 4;(gg) the amino acid sequence set forth as SEQ ID NO: 5;(hh) the amino acid sequence set forth as SEQ ID NO: 6;(ii) the amino acid sequence set forth as SEQ ID NO: 7;(jj) the amino acid sequence set forth as SEQ ID NO: 8;(kk) the amino acid sequence set forth as SEQ ID NO: 9;(11) the amino acid sequence set forth as SEQ ID NO: 10;(mm) the amino acid sequence set forth as SEQ ID NO: 11;(nn) the amino acid sequence set forth as SEQ ID NO: 12;(oo) the amino acid sequence set forth as SEQ ID NO: 13;(pp) the amino acid sequence set forth as SEQ ID NO: 14;(qq) the amino acid sequence set forth as SEQ ID NO: 15; and(rr) the amino acid sequence set forth as SEQ ID NO: 16.

38. The RGN polypeptide of any one of claims 32-37, wherein said RGN polypeptide is an isolated RGN polypeptide.

39. The RGN polypeptide of any one of claims 32-38, wherein said RGN polypeptide is capable of binding a target polynucleotide sequence of a DNA molecule in an RNA-guided sequence specific manner when bound to a guide RNA (gRNA) capable of hybridizing to said target polynucleotide sequence.

40. The RGN polypeptide of claim 39, wherein said RGN polypeptide recognizes a protospacer adjacent motif (PAM) that is 3' of said target polynucleotide sequence.

41. The RGN polypeptide of claim 40, wherein said RGN polypeptide recognizes a PAM having a consensus nucleotide sequence set forth as NNGG.

42. The RGN polypeptide of claim 40 or 41, wherein the RGN polypeptide comprises a PAM-interacting domain comprising the amino acid sequence set forth as SEQ ID NO: 253.

43. The RGN polypeptide of any one of claims 39-42, wherein said RGN polypeptide is capable of cleaving said target polynucleotide sequence upon binding.

44. The RGN polypeptide of claim 43, wherein cleavage by said RGN polypeptide generates a double-stranded break.

45. The RGN polypeptide of claim 43, wherein cleavage by said RGN polypeptide generates a single -stranded break.

46. The RGN polypeptide of any one of claims 32-42, wherein said RGN polypeptide is nuclease inactive or a nickase.

47. The RGN polypeptide of claim 46, wherein said RGN polypeptide comprises a D16A and / or a H611 A mutation(s) .

48. The RGN polypeptide of claim 46 or 47, wherein said RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 182-196, and 271-285.

49. The RGN polypeptide of any one of claims 32-48, wherein said RGN polypeptide is operably fused to a heterologous polypeptide.

50. The RGN polypeptide of claim 49, wherein said heterologous polypeptide is operably fused to the N-terminus, to the C-terminus, or to an internal location of said RGN polypeptide.

51. The RGN polypeptide of claim 49 or 50, wherein said heterologous polypeptide is a polymerase editing polypeptide.

52. The RGN polypeptide of claim 51, wherein said polymerase editing polypeptide comprises a DNA polymerase.

53. The RGN polypeptide of claim 51, wherein said polymerase editing polypeptide comprises a reverse transcriptase.

54. The RGN polypeptide of claim 56, wherein said reverse transcriptase has at least 90% sequence identity to the amino acid sequence of SEQ ID NO: 254 or 255.

55. The RGN polypeptide of claim 49 or 50, wherein the heterologous polypeptide is a base-editing polypeptide.

56. The RGN polypeptide of claim 55, wherein the base-editing polypeptide is a deaminase.

57. The RGN polypeptide of claim 56, wherein the deaminase is a cytosine deaminase or an adenine deaminase.

58. The RGN polypeptide of claim 56 or 57, wherein the deaminase has at least 90% sequence identity to an amino acid sequence of any one of SEQ ID NOs: 42-113, and 257.

59. The RGN polypeptide of claim 49 or 50, wherein the heterologous polypeptide is an effector domain, a detectable label, or a purification tag.

60. The RGN polypeptide of claim 59, wherein the effector domain is a cleavage domain, a deaminase domain, or an expression modulator domain.

61. The RGN polypeptide of any one of claims 32-60, wherein the RGN polypeptide comprises one or more nuclear localization signals.

62. A ribonucleoprotein (RNP) complex comprising the RGN polypeptide of any one of claims 32-61 and the guide RNA bound to the RGN polypeptide.

63. A system, said system comprising:A) one or more guide RNAs (gRNAs), or one or more polynucleotides comprising one or more nucleotide sequences encoding the one or more gRNAs; andB) an RNA-guided nuclease (RGN) polypeptide, or a polynucleotide comprising a nucleotide sequence encoding the RGN polypeptide, wherein said RGN polypeptide comprises an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, and wherein the amino acid residue at one or more of amino acid positions 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975 of said amino acid sequence differs from the corresponding amino acid residue in SEQ ID NO: 1.

64. The system of claim 63, wherein said RGN polypeptide comprises an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, and wherein the amino acid residue at one or more of amino acid positions 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975of said amino acid sequence is a positively charged amino acid residue.

65. The system of claim 63 or 64, wherein said RGN polypeptide comprises an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, and wherein the amino acid residue at one or more of amino acid positions 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975 of said amino acid sequence is an R.

66. The system of claim 65, wherein said RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence selected from:(a) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 52 is an R;(b) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 55 is an R;(c) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 86(d) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 472(e) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 533(f) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 541(g) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 643tuted by an R;(h) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 647(i) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 653(j) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 745(k) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position 774(l) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 778(m) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 780(n) the amino acid sequence set forth as SEQ ID NO: 1, wherein A at amino acid position 795(o) the amino acid sequence set forth as SEQ ID NO: 1, wherein Q at amino acid position 822(p) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 843(q) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 856(r) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 871(s) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 872(t) the amino acid sequence set forth as SEQ ID NO: 1, wherein D at amino acid position 900(u) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position 911 is an R;(v) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 913 is an R;(w) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position 954 is an R;(x) the amino acid sequence set forth as SEQ ID NO: 1, wherein V at amino acid position 958 is an R;(y) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 968 is an R;(z) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 969 is an R;(aa) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position973 is an R;(bb) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position974 is an R;(cc) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position975 is an R;(dd) the amino acid sequence set forth as SEQ ID NO: 2, wherein Q at amino acid position 822 is an R, G at amino acid position 856 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 778 is an R;(ee) the amino acid sequence set forth as SEQ ID NO: 3, wherein Q at amino acid position 822 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 778 is substituted by an R;(ff) the amino acid sequence set forth as SEQ ID NO: 4, wherein Q at amino acid position 822 is an R, S at amino acid position 55 is an R, and E at amino acid position 778 is an R;(gg) the amino acid sequence set forth as SEQ ID NO: 5, wherein Q at amino acid position 822 is an R, S at amino acid position 647 is an R, and E at amino acid position 778 is an R;(hh) the amino acid sequence set forth as SEQ ID NO: 6, wherein Q at amino acid position 822 is an R, G at amino acid position 856 is an R, and E at amino acid position 778 is an R;(ii) the amino acid sequence set forth as SEQ ID NO: 7, wherein G at amino acid position 856 is an R, and E at amino acid position 778 is an R;(jj) the amino acid sequence set forth as SEQ ID NO: 8, wherein E at amino acid position 778 is an R, Q at amino acid position 822 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(kk) the amino acid sequence set forth as SEQ ID NO: 9, wherein E at amino acid position 778 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(11) the amino acid sequence set forth as SEQ ID NO: 10, wherein E at amino acid position 778 is an R, G at amino acid position 856 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(mm) the amino acid sequence set forth as SEQ ID NO: 11, wherein E at amino acid position 778 is an R, G at amino acid position 856 is an R, S at amino acid position 647 is an R, and E at amino acid position 969 is an R;(nn) the amino acid sequence set forth as SEQ ID NO: 12, wherein E at amino acid position 778 is an R, Q at amino acid position 822 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(oo) the amino acid sequence set forth as SEQ ID NO: 13, wherein E at amino acid position 778 is an R, Q at amino acid position 822 is an R, S at amino acid position 647 is an R, and E at amino acid position 969 is an R;(pp) the amino acid sequence set forth as SEQ ID NO: 14, wherein E at amino acid position 778 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(qq) the amino acid sequence set forth as SEQ ID NO: 15, wherein E at amino acid position 778 is an R, G at amino acid position 856 is an R, and E at amino acid position 969 is an R; and(rr) the amino acid sequence set forth as SEQ ID NO: 16, wherein E at amino acid position 778 is an R, and E at amino acid position 969 is an R.

67. The system of claim 66, wherein the RGN polypeptide comprises an amino acid sequence having at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to the amino acid sequence selected from:(a) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 52 is an R;(b) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 55 is an R;(c) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 86 is an R;(d) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 472 is an R;(e) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 533 is an R;(f) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 541 is an R;(g) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 643tuted by an R;(h) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 647(i) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 653(j) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 745(k) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position 774(l) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 778(m) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 780(n) the amino acid sequence set forth as SEQ ID NO: 1, wherein A at amino acid position 795(o) the amino acid sequence set forth as SEQ ID NO: 1, wherein Q at amino acid position 822(p) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 843(q) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 856(r) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 871(s) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 872(t) the amino acid sequence set forth as SEQ ID NO: 1, wherein D at amino acid position 900(u) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position 911(v) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 913(w) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position R;(x) the amino acid sequence set forth as SEQ ID NO: 1, wherein V at amino acid position 958(y) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 968 is an R;(z) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 969 is an R;(aa) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position973 is an R;(bb) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position974 is an R;(cc) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position975 is an R;(dd) the amino acid sequence set forth as SEQ ID NO: 2, wherein Q at amino acid position 822 is an R, G at amino acid position 856 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 778 is an R;(ee) the amino acid sequence set forth as SEQ ID NO: 3, wherein Q at amino acid position 822 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 778 is substituted by an R;(ff) the amino acid sequence set forth as SEQ ID NO: 4, wherein Q at amino acid position 822 is an R, S at amino acid position 55 is an R, and E at amino acid position 778 is an R;(gg) the amino acid sequence set forth as SEQ ID NO: 5, wherein Q at amino acid position 822 is an R, S at amino acid position 647 is an R, and E at amino acid position 778 is an R;(hh) the amino acid sequence set forth as SEQ ID NO: 6, wherein Q at amino acid position 822 is an R, G at amino acid position 856 is an R, and E at amino acid position 778 is an R;(ii) the amino acid sequence set forth as SEQ ID NO: 7, wherein G at amino acid position 856 is an R, and E at amino acid position 778 is an R;(jj) the amino acid sequence set forth as SEQ ID NO: 8, wherein E at amino acid position 778 is an R, Q at amino acid position 822 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(kk) the amino acid sequence set forth as SEQ ID NO: 9, wherein E at amino acid position 778 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(11) the amino acid sequence set forth as SEQ ID NO: 10, wherein E at amino acid position 778 is an R, G at amino acid position 856 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(mm) the amino acid sequence set forth as SEQ ID NO: 11, wherein E at amino acid position 778 is an R, G at amino acid position 856 is an R, S at amino acid position 647 is an R, and E at amino acid position 969 is an R;(nn) the amino acid sequence set forth as SEQ ID NO: 12, wherein E at amino acid position 778 is an R, Q at amino acid position 822 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(oo) the amino acid sequence set forth as SEQ ID NO: 13, wherein E at amino acid position 778 is an R, Q at amino acid position 822 is an R, S at amino acid position 647 is an R, and E at amino acid position 969 is an R;(pp) the amino acid sequence set forth as SEQ ID NO: 14, wherein E at amino acid position 778 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(qq) the amino acid sequence set forth as SEQ ID NO: 15, wherein E at amino acid position 778 is an R, G at amino acid position 856 is an R, and E at amino acid position 969 is an R; and(rr) the amino acid sequence set forth as SEQ ID NO: 16, wherein E at amino acid position 778 is an R, and E at amino acid position 969 is an R.

68. The system of claim 67, wherein said RGN polypeptide comprises an amino acid sequence having 100% sequence identity to the amino acid sequence selected from:(a) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 52 is an R;(b) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 55 is an R;(c) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 86 is an R;(d) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 472 is an R;(e) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 533 is an R;(f) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 541 is an R;(g) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 643 is substituted by an R;(h) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 647 is an R;(i) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 653 is an R;(j) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 745 is an R;(k) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position 774 is an R;(1) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 778(m) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 780(n) the amino acid sequence set forth as SEQ ID NO: 1, wherein A at amino acid position 795(o) the amino acid sequence set forth as SEQ ID NO: 1, wherein Q at amino acid position 822(p) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 843(q) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 856(r) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 871(s) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 872(t) the amino acid sequence set forth as SEQ ID NO: 1, wherein D at amino acid position 900(u) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position 911(v) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 913(w) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position R;(x) the amino acid sequence set forth as SEQ ID NO: 1, wherein V at amino acid position 958(y) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 968(z) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 969(aa) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position R;(bb) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position R;(cc) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position R;(dd) the amino acid sequence set forth as SEQ ID NO: 2;(ee) the amino acid sequence set forth as SEQ ID NO: 3;(ff) the amino acid sequence set forth as SEQ ID NO: 4;(gg) the amino acid sequence set forth as SEQ ID NO: 5;(hh) the amino acid sequence set forth as SEQ ID NO: 6;(ii) the amino acid sequence set forth as SEQ ID NO: 7;(jj) the amino acid sequence set forth as SEQ ID NO: 8;(kk) the amino acid sequence set forth as SEQ ID NO: 9;(11) the amino acid sequence set forth as SEQ ID NO: 10;(mm) the amino acid sequence set forth as SEQ ID NO: 11;(nn) the amino acid sequence set forth as SEQ ID NO: 12;(oo) the amino acid sequence set forth as SEQ ID NO: 13;(pp) the amino acid sequence set forth as SEQ ID NO: 14;(qq) the amino acid sequence set forth as SEQ ID NO: 15; and(rr) the amino acid sequence set forth as SEQ ID NO: 16.

69. The system of any one of claims 63-68, wherein the polynucleotide comprising a nucleotide sequence encoding the RGN polypeptide comprises an mRNA comprising a 5 ’ untranslated region (UTR) and / or a 3 ’ UTR, wherein the 5 ’ UTR, the 3 ’ UTR, or both are heterologous to the mRNA.

70. The system of any one of claims 63-68, wherein at least one of said nucleotide sequences encoding the one or more gRNAs and said nucleotide sequence encoding the RGN polypeptide is operably linked to a promoter heterologous to said nucleotide sequence.

71. The system of any one of claims 63-70, wherein said gRNA is a single guide RNA.

72. The system of claim 71, wherein said single guide RNA further comprises an extension comprising an edit template for polymerase editing.

73. The system of any one of claims 63-70, wherein said gRNA is a dual-guide RNA.

74. The system of any one of claims 63-73, wherein said gRNA comprises a CRISPR RNA (crRNA) comprising a CRISPR repeat comprising a nucleotide sequence set forth as SEQ ID NO: 33, 244, or 245, or that differs from SEQ ID NO: 33, 244, or 245 by 1 to 5 nucleotides.

75. The system of any one of claims 63-74, wherein the guide RNA comprises a a tracrRNA having at least 90% sequence identity to SEQ ID NO: 34, 246, 247, or 248.

76. The system of any one of claims 63-75, wherein the one or more gRNAs are capable of hybridizing to a target polynucleotide sequence of a DNA molecule, and wherein the one or more guide RNAs are capable of forming a complex with the RGN polypeptide in order to direct said RGN polypeptide to bind to said target polynucleotide sequence of the DNA molecule.

77. The system of claim 76, wherein said target polynucleotide sequence is a eukaryotic target polynucleotide sequence.

78. The system of claim 76 or 77, wherein the target polynucleotide sequence is within a cell.

79. The system of any one of claims 76-78, wherein the complex comprising the one or more gRNAs and the RGN polypeptide directs cleavage of the target polynucleotide sequence.

80. The system of any one of claims 63-78, wherein said RGN polypeptide is nuclease inactive or is a nickase.

81. The system of claim 80, wherein said RGN polypeptide comprises a D16A and / or a H611A mutation(s).

82. The system of claim 80 or 81, wherein said RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 182-196, and 271- 285.

83. The system of any one of claims 63-82, wherein said RGN polypeptide is operably linked to a heterologous polypeptide.

84. The system of claim 83, wherein said heterologous polypeptide is operably fused to the N-terminus, to the C-terminus, or to an internal location of said RGN polypeptide.

85. The system of claim 83 or 84, wherein the heterologous polypeptide is a polymerase editing polypeptide.

86. The system of claim 85, wherein the polymerase editing polypeptide comprises a DNA polymerase.

87. The system of claim 85, wherein the polymerase editing polypeptide comprises a reverse transcriptase.

88. The system of claim 87, wherein said reverse transcriptase has at least 90% sequence identity to the amino acid sequence of SEQ ID NO: 254 or 255.

89. The system of claim 83 or 84, wherein the heterologous polypeptide is a base-editing polypeptide.

90. The system of claim 89, wherein the base-editing polypeptide is a deaminase.

91. The system of claim 90, wherein the deaminase is a cytosine deaminase or an adenine deaminase.

92. The system of claim 90 or 91, wherein the deaminase has at least 90% sequence identity to an amino acid sequence of any one of SEQ ID NOs: 42-113, and 257.

93. The system of claim 83 or 84, wherein the heterologous polypeptide is an effector domain, a detectable label, or a purification tag.

94. The system of claim 93, wherein the effector domain is a cleavage domain, a deaminase domain, or an expression modulator domain.

95. The system of any one of claims 63-94, wherein said system further comprises one or more donor polynucleotides.

96. A cell comprising the RGN polypeptide of any one of claims 32-61, the RNP complex of claim 62, or the system of any one of claims 63-95.

97. A pharmaceutical composition comprising the nucleic acid molecule of any one of claims 1-22, the vector of any one of claims 23-29, the cell of claim 30 or 96, the RGN polypeptide of any one of claims 32-61, the RNP complex of claim 62, or the system of any one of claims 63-95, and a pharmaceutically acceptable carrier.

98. A method for binding a target polynucleotide sequence of a nucleic acid molecule comprising delivering a system according to any one of claims 63-95, to said target polynucleotide sequence or a cell comprising the target polynucleotide sequence.

99. A method for cleaving and / or modifying a target polynucleotide sequence of a nucleic acid molecule comprising delivering a system according to any one of claims 63-95 to said target polynucleotide sequence or a cell comprising the nucleic acid molecule, wherein cleavage or modification of said target polynucleotide sequence occurs.

100. A method for binding a target polynucleotide sequence of a nucleic acid molecule comprising: a) assembling a ribonucleoprotein (RNP) complex under conditions suitable for formation of the RNP complex by combining: i) one or more guide RNAs (gRNAs); and ii) an RNA-guided nuclease (RGN) polypeptide comprising an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, and wherein the amino acid residue at one or more of amino acid positions 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975 of said amino acid sequence differs from the corresponding amino acid residue in SEQ ID NO: 1; and b) contacting said target polynucleotide sequence or a cell comprising said target polynucleotide sequence with the assembled RNP complex, thereby binding said target polynucleotide sequence with said RNP complex.

101. The method of claim 100, wherein said RGN polypeptide comprises an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, and wherein the amino acid residue at one or more of amino acid positions 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975 of said amino acid sequence is a positively charged amino acid residue.

102. The method of claim 100 or 101, wherein said RGN polypeptide comprises an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, and wherein the amino acidresidue at one or more of amino acid positions 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975 of said amino acid sequence is an R.

103. The method of claim 102, wherein said RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence selected from:(a) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 52 is an R;(b) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 55 is an R;(c) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 86 is an R;(d) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 472 is an R;(e) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 533 is an R;(f) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 541 is an R;(g) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 643 is substituted by an R;(h) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 647 is an R;(i) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 653 is an R;(j) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 745 is an R;(k) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position 774 is an R;(l) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 778 is an R;(m) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 780 is an R;(n) the amino acid sequence set forth as SEQ ID NO: 1, wherein A at amino acid position 795 is an R;(o) the amino acid sequence set forth as SEQ ID NO: 1, wherein Q at amino acid position 822 is an R;(p) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 843 is an R;(q) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 856 is an R;(r) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 871 is an R;(s) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 872 is an R;(t) the amino acid sequence set forth as SEQ ID NO: 1, wherein D at amino acid position 900 is an R;(u) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position 911 is an R;(v) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 913 is an R;(w) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position 954 is an R;(x) the amino acid sequence set forth as SEQ ID NO: 1, wherein V at amino acid position 958 is an R;(y) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 968 is an R;(z) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 969 is an R;(aa) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position973 is an R;(bb) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position974 is an R;(cc) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position975 is an R;(dd) the amino acid sequence set forth as SEQ ID NO: 2, wherein Q at amino acid position 822 is an R, G at amino acid position 856 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 778 is an R;(ee) the amino acid sequence set forth as SEQ ID NO: 3, wherein Q at amino acid position 822 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 778 is substituted by an R;(ff) the amino acid sequence set forth as SEQ ID NO: 4, wherein Q at amino acid position 822 is an R, S at amino acid position 55 is an R, and E at amino acid position 778 is an R;(gg) the amino acid sequence set forth as SEQ ID NO: 5, wherein Q at amino acid position 822 is an R, S at amino acid position 647 is an R, and E at amino acid position 778 is an R;(hh) the amino acid sequence set forth as SEQ ID NO: 6, wherein Q at amino acid position 822 is an R, G at amino acid position 856 is an R, and E at amino acid position 778 is an R;(ii) the amino acid sequence set forth as SEQ ID NO: 7, wherein G at amino acid position 856 is an R, and E at amino acid position 778 is an R;(jj) the amino acid sequence set forth as SEQ ID NO: 8, wherein E at amino acid position 778 is an R, Q at amino acid position 822 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(kk) the amino acid sequence set forth as SEQ ID NO: 9, wherein E at amino acid position 778 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(11) the amino acid sequence set forth as SEQ ID NO: 10, wherein E at amino acid position 778 is an R, G at amino acid position 856 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(mm) the amino acid sequence set forth as SEQ ID NO: 11, wherein E at amino acid position 778 is an R, G at amino acid position 856 is an R, S at amino acid position 647 is an R, and E at amino acid position 969 is an R;(nn) the amino acid sequence set forth as SEQ ID NO: 12, wherein E at amino acid position 778 is an R, Q at amino acid position 822 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(oo) the amino acid sequence set forth as SEQ ID NO: 13, wherein E at amino acid position 778 is an R, Q at amino acid position 822 is an R, S at amino acid position 647 is an R, and E at amino acid position 969 is an R;(pp) the amino acid sequence set forth as SEQ ID NO: 14, wherein E at amino acid position 778 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(qq) the amino acid sequence set forth as SEQ ID NO: 15, wherein E at amino acid position 778 is an R, G at amino acid position 856 is an R, and E at amino acid position 969 is an R; and(rr) the amino acid sequence set forth as SEQ ID NO: 16, wherein E at amino acid position 778 is an R, and E at amino acid position 969 is an R.

104. The method of claim 103, wherein said RGN polypeptide comprises an amino acid sequence having 100% sequence identity to the amino acid sequence selected from:(a) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 52 is an R;(b) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 55 is an R;(c) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 86(d) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 472(e) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 533(f) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 541(g) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 643tuted by an R;(h) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 647(i) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 653(j) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 745(k) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position 774(l) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 778(m) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 780(n) the amino acid sequence set forth as SEQ ID NO: 1, wherein A at amino acid position 795(o) the amino acid sequence set forth as SEQ ID NO: 1, wherein Q at amino acid position 822(p) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 843(q) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 856(r) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 871(s) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 872(t) the amino acid sequence set forth as SEQ ID NO: 1, wherein D at amino acid position 900(u) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position 911 is an R;(v) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 913 is an R;(w) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position 954 is an R;(x) the amino acid sequence set forth as SEQ ID NO: 1, wherein V at amino acid position 958 is an R;(y) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 968 is an R;(z) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 969 is an R;(aa) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position973 is an R;(bb) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position974 is an R;(cc) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position975 is an R;(dd) the amino acid sequence set forth as SEQ ID NO: 2;(ee) the amino acid sequence set forth as SEQ ID NO: 3;(ff) the amino acid sequence set forth as SEQ ID NO: 4;(gg) the amino acid sequence set forth as SEQ ID NO: 5;(hh) the amino acid sequence set forth as SEQ ID NO: 6;(ii) the amino acid sequence set forth as SEQ ID NO: 7;(jj) the amino acid sequence set forth as SEQ ID NO: 8;(kk) the amino acid sequence set forth as SEQ ID NO: 9;(11) the amino acid sequence set forth as SEQ ID NO: 10;(mm) the amino acid sequence set forth as SEQ ID NO: 11;(nn) the amino acid sequence set forth as SEQ ID NO: 12;(oo) the amino acid sequence set forth as SEQ ID NO: 13;(pp) the amino acid sequence set forth as SEQ ID NO: 14;(qq) the amino acid sequence set forth as SEQ ID NO: 15; and(rr) the amino acid sequence set forth as SEQ ID NO: 16.

105. The method of any one of claims 100-104, wherein said RGN polypeptide is capable of cleaving said target polynucleotide sequence, thereby allowing for the cleaving and / or modifying of said target polynucleotide sequence.

106. A method for cleaving and / or modifying a target polynucleotide sequence of a nucleic acid molecule, comprising contacting the nucleic acid molecule with: a) an RNA-guided nuclease (RGN) polypeptide, wherein said RGN polypeptide comprises an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, and wherein the amino acid residue at one or more of amino acid positions 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975 of said amino acid sequence differs from the corresponding amino acid residue in SEQ ID NO: 1; and b) one or more guide RNAs (gRNAs) capable of targeting the RGN of (a) to the target polynucleotide sequence, thereby cleaving and / or modifying said target polynucleotide sequence to generate a modified target polynucleotide sequence.

107. The method of claim 106, wherein the one or more gRNAs is capable of hybridizing to the target polynucleotide sequence and binding to said RGN polypeptide.

108. The method of claim 106 or 107, wherein said RGN polypeptide comprises an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, and wherein the amino acid residue at one or more of amino acid positions 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975 of said amino acid sequence is a positively charged amino acid residue.

109. The method of any one of claims 106-108, wherein said RGN polypeptide comprises an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 1, and wherein the amino acid residue at one or more of amino acid positions 52, 55, 86, 472, 533, 541, 643, 647, 653, 745, 774, 778, 780, 795, 822, 843, 856, 871, 872, 900, 911, 913, 954, 958, 968, 969, 973, 974, and 975 of said amino acid sequence is an R.

110. The method of claim 109, wherein said RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to the amino acid sequence selected from:(a) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 52 is an R;(b) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 55 is an R;(c) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 86 is an R;(d) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 472 is an R;(e) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 533 is an R;(f) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 541(g) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 643tuted by an R;(h) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 647(i) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 653(j) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 745(k) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position 774(l) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 778(m) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 780(n) the amino acid sequence set forth as SEQ ID NO: 1, wherein A at amino acid position 795(o) the amino acid sequence set forth as SEQ ID NO: 1, wherein Q at amino acid position 822(p) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 843(q) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 856(r) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 871(s) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 872(t) the amino acid sequence set forth as SEQ ID NO: 1, wherein D at amino acid position 900(u) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position 911(v) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 913(w) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position R;(x) the amino acid sequence set forth as SEQ ID NO: 1, wherein V at amino acid position 958 is an R;(y) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 968 is an R;(z) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 969 is an R;(aa) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position973 is an R;(bb) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position974 is an R;(cc) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position975 is an R;(dd) the amino acid sequence set forth as SEQ ID NO: 2, wherein Q at amino acid position 822 is an R, G at amino acid position 856 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 778 is an R;(ee) the amino acid sequence set forth as SEQ ID NO: 3, wherein Q at amino acid position 822 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 778 is substituted by an R;(ff) the amino acid sequence set forth as SEQ ID NO: 4, wherein Q at amino acid position 822 is an R, S at amino acid position 55 is an R, and E at amino acid position 778 is an R;(gg) the amino acid sequence set forth as SEQ ID NO: 5, wherein Q at amino acid position 822 is an R, S at amino acid position 647 is an R, and E at amino acid position 778 is an R;(hh) the amino acid sequence set forth as SEQ ID NO: 6, wherein Q at amino acid position 822 is an R, G at amino acid position 856 is an R, and E at amino acid position 778 is an R;(ii) the amino acid sequence set forth as SEQ ID NO: 7, wherein G at amino acid position 856 is an R, and E at amino acid position 778 is an R;(jj) the amino acid sequence set forth as SEQ ID NO: 8, wherein E at amino acid position 778 is an R, Q at amino acid position 822 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(kk) the amino acid sequence set forth as SEQ ID NO: 9, wherein E at amino acid position 778 is an R, S at amino acid position 647 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(11) the amino acid sequence set forth as SEQ ID NO: 10, wherein E at amino acid position 778 is an R, G at amino acid position 856 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(mm) the amino acid sequence set forth as SEQ ID NO: 11, wherein E at amino acid position 778 is an R, G at amino acid position 856 is an R, S at amino acid position 647 is an R, and E at amino acid position 969 is an R;(nn) the amino acid sequence set forth as SEQ ID NO: 12, wherein E at amino acid position 778 is an R, Q at amino acid position 822 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(oo) the amino acid sequence set forth as SEQ ID NO: 13, wherein E at amino acid position 778 is an R, Q at amino acid position 822 is an R, S at amino acid position 647 is an R, and E at amino acid position 969 is an R;(pp) the amino acid sequence set forth as SEQ ID NO: 14, wherein E at amino acid position 778 is an R, S at amino acid position 55 is an R, and E at amino acid position 969 is an R;(qq) the amino acid sequence set forth as SEQ ID NO: 15, wherein E at amino acid position 778 is an R, G at amino acid position 856 is an R, and E at amino acid position 969 is an R; and(rr) the amino acid sequence set forth as SEQ ID NO: 16, wherein E at amino acid position 778 is an R, and E at amino acid position 969 is an R.

111. The method of claim 110, wherein said RGN polypeptide comprises an amino acid sequence having 100% sequence identity to the amino acid sequence selected from:(a) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 52 is an R;(b) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 55 is an R;(c) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 86 is an R;(d) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 472 is an R;(e) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 533 is an R;(f) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 541 is an R;(g) the amino acid sequence set forth as SEQ ID NO: 1, wherein Y at amino acid position 643 is substituted by an R;(h) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position 647 is an R;(i) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 653 is an R;(j) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 745(k) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position 774(l) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 778(m) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 780(n) the amino acid sequence set forth as SEQ ID NO: 1, wherein A at amino acid position 795(o) the amino acid sequence set forth as SEQ ID NO: 1, wherein Q at amino acid position 822(p) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 843(q) the amino acid sequence set forth as SEQ ID NO: 1, wherein G at amino acid position 856(r) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 871(s) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 872(t) the amino acid sequence set forth as SEQ ID NO: 1, wherein D at amino acid position 900(u) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position 911(v) the amino acid sequence set forth as SEQ ID NO: 1, wherein T at amino acid position 913(w) the amino acid sequence set forth as SEQ ID NO: 1, wherein N at amino acid position R;(x) the amino acid sequence set forth as SEQ ID NO: 1, wherein V at amino acid position 958(y) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position 968(z) the amino acid sequence set forth as SEQ ID NO: 1, wherein E at amino acid position 969(aa) the amino acid sequence set forth as SEQ ID NO: 1, wherein K at amino acid position R;(bb) the amino acid sequence set forth as SEQ ID NO: 1, wherein S at amino acid position974 is an R;(cc) the amino acid sequence set forth as SEQ ID NO: 1, wherein L at amino acid position975 is an R;(dd) the amino acid sequence set forth as SEQ ID NO: 2;(ee) the amino acid sequence set forth as SEQ ID NO: 3;(ff) the amino acid sequence set forth as SEQ ID NO: 4;(gg) the amino acid sequence set forth as SEQ ID NO: 5;(hh) the amino acid sequence set forth as SEQ ID NO: 6;(ii) the amino acid sequence set forth as SEQ ID NO: 7;(jj) the amino acid sequence set forth as SEQ ID NO: 8;(kk) the amino acid sequence set forth as SEQ ID NO: 9;(11) the amino acid sequence set forth as SEQ ID NO: 10;(mm) the amino acid sequence set forth as SEQ ID NO: 11;(nn) the amino acid sequence set forth as SEQ ID NO: 12;(oo) the amino acid sequence set forth as SEQ ID NO: 13;(pp) the amino acid sequence set forth as SEQ ID NO: 14;(qq) the amino acid sequence set forth as SEQ ID NO: 15; and(rr) the amino acid sequence set forth as SEQ ID NO: 16.

112. The method of any one of claims 99, and 105-111, wherein said modified target DNA sequence comprises insertion of heterologous DNA into the target DNA sequence or deletion of at least one nucleotide from the target DNA sequence.

113. The method of any one of claims 99, and 105-111, wherein said modified target DNA sequence comprises mutation of at least one nucleotide in the target DNA sequence.

114. The method of any one of claims 95-113, wherein said method is performed in vitro, in vivo, or ex vivo.

115. The method of any one of claims 98-114, wherein said target polynucleotide sequence is a eukaryotic target DNA sequence.

116. The method of any one of claims 98-115, wherein the target polynucleotide sequence is within a cell.

117. The method of claim 116, wherein the cell is a eukaryotic cell.

118. The method of claim 117, wherein the eukaryotic cell is a mammalian cell.

119. The method of any one of claims 116-118, further comprising culturing the cell under conditions in which the RGN polypeptide is expressed and cleaves and modifies the target polynucleotide sequence to produce a nucleic acid molecule comprising a modified targetpolynucleotide sequence; and selecting a cell comprising said modified target polynucleotide sequence.

120. A cell comprising a modified target polynucleotide sequence produced according to the method of claim 119.

121. A pharmaceutical composition comprising the cell of any one of claims 30, 96, and 120, and a pharmaceutically acceptable carrier.

122. A method for treating a subject having or at risk of developing a disease, disorder, or condition, the method comprising: administering to the subject the nucleic acid molecule of any one of claims 1-22, the vector of any one of claims 23-29, the RGN polypeptide of any one of claims 32-61, the RNP complex of claim 62, the system of any one of claims 63-95, the cell of any one of claims 30, 96 and 120, or the pharmaceutical composition of claim 97 or 121.

123. The method of claim 122, wherein said disease, disorder, or condition is associated with a mutation and said treating comprises correcting said mutation.

124. Use of the nucleic acid molecule of any one of claims 1-22, the vector of any one of claims 23-29, the RGN polypeptide of any one of claims 32-61, the RNP complex of claim 62, the system of any one of claims 63-95, the cell of any one of claims 30, 96 and 120, or the pharmaceutical composition of claim 97 or 121 in the manufacture of a medicament useful for treating a disease, disorder, or condition.

125. The use of claim 124, wherein said disease, disorder, or condition is associated with a mutation and an effective amount of said medicament corrects said mutation.

126. The RNP complex of claim 62 or the system of any one of claims 63-95 for use in binding a target polynucleotide sequence of a nucleic acid molecule.

127. The RNP complex of claim 62 or the system of any one of claims 63-95 for use in cleaving and / or modifying a target polynucleotide sequence of a nucleic acid molecule.

128. The nucleic acid molecule of any one of claims 1-22, the vector of any one of claims 23-29, the RGN polypeptide of any one of claims 32-61, the RNP complex of claim 62, the system of any one of claims 63-95, the cell of any one of claims 30, 96 and 120, or the pharmaceutical composition of claim 97 or 121 for use in treating a disease, disorder, or condition.

129. A method of increasing efficiency of cleaving and / or modifying a nucleic acid molecule comprising a target sequence, the method comprising delivering the system of any one of claims 63-95 or the RNP complex of claim 62 to the target sequence or to a cell comprising the target sequence, wherein cleavage or modification of the nucleic acid molecule occurs at greater efficiency as compared to cleavage or modification of the nucleic acid molecule by a method comprising delivering to the target sequence or to a cell comprising the target sequence a reference RGN systemor RNP complex, wherein the reference RGN system or RNP complex does not comprise said RGN polypeptide.

130. The method of claim 129, wherein the efficiency of cleaving and / or modifying the target sequence is increased by at least 15%.

131. The method of claim 129 or 130, wherein the efficiency of cleaving and / or modifying the target sequence is measured by next generation sequencing, Tracking of Indels by DEcomposition (TIDE) analysis, flow cytometry, or a combination thereof.

132. A fusion polypeptide comprising:(a) the RGN polypeptide of any one of claims 32-48; and(b) a heterologous polypeptide.

133. The fusion polypeptide of claim 132, wherein said heterologous polypeptide is operably fused to the N-terminus, to the C-terminus, or to an internal location of said RGN polypeptide.

134. The fusion polypeptide of claim 132 or 133, wherein said heterologous polypeptide is a polymerase editing polypeptide.

135. The fusion polypeptide of claim 132 or 133, wherein the heterologous polypeptide is a base-editing polypeptide.

136. The fusion polypeptide of claim 132 or 133, wherein the heterologous polypeptide is an effector domain, a detectable label, or a purification tag.

137. The fusion polypeptide of claim 136, wherein the effector domain is a cleavage domain, a deaminase domain, or an expression modulator domain.

138. A polynucleotide comprising a nucleotide sequence encoding the fusion polypeptide of any one of claims 132-137.

139. One or more polynucleotides encoding a polymerase editor (PE) comprising a polymerase and an RNA-guided nuclease (RGN) polypeptide of any one of claims 32-48.

140. A polymerase editor (PE) comprising a polymerase and an RNA-guided nuclease (RGN) polypeptide of any one of claims 32-48.

141. The PE of claim 140, wherein said polymerase is a reverse transcriptase.

142. The PE of claim 141, wherein said reverse transcriptase has at least 90% sequence identity to the amino acid sequence of SEQ ID NO: 254 or 255.

143. The PE of any one of claims 140-142, wherein said polymerase is operably fused to the N-terminus of said RGN polypeptide.

144. The PE of any one of claims 140-142, wherein said polymerase is operably fused to the C-terminus of said RGN polypeptide.

145. The PE of any one of claims 140-142, wherein said polymerase is operably fused to an internal location of said RGN polypeptide.

146. The PE of claim 145, wherein said polymerase is operably fused within said RGN polypeptide immediately after an amino acid at a position selected from the group consisting of: a) an amino acid position corresponding to position 666 of SEQ ID NO: 1; b) an amino acid position corresponding to position 785 of SEQ ID NO: 1; and c) an amino acid position corresponding to position 910 of SEQ ID NO: 1.

147. The PE of any one of claims 140-146, wherein said PE comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 135, and 162-181.

148. A PE system for modifying one or more target sequences in a target polynucleotide, said system comprising: a) one or more polymerase editing guide RNAs (PEgRNAs), or one or more polynucleotides comprising one or more nucleotide sequences encoding the one or more PEgRNAs, wherein the one or more PEgRNAs comprise an extension arm, wherein the extension arm comprises a primer binding site (PBS) and a DNA synthesis template sequence; and b) a PE of any one of claims 140-147; wherein said one or more PEgRNAs are capable of binding to said RGN polypeptide of said PE.

149. A method for modifying a target polynucleotide comprising a target sequence, said method comprising delivering a PE system according to claim 148 to said target sequence or a cell comprising the target sequence, wherein said method generates a modified target polynucleotide, and wherein components of said PE system are delivered simultaneously or sequentially to said target sequence or said cell comprising the target sequence.

150. One or more polynucleotides encoding a base editor comprising a base editing polypeptide and an RNA-guided nuclease (RGN) polypeptide of any one of claims 32-48.

151. A base editor comprising a base editing polypeptide and an RNA-guided nuclease (RGN) polypeptide of any one of claims 32-48.

152. The base editor of claim 151, wherein said base editing polypeptide is a deaminase.

153. The base editor of claim 152, wherein the deaminase has at least 90% sequence identity to an amino acid sequence of any one of SEQ ID NOs: 42-113, and 257.

154. The base editor of claim 151, wherein said base editing polypeptide is operably fused to the N-terminus of said RGN polypeptide.

155. The base editor of claim 151, wherein said base editing polypeptide is operably fused to the C-terminus of said RGN polypeptide.

156. The base editor of claim 151, wherein said base editing polypeptide is operably fused to an internal location of said RGN polypeptide.

157. The base editor of claim 156, wherein said base editing polypeptide is operably fused within said RGN polypeptide immediately after an amino acid at a position selected from the group consisting of:a) an amino acid position corresponding to position 666 of SEQ ID NO: 1; b) an amino acid position corresponding to position 785 of SEQ ID NO: 1; and c) an amino acid position corresponding to position 910 of SEQ ID NO: 1.

158. The base editor of any one of claims 151-157, wherein said base editor comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 258-270.

159. A base editing system for modifying one or more target sequences in a target polynucleotide, said system comprising: a) one or more guide RNAs (gRNAs), or one or more polynucleotides comprising one or more nucleotide sequences encoding the one or more gRNAs; and b) a base editor of any one of claims 151-158; wherein said one or more gRNAs is capable of binding to said RGN polypeptide of said base editor.

160. A method for modifying a target polynucleotide comprising a target sequence, said method comprising delivering a base editing system according to claim 159 to said target sequence or a cell comprising the target sequence, wherein said method generates a modified target polynucleotide, and wherein components of said base editing system are delivered simultaneously or sequentially to said target sequence or said cell comprising the target sequence.