Novel compact crispr / cas12f1 system

A compact Cas12f polypeptide with specific amino acid substitutions addresses the delivery challenges of CRISPR-Cas systems by enhancing targeting and editing efficiency, enabling efficient gene editing and epigenetic modifications within cells.

WO2025117934A1PCT designated stage expired Publication Date: 2025-06-05EPITOR THERAPEUTICS
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/058047
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-07
Filing Date
2024-12-02
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Current CRISPR-Cas systems face challenges in delivering large Cas nuclease proteins and guide RNAs into cells efficiently, particularly when combined with other effector proteins, due to size constraints and delivery limitations.

Method used

Development of a compact Cas12f polypeptide with specific amino acid substitutions, such as N195K, Y306K, and T307K, which improves targeting and editing efficiency, allowing for efficient delivery within an adenovirus-associated virus (AAV) and enabling epigenetic editing.

Benefits of technology

The compact Cas12f polypeptide demonstrates enhanced editing activity and reduced nuclease activity, facilitating efficient gene editing and epigenetic modifications within cells, while being compatible with AAV delivery systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024058047_05062025_PF_FP_ABST
    Figure US2024058047_05062025_PF_FP_ABST
Patent Text Reader

Abstract

Provided herein are novel compact Cas12f nucleases and polypeptides, improved gRNAs and compositions and methods of use thereof.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No: 124540-829360 NOVEL COMPACT CRISPR / Cas12f1 SYSTEM CROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of U.S. Provisional Application Ser. No. 63 / 604,969 filed December 1, 2023, and U.S. Provisional Application Ser. No.63 / 643,752 filed May 7, 2024, the contents of each of which are hereby incorporated by reference in their entirety. INCORPORATION BY REFERENCE OF SEQUENCE LISTING

[0002] This application contains a Sequence Listing that has been submitted in XML format via Patent Center and is hereby incorporated by reference in its entirety. The XML file, created on November 27, 2024, is named 124540-829360_SequenceListing, and is about 1,085,440 bytes in size. FIELD OF THE INVENTION

[0003] The present disclosure relates to the field of gene editing and, in some embodiments, to compact nuclease polypeptides for use in editing or modifying nucleic acids. BACKGROUND

[0004] In the field of clustered regularly interspaced palindromic repeats (CRISPR) and CRISPR-associated (Cas) discovery, prokaryotic CRISPR-Cas systems have been leveraged for uses as a genome engineering tools. The CRISPR-Cas system comprises a CRISPR associated (Cas) nuclease which works in combination with a guide RNA (gRNA) to specifically target a nucleic acid of interest. The challenge with the system is in delivering the Cas nuclease (or a fusion protein comprising the Cas nuclease), which may be very large, along with the target gRNA into a cell. To address this, various compact Cas nucleases have been developed, but very few are capable of delivery in a single adenovirus associated virus (AAV), especially when the Cas protein is fused with other effector proteins. New compact Cas nucleases are needed, especially with improved targeting and / or editing efficiency. SUMMARY

[0005] In some aspects, the present disclosure is directed to a Cas12f polypeptide having one or more amino acid substitutions relative to an amino acid sequence of a wildtype Cas12f nuclease obtained from a Blautia species, wherein the Cas12f polypeptide has improved targeting and / or editing efficiency relative to the wildtype bsCas12f (bsCas12f). -1- 99975521.7Attorney Docket No: 124540-829360

[0006] In an aspect, the wildtype Cas12f nuclease has an amino acid sequence of SEQ ID NO:1, and the one or more amino acid substitutions are selected from a substitution at any one or more of I10, Y48, A50, T66, D76, S98, T99, E109, I110, G111, N114, F119, N103, M122R, S131, W143, I159, G176, E181, D183, K184, N195, I196, I201, P209, K220, A223, G266, A276, Y306, T307, N312, K318, A355, and D353 in SEQ ID NO: 1.

[0007] In another aspect, the one or more amino acid substitutions are selected from a substitution at any one or more of I10, Y48, A50, T66, S98, T99, E109, N114, F119, M122R, S131, W143, G176, E181, D183, K184, N195, I196, K220, A223, G266, A276, Y306, T307, N312, K318, and D353 in SEQ ID NO: 1. For instance, in an aspect, the one or more amino acid substitutions are at one or more residues of N195, Y306, T307, E109, S131, E181, D183, D353, D183, I10, D76, G111, I159, I196 in SEQ ID NO: 1. In another aspect, the one or more amino acid substitutions are selected from I10R, Y48F, Y48R, A50R, A50Y, T66Q, D76K, S98H, S98R, T99R, E109M, E109R, G111N, N114Q, N114A, F119Y, F119H, M122R, S131R, S131H, S131C, W143T, I159R, G176R, G176H, E181Q, D183K, D183T, K184R, N195Q, N195R, N195K, N195H, I196V, K220R, A223R, G266R, A276R, Y306F, Y306M, Y306K, T307K, N312V, N312Y, K318R, K318G and D353R..

[0008] In an aspect, the Cas12f polypeptide comprises two or more amino acid substitutions selected from: N195K, N195H, E109H, S131W, S131C, E181Q, D183T, Y306K, T307K, I10R, D76K, G111N, I159R, I196V, and D353K. In another aspect, the Cas12f polypeptide comprises three or more amino acid substitutions selected from: N195K, N195H, Y306K, D353K, S131Q, E181Q, D183T, I10R, D76K, G111N, I159R, I196V, and T307K. In another aspect, the Cas12f polypeptide comprises four or more amino acid substitutions selected from N195K, Y306K, T307K, E109H, S131C, S131Q, E181Q, D183T, I10R, D76K, G111N, I159R, I196V, and D353K. In another aspect, the Cas12f polypeptide comprises five or more amino acid substitutions selected from N195K, Y306K, T307K, S131C, E181Q I10R, D76K, G111N, I159R, I196V, and D183T. In still another aspect, the Cas12f polypeptide comprises six or more amino acid substitutions selected from N195K, Y306K, T307K, S131C, E181Q, I10R, D76K, G111N, I159R, I196V, and D183T. In another aspect, the Cas12f polypeptide comprises seven or more amino acid substitutions selected from N195K, Y306K, T307K, S131C, E181Q, D183T, and I10R. In another aspect, the Cas12f polypeptide comprises eight or more amino acid substitutions selected from N195K, Y306K, T307K, S131C, E181Q, D183T, I10R, and D76K. In another aspect, the Cas12f polypeptide comprises nine or more amino acid substitutions selected from N195K, -2- 99975521.7Attorney Docket No: 124540-829360 Y306K, T307K, S131C, E181Q, D183T, I10R, D76K, and G111N. In another aspect, the Cas12f polypeptide comprises ten or more amino acid substitutions selected from N195K, Y306K, T307K, S131C, E181Q, D183T, I10R, D76K, G111N, and I159R. In another aspect, the Cas12f polypeptide comprises eleven or more amino acid substitutions selected from N195K, Y306K, T307K, S131C, E181Q, D183T, I10R, D76K, G111N, I159R, and I196V.

[0009] In various aspects, the Cas12f polypeptide has increased editing activity relative to the wildtype Cas12f nuclease. In various aspects, the Cas12f polypeptide has reduced nuclease activity relative to the wildtype Cas12f nuclease. For instance, in some aspects, the Cas12f polypeptide further comprising one or more amino acid substitutions selected from D216A and D390A.

[0010] In various aspects, the Cas12f polypeptide comprises or consists of an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs: 16-391 and SEQ ID NOs: 803-807. For instance, in an aspect, the Cas12f polypeptide may comprise or consist of an amino acid sequence of SEQ ID NO: 329- 338 and SEQ ID NOs: 803-807. In some aspects, the Cas12f polypeptide comprises or consists of an amino acid sequence of SEQ ID NO: 338. In some aspects, the Cas12f polypeptide comprises or consists of an amino acid sequence of SEQ ID NO: 807.

[0011] Further aspects of the present disclosure provide for a fusion protein comprising a Cas12f polypeptide of any one of claims 1 to 16 and a functional polypeptide.

[0012] In an aspect, the functional polypeptide can comprise a base editing domain, for example, a deaminase or a catalytic domain thereof, a base excising domain, an uracil glycosylase inhibitor (UGI) or a catalytic domain thereof, an uracil glycosylase (UNG) or a catalytic domain thereof, a methylpurine glycosylase (MPG) or a catalytic domain thereof, a methylase or a catalytic domain thereof, a demethylase or a catalytic domain thereof, an transcription activating domain (e.g., VP64 or VPR) , an transcription inhibiting domain (e.g., KRAB moiety or SID moiety) , a reverse transcriptase or a catalytic domain thereof, an exonuclease (e.g., T5E) or a catalytic domain thereof, a destabilized domain (e.g., destabilized domains (DD) of E. coli dihydrofolate reductase (ecDHFR) ) , a histone residue modification domain, a nuclease catalytic domain (e.g., FokI) , a transcription modification factor, a light gating factor, a chemical inducible factor, a chromatin visualization factor, a targeting polypeptide for providing binding to a cell surface portion on a target cell or a target -3- 99975521.7Attorney Docket No: 124540-829360 cell type, a reporter (e.g., fluorescent) polypeptide or a detection label (e.g., GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, BFP) , a localization signal, a polypeptide targeting moiety, a DNA binding domain (e.g., MBP, Lex A DBD, Gal4 DBD) , an epitope tag (e.g., His, myc, V5, FLAG, HA, VSV-G, Trx, etc) , a transcription release factor, an HDAC, a moiety having ssRNA cleavage activity, a moiety having dsRNA cleavage activity, a moiety having ssDNA cleavage activity, a moiety having dsDNA cleavage activity, a DNA or RNA ligase, a functional domain exhibiting activity to modify a target DNA, selected from the group consisting of: methyltransferase activity, DNA repair activity, DNA damage activity, dismutase activity, alkylation activity, dealkylation activity, depurination activity, oxidation activity, deoxidation activity, pyrimidine dimer forming activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, glycosylase activity, acetyl transferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, deSUMOylation activity, ribosylation activity, deribosylation activity, myristoylation activity, demyristoylation activity, glycosylation activity (e.g., from O-GlcNAc transferase) , deglycosylation activity, reverse transcriptase activity, and a catalytic domain thereof, and a functional fragment (e.g., a functional truncation) thereof, or any combination thereof.

[0013] In an aspect, the functional polypeptide comprises a Ten-eleven translocation methylcytosine dioxygenase (Tet, Tet2 or Tet3, or a functional or catalytic domain thereof), DNA methyl transferase (e.g. DNMT3A and / or DNMT3L, or a functional or catalytic domain thereof), VP16, VP64, VPR, ABE, CBE, or a Krüppel-associated box (KRAB).

[0014] In an aspect, the functional polypeptide comprises a catalytically active truncated Tet polypeptide comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 408 or 409.

[0015] In an aspect, the functional polypeptide comprises a DNMT3A protein.

[0016] In still further aspects, any Cas12f polypeptide or fusion protein provided herein can further comprise a nuclear localization signal (NLS).

[0017] In various aspects, a nucleic acid is provided encoding any Cas12f polypeptide or fusion protein provided herein. -4- 99975521.7Attorney Docket No: 124540-829360

[0018] Further aspects of the present disclosure provide for a single molecule guide RNA (sgRNA), comprising a scaffold sequence capable of forming a complex with any Cas12f polypeptide provided herein and / or any fusion protein provided herein and a spacer sequence capable of hybridizing to a target sequence of a target DNA, thereby guiding the complex to the target DNA.

[0019] In an aspect, the sgRNA may comprise a scaffold sequence comprising a nucleic acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to any one of SEQ ID NOs: 438-515 and SEQ ID NOs: 673-802.

[0020] In an aspect, the single-molecule guide RNA may be chemically modified. In various aspects, the sgRNA may be pre-complexed with any Cas12f polypeptide or a fusion protein provided herein.

[0021] Also provided are nucleic acids encoding any sgRNA provided herein.

[0022] Further aspects of the present disclosure are directed to a vector comprising one or more expression constructs comprising one or more polynucleotides encoding the Cas12f polypeptide, fusion protein and / or sgRNA provided herein. In some aspects, the vector comprises both a polynucleotide encoding a Cas12f polypeptide and / or a fusion protein comprising it and an sgRNA. In various aspects, the vector is an adenoviral, adeno-associated viral (AAV), or lentiviral vector.

[0023] In various aspects, the vector provided herein comprises a nucleic acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 558-605.

[0024] In still further aspects, a cell is provided comprising any polynucleotide or vector provided herein encoding one or more of the Cas12f polypeptide, the fusion protein and / or the sgRNA, or any combination thereof. In some aspects, the cell further comprises one or more TurboRFP expressing constructs under control of a TRE operator comprising seven TetO operators, wherein the cell expresses RFP only in the presence of a targeted transcriptional activator binding to the TetO operator of at least one TurboRFP expressing construct. In still further aspects, each TetO operator in each TurboRFP expressing construct may comprise a PAM sequence selected from TTTA, TTTC or TCCA preceding each TetO operator sequence. For instance, such a cell may comprise a vector having at least 70%, at -5- 99975521.7Attorney Docket No: 124540-829360 least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 593-595. In any of these aspects the cell may be a mammalian cell (e.g., a human cell).

[0025] Further aspects of the present disclosure are also directed to a gene editing system comprising (1) a Cas12f polypeptide as provided herein or an encoding polynucleotide thereof or a fusion protein provided herein or the polynucleotide thereof and (2) a single molecule guide RNA (sgRNA) provided herein or an encoding polynucleotide thereof, and optionally, (3) a transgene encoding a therapeutic protein of interest, which is optionally comprised in an episome. In an aspect, the system is packaged into an adenoviral, adeno- associated viral (AAV), or lentiviral vector. In an aspect, the system may be further packaged into a liposome or lipid nanoparticle.

[0026] Further aspects of the present disclosure are directed to a method for modifying a target DNA by contacting the target DNA with a complex formed between a Cas12f polypeptide described herein or a fusion protein described herein and a sgRNA described herein, wherein the sgRNA comprises a spacer sequence capable of hybridizing to a target sequence in the target DNA, and wherein the complex modifies the target DNA. In an aspect, the complex can modify the target DNA by introducing one or more single or double stranded breaks at the target sequence. In further aspects, the method can further comprise providing an exogenous nucleic acid, wherein the exogenous nucleic acid is inserted at the site of the double stranded break in the target sequence. In various aspects, the complex can modify the target DNA by editing a base at or near the target sequence.

[0027] Further aspects of the present disclosure are directed to a method for modifying expression of a target gene, the method comprising: contacting a nucleic acid comprising the target gene with a complex formed between a Cas12f polypeptide provided herein or a fusion protein provided herein and an sgRNA provided herein, wherein the sgRNA comprises a spacer sequence capable of hybridizing to a target sequence in the target gene, and wherein the complex modifies expression of the target gene.

[0028] In any of the methods provided herein, the complex can comprise a fusion protein comprising a functional polypeptide that modifies expression of the target gene. In further aspects, the functional polypeptide has methyltransferase activity, DNA repair activity, DNA damage activity, dismutase activity, alkylation activity, demethylation activity, dealkylation activity, depurination activity, oxidation activity, deoxidation activity, pyrimidine dimer -6- 99975521.7Attorney Docket No: 124540-829360 forming activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, glycosylase activity, acetyl transferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, deSUMOylation activity, ribosylation activity, deribosylation activity, myristoylation activity, demyristoylation activity, glycosylation activity (e.g., from O- GlcNAc transferase) , deglycosylation activity or any combination thereof.

[0029] In still further aspects, the functional polypeptide comprises a Ten-eleven translocation methylcytosine dioxygenase (Tet, Tet2 or Tet3, or a functional or catalytic domain thereof (e.g., TetMini)), DNA methyl transferase (e.g. DNMT3A and / or DNMT3L or a functional or catalytic domain thereof), VP16, VP64, VPR, ABE, CBE, or a Krüppel- associated box (KRAB).

[0030] In any of the methods provided herein, the target DNA and / or the target gene are in a cell, and the method further comprises delivering to the cell the Cas12f polypeptide or a polynucleotide encoding the Cas12f polypeptide and the sgRNA or a polynucleotide encoding the sgRNA.

[0031] In any of the methods provided herein, the cell may be a human cell. In an aspect, the cell may be in vitro or in vivo.

[0032] Also provided herein is a method for diagnosing, preventing or treating a disease in a subject in need thereof, the method comprising: modifying a target DNA and / or modifying expression of a target gene in at least one cell of the subject according to any of the methods provided herein, wherein modifying the target DNA and / or modifying expression of the target gene diagnoses, prevents and / or treats the disease. In an aspect, the disease may comprise a hereditary genetic disease. In further aspects, the subject may be a human.

[0033] While the disclosure is susceptible to various modifications and alternative forms, specific embodiments thereof are shown by way of example in the drawings and will herein be described in detail. It should be understood, however, that the drawings and detailed description presented herein are not intended to limit the disclosure to the particular embodiments disclosed, but on the contrary, the intention is to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the present disclosure as defined by the appended claims. -7- 99975521.7Attorney Docket No: 124540-829360

[0034] Other features and advantages of this disclosure will become apparent in the following detailed description of embodiments, taken with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] FIG.1 shows the plasmid map of an empty pCS1-VPR vector (SEQ ID NO: 588).

[0036] FIG.2 shows a plasmid map of the pCS1-dBsCas12-VPR vector (SEQ ID NO: 589).

[0037] FIG.3 shows a plasmid map of a pMINI vector (SEQ ID NO: 590).

[0038] FIG.4 shows a plasmid map of the pMINI-BsCas12f gRNA (174-nt)-TetO vector (SEQ ID NO: 591).

[0039] FIG.5 shows a plasmid map of pMINI-BsCas12f gRNA (140-nt):-TetO vector (SEQ ID NO: 592).

[0040] FIG.6 depicts a schematic of the TurboRFP-expressing cassette under control of TRE.

[0041] FIG.7 shows a plasmid map of the TetO-TurboRFP reporter plasmid (TCCA version) (SEQ ID NO: 593).

[0042] FIG.8A-8B shows a table (FIG.8A) and plot (FIG.8B) quantifying RFP expression in TurboRFP reporter cells expressing the indicated dCas polypeptides in the presence of control (blue bars, left) or TetO gRNA (orange bars, right).

[0043] FIG.8C shows a plot of RFP expression in TurboRFP reporter cells expressing dCas proteins from different species.

[0044] FIG.9 shows a plot quantifying the mean fluorescent intensity (MFI) ratio of TurboRFP reporter cells expressing dBsCas12f polypeptides comprising the indicated mutations as compared to reporter cells expressing wildtype dBsCas12f polypeptides.

[0045] FIG.10 shows a plot quantifying RFP expression in TurboRFP reporter cells expressing wildtype dBsCas12f and TetO targeting gRNA comprising different scaffolds as indicated.

[0046] FIG.11 shows results from an AI generated predictions of different substitutions at each residue of wildtype BsCas12f. -8- 99975521.7Attorney Docket No: 124540-829360

[0047] FIG.12 shows summary of relative performance of different mutated Cas12f polypeptides having 1 to 6 (V1 to V6) substitutions relative to wildtype.

[0048] FIG.13 shows (at top) a diagram demonstrating the small size of CasNano compared to other known Cas nucleases and (at bottom) a plot of the total size of different gene editing systems using CasNano or other known Cas nucleases.

[0049] FIG.14 shows a diagram of the PAM site for CasNano compared to SpCas9.

[0050] FIG.15A shows a diagram summarizing a process for systemically engineering an improved CasNano.

[0051] FIG.15B shows a diagram of results from systemic and targeted mutagenesis of a wildtype Cas showing improved efficiency by combining substitutions.

[0052] FIG.15C shows a table showing that optimized gRNA scaffolds improve targeting efficiency.

[0053] FIG.16D shows results demonstrating that CasNano (an optimized Cas polypeptide disclosed herein) has low tolerance for spacer mismatches.

[0054] FIG.16A shows results of editing efficiency (measured as reporter cell line fluorescence) for different editing systems using different versions of an optimized Cas protein of the instant disclosure (CasNano (V4, V5 and V6)), vs SpCas9 and in combination with two different gRNAs.

[0055] FIG.16B shows the results of endogenous gene activation triggered by an optimized Cas protein of the instant disclosure (dCasNano) in construct with a transcriptional activator (VPR) in HEK293 cells using two different sets of gRNAs.

[0056] FIG.17A shows a diagram of different epigenetic editors that were tested to compare demethylation efficiency.

[0057] FIG.17B shows a plot of demethylation of MeCP2 promoter region using a construct of an optimized Cas protein (dCasNano) and TETmini in HEK293 cells.

[0058] FIG.18A shows a diagram of C9Orf72(GGGCC)n Repeats reporters that were used to test the methylation efficiency of different epigenetic editors.

[0059] FIG.18B shows a plot of % DNA methylation of different epigenetic editors at the PGK promoter region. -9- 99975521.7Attorney Docket No: 124540-829360

[0060] FIG.18C shows a plot of % DNA methylation of different epigenetic editors at the G4C2 repeat region.

[0061] FIG.19A shows a plasmid map of a pRP[Exp]-UBC-dSpCas9-TET2 (1.6K)-Puro vector (SEQ ID NO: 596).

[0062] FIG.19B shows a plasmid map of a pCS1-dCasNano-Tet1-GFP-Puro vector (SEQ ID NO: 597).

[0063] FIG.19C shows a plasmid map of a pCS1-dCasNano-Tet2-Puro vector (SEQ ID NO: 598).

[0064] FIG.19D shows a plasmid map of a pCS1-Tet2-dCasNano-Puro vector (SEQ ID NO: 599).

[0065] FIG.20A shows a plasmid map of a pCS3-dCasNano-hDNMT3A-3L (CD) vector (SEQ ID NO: 600).

[0066] FIG.20B shows a plasmid map of a pCS3-dCasNano-hDNMT3A(CD) vector (SEQ ID NO: 601).

[0067] FIG.20C shows a plasmid map of a pCS3-dCasNano-hDNMT3L(CD) vector (SEQ ID NO: 602).

[0068] FIG.20D shows a plasmid map of a pCS3-dCasNano-hH3-Head-DNMT3A (CD) vector (SEQ ID NO: 603).

[0069] FIG.20E shows a plasmid map of a pCS3-dCasNano-mDNMT3A-3L(CD) vector (SEQ ID NO: 604).

[0070] FIG.21 shows a plasmid map of a pCS1-dBsCas12F-V333 (CasNano)-VPR vector (SEQ ID NO: 605).

[0071] FIG.22A-22B show plots of TurboRFP+ population (%) (FIG.22A) and a Geo MFI plot (FIG.22B) of CasNano editing systems using different spacers having lengths from 16 to 24 nucleotides.

[0072] FIG.23A shows a diagram of results from systemic and targeted mutagenesis of a dCasNano showing editing efficiency and

[0073] FIG.23B shows a diagram of results from systemic and targeted mutagenesis of a dCasNano showing editing efficiency and -10- 99975521.7Attorney Docket No: 124540-829360

[0074] FIG.24A shows a bar graph displaying the performance of CasNano variants on TCCA-NGG reporter cells. V1 corresponds to N195K; V2 corresponds to N195K and Y306K; V3 corresponds to N195K, Y306K, and T307K; V4 corresponds to N195K, Y306K, T307K, and S131C; V5 corresponds to N195K, Y306K, T307K, S131C, and E181Q; V6 corresponds to N195K, Y306K, T307K, S131C, E181Q, and D183T.

[0075] FIG.24B shows a bar graph displaying the performance of CasNano variants on TCCA-NGG reporter cells. V7 corresponds to N195K, Y306K, T307K, S131C, E181Q, D183T, and I10R; V8 corresponds to N195K, Y306K, T307K, S131C, E181Q, D183T, I10R, and D76K; V9 corresponds to N195K, Y306K, T307K, S131C, E181Q, D183T, I10R, D76K, and G111N; V10 corresponds to N195K, Y306K, T307K, S131C, E181Q, D183T, I10R, D76K, G111N, and I159R; V11 corresponds to N195K, Y306K, T307K, S131C, E181Q, D183T, I10R, D76K, G111N, I159R, and I196V.

[0076] FIG.24C is a matrix diagram displaying the activities observed when pairing different CasNano variants with different gRNA scaffold variants. V8 corresponds to N195K, Y306K, T307K, S131C, E181Q, D183T, I10R, and D76K; V9 corresponds to N195K, Y306K, T307K, S131C, E181Q, D183T, I10R, D76K, and G111N; V10 corresponds to N195K, Y306K, T307K, S131C, E181Q, D183T, I10R, D76K, G111N, and I159R; V11 corresponds to N195K, Y306K, T307K, S131C, E181Q, D183T, I10R, D76K, G111N, I159R, and I196V. S001 corresponds to SEQ ID NO: 673, S003 corresponds to SEQ ID NO: 675, S075 corresponds to SEQ ID NO: 735, S077 corresponds to SEQ ID NO: 737, S078 corresponds to SEQ ID NO: 738.

[0077] FIG.25A is a schematic diagram depicting the sense and antisense strands of spCas9 and CasNano.

[0078] FIG.25B is a bar graph depicting RFP fluorescence when either spCas9 or CasNano were targeted to the sense strand or the antisense strand to activate RFP expression.

[0079] FIG.26 is a bar graph depicting the activity of CasNano using the following 13 variations of the PAM sequence: TCCA, ACCA, GCCA, CCCA, TACA, TTCA, TGCA, TCAA, TCTA, TCGA, TCCT, TCCG, TCCC.

[0080] FIG.27 shows flow cytometry results depicting RFP fluorescence levels when the CasNano variants CasNano, CasNano D216A, CasNano D390A, and CasNano D216A and D390A were targeted to an RFP expression cassette in RFP expressing reporter cells. -11- 99975521.7Attorney Docket No: 124540-829360

[0081] FIG.28A is a line graph depicting the methylation levels of the MeCP2 locus in HEK293 cells expressing dCasNano-TETmini epigenetic editors.

[0082] FIG.28B is a violin plot depicting the methylation levels of the MeCP2 locus in HEK293 cells expressing dCasNano-TETmini epigenetic editors.

[0083] FIG.29A is a line graph depicting the methylation levels of the MeCP2 locus in HEK293 cells expressing dCasNano-DMNT epigenetic editors.

[0084] FIG.29B a violin plot depicting the methylation levels of the MeCP2 locus in HEK293 cells expressing dCasNano- DMNT epigenetic editors.

[0085] FIG.30 is a representative structure diagram annotating the predicted CasNano gRNA scaffold structure. DETAILED DESCRIPTION

[0086] Gene therapy and gene editing hold great potential for treating genetic disorders, but delivery challenges remain. Adeno-associated viruses (AAVs) are promising delivery vehicles; however, their limited packaging capacity of ~4.5 kb hinders the delivery of many CRISPR-based gene editing tools, particularly those involving large epigenetic effectors like TET enzymes, DNMTs or VPR. Epigenetic editing enables precise modulation of gene expression without altering the DNA sequence, but the size of effector proteins combined with conventional CRISPR / Cas systems exceeds AAV packaging limits. To address this, the development of ultracompact CRISPR / Cas systems accommodating epigenetic effectors within AAV capacity is crucial. These compact systems would enable efficient delivery of epigenetic editing tools and expand possibilities for multiplex editing and tissue-specific targeting. The present disclosure describes how overcoming delivery bottlenecks enables the identification and engineering of an ultracompact CRISPR / Cas system originated from a commensal bacteria species which is called CasNano (see FIG.13). CasNano has the potential to revolutionize gene therapy and gene editing, enabling novel therapeutic strategies for a wide range of genetic disorders.

[0087] Accordingly, the present disclosure is directed to novel proteins for the use of DNA targeting, editing and / or binding with or without additional effector proteins. In the field of clustered regularly interspaced palindromic repeats (CRISPR) and CRISPR- associated (Cas) discovery, prokaryotic CRISPR-Cas systems have been leveraged for uses as genome engineering tools. The novel proteins described below are related to the BsCas12f1, which is a CRISPR Type V-F endonuclease. Due to the compact size of this Cas protein, it -12- 99975521.7Attorney Docket No: 124540-829360 can be packaged in delivery vehicles such as adeno-associated viruses (AAVs) that are known for their small cargo carrying capacity and delivery to cells.

[0088] The Cas12f polypeptides provided herein can be used directly for its nuclease activity, or DNA cleavage. Additionally, the Cas12f polypeptides further comprise other functional domains to modify gene expression or a target nucleic acid. In other words, the novel Cas12f polypeptides provided herein can be used alone or with other protein effectors to edit DNA (e.g., nucleotide base editing), or transcriptionally regulate target genes either through epigenetic modifications (e.g., methylation / demethylation) or transcriptional modulators (e.g., activators or repressors).

[0089] Also provided herein are improved guide RNAs designed to improve the targeting and editing efficiency of the Cas12f polypeptides herein. Nucleic acids and vectors encoding these improved constructs are also provided along with methods of using these Cas12f polypeptides and gRNAs to modify target nucleic acids for therapeutic purposes. I. Compositions

[0090] Various aspects of the present disclosure relate to improved, compact CRISPR associated (Cas) nucleases. These Cas nucleases are derived from Cas12f nucleases and are referred to herein as “Cas12f polypeptides”. The Cas12f polypeptides of the present disclosure generally comprise one or more amino acid substitutions relative to a wildtype Cas12f nuclease and have improved targeting and / or editing efficiency compared to the wildtype nuclease. Cas12f Polypeptides

[0091] The Cas12f Polypeptides as used herein can comprise an amino acid sequence having at least 10%, at least 15%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to a wild-type Cas12f nuclease, e.g., Cas12f from Blautia species, Clostridium novyi, Parageobacillus thermoglucosidasius, Oscillospiraceae bacterium, Syntrophomonas palmitatica, Oscillospiraceae bacterium, Clostridia bacterium, Eubacterium siraeum, Ruminiclostridium hungatei, Clostridium botulinum, Cellulosilyticum ruminicola, Clostridium hiranonis strain DSM 13275, Acidibacillus sulfuroxidans.

[0092] Advantageously, the Cas12f polypeptides described herein have improved targeting and / or editing efficiency compared to a wildtype Cas12f nuclease. Accordingly, in -13- 99975521.7Attorney Docket No: 124540-829360 various aspects, the Cas12f polypeptides further comprise one or more amino acid substitutions relative to a wildtype Cas12f nuclease. Exemplary wildtype Cas12f nucleases from various sources are provided in Table 1 below. Table 1 Wildtype Cas12f cies) Source S SEQ Nucleases (Spe equence ID NO: MITVRKVKLIVNSEEAEEINRTYK FIRDSMYAQYQGLNRCMGYLLS GYYANGMDIKSDGFKNHMKTIK NSLNIFDDINFGIGIDSKSAITQKV KKDFSTSLKNGLAKGERGATNY KRNFPLMTRGRDVKISYLEDTNT FVIKWVNKIEFKVILGQKDNIELS HTLHKIINKEYTLGQCTFEFDKN BsCas12f Blautia species NKLLLALNINIPDNLISKNKEIIPG RVLGVDLGVKVPAMICLNDNTFI 1 KKSIGSYNEFFKVRSQFKARRER LYKQLESSNGGKGRKHKLKATM QFRDKEKNFARTYNHFLSKNIIEF AQKYTCETINLEELNKKGFDNNL LGKWGYYQLQSMIEYKAERVGI KVKYVDPAFTSQTCSKCGYVDE ENRITQDKFECQKCGFTLNADHN AAINIARK MITVRKIKLTIMGDKDTRNSQYK 2 WIRDEQYNQYRALNMGMTYLA VNDILYMNESGLEIRTIKDLKDCE KDIDKNKKEIEKLTARLEKEQNK KNSSSEKLDEIKYKISLVENKIED YKLKIVELNKILEETQKERMDIQ KEFKEKYVDDLYQVLDKIPFKHL DNKSLVTQRIKADIKSDKSNGLL KGERSIRNYKRNFPLMTRGRDLK FKYDDNDDIEIKWMEGIKFKVIL CnCas12f1 Clostridium novyi GNRIKNSLELRHTLHKVIEGKYKI CDSSLQFDKNNNLILNLTLDIPIDI VNKKVSGRVVGVDLGLKIPAYC ALNDVEYIKKSIGRIDDFLKVRTQ MQSRRRRLQIAIQSAKGGKGRVN KLQALERFAEKEKNFAKTYNHFL SSNIVKFAVSNQAEQINMELLSL KETQNKSILRNWSYYQLQTMIEY KAQREGIKVKYIDPYHTSQTCSK CGNYEEGQRESQADFICKKCGY KVNADYNAARNIAMSNKYITKK EESKYYKIKESMV -14- 99975521.7Attorney Docket No: 124540-829360 MKYTKVMRYQIIKPLNAEWDEL 3 GMVLRDIQKETRAALNKTIQLC WEYQGFSADYKQIHGQYPKPKD VLGYTSMHGYAYDRLKNEFSKI ASSNLSQTIKRAVDKWNSDLKEI LRGDRSIPNFRKDCPIDIVKQSTKI QKCNDGYVLSLGLINREYKNELG RKNGVFDVLIKANDKTQQTILER Par IINGDYTYTASQIINHKNKWFINL Pt1Cas12f1 ageobacillus thermoglucosidasius TYQFETKETALDPNNVMGVALGI VYPVYIAFNNSLHRYHIKGGEIER FRRQVEKRKRELLNQGKYCGDG RKGHGYATRTKSIESISDKIARFR DTCNHKYSRFIVDMALKHNCGII QMEDLTGISKESTFLKNWTYYDL QQKIEYKAREAGIQVIKIEPQYTS QRCSKCGYIDKENRQEQATFKCI ECGFKTNAAYNAARNIAIPNIDKI IRKTLKMQ MGESVKAIKLKILDMFLDPECTK 4 QDDNWRKDLSTMSRFCAEAGN MCLRDLYNYFSMPKEDRISSKDL YNAMYHKTKLLHPELPGKVANQ IVNHAKDVWKRNAKLIYRNQIS MPTYKITTAPIRLQNNIYKLIKNK NKYIIDVQLYSKEYSKDSGKGTH RYFLVAVRDSSTRMIFDRIMSKD HIDSSKSYTQGQLQIKKDHQGK WYCIIPYTFPTHETVLDPDKVMG SpCas12f1 Syntrophomonas VALGVAKAVYWAFNSSYKRGCI palmitatica DGGEIEHFRKMIRARRVSIQNQIK HSGDARKGHGRKRALKPIETLSE KEKNFRDTINHRYANRIVEAAIK QGCGTIQIENLEGIADTTGSKFLK NWPYYDLQTKIVNKAKEHGITV VAINPQYTSQRCSMCGYIEKTNR SSQAVFECKQCGYGSRTICINCR HVQVSGDVCEECGGIVKKENVN AAYNAAKNISTPYIDQIIMEKCLE LGIPYRSITCKECGHIQASGNTCE VCGSTNILKPKKIRKAK MAKNTITKTLKLRIVRPYNSAEV 5 EKIVADEKNNREKIALEKNKDKV KEACSKHLKVAAYCTTQVERNA CLFCKARKLDDKFYQKLRGQFP UnCas12f1 Uncultured bacterium DAVFWQEISEIFRQLQKQAAEIY NQSLIELYYEIFIKGKGIANASSVE HYLSRVCYRRAAELFKNAAIASG LRSKIKSNFRLKELKNMKSGLPT TKSDNFPIPLVKQKGGQYTGFEIS -15- 99975521.7Attorney Docket No: 124540-829360 NHNSDFIIKIPFGRWQVKKEIDKY RPWEKFDFEQVQKSPKPISLLLST QRRKRNKGWSKDEGTEAEIKKV MNGDYQTSYIEVKRGSKIGEKSA WMLNLSIDVPKIDKGVDPSIIGGI AVGVRSPLVCAINNAFSRYSISDN DLFHFNKKMFARRRILLKKNRH KRAGHGAKNKLKPITILTEKSERF RKKLIERWACEIADFFIKNKVGT VQMENLESMKRKEDSYFNIRLR GFWPYAEMQNKIEFKLKQYGIEI RKVAPNNTSKTCSKCGHLNNYF NFEYRKKNKFPHFKCEKCNFKEN AAYNAALNISNPKLKSTKERP MEVQKTVMKTLSLRILRPLYSQE 6 IEKEIKEEKERRKQAGGTGELDG GFYKKLEKKHSEMFSFDRLNLLL NQLQREIAKVYNHAISELYIATIA QGNKSNKHYISSIVYNRAYGYFY NAYIALGICSKVEANFRSNELLTQ QSALPTAKSDNFPIVLHKQKGAE GEDGGFRISTEGSDLIFEIPIPFYE YNGENRKEPYKWVKKGGQKPV LKLILSTFRRQRNKGWAKDEGTD Un2Cas12f1 Uncultured bacteriumAEIRKVTEGKYQVSQIEINRGKK LGEHQKWFANFSIEQPIYERKPN RSIVGGLAVGIRSPLVCAINNSFS RYSVDSNDVFKFSKQVFAFRRRL LSKNSLKRKGHGAAHKLEPITEM TEKNDKFRKKIIERWAKEVTNFF VKNQVGIVQIEDLSTMKDREDHF FNQYLRGFWPYYQMQTLIENKL KEYGIEVKRVQAKYTSQLCSNPN CRYWNNYFNFEYRKVNKFPKFK CEKCNLEISAAYNAARNLSTPDIE KFVAKATKGINLPEK Ob3Cas12f1 OscillospiraceaeMGKGEISKVMKYELRYLDGSGS 7 bacterium FEEMQQRVWALQRKTREIQNRT VQIAFHWDYINREHFIQTGNNLN VLQETGYKRLDGYIYDRLKGQS AEMSGANLNATIQTAWKKYNSA KPKVLSGTMSVPSFKRDQPLIINS NCVKFSRSESECLAELTLFSREYK KEHDLSSNVRFAIRLHDSTQRSIL ERVLSGAYRKGQCQLVYQRPKW FLFLTYSFFPMQHDLDPEKYLGV DLGECCALYASSVGEYGSLKLEG GEITAFAKQLEARKRSMQKQAA YCGEGRIGHGTKTRVADVYKME NRIANFRDTVNHRYSKALIDYAV -16- 99975521.7Attorney Docket No: 124540-829360 KHQYGTIQMEDLSGIKNDTGFPK FLRHWTYFDLQEKIDAKAREHGI HVVKVNPQYTSQRCSKCGSIDSR NRKSQKEFCCLNCGYKVNADFN ASQNLSIKGIDVIIQKYIGAKSKQ TENNG* Cb1Cas12f1 Clostridia bacterium MAKGTVTKVMKYELRYLSGFSD8 FHAMQQAVWGLQRQSREILNKT IQMAFHWDYISRENFNANGVYL DVKAETGYKTYDGYIYNSLKSA YADMAAANLNAAIQKAWKKYK DAKMEVLRGTMSTPSYRSDQPV LINKNCVKLFDGGVRLTLFSDRF KRENNLNGNLEFAVQLHDGTQR SIFANLLNGTYALGQCQLVYDKR KWFLLVTYIFTPEKHELDPEKILG VDLGQTYALYASSVCARGTFRIE GGEAAECAHRLEQRKRSLQQQA RFCGEGRVGHGTKTRVAAVYSA GDKIASYRDSINHRYSKALVEYA VKNGYGTIQMEDLTGIQNDLDHP KRLQHWTYYDLQTKIENKAKEH GVGVVKVNPRYTSQRCSRCGHIE RENRPTQKVFCCKACGFEGNAD YNASQNLSMRNIDKIIEKELSAK GE* EsCas12f1 Eubacterium siraeum MVCNKVVKIALICDQIDKDGKD9 VNYNDIYKLLWDLQKQTREAKN KVIRLCWEWSGYSSEYFKTHEEY PKDKEILGISLRSYLYNRIKGDYN LYSGNLSQSAKIAYIEYKNSLTD VLRGDKSIINYRENQPLDIKNKAI QLLYENDNFFVRVALINKDKRKE LNFKDCSVRFKLLVKDDSTRTIL ERCFDEVYTITASKIMYNKKKKQ WYINLGYKFTKEIDKTLDKDRIL GVDLGVINPLVASVYGSYDRLIIG GGEIDKFRKRVEANKVQMLKQG KYCGDGRIGHGVNTRNKPAYNIE DKISRFRDTVNHKYSKAVVDYA VKNNCGTIQMEDLKGITQNKNE RYLKNWTYFDLQTKIEYKAKAL GIEVKYKNPKYTSQRCSKCGHIA EENRPEQKTFKCVKCGFKVNAD YNASQNLAIKDIDKIIEQYYNKG* RhgCas12f1 RuminiclostridiumMATKVMRYQIIKPIDCNWDLFG 10 hungatei KVLRDIQYDTRQIMNRTIQYCWE WQGYSSDYKIAKGEYPKTRETFG YSDMRGYAYDKLKSIYQRLNTA NLTTSITRAVQRWKTDTKDVIRG -17- 99975521.7Attorney Docket No: 124540-829360 DKSIACFRADVPIDLHNKSMNIE KSDDGYIVALSLASNIYKKELDR NSGQFSVLINEGNKSNRDVLDRC IAGQYKISASQILREKNKWFLNLS YSFEISKPDKSRDNILGIDVGIVHP VYMAVYNSPARRSISGGEIDNFR KQVQKRIKELQLQGKQCGEGRIG HGIKTRVKPIEFAKDKVANFRNTI NHKYSKAIVEFAIKNGCGIIQME DLKGINTDNVFLKNWTYYDLQQ KVKYKAELEGIEVKLIDPQYTSQ RCCKCGYIHRDNRPEQAKFKCID CGFEVNADYNASLNIATPDIDKII LEFLKCET* Cb3Cas12f1 ClostridiumMNTVRKIKIIINNENNELRKEQYK 11 botulinum FIRDSQYAQYQGLNRCMGYLMS GFYVNNMDIKSEEFKTWQKGVT NSANFFQEISFGKGIDSKSSITQK VKKDFSIALKNGLAKGERNINNY KRIAPLMTRGRNLKFKYDDNEL DILINWVNKIQFKCVLGEHKNSL ELQHTLHKVINNEYKIGQSSLYF NKKNELILILTIDIPTAKSSYEPIK DRILGVDLGMAVPVYMSINDNS YIKKSLGSYSEFAKVRKQFKERR NRLYKQLEACKGGRGRKDKLKA MNQFKEKEKNFAKTYNHFLSKNI VEFALKNKCEFIHLEKIESKGLEN SVLANWTYYDLQEKIIYKAKREG IGIKFVNSSYTSQTCSKCNYVDKE NRKTQAKFICKNCGFKANADYN ASQNISKSKEFIK* Pt2Cas12f1 ParageobacillusMIAVKKLKLTIVEEEEKRKEQYK 12 thermoglucosidasius FIRDSQYAQYQGLNLAMGILTSA YLVSGRDIKSDLFKDSQKSLTNS NEIFNGINFGKGIDTKSSITQKVK KDFSTSLKNGLAKGERGFTNYKR DFPLMTRGRDLKFYEEDKEFYIK WVNKIVFKILIGRKDKNKVELIH TLNKVLNKEYKVSQSSLQFDKN NKLILNLTIDIPYKKVDEIVKDRV CGVDMGIAIPIYVALNDVSYVRE GMGTIDEFMKQRLQFQSRRRRL QQQLKNVNGGKGRKDKLKGLES LREKEKSWVKTYNHALSKRVVE FAKKNKCEYIHLEKLTKDGFGDR LLRNWSYYELQEMIKYKADRVG IKVKHVNPAYTSQTCSECGHADK ENRETQAKFKCLECGFEANADY NAARNIAKSDKFVK* -18- 99975521.7Attorney Docket No: 124540-829360 CrCas12f1 Cellulosilyticum MIAVRKLKIMVLCDDESKKNEQ 13 ruminicola YKFLRDSQYAQYLGLNRAMSFL AKEYLSGDKERFKEAKKKLTNT CECYQNINFGTGIDSKSQITQKVK KDLQADIKNGLARGERSIRNYRR TFPLITRGRDLKFSYNGDEIIIKW VNKIYFKVLIGRKDKNYLELMHT LEKIINGEYKVCTSSIQIDKKLILN LTLEIPDKVKKEFQENRVLGVDL GIKFPAYACVSDNTYVRRSFGSID EFLKVRIQFDKRRKRIQQQLQNV KGGKGRKDKLQALDRMRDCER KWVRNYNHALSKRIIDFAFRNKC GIIHLEKLEKDGFKNKLLRNWSY YELQDMIGYKAEREGIVVKYVEP AYTSQTCSKCGYVDRENRPSQEH FLCKECGFEINADHNAAINIARSN KVIVDK* ChCas12f1 Clostridium MITVRKLKLTIINDDETKRNEQY 14 hiranonis strain DSM KFIRDSQYAQYQGLNLAMSVLT 13275 NAYLSSNRDIKSDLFKETQKNLK NSSHIFDDITFGKGTDNKSLINQK VKKDFNSAIKNGLARGERNITNY KRTFPLMTRGTALKFSYKDDCSD EIIIKWVNKIVFKVVIGRKDKNYL ELMHTLNKVINGEYKVGQSSIYF DKSNKLILNLTLYIPEKKDDDAIN GRTLGVDLGIKYPAYVCLNDDTF IRQHIGESLELSKQREQFRNRRKR LQQQLKNVKGGKGREKKLAALD KVAVCERNFVKTYNHTISKRIIDF AKKNKCEFINLEQLTKDGFDNIIL SNWSYYELQNMIKYKADREGIK VRYVNPAYTSQKCSKCGYIDKE NRPTQEKFKCIKCGFELNADHNA AINISRLEE* AsCas12f1 Acidibacillus MIKVYRYEIVKPLDLDWKEFGTI 15 sulfuroxidans LRQLQQETRFALNKATQLAWEW MGFSSDYKDNHGEYPKSKDILG YTNVHGYAYHTIKTKAYRLNSG NLSQTIKRATDRFKAYQKEILRG DMSIPSYKRDIPLDLIKENISVNR MNHGDYIASLSLLSNPAKQEMN VKRKISVIIIVRGAGKTIMDRILSG EYQVSASQIIHDDRKNKWYLNIS YDFEPQTRVLDLNKIMGIDLGVA VAVYMAFQHTPARYKLEGGEIE NFRRQVESRRISMLRQGKYAGG ARGGHGRDKRIKPIEQLRDKIAN FRDTTNHRYSRYIVDMAIKEGCG -19- 99975521.7Attorney Docket No: 124540-829360 TIQMEDLTNIRDIGSRFLQNWTY YDLQQKIIYKAEEAGIKVIKIDPQ YTSQRCSECGNIDSGNRIGQAIFK CRACGYEANADYNAARNIAIPNI DKIIAESIK

[0093] In various aspects, the improved Cas12f polypeptide provided herein has one or more amino acid substitutions relative to a Cas12f nuclease derived from the Blautia species (e.g., SEQ ID NO: 1). Exemplary substitutions are described herein below in reference to SEQ ID NO: 1, but equivalent substitutions in Cas12f nucleases from different sources may be contemplated.

[0094] In view of the foregoing, in various aspects, the Cas12f polypeptide may comprise one or more substitutions at any one or more of the following residues: I10, Y48, A50, T66, D76, S98, T99, E109, I110, G111, N114, F119, N103, M122R, S131, W143, I159, G176, E181, D183, K184, N195, I196, I201, P209, K220, A223, G266, A276, Y306, T307, N312, K318, A355, and D353, according to the numbering of SEQ ID NO: 1. In various aspects, the Cas12f polypeptide may comprise one or more substitutions at any one or more of the following residues: I10, Y48, A50, T66, S98, T99, E109, N114, F119, M122R, S131, W143, G176, E181, D183, K184, N195, I196, K220, A223, G266, A276, Y306, T307, N312, K318, and D353 according to the numbering of SEQ ID NO: 1. In some aspects, the BsCas12f polypeptide may comprise one or more substitutions at any one or more of T66, E109, S131, D183, K184, N195, K220, A276, Y306, K318 or D353 in SEQ ID NO: 1. In some aspects, the BsCas12f polypeptide may comprise one or more substitutions at any one or more of E109, S131, E181, D183, K184, N195, K220, Y306, K318, and D353 in SEQ ID NO: 1 In some aspects, the BsCas12f polypeptide may comprise one or more substitutions at any one or more of N195, Y306, T307, S131 E181, and D183 in SEQ ID NO: 1. In some aspects, the BsCas12f polypeptide may comprise one or more substitutions at T307, N312 or W143 in SEQ ID NO: 1.

[0095] While any substitutions are contemplated at one or more of the residues described above, certain substitutions are specifically contemplated. For instance, in some embodiments, the one or more amino acid substitutions may be selected from I10R, I10H, I10A, I10C, I10D, I10E, I10F, I10G, I10K, I10L, I10M, I10N, I10P, I10Q, I10T, I10V, I10W, I10Y, I10S, Y48F, Y48R, A50R, A50Y, T66Q, T66F, T66F, T66L, T66I, T66M, T66V, T66S, T66P, T66A, T66Y, T66H, T66N, T66K, T66D, T66E, T66C, T66W, T66R, -20- 99975521.7Attorney Docket No: 124540-829360 T66G, D76R, D76G, S98H, S98R, T99R, N103R, G104R, G104M, E109M, E109R, E109A, E109C, E109D, E109F, E109G, E109H, E109I, E109K, E109L, E109N, E109P, E109Q, E109S, E109T, E109V, E109W, E109Y, R110W , G111R, G111H, A112V, N114Q, N114A, F119Y, F119H, M122R, R124M, S131R, S131H, S131A, S131C, S131D, S131E, S131F, S131G, S131I, S131K, S131L, S131M, S131N, S131P, S131Q, S131T, S131V, S131W, S131Y, W143T, I159A, I159R, L175H, G176R, G176H, E181A, E181C, E181D, E181F, E181G, E181H, E181I, E181K, E181L, E181M, E181N, E181P, E181Q, E181R, E181S, E181T, E181V, E181W, E181Y, D183K, D183A, D183C, D183E, D183F, D183G, D183H, D183I, D183L, D183M, D183N, D183P, D183Q, D183R, D183S, D183T, D183V, D183W, D183Y, K184R, K184A, K184C, K184D, K184E, K184F, K184G, K184H, K184I, K184L, K184M, K184N, K184P, K184Q, K184S, K184T, K184V, K184W, K184Y, S162G, N195Q, N195R, N195A, N195C, N195D, N195E, N195F, N195G, N195H, N195I, N195K, N195L, N195M, N195P, N195S, N195T, N195V, N195W, N195Y, I196A, I196R, I201A, I201C, I201D, I201E, I201F, I201G, I201H, I201K, I201L, I201M, I201N, I201P, I201Q, I201R, I201S, I201T, I201V, I201W, I201Y, P209A, P209C, P209D, P209E, P209F, P209G, P209H, P209I, P209K, P209L, P209M, P209N, P209Q, P209R, P209S, P209T, P209V, P209W, P209Y, V219A, K220R, K220A, K220C, K220D, K220E, K220F, K220G, K220H, K220I, K220L, K220M, K220N, K220P, K220Q, K220S, K220T, S263Y, K220V, K220W, K220Y, A223R, R247H, Q249K, Y258H, G266R, A276R, A276M, A276F, A276L, A276I, A276V, A276S, A276P, A276T, A276Y, A276H, A276Q, A276N, A276K, A276D, A276E, A276C, A276W, A276G, Y306F, Y306M, Y306A, Y306C, Y306D, Y306E, Y306G, Y306H, Y306I, Y306K, Y306L, Y306N, Y306P, Y306Q, Y306R, Y306S, Y306T, Y306V, Y306W, T307K, T307N, T307G, N312V, N312Y, N317D, K318R, K318G, K318A, K318C, K318D, K318E, K318F, K318H, K318I, K318L, K318M, K318N, K318P, K318Q, K318S, K318T, K318V, K318W, K318Y, D322Y, D353R, D353A, D353C, D353E, D353F, D353G, D353H, D353I, D353K, D353L, D353M, D353N, D353P, D353Q, D353S, D353T, D353V, D353W, D353Y, A355K, A355T, A355R, D368Y, and A389V. In an aspect, the one or more amino acid substitutions may be selected from: Y48F, Y48R, A50R, A50Y, T66F, T66Q, S98H, S98R, T99R, E109M, E109R, N114Q, N114A, F119Y, F119H, M122R, S131H, S131R, G176R, G176H, D183K, K184R, N195Q, N195R, K220R, A223R, G266R, A276R, Y306F, Y306M, K318R, K318G, D353R, E109A, E109C, E109D, E109F, E109G, E109H, E109I, E109K, E109L, E109N, E109P, E109Q, E109S, E109T, E109V, E109W, E109Y, E181A, E181C, E181D, E181F, E181G, E181H, E181I, E181K, E181L, E181M, E181N, E181P, E181Q, E181R, E181S, E181T, E181V, E181W, E181Y, K184A, K184C, K184D, K184E, K184F, -21- 99975521.7Attorney Docket No: 124540-829360 K184G, K184H, K184I, K184L, K184M, K184N, K184P, K184Q, K184S, K184T, K184V, K184W, K184Y, N195A, N195C, N195D, N195E, N195F, N195G, N195H, N195I, N195K, N195L, N195M, N195P, N195S, N195T, N195V, N195W, N195Y, K318A, K318C, K318D, K318E, K318F, K318H, K318I, K318L, K318M, K318N, K318P, K318Q, K318S, K318T, K318V, K318W, K318Y, D353A, D353C, D353E, D353F, D353G, D353H, D353I, D353K, D353L, D353M, D353N, D353P, D353Q, D353S, D353T, D353V, D353W, D353Y, S131A, S131C, S131D, S131E, S131F, S131G, S131I, S131K, S131L, S131M, S131N, S131P, S131Q, S131T, S131V, S131W, S131Y, D183A, D183C, D183E, D183F, D183G, D183H, D183I, D183L, D183M, D183N, D183P, D183Q, D183R, D183S, D183T, D183V, D183W, D183Y, K220A, K220C, K220D, K220E, K220F, K220G, K220H, K220I, K220L, K220M, K220N, K220P, K220Q, K220S, K220T, K220V, K220W, K220Y, Y306A, Y306C, Y306D, Y306E, Y306G, Y306H, Y306I, Y306K, Y306L, Y306N, Y306P, Y306Q, Y306R, Y306S, Y306T, Y306V, Y306W, T66F, T66L, T66I, T66M, T66V, T66S, T66P, T66A, T66Y, T66H, T66N, T66K, T66D, T66E, T66C, T66W, T66R, T66G, A276F, A276L, A276I, A276M, A276V, A276S, A276P, A276T, A276Y, A276H, A276Q, A276N, A276K, A276D, A276E, A276C, A276W, A276G, T307G, T307K, T307N, N312V, N312Y, and W143T. For instance, in some embodiments, the one or more amino acid substitutions may be selected from Y48F, Y48R, A50R, A50Y, T66Q,T66F, S98H, S98R, T99R, E109M, E109R, N114Q, N114A, F119Y, F119H, M122R, S131R, S131H, W143T, G176R, G176H, E181G, D183K, K184R, N195Q, N195R, K220R, A223R, G266R, A276R, A276M, Y306F, Y306M, T307K, T307N, N312V, N312Y, K318R, K318G, and D353R. Exemplary Cas12f polypeptides comprising at least one substitution relative to SEQ ID NO: 1 (WT) are provided as SEQ ID NOs: 16-266.

[0096] In some aspects, the Cas12f polypeptide may comprise substitutions at two or more residues selected from Y48, A50, T66, D76, S98, T99, E109, I110, G111, N114, F119, N103, M122R, S131, W143, I159, G176, E181, D183, K184, N195, I201, P209, K220, A223, G266, A276, Y306, T307, N312, K318, A355, and D353. In some aspects the Cas12f polypeptide may comprise substitutions at two or more residues selected from Y48, A50, T66, S98, T99, E109, N114, F119, M122R, S131, W143, G176, E181, D183, K184, N195, K220, A223, G266, A276, Y306, T307, N312, K318, and D353. While any substitutions are contemplated at these two or more residues, certain substitutions and combinations of substitutions are specifically contemplated. For instance, in an embodiment, at least one of the two or more amino acid substitutions may comprise N195K or N195H. In an aspect, the -22- 99975521.7Attorney Docket No: 124540-829360 two or more amino acid substitutions may comprise N195K and a second amino acid substitution selected from: E109H, S131Q, S131C, E181Q, D183T, Y306K, T307K, and D353K. In an aspect, the two or more amino acid substitutions may comprise N195H and a second amino acid substitution selected from: E109H, S131Q, S131C, E181Q, D183T, Y306K, T307K, and D353K. In an aspect, the two or more amino acid substitutions may comprise N195K and Y306K. In an aspect, the two or more amino acid substitutions may comprise N195H and Y306K. In an aspect, the two or more amino acid substitutions may comprise N195K and D353K. In an aspect, the two or more amino acid substitutions may comprise N195H and D353K. In an aspect, the two or more amino acid substitutions may comprise any one of the following combinations: N195K and E109H, N195K and E181Q, N195K and D353K, N195K and S131Q, N195K and Y306K, N195K and T307K, N195H and E109H, N195H and E181Q, N195H and D353K, N195H and S131Q, N195H and Y306K, N195H and T307K, T66F and A112V, E109I and S162G, E181N and L175H, E181S and Q249K, E181Y and R247H, K184G and A355T, K184L and D322Y, K184P and V219A, K184P and Y258H, K184V and A112V, K318M and S263Y, D353A and R124M, S131G and A389V, S131G and R110W, D183Y and D368Y, N195K and D183T, N195H and D183T, N195K and S131C, or N195H and S131C. Exemplary Cas12f polypeptides comprising at least two substitutions relative to SEQ ID NO: 1 (WT) are provided as SEQ ID NOs: 267-297.

[0097] In some aspects, the Cas12f polypeptide may comprise substitutions at three or more residues selected from I10, Y48, A50, T66, D76, S98, T99, E109, I110, G111, N114, F119, N103, M122R, S131, W143, I159, G176, E181, D183, K184, N195, I201, P209, K220, A223, G266, A276, Y306, T307, N312, K318, A355, and D353. In some aspects, the Cas12f polypeptide may comprise substitutions at three or more residues selected from I10, Y48, A50, T66, S98, T99, E109, N114, F119, M122R, S131, W143, G176, E181, D183, K184, N195, I196, K220, A223, G266, A276, Y306, T307, N312, K318, and D353. While any substitutions are contemplated at these three or more residues, certain substitutions and combinations of substitutions are specifically contemplated. For instance, in an embodiment, at least one of the three or more amino acid substitutions may be selected from N195K, N195H, D353K, Y306K. In an aspect, at least two of the three or more amino acid substitutions may be N195K and Y306K. In an aspect, at least two of the three or more amino acid substitutions may be N195H and Y306K. In an aspect, at least two of the three or more amino acid substitutions may be N195K and D353K. In an aspect, at least two of the three or -23- 99975521.7Attorney Docket No: 124540-829360 more amino acid substitutions may be N195H and D353K. In an aspect, the three or more amino acid substitutions may comprise N195K and Y306K and a third amino acid substitution selected from: E109H, S131Q, S131C, E181Q, E183T, T307K, and D353K. In an aspect, the three or more amino acid substitutions may comprise N195H and Y306K and a third amino acid substitution selected from: E109H, S131Q, S131C, E181Q, E183T, T307K, and D353K. In an aspect, the three or more amino acid substitutions may comprise N195K and D353K and a third amino acid substitution selected from: E109H, S131Q, S131C, E181Q, E183T, T307K, and D353K. In an aspect, the three or more amino acid substitutions may comprise N195H and D353K and a third amino acid substitution selected from: E109H, S131Q, S131C, E181Q, E183T, T307K, and D353K. In some aspects, the three or more amino acid substitutions may comprise N195K, Y306K and T307K. In an aspect, the three or more amino acid substitutions may be selected from any of the following combinations: N195K, Y306K and D353K; E109H, N195H and Y306K; S131C, N195H and Y306K; S131Q, N195H and Y306K; E181Q, N195H and Y306K; D183T, N195H and Y306K; N195H, Y306K and T307K; N195H, Y306K and D353K; E109H, N195K and D353K; S131C, N195K and D353K; S131Q, N195K and D353K; E181Q, N195K and D353K; D183T, N195K and D353K; N195K, T307K and D353K; E109H, N195H and D353K; S131C, N195H and D353K; S131Q, N195H and D353K; E181Q, N195H and D353K; D183T, N195H and D353K; N195H, T307K and D353K; or N195H, S131Q and A342E. Exemplary Cas12f polypeptides comprising at least three substitutions relative to SEQ ID NO: 1 (WT) are provided as SEQ ID NOs: 298-327.

[0098] In some aspects, the Cas12f polypeptide may comprise substitutions at four or more residues selected from I10, Y48, A50, T66, D76, S98, T99, E109, I110, G111, N114, F119, N103, M122R, S131, W143, I159, G176, E181, D183, K184, N195, I196, I201, P209, K220, A223, G266, A276, Y306, T307, N312, K318, A355, and D353. In some aspects, the Cas12f polypeptide may comprise substitutions at four or more residues selected from Y48, A50, T66, S98, T99, E109, N114, F119, M122R, S131, W143, G176, E181, D183, K184, N195, K220, A223, G266, A276, Y306, T307, N312, K318, and D353. While any substitutions are contemplated at these four or more residues, certain substitutions and combinations of substitutions are specifically contemplated. For instance, in an embodiment, at least one of the four or more amino acid substitutions may be selected from N195K, Y306K and T307K. In an aspect, the four or more amino acid substitutions comprise N195K, Y306K and T307K and a fourth amino acid substitution selected from: E109H, S131C, -24- 99975521.7Attorney Docket No: 124540-829360 S131Q, E181Q, D183T and D353K. In an aspect, the four or more amino acid substitutions comprise N195K, Y306K, T307K and E109H. In an aspect, the four or more amino acid substitutions comprise N195K, Y306K, T307K and S131C. In an aspect, the four or more amino acid substitutions comprise N195K, Y306K, T307K and S131Q. In an aspect, the four or more amino acid substitutions comprise N195K, Y306K, T307K and E181Q. In an aspect, the four or more amino acid substitutions comprise N195K, Y306K, T307K and D183T. In an aspect, the four or more amino acid substitutions comprise N195K, Y306K, T307K and D353K. In an aspect, the four or more amino acid substitutions may be selected from any of the following combinations: E109M, E181G, N195R and D353R; N195K, Y306K, T307K and E109H; N195K, Y306K, T307K and S131C; N195K, Y306K, T307K and S131Q; N195K, Y306K, T307K and E181Q; N195K, Y306K, T307K and D183T; or N195K, Y306K, T307K and D353K. Exemplary Cas12f polypeptides comprising at least four substitutions relative to SEQ ID NO: 1 (WT) are provided as SEQ ID NOs: 328-334.

[0099] In some aspects, the Cas12f polypeptide may comprise substitutions at five or more residues selected from I10, Y48, A50, T66, D76, S98, T99, E109, I110, G111, N114, F119, N103, M122R, S131, W143, I159, G176, E181, D183, K184, N195, I196, I201, P209, K220, A223, G266, A276, Y306, T307, N312, K318, A355, and D353. In some aspects, the Cas12f polypeptide may comprise substitutions at five or more residues selected from I10, Y48, A50, T66, S98, T99, E109, N114, F119, M122R, S131, W143, G176, E181, D183, K184, N195, I196, K220, A223, G266, A276, Y306, T307, N312, K318, and D353. While any substitutions are contemplated at these five or more residues, certain substitutions and combinations of substitutions are specifically contemplated. For instance, in an embodiment, at least one of the five or more residues may be selected from N195K, Y306K, T307K, S131C, E181Q, D183T. In an aspect, the five or more substitutions may comprise N195K, Y306K, T307K and two substitutions selected from S131C, E181Q, D183T. In an aspect, the five or more substitutions may comprise N195K, Y306K, T307K, S131C, and E181Q. In an aspect, the five or more substitutions may comprise N195K, Y306K, T307K, S131C, and D183T. In an aspect, the five or more substitutions may comprise N195K, Y306K, T307K, E181Q, and D183T. Exemplary Cas12f polypeptides comprising at least five substitutions relative to SEQ ID NO: 1 (WT) are provided as SEQ ID NOs: 335-337.

[0100] In some aspects, the Cas12f polypeptide may comprise substitutions at six or more residues selected I10, Y48, A50, T66, D76, S98, T99, E109, I110, G111, N114, F119, N103, M122R, S131, W143, I159, G176, E181, D183, K184, N195, I196, I201, P209, K220, A223, -25- 99975521.7Attorney Docket No: 124540-829360 G266, A276, Y306, T307, N312, K318, A355, and D353. While any substitutions are contemplated at these six or more residues, certain substitutions and combinations of substitutions are specifically contemplated. For instance, in an embodiment, the six or more substitutions may comprise N195K, Y306K, T307K, S131C, E181Q, and D183T. An exemplary Cas12f polypeptides comprising at least six substitutions relative to SEQ ID NO: 1 (WT) are provided as SEQ ID NO: 338.

[0101] In some aspects, the Cas12f polypeptide may comprise substitutions at seven or more residues selected from I10, Y48, A50, T66, D76, S98, T99, E109, G111, N114, F119, M122R, S131, W143, G176, E181, D183, K184, N195, I196, K220, A223, G266, A276, Y306, T307, N312, K318, and D353. While any substitutions are contemplated at these seven or more residues, certain substitutions and combinations of substitutions are specifically contemplated. For instance, in an embodiment, the seven or more substitutions may comprise N195K, Y306K, T307K, S131C, E181Q, D183T and at least one substitution selected from I10R, I10H, D76R, D76G, N103R, G104R, G104M, G111R, G111H, I201A, I201C, I201D, I201E, I201F, I201G, I201H, I201K, I201L, I201M, I201N, I201P, I201Q, I201R, I201S, I201T, I201V, I201W, I201Y, P209A, P209C, P209D, P209E, P209F, P209G, P209H, P209I, P209K, P209L, P209M, P209N, P209Q, P209R, P209S, P209T, P209V, P209W, P209Y, I159A, I159R, I196A, I196R, A355K, and A355R. For instance, in an aspect, any of the following substitution combinations are contemplated: N195K, Y306K, T307K, S131C, E181Q, D183T and I10R; N195K, Y306K, T307K, S131C, E181Q, D183T and I10H; N195K, Y306K, T307K, S131C, E181Q, D183T and D76R; N195K, Y306K, T307K, S131C, E181Q, D183T and D76G; N195K, Y306K, T307K, S131C, E181Q, D183T and N103R; N195K, Y306K, T307K, S131C, E181Q, D183T and G104R; N195K, Y306K, T307K, S131C, E181Q, D183T and G104M; N195K, Y306K, T307K, S131C, E181Q, D183T and G111R; N195K, Y306K, T307K, S131C, E181Q, D183T and G111H; N195K, Y306K, T307K, S131C, E181Q, D183T and I201A; N195K, Y306K, T307K, S131C, E181Q, D183T and I201C; N195K, Y306K, T307K, S131C, E181Q, D183T and I201D; N195K, Y306K, T307K, S131C, E181Q, D183T and I201E; N195K, Y306K, T307K, S131C, E181Q, D183T and I201F; N195K, Y306K, T307K, S131C, E181Q, D183T and I201G; N195K, Y306K, T307K, S131C, E181Q, D183T and I201H; N195K, Y306K, T307K, S131C, E181Q, D183T and I201K; N195K, Y306K, T307K, S131C, E181Q, D183T and I201L; N195K, Y306K, T307K, S131C, E181Q, D183T and I201M; N195K, Y306K, T307K, S131C, E181Q, D183T and I201N; N195K, Y306K, T307K, S131C, E181Q, D183T -26- 99975521.7Attorney Docket No: 124540-829360 and I201P; N195K, Y306K, T307K, S131C, E181Q, D183T and I201Q; N195K, Y306K, T307K, S131C, E181Q, D183T and I201R; N195K, Y306K, T307K, S131C, E181Q, D183T and I201S; N195K, Y306K, T307K, S131C, E181Q, D183T and I201T; N195K, Y306K, T307K, S131C, E181Q, D183T and I201V; N195K, Y306K, T307K, S131C, E181Q, D183T and I201W; N195K, Y306K, T307K, S131C, E181Q, D183T and I201Y; N195K, Y306K, T307K, S131C, E181Q, D183T and P209A; N195K, Y306K, T307K, S131C, E181Q, D183T and P209C; N195K, Y306K, T307K, S131C, E181Q, D183T and P209D; N195K, Y306K, T307K, S131C, E181Q, D183T and P209E; N195K, Y306K, T307K, S131C, E181Q, D183T and P209F; N195K, Y306K, T307K, S131C, E181Q, D183T and P209G; N195K, Y306K, T307K, S131C, E181Q, D183T and P209H; N195K, Y306K, T307K, S131C, E181Q, D183T and P209I; N195K, Y306K, T307K, S131C, E181Q, D183T and P209K; N195K, Y306K, T307K, S131C, E181Q, D183T and P209L; N195K, Y306K, T307K, S131C, E181Q, D183T and P209M; N195K, Y306K, T307K, S131C, E181Q, D183T and P209N; N195K, Y306K, T307K, S131C, E181Q, D183T and P209Q; N195K, Y306K, T307K, S131C, E181Q, D183T and P209R; N195K, Y306K, T307K, S131C, E181Q, D183T and P209S; N195K, Y306K, T307K, S131C, E181Q, D183T and P209T; N195K, Y306K, T307K, S131C, E181Q, D183T and P209V; N195K, Y306K, T307K, S131C, E181Q, D183T and P209W; N195K, Y306K, T307K, S131C, E181Q, D183T and P209Y; N195K, Y306K, T307K, S131C, E181Q, D183T and I159A; N195K, Y306K, T307K, S131C, E181Q, D183T and I159R; N195K, Y306K, T307K, S131C, E181Q, D183T and I196A; N195K, Y306K, T307K, S131C, E181Q, D183T and I196R; N195K, Y306K, T307K, S131C, E181Q, D183T and A355K; and N195K, Y306K, T307K, S131C, E181Q, D183T and A355R. Exemplary Cas12f polypeptides comprising at least seven substitutions relative to SEQ ID NO: 1 (WT) are provided as SEQ ID NOs: 339-391. In some aspects, the Cas12f polypeptide may comprise substitutions at seven or more residues selected from N195, Y306, T307, S131, E181, D183, and I10. In some aspects, the seven or more substitutions may comprise N195K, Y306K, T307K, S131C, E181Q, D183T, and I10R. An exemplary Cas12f polypeptides comprising at least seven substitutions relative to SEQ ID NO: 1 (WT) are provided as SEQ ID NO: 803.

[0102] In some aspects, the Cas12f polypeptide may comprise substitutions at eight or more residues selected from N195, Y306, T307, S131, E181, D183, I10, and D76. In some aspects, the eight or more substitutions may comprise N195K, Y306K, T307K, S131C, E181Q, D183T, I10R, and D76K. An exemplary Cas12f polypeptides comprising at least eight substitutions relative to SEQ ID NO: 1 (WT) are provided as SEQ ID NO: 804. -27- 99975521.7Attorney Docket No: 124540-829360

[0103] In some aspects, the Cas12f polypeptide may comprise substitutions at nine or more residues selected from N195, Y306, T307, S131, E181, D183, I10, D76, and G111. In some aspects, the nine or more substitutions may comprise N195K, Y306K, T307K, S131C, E181Q, D183T, I10R, D76K, and G111N. An exemplary Cas12f polypeptides comprising at least nine substitutions relative to SEQ ID NO: 1 (WT) are provided as SEQ ID NO: 805.

[0104] In some aspects, the Cas12f polypeptide may comprise substitutions at ten or more residues selected from N195, Y306, T307, S131, E181, D183, I10, D76, G111, and I159. In some aspects, the ten or more substitutions may comprise N195K, Y306K, T307K, S131C, E181Q, D183T, I10R, D76K, G111N, and I159R. An exemplary Cas12f polypeptides comprising at least ten substitutions relative to SEQ ID NO: 1 (WT) are provided as SEQ ID NO: 806.

[0105] In some aspects, the Cas12f polypeptide may comprise substitutions at eleven or more residues selected from N195, Y306, T307, S131, E181, D183, I10, D76, G111, I159, and I196. In some aspects, the eleven or more substitutions may comprise N195K, Y306K, T307K, S131C, E181Q, D183T, I10R, D76K, G111N, I159R, and I196V. An exemplary Cas12f polypeptides comprising at least eleven substitutions relative to SEQ ID NO: 1 (WT) are provided as SEQ ID NO: 807.

[0106] In view of the foregoing, various exemplary Cas12f polypeptides comprising one or more of the discussed substitutions above are provided herein as SEQ ID NOs: 16-391 and are discussed further in the Examples below.

[0107] In various aspects, accordingly, the Cas12f polypeptide may comprise or consist of an amino acid sequence having 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs: 16-391 and SEQ ID NOs: 803-807. In various aspects, the Cas12f polypeptide may comprise or consist of an amino acid sequence of any one of SEQ ID NOs: 329-338 and SEQ ID NOs: 803-807. In various aspects, the Cas12f polypeptide may comprise or consist of an amino acid sequence of SEQ ID NO: 338. For ease of reference, the sequences of illustrative Cas12f polypeptides comprising combinations of substitutions shown to have surprisingly improved targeting and / or editing efficiency are provided in Table 2 below. Table 2 – Illustrative Optimized Cas12f Polypeptides -28- 99975521.7Attorney Docket No: 124540-829360 Cas12f Substitutions Sequence SEQ Polypeptide relative to WT ID (SEQ ID NO: 1)* NO: dBsCas12f1 N195K, Y306K, MITVRKVKLIVNSEEAEEINRTYKFIRDSMYA 329 variant T307K, E109H QYQGLNRCMGYLLSGYYANGMDIKSDGFK (V4.1) NHMKTIKNSLNIFDDINFGIGIDSKSAITQKV KKDFSTSLKNGLAKGHRGATNYKRNFPLMT RGRDVKISYLEDTNTFVIKWVNKIEFKVILG QKDNIELSHTLHKIINKEYTLGQCTFEFDKNN KLLLALNIKIPDNLISKNKEIIPGRVLGVALGV KVPAMICLNDNTFIKKSIGSYNEFFKVRSQFK ARRERLYKQLESSNGGKGRKHKLKATMQFR DKEKNFARTYNHFLSKNIIEFAQKKKCETINL EELNKKGFDNNLLGKWGYYQLQSMIEYKAE RVGIKVKYVDPAFTSQTCSKCGYVDEENRIT QDKFECQKCGFTLNAAHNAAINIARK dBsCas12f1 N195K, Y306K, MITVRKVKLIVNSEEAEEINRTYKFIRDSMYA 330 variant T307K, S131C QYQGLNRCMGYLLSGYYANGMDIKSDGFK (V4.2) NHMKTIKNSLNIFDDINFGIGIDSKSAITQKV KKDFSTSLKNGLAKGERGATNYKRNFPLMT RGRDVKICYLEDTNTFVIKWVNKIEFKVILG QKDNIELSHTLHKIINKEYTLGQCTFEFDKNN KLLLALNIKIPDNLISKNKEIIPGRVLGVALGV KVPAMICLNDNTFIKKSIGSYNEFFKVRSQFK ARRERLYKQLESSNGGKGRKHKLKATMQFR DKEKNFARTYNHFLSKNIIEFAQKKKCETINL EELNKKGFDNNLLGKWGYYQLQSMIEYKAE RVGIKVKYVDPAFTSQTCSKCGYVDEENRIT QDKFECQKCGFTLNAAHNAAINIARK dBsCas12f1 N195K, Y306K, MITVRKVKLIVNSEEAEEINRTYKFIRDSMYA 331 variant T307K, S131Q QYQGLNRCMGYLLSGYYANGMDIKSDGFK (V4.3) NHMKTIKNSLNIFDDINFGIGIDSKSAITQKV KKDFSTSLKNGLAKGERGATNYKRNFPLMT RGRDVKIQYLEDTNTFVIKWVNKIEFKVILG QKDNIELSHTLHKIINKEYTLGQCTFEFDKNN KLLLALNIKIPDNLISKNKEIIPGRVLGVALGV KVPAMICLNDNTFIKKSIGSYNEFFKVRSQFK ARRERLYKQLESSNGGKGRKHKLKATMQFR DKEKNFARTYNHFLSKNIIEFAQKKKCETINL EELNKKGFDNNLLGKWGYYQLQSMIEYKAE RVGIKVKYVDPAFTSQTCSKCGYVDEENRIT QDKFECQKCGFTLNAAHNAAINIARK dBsCas12f1 N195K, Y306K, MITVRKVKLIVNSEEAEEINRTYKFIRDSMYA 332 variant T307K, E181Q QYQGLNRCMGYLLSGYYANGMDIKSDGFK (V4.4) NHMKTIKNSLNIFDDINFGIGIDSKSAITQKV KKDFSTSLKNGLAKGERGATNYKRNFPLMT RGRDVKISYLEDTNTFVIKWVNKIEFKVILG QKDNIELSHTLHKIINKEYTLGQCTFQFDKN -29- 99975521.7Attorney Docket No: 124540-829360 NKLLLALNIKIPDNLISKNKEIIPGRVLGVALG VKVPAMICLNDNTFIKKSIGSYNEFFKVRSQF KARRERLYKQLESSNGGKGRKHKLKATMQF RDKEKNFARTYNHFLSKNIIEFAQKKKCETIN LEELNKKGFDNNLLGKWGYYQLQSMIEYKA ERVGIKVKYVDPAFTSQTCSKCGYVDEENRI TQDKFECQKCGFTLNAAHNAAINIARK dBsCas12f1 N195K, Y306K, MITVRKVKLIVNSEEAEEINRTYKFIRDSMYA 333 variant T307K, D183T QYQGLNRCMGYLLSGYYANGMDIKSDGFK (V4.5) NHMKTIKNSLNIFDDINFGIGIDSKSAITQKV KKDFSTSLKNGLAKGERGATNYKRNFPLMT RGRDVKISYLEDTNTFVIKWVNKIEFKVILG QKDNIELSHTLHKIINKEYTLGQCTFEFTKNN KLLLALNIKIPDNLISKNKEIIPGRVLGVALGV KVPAMICLNDNTFIKKSIGSYNEFFKVRSQFK ARRERLYKQLESSNGGKGRKHKLKATMQFR DKEKNFARTYNHFLSKNIIEFAQKKKCETINL EELNKKGFDNNLLGKWGYYQLQSMIEYKAE RVGIKVKYVDPAFTSQTCSKCGYVDEENRIT QDKFECQKCGFTLNAAHNAAINIARK dBsCas12f1 N195K, Y306K, MITVRKVKLIVNSEEAEEINRTYKFIRDSMYA 334 variant T307K, D353K QYQGLNRCMGYLLSGYYANGMDIKSDGFK (V4.6) NHMKTIKNSLNIFDDINFGIGIDSKSAITQKV KKDFSTSLKNGLAKGERGATNYKRNFPLMT RGRDVKISYLEDTNTFVIKWVNKIEFKVILG QKDNIELSHTLHKIINKEYTLGQCTFEFDKNN KLLLALNIKIPDNLISKNKEIIPGRVLGVALGV KVPAMICLNDNTFIKKSIGSYNEFFKVRSQFK ARRERLYKQLESSNGGKGRKHKLKATMQFR DKEKNFARTYNHFLSKNIIEFAQKKKCETINL EELNKKGFDNNLLGKWGYYQLQSMIEYKAE RVGIKVKYVKPAFTSQTCSKCGYVDEENRIT QDKFECQKCGFTLNAAHNAAINIARK dBsCas12f1 N195K, Y306K, MITVRKVKLIVNSEEAEEINRTYKFIRDSMYA 335 variant T307K, S131C, QYQGLNRCMGYLLSGYYANGMDIKSDGFK (V5.1) E181Q NHMKTIKNSLNIFDDINFGIGIDSKSAITQKV KKDFSTSLKNGLAKGERGATNYKRNFPLMT RGRDVKICYLEDTNTFVIKWVNKIEFKVILG QKDNIELSHTLHKIINKEYTLGQCTFQFDKN NKLLLALNIKIPDNLISKNKEIIPGRVLGVALG VKVPAMICLNDNTFIKKSIGSYNEFFKVRSQF KARRERLYKQLESSNGGKGRKHKLKATMQF RDKEKNFARTYNHFLSKNIIEFAQKKKCETIN LEELNKKGFDNNLLGKWGYYQLQSMIEYKA ERVGIKVKYVDPAFTSQTCSKCGYVDEENRI TQDKFECQKCGFTLNAAHNAAINIARK -30- 99975521.7Attorney Docket No: 124540-829360 dBsCas12f1 N195K, Y306K, MITVRKVKLIVNSEEAEEINRTYKFIRDSMYA 336 variant T307K, S131C, QYQGLNRCMGYLLSGYYANGMDIKSDGFK (V5.2) D183T NHMKTIKNSLNIFDDINFGIGIDSKSAITQKV KKDFSTSLKNGLAKGERGATNYKRNFPLMT RGRDVKICYLEDTNTFVIKWVNKIEFKVILG QKDNIELSHTLHKIINKEYTLGQCTFEFTKNN KLLLALNIKIPDNLISKNKEIIPGRVLGVALGV KVPAMICLNDNTFIKKSIGSYNEFFKVRSQFK ARRERLYKQLESSNGGKGRKHKLKATMQFR DKEKNFARTYNHFLSKNIIEFAQKKKCETINL EELNKKGFDNNLLGKWGYYQLQSMIEYKAE RVGIKVKYVDPAFTSQTCSKCGYVDEENRIT QDKFECQKCGFTLNAAHNAAINIARK dBsCas12f1 N195K, Y306K, MITVRKVKLIVNSEEAEEINRTYKFIRDSMYA 337 variant T307K, E181Q, QYQGLNRCMGYLLSGYYANGMDIKSDGFK (V5.3) D183T NHMKTIKNSLNIFDDINFGIGIDSKSAITQKV KKDFSTSLKNGLAKGERGATNYKRNFPLMT RGRDVKISYLEDTNTFVIKWVNKIEFKVILG QKDNIELSHTLHKIINKEYTLGQCTFQFTKNN KLLLALNIKIPDNLISKNKEIIPGRVLGVALGV KVPAMICLNDNTFIKKSIGSYNEFFKVRSQFK ARRERLYKQLESSNGGKGRKHKLKATMQFR DKEKNFARTYNHFLSKNIIEFAQKKKCETINL EELNKKGFDNNLLGKWGYYQLQSMIEYKAE RVGIKVKYVDPAFTSQTCSKCGYVDEENRIT QDKFECQKCGFTLNAAHNAAINIARK dBsCas12f1 N195K, Y306K, MITVRKVKLIVNSEEAEEINRTYKFIRDSMYA 338 variant (V6) T307K, S131C, QYQGLNRCMGYLLSGYYANGMDIKSDGFK E181Q, D183T NHMKTIKNSLNIFDDINFGIGIDSKSAITQKV KKDFSTSLKNGLAKGERGATNYKRNFPLMT RGRDVKICYLEDTNTFVIKWVNKIEFKVILG QKDNIELSHTLHKIINKEYTLGQCTFQFTKNN KLLLALNIKIPDNLISKNKEIIPGRVLGVALGV KVPAMICLNDNTFIKKSIGSYNEFFKVRSQFK ARRERLYKQLESSNGGKGRKHKLKATMQFR DKEKNFARTYNHFLSKNIIEFAQKKKCETINL EELNKKGFDNNLLGKWGYYQLQSMIEYKAE RVGIKVKYVDPAFTSQTCSKCGYVDEENRIT QDKFECQKCGFTLNAAHNAAINIARK dBsCas12f1 N195K, Y306K, MITVRKVKLRVNSEEAEEINRTYKFIRDSMY 803 variant (V7) T307K, S131C, AQYQGLNRCMGYLLSGYYANGMDIKSDGF E181Q, D183T, KNHMKTIKNSLNIFDDINFGIGIDSKSAITQK I10R VKKDFSTSLKNGLAKGERGATNYKRNFPLM TRGRDVKICYLEDTNTFVIKWVNKIEFKVIL GQKDNIELSHTLHKIINKEYTLGQCTFQFTKN NKLLLALNIKIPDNLISKNKEIIPGRVLGVALG VKVPAMICLNDNTFIKKSIGSYNEFFKVRSQF KARRERLYKQLESSNGGKGRKHKLKATMQF RDKEKNFARTYNHFLSKNIIEFAQKKKCETIN -31- 99975521.7Attorney Docket No: 124540-829360 LEELNKKGFDNNLLGKWGYYQLQSMIEYKA ERVGIKVKYVDPAFTSQTCSKCGYVDEENRI TQDKFECQKCGFTLNAAHNAAINIARK dBsCas12f1 N195K, Y306K, MITVRKVKLRVNSEEAEEINRTYKFIRDSMY 804 variant (V8) T307K, S131C, AQYQGLNRCMGYLLSGYYANGMDIKSDGF E181Q, D183T, KNHMKTIKNSLNIFDKINFGIGIDSKSAITQK I10R, D76K VKKDFSTSLKNGLAKGERGATNYKRNFPLM TRGRDVKICYLEDTNTFVIKWVNKIEFKVIL GQKDNIELSHTLHKIINKEYTLGQCTFQFTKN NKLLLALNIKIPDNLISKNKEIIPGRVLGVALG VKVPAMICLNDNTFIKKSIGSYNEFFKVRSQF KARRERLYKQLESSNGGKGRKHKLKATMQF RDKEKNFARTYNHFLSKNIIEFAQKKKCETIN LEELNKKGFDNNLLGKWGYYQLQSMIEYKA ERVGIKVKYVDPAFTSQTCSKCGYVDEENRI TQDKFECQKCGFTLNAAHNAAINIARK dBsCas12f1 N195K, Y306K, MITVRKVKLRVNSEEAEEINRTYKFIRDSMY 805 variant (V9) T307K, S131C, AQYQGLNRCMGYLLSGYYANGMDIKSDGF E181Q, D183T, KNHMKTIKNSLNIFDKINFGIGIDSKSAITQK I10R, D76K, VKKDFSTSLKNGLAKGERNATNYKRNFPLM G111N TRGRDVKICYLEDTNTFVIKWVNKIEFKVIL GQKDNIELSHTLHKIINKEYTLGQCTFQFTKN NKLLLALNIKIPDNLISKNKEIIPGRVLGVALG VKVPAMICLNDNTFIKKSIGSYNEFFKVRSQF KARRERLYKQLESSNGGKGRKHKLKATMQF RDKEKNFARTYNHFLSKNIIEFAQKKKCETIN LEELNKKGFDNNLLGKWGYYQLQSMIEYKA ERVGIKVKYVDPAFTSQTCSKCGYVDEENRI TQDKFECQKCGFTLNAAHNAAINIARK dBsCas12f1 N195K, Y306K, MITVRKVKLRVNSEEAEEINRTYKFIRDSMY 806 variant T307K, S131C, AQYQGLNRCMGYLLSGYYANGMDIKSDGF (V10) E181Q, D183T, KNHMKTIKNSLNIFDKINFGIGIDSKSAITQK I10R, D76K, VKKDFSTSLKNGLAKGERNATNYKRNFPLM G111N, I159R TRGRDVKICYLEDTNTFVIKWVNKIEFKVIL GQKDNRELSHTLHKIINKEYTLGQCTFQFTK NNKLLLALNIKIPDNLISKNKEIIPGRVLGVAL GVKVPAMICLNDNTFIKKSIGSYNEFFKVRS QFKARRERLYKQLESSNGGKGRKHKLKATM QFRDKEKNFARTYNHFLSKNIIEFAQKKKCE TINLEELNKKGFDNNLLGKWGYYQLQSMIE YKAERVGIKVKYVDPAFTSQTCSKCGYVDE ENRITQDKFECQKCGFTLNAAHNAAINIARK dBsCas12f1 N195K, Y306K, MITVRKVKLRVNSEEAEEINRTYKFIRDSMY 807 variant T307K, S131C, AQYQGLNRCMGYLLSGYYANGMDIKSDGF (V11) E181Q, D183T, KNHMKTIKNSLNIFDKINFGIGIDSKSAITQK I10R, D76K, VKKDFSTSLKNGLAKGERNATNYKRNFPLM G111N, I159R, TRGRDVKICYLEDTNTFVIKWVNKIEFKVIL I196V GQKDNRELSHTLHKIINKEYTLGQCTFQFTK -32- 99975521.7Attorney Docket No: 124540-829360 NNKLLLALNIKVPDNLISKNKEIIPGRVLGVA LGVKVPAMICLNDNTFIKKSIGSYNEFFKVRS QFKARRERLYKQLESSNGGKGRKHKLKATM QFRDKEKNFARTYNHFLSKNIIEFAQKKKCE TINLEELNKKGFDNNLLGKWGYYQLQSMIE YKAERVGIKVKYVDPAFTSQTCSKCGYVDE ENRITQDKFECQKCGFTLNAAHNAAINIARK *All variants shown further comprise D216A D390A substitutions relative to SEQ ID NO: 1 as discussed below.

[0108] Also contemplated herein are variants of Cas12f polypeptides comprising deletions or insertions relative to wildtype. In an aspect, a variant having at least one deletion or insertion relative to wildtype may have a ∆G214-G218 deletion relative to SEQ ID NO: 1 (WT). This variant is provided herein as SEQ ID NO: 392.

[0109] In any of the foregoing or related aspects, the Cas12f polypeptides may be a functional nuclease and may, in some aspects, have improved or increased nuclease activity relative to a wildtype Cas12f nuclease. In other aspects, however, the Cas12f polypeptide may comprise one or more further substitutions that deactivate its native nuclease activity. In these aspects, the Cas12f polypeptide may have decreased nuclease activity relative to a wildtype Cas12f nuclease.

[0110] In various aspects, the Cas12f polypeptide may further comprise a D216A and / or a D390A substitution relative to SEQ ID NO: 1. For ease of reference, representative polypeptides comprising these substitutions are provided herein in Table 3 as SEQ ID NO: 393-407. These “deactivating” mutations may be combined with any combination of the aforementioned mutations as desired depending on whether nuclease activity is desired. Table 3 – Illustrative deactivated Cas12f polypeptides from various species SEQ dCas12f Polypeptide Substitutions ID NO: dBsCas12f1 D216A D390A relative to SEQ ID NO: 1 393 dCnCas12f1 D292A, D467A relative to SEQ ID NO: 2 394 dPt1Cas12f1 D226A, D401A relative to SEQ ID NO: 3 395 dSpCas12f1 D228A, D434A relative to SEQ ID NO: 4 396 dUnCas12f1 D326A, D510A relative to SEQ ID NO: 5 397 dUn2Cas2f1 D286A, D472A relative to SEQ ID NO: 6 398 -33- 99975521.7Attorney Docket No: 124540-829360 dOb3Cas12f1 D229A, D407A relative to SEQ ID NO: 7 399 dCb1Cas12f1 D225A, D403A relative to SEQ ID NO: 8 400 dEsCas12f1 D233A, D410A relative to SEQ ID NO: 9 401 dRhgCas12f1 D225A, D400A relative to SEQ ID NO: 10 402 dCb3Cas12f1 D215A, D389A relative to SEQ ID NO: 11 403 dPt2Cas12f1 D216A, D386A relative to SEQ ID NO: 12 404 dCrCas12f1 D207A, D381A relative to SEQ ID NO: 13 405 dChCas12f1 D216A, D390A relative to SEQ ID NO: 14 406 dAsCas12f1 D225A, D401A relative to SEQ ID NO: 15 407

[0111] In view of the foregoing, in various aspects, the Cas12f polypeptide may comprise or consist of an amino acid sequence having 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs: 393-407.

[0112] In any of the foregoing aspects, the Cas12f polypeptide is advantageously compact, relative to other Cas proteins. Accordingly, in various aspects, the Cas12f polypeptide of the present disclosure may be less than 1000 amino acids in length (e.g., less than 900, less than 800, less than 700, less than 600, less than 500 amino acids or less than 400 amino acids in length). In some aspects, the Cas12f polypeptide of the present disclosure may be about 200 to 1000 amino acids in length. In some aspects, the Cas12f polypeptide is about 200 to about 1000 amino acids, about 200 to about 950 amino acids, about 200 to about 900 amino acids, about 200 to about 850 amino acids, about 200 to about 800 amino acids, about 200 to about 750 amino acids, about 200 to about 700 amino acids, about 200 to about 650 amino acids, about 200 to about 600 amino acids, about 200 to about 550 amino acids, about 200 to about 500 amino acids, about 200 to about 450 amino acids, about 200 to about 400 amino acids, about 250 to about 1000 amino acids, about 250 to about 950 amino acids, about 250 to about 900 amino acids, about 250 to about 850 amino acids, about 250 to about 800 amino acids, about 250 to about 750 amino acids, about 250 to about 700 amino acids, about 250 to about 650 amino acids, about 250 to about 600 amino acids, about 250 to about 550 amino acids, about 250 to about 500 amino acids, about 250 to about 450 amino acids, about 250 to about 400 amino acids, about 300 to about 1000 amino acids, about 300 to about 950 amino acids, about 300 to about 900 amino acids, about 300 to about 850 amino acids, about 300 to about 800 amino acids, about 300 to about 750 amino acids, about 300 to about 700 amino acids, -34- 99975521.7Attorney Docket No: 124540-829360 about 300 to about 650 amino acids, about 300 to about 600 amino acids, about 300 to about 550 amino acids, about 300 to about 500 amino acids, about 300 to about 450 amino acids, about 300 to about 400 amino acids, about 350 to about 1000 amino acids, about 350 to about 950 amino acids, about 350 to about 900 amino acids, about 350 to about 850 amino acids, about 350 to about 800 amino acids, about 350 to about 750 amino acids, about 350 to about 700 amino acids, about 350 to about 650 amino acids, about 350 to about 600 amino acids, about 350 to about 550 amino acids, about 350 to about 500 amino acids, about 350 to about 450 amino acids, or about 350 to about 400 amino acids. In some embodiments, the Cas12f polypeptide is between 350 and 400 amino acids in length (i.e., 350, 351, 352, 353, 354, 355, 356, 357, 358, 359, 360, 361, 362, 363, 364, 365, 366, 367, 368, 369, 370, 371, 372, 373, 374, 375, 376, 377, 378, 379, 380, 381, 382, 383, 384, 385, 386, 387, 388, 389, 390, 391, 392, 393, 394, 395, 396, 397, 398, 399, or 400 amino acids in length). In all of these and related embodiments, the Cas12f polypeptide is active (i.e., can interact with an associated sgRNA and target a DNA) despite its small size. Fusion Proteins comprising Cas12f polypeptides.

[0113] In another aspect of the present disclosure, fusion proteins comprising a Cas12f polypeptide and one or more functionally active polypeptides are provided.

[0114] In an aspect, the functional polypeptide may be capable of modifying a target nucleic acid. For instance, the functional polypeptide or moiety may have methyltransferase activity, DNA repair activity, DNA damage activity, dismutase activity, alkylation activity, dealkylation activity, depurination activity, oxidation activity, deoxidation activity, pyrimidine dimer forming activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, glycosylase activity, acetyl transferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, deSUMOylation activity, ribosylation activity, deribosylation activity, myristoylation activity, demyristoylation activity, glycosylation activity (e.g., from O-GlcNAc transferase), deglycosylation activity, reverse transcriptase activity, or a combination of any thereof.

[0115] In various embodiments, therefore, the functional polypeptide may comprise a base editing domain (e.g., a deaminase), a base excising domain, an uracil glycosylase inhibitor (UGI), an uracil glycosylase (UNG), a methylpurine glycosylase (MPG), a methylase, a -35- 99975521.7Attorney Docket No: 124540-829360 demethylase, an transcription activating domain (e.g., VP16, VP64 or VPR), a transcription inhibiting domain (e.g., KRAB moiety or SID moiety) , a reverse transcriptase, an exonuclease (e.g., T5E), a destabilized domain (e.g., destabilized domains (DD) of E. coli dihydrofolate reductase (ecDHFR), a histone residue modification domain, a nuclease catalytic domain (e.g., FokI) , a transcription modification factor, a light gating factor, a chemical inducible factor, a chromatin visualization factor, a transcription release factor, an HDAC, a moiety having single strand RNA (ssRNA) cleavage activity, a moiety having double strand (dsRNA) cleavage activity, a moiety having ssDNA cleavage activity, a moiety having dsDNA cleavage activity, a DNA or RNA ligase, any catalytic domain thereof, or any combination thereof.

[0116] In various aspects, the functional polypeptide comprises a Ten-eleven translocation methylcytosine dioxygenase (i.e., Tet, Tet1, Tet2 or mini Tet2), DNA methyl transferase (e.g. DNMT3A and / or DNMT3L), VP16, VP64, VPR, CBE, ABE, a Krüppel-associated box (KRAB), or any catalytically active domain thereof.

[0117] In an aspect, the functional polypeptide comprises a Ten-eleven translocation methylcytosine dioxygenase (Tet) polypeptide. Tet polypeptides are Ten-eleven translocation proteins (TET1-3) are dioxygenases that oxidize 5-methyldeoxycytosine, thus taking part in passive and active demethylation. In an aspect, the Tet polypeptide may comprise a full length (wildtype) Tet polypeptide. For instance, the Tet polypeptide may comprise Tet1, Tet2 or Tet3. In an aspect, the Tet polypeptide may comprise a catalytically active Tet polypeptide that is truncated or otherwise modified. Exemplary truncated Tet proteins include truncated versions of Tet1 (i.e., Tet1-CD) or Tet2 (called TETmini). Illustrative sequences of these truncated versions of Tet include SEQ ID NOs: 408 or 409, as depicted in Table 4 below. In an aspect, the Tet polypeptide may comprise an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 408 or 409. In an aspect, the Tet polypeptide may comprise an amino acid sequence comprising or consisting of SEQ ID NO: 408 or 409. In an aspect, the Tet polypeptide may comprise an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 408. In an aspect, the Tet polypeptide may comprise an amino acid sequence comprising or consisting of SEQ ID NO: 408. Table 4 – Illustrative Tet Polypeptides -36- 99975521.7Attorney Docket No: 124540-829360 Tet Protein Sequence SEQ ID NO: DTPGEQSQNGKCEGCNPDKDEAPYYTH LGAGPDVAAIRTLMEERYGEKGKAIRIE KVIYTGKEGKSSQGCPIAKWVYRRSSEE EKLLCLVRVRPNHTCETAVMVIAIMLW DGIPKLLASELYSELTDILGKCGICTNRR CSQNETKKKQSPPSRNCCCQGENPETCG ASFSFGCSWSMYYNGCKFARSKKPRKF RLHGAEPKEEERLGSHLQNLATVIAPIY KKLAPDAYNNQVEFEHQAPDCCLGLKE Tet2 -truncated (TET-mini) GRPFSGVTACLDFSAHSHRDQQNMPNG 408 STVVVTLNREDNREVGAKPEDEQFHVL PMYIIAPEDEFGSTEGQEKKIRMGSIEVL QSFRRRRVIRIGELPKPGKKLLPGLAAV QEIEYWSDSEHNFQDPCIGGVAIAPTHG SILIECAKCEVHATTKVNDPDRNHPTRIS LVLYRHKNLFLPKHCLALWEAKMAEK ARKEEECGKNGSDHVSQKNHGKQEKR EPTGPQEPSYLRFIQSLAENTGSVTTDST VTTSPYAFTQVTGPYNTFV LPTCSCLDRVIQKDKGPYYTHLGAGPSV AAVREIMENRYGQKGNAIRIEIVVYTGK EGKSSHGCPIAKWVLRRSSDEEKVLCLV RQRTGHHCPTAVMVVLIMVWDGIPLPM ADRLYTELTENLKSYNGHPTDRRCTLN ENRTCTCQGIDPETCGASFSFGCSWSMY FNGCKFGRSPSPRRFRIDPSSPLHEKNLE DNLQSLATRLAPIYKQYAPVAYQNQVE YENVARECRLGSKEGRPFSGVTACLDF CAHPHRDIHNMNNGSTVVCTLTREDNR SLGVIPQDEQLHVLPLYKLSDTDEFGSK EGMEAKIKSGAIEVLAPRRKKRTCFTQP Tet1-Catalytic Domain (CD) VPRSGKKRAAMMTEVLAHKIRAVEKKP IPRIKRKNNSTTTNNSKPSSLPTLGSNTE 409 TVQPEVKSETEPHFILKSSDNTKTYSLM PSAPHPVKEASPGFSWSPKTASATPAPL KNDATASCGFSERSSTPHCTMPSGRLSG ANAAAADGPGISQLGEVAPLPTLSAPV MEPLINSEPSTGVTEPLTPHQPNHQPSFL TSPQDLASSPMEEDEQHSEADEPPSDEP LSDDPLSPAEEKLPHIDEYWSDSEHIFLD ANIGGVAIAPAHGSVLIECARRELHATT PVEHPNRNHPTRLSLVFYQHKNLNKPQ HGFELNKIKFEAKEAKNKKMKASEQKD QAANEGPEQSSEVNELNQIPSHKALTLT HDNVVTVSPYALTHVAGPYNHWV -37- 99975521.7Attorney Docket No: 124540-829360

[0118] In an aspect, the functional polypeptide comprises a DNA methyltransferase 3 alpha (DNMT3A) polypeptide. This polypeptide, also referred to interchangeably as DNA (cytosine-5)-methyltransferase 3A facilitates transfer of methyl groups to specific CpG structures in DNA, which contributes to gene silencing. Exemplary DNMT3A or related polypeptides that may be combined with any of the Cas12f polypeptides described above include hDNMT3A-L, hDNMT3A, hDNMT3L, H3-hDNMT3L, mDNMT3A-L. For ease of reference, illustrative sequences of these DNMT3A polypeptides are provided in Table 5 below. Table 5 – Illustrative DNMT3A polypeptides DNMT3 Polypeptide Sequence SEQ ID NO: hDNMT3A-L NHDQEFDPPKVYPPVPAEKRKPIRVLSL 410 FDGIATGLLVLKDLGIQVDRYIASEVCE DSITVGMVRHQGKIMYVGDVRSVTQK HIQEWGPFDLVIGGSPCNDLSIVNPARK GLYEGTGRLFFEFYRLLHDARPKEGDD RPFFWLFENVVAMGVSDKRDISRFLESN PVMIDAKEVSAAHRARYFWGNLPGMN RPLASTVNDKLELQECLEHGRIAKFSKV RTITTRSNSIKQGKDQHFPVFMNEKEDIL WCTEMERVFGFPVHYTDVSNMSRLAR QRLLGRSWSVPVIRHLFAPLKEYFACVS SGNSNANSRGPSFSSGLVPLSLRGSHNP LEMFETVPVWRRQPVRVLSLFEDIKKEL TSLGFLESGSDPGQLKHVVDVTDTVRK DVEEWGPFDLVYGATPPLGHTCDRPPS WYLFQFHRLLQYARPKPGSPRPFFWMF VDNLVLNKEDLDVASRFLEMEPVTIPD VHGGSLQNAVRVWSNIPAIRSSRHWAL VSEEELSLLAQNKQSSKLAAKWPTKLV KNCFLPLREYFKYFSTELTSSL hDNMT3A NHDQEFDPPKVYPPVPAEKRKPIRVLSL 411 FDGIATGLLVLKDLGIQVDRYIASEVCE DSITVGMVRHQGKIMYVGDVRSVTQK HIQEWGPFDLVIGGSPCNDLSIVNPARK GLYEGTGRLFFEFYRLLHDARPKEGDD RPFFWLFENVVAMGVSDKRDISRFLESN PVMIDAKEVSAAHRARYFWGNLPGMN RPLASTVNDKLELQECLEHGRIAKFSKV RTITTRSNSIKQGKDQHFPVFMNEKEDIL WCTEMERVFGFPVHYTDVSNMSRLAR QRLLGRSWSVPVIRHLFAPLKEYFACV -38- 99975521.7Attorney Docket No: 124540-829360 hDNMT3L NPLEMFETVPVWRRQPVRVLSLFEDIKK 412 ELTSLGFLESGSDPGQLKHVVDVTDTVR KDVEEWGPFDLVYGATPPLGHTCDRPP SWYLFQFHRLLQYARPKPGSPRPFFWM FVDNLVLNKEDLDVASRFLEMEPVTIPD VHGGSLQNAVRVWSNIPAIRSSRHWAL VSEEELSLLAQNKQSSKLAAKWPTKLV KNCFLPLREYFKYFSTELTSSL H3-hDNMT3L ARTKQTARKSTGGKAPRKQLSSGNSNA 413 NSRGPSFSSGLVPLSLRGSHNPLEMFET VPVWRRQPVRVLSLFEDIKKELTSLGFL ESGSDPGQLKHVVDVTDTVRKDVEEW GPFDLVYGATPPLGHTCDRPPSWYLFQF HRLLQYARPKPGSPRPFFWMFVDNLVL NKEDLDVASRFLEMEPVTIPDVHGGSLQ NAVRVWSNIPAIRSSRHWALVSEEELSL LAQNKQSSKLAAKWPTKLVKNCFLPLR EYFKYFSTELTSSL mDNMT3A-L NHDQEFDPPKVYPPVPAEKRKPIRVLSL 414 FDGIATGLLVLKDLGIQVDRYIASEVCE DSITVGMVRHQGKIMYVGDVRSVTQK HIQEWGPFDLVIGGSPCNDLSIVNPARK GLYEGTGRLFFEFYRLLHDARPKEGDD RPFFWLFENVVAMGVSDKRDISRFLESN PVMIDAKEVSAAHRARYFWGNLPGMN RPLASTVNDKLELQECLEHGRIAKFSKV RTITTRSNSIKQGKDQHFPVFMNEKEDIL WCTEMERVFGFPVHYTDVSNMSRLAR QRLLGRSWSVPVIRHLFAPLKEYFACVS SGNSNANSRGPSFSSGLVPLSLRGSHMG PMEIYKTVSAWKRQPVRVLSLFRNIDK VLKSLGFLESGSGSGGGTLKYVEDVTN VVRRDVEKWGPFDLVYGSTQPLGSSCD RCPGWYMFQFHRILQYALPRQESQRPFF WIFMDNLLLTEDDQETTTRFLQTEAVTL QDVRGRDYQNAMRVWSNIPGLKSKHA PLTPKEEEYLQAQVRSRSKLDAPKVDLL VKNCLLPLREYFKYFSQNSLPL

[0100] In an aspect, the functional polypeptide comprises a VPR polypeptide. Thispolypeptide, also referred to interchangeably VP64-p65-Rta is a transcriptional activator created by combining the transcription factors of p65, Rta and VP64. An illustrative amino acid sequence of a VPR polypeptide contemplated to be combined with any of the optimized Cas12f polypeptides herein is provided as SEQ ID NO: 415 and is shown in Table 6 below. -39- 99975521.7Attorney Docket No: 124540-829360 Table 6 – Illustrative VPR Polypeptide VPR Polypeptide Sequence SEQ ID NO: MPKKKRKVEASGSGRADALDDFDLDM LGSDALDDFDLDMLGSDALDDFDLDM LGSDALDDFDLDMLINSRSSGSPKKKRK VGSQYLPDTDDRHRIEEKRKRTYETFKS IMKKSPFSGPTDPRPPPRRIAVPSRSSAS VPKPAPQPYPFTSSLSTINYDEFPTMVFP SGQISQASALAPAPPQVLPQAPAPAPAP AMVSALAQAPAPVPVLAPGPPQAVAPP APKPTQAGEGTLSEALLQLQFDDEDLG VP64-p65-Rta ALLGNSTDPAVFTDLASVDNSEFQQLLN 415 QGIPVAPHTTEPMLMEYPEAITRLVTGA QRPPDPAPAPLGAPGLPNGLLSGDEDFS SIADMDFSALLGSGSGSRDSREGMFLPK PEAGSAISDVFEGREVCQPKRIRPFHPPG SPWANRPLPASLAPTPTGPVHEPVGSLT PAPVPQPLDPAPAVTPEASHLLEDPDEE TSQAVKALREMADTVIPQKEEAAICGQ MDLSHPPPRGHLDELTTTLESMTEDLNL DSPLTPELNEILDTFLNDECLLHAMHIST GLSIFDTSLF

[0101] In various aspects, the fusion proteins provided herein may comprise a flexiblelinker that connects the Cas12f polypeptide with one or more functional polypeptides listed above. In some aspects, the flexible linker may be an XTEN linker (e.g., SGGSSGGSSGSETPGTSESATPESSGGSSGGSS, SEQ ID NO: 608), a GS linker, or a GSSG (or SGGS) linker.

[0102] In various aspects, the Cas12f polypeptide and / or fusion proteins comprising theCas12f polypeptide may further comprise a localization signal such as a nuclear localization signal (NLS), a nuclear export signal (NES), a targeting polypeptide for providing binding to a cell surface portion on a target cell or a target cell type, a polypeptide targeting moiety, a DNA binding domain (e.g., MBP, Lex A DBD, Gal4 DBD), or any other localization signal. In some aspects, the Cas12f polypeptide or fusion protein thereof comprises an NLS. In some aspects, the NLS can be located at or within 50 amino acids of the amino-terminus of the Cas12f polypeptide and / or be located at or within 50 amino acids of the carboxy-terminus of the Cas12f polypeptide. In some aspects, particularly when the NLS is included in a fusion protein comprising a functional polypeptide, the NLS may be located at or within 50 amino -40- 99975521.7Attorney Docket No: 124540-829360 acids of the amino terminus of the functional polypeptide and / or be located at or within 50 amino acids of the carboxyl terminus of the functional polypeptide.

[0103] In various aspects, the Cas12f polypeptide and / or fusion proteins comprising theCas12f polypeptide may further comprise a moiety to aid in the purification, visualization or detection of the Cas12f polypeptide and / or fusion protein. For instance, such a moiety may comprise a reporter (e.g., fluorescent) polypeptide or a detection label (e.g., GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, BFP), or an epitope tag (e.g., His, myc, V5, FLAG, HA, VSV-G, Trx, etc).

[0104] One exemplary fusion protein provided herein comprises a Cas12f polypeptidelinked to a VPR transcriptional activator. In some aspects, the fusion protein comprises from the N terminus to C terminus, a nuclear localization sequence (NLS, SEQ ID NO: 606), a Cas12f polypeptide (e.g., a BsCas12f polypeptide comprising SEQ ID NO: 393), a flexible XTEN linker (SEQ ID NO: 608) and a VPR transcriptional activator (VP64-p65-Rta, SEQ ID NO: 415) and a second NLS (SEQ ID NO: 606). The amino acid sequence of this Cas12f polypeptide is provided herein as SEQ ID NO: 416 and a nucleic acid encoding it is provided herein as SEQ ID NO: 417. Accordingly, in some aspects, a fusion protein comprising a Cas12f polypeptide linked to a VPR transcriptional activator is provided comprising at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 416. In some aspects, the fusion protein comprising a Cas12f polypeptide linked to a VPR transcriptional activator comprises an amino acid sequence comprising or consisting of SEQ ID NO: 416.

[0105] One exemplary fusion protein provided herein comprises a Cas12f polypeptidelinked to a VPR transcriptional activator. In some aspects, the fusion protein comprises from the N terminus to C terminus, a nuclear localization sequence (NLS, SEQ ID NO: 606), a Cas12f polypeptide (e.g., a BsCas12f polypeptide comprising SEQ ID NO: 338), a flexible XTEN linker (SEQ ID NO: 608), a VPR transcriptional activator (VP64-p65-Rta, SEQ ID NO: 415) and a second NLS (SEQ ID NO: 606). The amino acid sequence of this Cas12f polypeptide is provided herein as SEQ ID NO: 418 and a nucleic acid encoding it is provided herein as SEQ ID NO: 419. Accordingly, in some aspects, a fusion protein comprising a Cas12f polypeptide linked to a VPR transcriptional activator is provided comprising at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 418. In some -41- 99975521.7Attorney Docket No: 124540-829360 aspects, the fusion protein comprising a Cas12f polypeptide linked to a VPR transcriptional activator comprises an amino acid sequence comprising or consisting of SEQ ID NO: 418.

[0106] One exemplary fusion protein provided herein comprises a Cas12f polypeptidelinked to a Tet polypeptide. In some aspects, the fusion protein comprises from the N terminus to C terminus, a nuclear localization sequence (NLS, SEQ ID NO: 606), a Cas12f polypeptide (e.g., a BsCas12f polypeptide comprising SEQ ID NO: 338) a flexible XTEN linker (SEQ ID NO: 608), a Tet polypeptide (TetMini, SEQ ID NO: 408) and a second NLS (SEQ ID NO: 606). The amino acid sequence of this fusion protein is provided herein as SEQ ID NO: 420 and a nucleic acid encoding it is provided herein as SEQ ID NO: 421. Accordingly, in some aspects, a fusion protein comprising a Cas12f polypeptide linked to a Tet polypeptide is provided comprising at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 420. In some aspects, the fusion protein comprising a Cas12f polypeptide linked to a Tet polypeptide comprises an amino acid sequence comprising or consisting of SEQ ID NO: 420.

[0107] One exemplary fusion protein provided herein comprises a Cas12f polypeptidelinked to a Tet polypeptide. In some aspects, the fusion protein comprises from the N terminus to C terminus, a nuclear localization sequence (NLS, SEQ ID NO: 606), a Cas12f polypeptide (e.g., a BsCas12f polypeptide comprising SEQ ID NO: 338) a flexible XTEN linker (SEQ ID NO: 608), a catalytic domain of a Tet1 polypeptide (Tet1-CD, SEQ ID NO: 409) and a second NLS (SEQ ID NO: 606). The amino acid sequence of this fusion protein is provided herein as SEQ ID NO: 422 and a nucleic acid encoding it is provided herein as SEQ ID NO: 423. Accordingly, in some aspects, a fusion protein comprising a Cas12f polypeptide linked to a Tet polypeptide is provided comprising at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 422. In some aspects, the fusion protein comprising a Cas12f polypeptide linked to a Tet polypeptide comprises an amino acid sequence comprising or consisting of SEQ ID NO: 422.

[0108] Another exemplary fusion protein provided herein comprises a Cas12f polypeptidelinked to a Tet polypeptide. In some aspects, the fusion protein comprises from the N terminus to C terminus, a nuclear localization sequence (NLS, SEQ ID NO: 606), a truncated Tet2 polypeptide (Tet2-Mini, SEQ ID NO: 408), a flexible XTEN linker (SEQ ID NO: 608), a Cas12f polypeptide (e.g., a BsCas12f polypeptide comprising SEQ ID NO: 338), and a -42- 99975521.7Attorney Docket No: 124540-829360 second NLS (SEQ ID NO: 606). The amino acid sequence of this fusion protein is provided herein as SEQ ID NO: 424 and a nucleic acid encoding it is provided herein as SEQ ID NO: 425. Accordingly, in some aspects, a fusion protein comprising a Cas12f polypeptide linked to a polypeptide is provided comprising at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 424. In some aspects, the fusion protein comprising a Cas12f polypeptide linked to a Tet polypeptide comprises an amino acid sequence comprising or consisting of SEQ ID NO: 424.

[0109] Also provided are fusion proteins comprising a Tet polypeptide provide herein(i.e., Tet1-CD (SEQ ID NO: 409) or Tet2-Mini (SEQ ID NO: 408)) linked to another Cas protein (i.e., a Cas9). One such fusion protein may comprise, for example, from the N terminus to C terminus, a nuclear localization sequence (NLS, SEQ ID NO: 606), a Cas9 nuclease (e.g., SpCas9, SEQ ID NO: 652), a flexible XTEN linker (SEQ ID NO: 608), a Tet polypeptide (TetMini, SEQ ID NO: 408) and a second NLS (SEQ ID NO: 606). The amino acid sequence of this fusion protein is provided herein as SEQ ID NO: 426 and a nucleic acid encoding it is provided herein as SEQ ID NO: 427. Accordingly, a fusion protein comprising a Cas9 nuclease and a truncated, catalytically active, Tet polypeptide can comprise an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 426. In an aspect, a fusion protein comprising a Cas9 nuclease and a truncated, catalytically active, Tet polypeptide can comprise or consist of SEQ ID NO: 426.

[0110] Another exemplary fusion protein provided herein comprises a Cas12f polypeptidelinked to a DNMT3A polypeptide. In some aspects, the fusion protein comprises from the N terminus to C terminus, a nuclear localization sequence (NLS, SEQ ID NO: 606), a Cas12f polypeptide (e.g., a BsCas12f polypeptide comprising SEQ ID NO: 338), a flexible XTEN linker (SEQ ID NO: 608) and a DNMT3A polypeptide (i.e., any one of 410-414). An amino acid sequence of such a fusion protein comprising a hDNMT3A-L polypeptide (SEQ ID NO: 410) is provided herein as SEQ ID NO: 428 and a nucleic acid encoding it is provided herein as SEQ ID NO: 429. An amino acid sequence of another such a fusion protein comprising a hDNMT3A polypeptide (SEQ ID NO: 411) is provided herein as SEQ ID NO: 430 and a nucleic acid encoding it is provided herein as SEQ ID NO: 431. An amino acid sequence of another such a fusion protein comprising a hDNMT3L polypeptide (SEQ ID NO: 412) is provided herein as SEQ ID NO: 432 and a nucleic acid encoding it is provided herein as SEQ -43- 99975521.7Attorney Docket No: 124540-829360 ID NO: 433. An amino acid sequence of another such a fusion protein comprising a H3- hDNMT3L polypeptide (SEQ ID NO: 413) is provided herein as SEQ ID NO: 434 and a nucleic acid encoding it is provided herein as SEQ ID NO: 435. An amino acid sequence of another such a fusion protein comprising a mDNMT3A-L polypeptide (SEQ ID NO: 414) is provided herein as SEQ ID NO: 436 and a nucleic acid encoding it is provided herein as SEQ ID NO: 437. Accordingly, in some aspects, a fusion protein comprising a Cas12f polypeptide linked to a DNMT3A polypeptide is provided comprising at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 428, 430, 432, 434, and 436. In some aspects, the fusion protein comprising a Cas12f polypeptide linked to a DNMT3A comprises an amino acid sequence comprising or consisting of any one of SEQ ID NOs: 428, 430, 432, 434, and 436. gRNAs

[0111] The present disclosure also provides guide RNAs (gRNAs) that can direct theCas12f polypeptides (or fusion proteins thereof) provided herein to a specific target site within a nucleic acid. In an aspect, the gRNAs provided herein are single guide RNAs which comprise a crRNA and tracrRNA linked by a short DNA linker. As discussed further below, the crRNA comprises a spacer sequence that hybridizes to a target nucleic acid and the tracrRNA forms a scaffold structure that allows for the formation of a complex with the Cas12f polypeptide – thereby guiding the complex to the target DNA. In an aspect, the tracrRNA, flexible linker and any portion of the crRNA that is not the spacer sequence may be understood to be a “scaffold sequence.” In other words, the gRNAs of the present disclosure may be understood to comprise a spacer sequence that hybridizes to a target nucleic acid, and a scaffold sequence capable of forming a complex with the Cas12f polypeptide – thereby guiding the complex to the target DNA. Advantageously, the gRNAs provided herein have improved scaffold sequences that improve the efficiency and accuracy of the associated Cas12f polypeptides in this disclosure. Exemplary spacer sequences and scaffold sequences are discussed further herein below.

[0112] In view of the foregoing, the guide RNAs provided herein are provided as single-molecule guide RNAs (sgRNAs). These single-molecule guide RNA (sgRNA) can comprise, in the 5' to 3' direction, the scaffold sequence (e.g., any of SEQ ID NOs: 438-515, provided below) and a spacer sequence. In various aspects, the single-molecule guide RNA (sgRNA) can further comprise elements that contribute additional functionality (e.g., stability) to the -44- 99975521.7Attorney Docket No: 124540-829360 guide RNA. These additional elements are encompassed herein in a “scaffold”. The scaffold itself may further comprise additional components (i.e., hairpins, linkers, CRISPR repeat sequences, Tracr regions, and / or spacer extension regions) that all improve function and stability of the guide RNA. For instance, the scaffold may comprise a MS1 stem loop at either the 5’- or 3’ end for recruiting endogenous or exogenous gene regulators. Spacer Sequences

[0113] As noted above, the guide RNAs of the present disclosure further comprise aspacer sequence targeting a genome location of interest. As is understood by a person of ordinary skill in the art, the guide RNAs of the present disclosure may include spacer sequence complementary to its genomic target site or region. See Jinek et al., Science, 2012, 337, 816-821 and Deltcheva et al., Nature, 2011, 471, 602-607, which are each incorporated herein by reference in their entirety.

[0114] The spacer sequence may in certain embodiments comprise between 15 and 200nucleotides, but preferably comprises between about 15 and 30 nucleotides (e.g., 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides). In some aspects, the spacer sequence comprises about 16 to 24 nucleotides.

[0115] In some embodiments, the sgRNA comprises a spacer sequence that hybridizes to asequence in a target polynucleotide. The spacer of the sgRNA can interact with a target polynucleotide in a sequence-specific manner via hybridization (i.e., base pairing). The nucleotide sequence of the spacer can vary depending on the sequence of the target nucleic acid of interest.

[0116] In some embodiments, a sgRNA comprises a 20-nucleotide spacer sequence. Insome embodiments, a sgRNA comprises a less than a 20-nucleotide spacer sequence. In some embodiments, a sgRNA comprises a more than 20 nucleotide spacer sequence. In some embodiments, a sgRNA comprises a variable length spacer sequence with 17-30 nucleotides.

[0117] In a CRISPR-endonuclease system described herein, a spacer sequence can bedesigned to hybridize to a target polynucleotide that is located 3' of a protospacer adjacent motif (PAM) of the endonuclease used in the system. The spacer may perfectly match the target sequence or may have mismatches. Each endonuclease, e.g., a Cas12f nuclease, has a particular PAM sequence that it recognizes in a target DNA. For example, the Cas12f polypeptides of the present disclosure recognize a PAM that comprises the sequence 5'-CCN- 3', where CCN is immediately 5’ to the target nucleic acid sequence targeted by the spacer -45- 99975521.7Attorney Docket No: 124540-829360 sequence. For example, the target locus can have 5’-CCA [N]16-24-3’ sequence, where CCA = CCN is the pam, the [N]16-24 is a 16-24 nucleotide long target polynucleotide sequence (equivalent to the spacer sequence of the gRNA). This PAM target is G / C rich, which further distinguishes these polypeptides from other Cas proteins in the Cas12f family, which usually prefer an A / T rich PAM. For example, UnCas12f uses TTTA as its PAM.

[0118] A target polynucleotide sequence can comprise 16 to 24 nucleotides. The targetpolynucleotide can comprise less than 20 nucleotides. The target polynucleotide can comprise more than 20 nucleotides. The target polynucleotide can comprise at least: 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30 or more nucleotides. The target polynucleotide can comprise at most: 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30 or more nucleotides. The target polynucleotide sequence can comprise 20 bases immediately 3' of the first nucleotide of the PAM. The target polynucleotide sequence can comprise 16 to 24 bases (i.e., 16, 17, 18, 19, 20, 21, 22, 23, or 24 bases) immediately 3’ of the first nucleotide of the PAM.

[0119] A spacer sequence that hybridizes to a target polynucleotide can have a length of atleast about 6 nucleotides (nt). The spacer sequence can be at least about 6 nt, at least about 10 nt, at least about 15 nt, at least about 18 nt, at least about 19 nt, at least about 20 nt, at least about 25 nt, at least about 30 nt, at least about 35 nt or at least about 40 nt, from about 6 nt to about 80 nt, from about 6 nt to about 50 nt, from about 6 nt to about 45 nt, from about 6 nt to about 40 nt, from about 6 nt to about 35 nt, from about 6 nt to about 30 nt, from about 6 nt to about 25 nt, from about 6 nt to about 20 nt, from about 6 nt to about 19 nt, from about 10 nt to about 50 nt, from about 10 nt to about 45 nt, from about 10 nt to about 40 nt, from about 10 nt to about 35 nt, from about 10 nt to about 30 nt, from about 10 nt to about 25 nt, from about 10 nt to about 20 nt, from about 10 nt to about 19 nt, from about 19 nt to about 25 nt, from about 19 nt to about 30 nt, from about 19 nt to about 35 nt, from about 19 nt to about 40 nt, from about 19 nt to about 45 nt, from about 19 nt to about 50 nt, from about 19 nt to about 60 nt, from about 20 nt to about 25 nt, from about 20 nt to about 30 nt, from about 20 nt to about 35 nt, from about 20 nt to about 40 nt, from about 20 nt to about 45 nt, from about 20 nt to about 50 nt, or from about 20 nt to about 60 nt. In some examples, the spacer sequence can comprise from about 16 nt to about 24 nt. In some examples, the spacer sequence can comprise from about 17 nt to about 23 nt. In some examples, the spacer sequence can comprise 16 nt. In some examples, the spacer sequence can comprise 17 nt. In some examples, the spacer sequence can comprise 18 nt. In some examples, the spacer sequence can comprise 19 nt. In some examples, the spacer sequence can comprise 20 nucleotides. In -46- 99975521.7Attorney Docket No: 124540-829360 some examples, the spacer sequence can comprise 21 nucleotides. In some examples, the spacer can comprise 22 nucleotides. In some examples, the spacer can comprise 23 nucleotides. In some examples, the spacer can comprise 24 nucleotides.

[0120] In some examples, the percent complementarity between the spacer sequence andthe target nucleic acid is at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 97%, at least about 98%, at least about 99%, or 100%. In some examples, the percent complementarity between the spacer sequence and the target nucleic acid is at most about 30%, at most about 40%, at most about 50%, at most about 60%, at most about 65%, at most about 70%, at most about 75%, at most about 80%, at most about 85%, at most about 90%, at most about 95%, at most about 97%, at most about 98%, at most about 99%, or 100%. In some examples, the percent complementarity between the spacer sequence and the target nucleic acid is 100% over the six contiguous 5'-most nucleotides of the target sequence of the complementary strand of the target nucleic acid. The percent complementarity between the spacer sequence and the target nucleic acid can be at least 60% over about 20 contiguous nucleotides. The length of the spacer sequence and the target nucleic acid can differ by 1 to 6 nucleotides, which may be thought of as a bulge or bulges.

[0121] In some embodiments, a sgRNA comprises a spacer extension sequence thatcomprises another moiety (e.g., a stability control sequence, an endoribonuclease binding sequence, or a ribozyme). The moiety can decrease or increase the stability of a nucleic acid targeting nucleic acid. The moiety can be a transcriptional terminator segment (i.e., a transcription termination sequence). The moiety can function in a eukaryotic cell. The moiety can function in a prokaryotic cell. The moiety can function in both eukaryotic and prokaryotic cells. Non-limiting examples of suitable moieties include: a 5' cap (e.g., a 7- methylguanylate cap (m7 G)), a riboswitch sequence (e.g., to allow for regulated stability and / or regulated accessibility by proteins and protein complexes), a sequence that forms a dsRNA duplex (i.e., a hairpin), a sequence that targets the RNA to a subcellular location (e.g., nucleus, mitochondria, chloroplasts, and the like), a modification or sequence that provides for tracking (e.g., direct conjugation to a fluorescent molecule, conjugation to a moiety that facilitates fluorescent detection, a sequence that allows for fluorescent detection, etc.), and / or a modification or sequence that provides a binding site for proteins (e.g., proteins that act on DNA, including transcriptional activators, transcriptional repressors, DNA -47- 99975521.7Attorney Docket No: 124540-829360 methyltransferases, DNA demethylases, histone acetyltransferases, histone deacetylases, and the like).

[0122] Exemplary non-limiting spacer sequences are provided herein. In non-limitingembodiments, spacer sequences are provided targeting a TetO sequence (as used in the reporter cell lines described further below), an HBB gene, an HBG1 gene, an MECP2 gene, a PGK gene or a G4C2 repeat (i.e., (GGGGCC)n), but additional genes may be targeted and contemplated herein. Exemplary gRNA spacer sequences for these genes are provided in the Tables 7A-7F herein below. Table 7A- Spacer Sequences targeting Tet0 Target Gene gRNA name Spacer Sequence SEQ ID NO:Tet0 g004 ctccctatcagtgatagaga 516Tet0 g030 tctctatcactgatagggag 517 Table 7B – Spacer sequences targeting HBB Target Gene gRNA name Spacer Sequence SEQ ID NO:HBB g051 gtgccagaagagccaaggac 518HBB g052 gaagagccaaggacaggta 519 HBB g053 gggctgggcataaaagtca 520 HBB g054 gtggagccacaccctagggt 521 HBB g055 atctactcccaggagcaggg 522 HBB g056 ttactgccctgtggggca 523 HBB g081 gtccttggctcttctggcac 524 HBB g082 tacctgtccttggctcttc 525 HBB g083 tgacttttatgcccagccc 526 HBB g084 accctagggtgtggctccac 527 HBB g085 ccctgctcctgggagtagat 528 HBB g086 tgccccacagggcagtaa 529 -48- 99975521.7Attorney Docket No: 124540-829360 Table 7C - Spacer sequences targeting HBP1 Target Gene gRNA name Spacer sequence SEQ ID NO:HBP1 g041 ggctaaactccacccatgggt 530 HBP1 g042 atagtcttagagtatccagtg 531 HBP1 g043 gtgaggccaggggccggcggc 532HBP1 g044 tgggtcatttcacagagg 533 HBP1 g071 acccatgggtggagtttagcc 534 HBP1 g072 cactggatactctaagactat 535 HBP1 g073 gccgccggcccctggcctcac 536 HBP1 g074 cctctgtgaaatgaccca 537 Table 7D– Spacer sequences targeting MeCP2 Target Gene gRNA name Spacer Sequence SEQ ID NO:MeCP2 g091 gagctctcacccccatctct 538 MeCP2 g101 tttccctggccgaaatggac 539 MeCP2 g113 cctacttgttcctgctagat 540 MeCP2 g114 tctagcaggaacaagtaggt 541MeCP2 g115 cagccctctctccgagagg 542 MeCP2 g116 tcacagccaatgacgggc 543 MeCP2 g117 tttcggccagggaaaagg 544 Table 7E - gRNAs targeting PGK Promoter Target Gene gRNA name Spacer Sequence SEQ ID NO:PGK g171 cgttgaccgaatcaccgacc 545 -49- 99975521.7Attorney Docket No: 124540-829360 PGK g172 ccggagcgcacgtcggcagt 546PGK g173 gttcctgcccgcgcggtgtt 547 Table 7F - Spacer sequences targeting GC2 ((GGGCC)n Repeats) Target Gene gRNA name Spacer Sequence SEQ ID NO:GC2 g026 gggccggggccggggccggg 548GC2 g027 cggccccggccccggccccg 549 GC2 g028 ggccccggccccggccccgg 550 GC2 g029 gccccggccccggccccggc 551 Scaffold Sequences

[0123] As discussed above, the sgRNAs of the instant disclosure comprise scaffoldsequences which may comprise all or part of a tracrRNA sequence, a flexible linker, and / or a portion of a crRNA sequence not including the spacer sequence. As discussed further below, the instant disclosure provides new scaffold sequences that were found to be surprisingly effective at improving the function and activity of the disclosed Cas12f polypeptides.

[0124] In an aspect, the scaffold sequence comprises a tracrRNA (trRNA) sequence. ThetracrRNA sequence may comprise a minimum tracrRNA sequence and, optionally, a tracrRNA extension sequence. The minimum tracrRNA sequence can have a length from about 7 nucleotides to about 100 nucleotides. For example, the minimum tracrRNA sequence can be from about 7 nucleotides (nt) to about 50 nt, from about 7 nt to about 40 nt, from about 7 nt to about 30 nt, from about 7 nt to about 25 nt, from about 7 nt to about 20 nt, from about 7 nt to about 15 nt, from about 8 nt to about 40 nt, from about 8 nt to about 30 nt, from about 8 nt to about 25 nt, from about 8 nt to about 20 nt, from about 8 nt to about 15 nt, from about 15 nt to about 100 nt, from about 15 nt to about 80 nt, from about 15 nt to about 50 nt, from about 15 nt to about 40 nt, from about 15 nt to about 30 nt or from about 15 nt to about 25 nt long. The tracrRNA sequence can be approximately 9 nucleotides in length. The minimum tracrRNA sequence can be approximately 12 nucleotides.

[0125] The minimum tracrRNA sequence can be at least about 60% identical to areference minimum tracrRNA (e.g., wildtype, tracrRNA from a Blautia species) sequence -50- 99975521.7Attorney Docket No: 124540-829360 over a stretch of at least 6, 7, or 8 contiguous nucleotides. For example, the minimum tracrRNA sequence can be at least about 65% identical, about 70% identical, about 75% identical, about 80% identical, about 85% identical, about 90% identical, about 95% identical, about 98% identical, about 99% identical or 100% identical to a reference minimum tracrRNA sequence over a stretch of at least 6, 7, or 8 contiguous nucleotides.

[0126] In some embodiments, a gRNA may comprise a tracrRNA extension sequence. AtracrRNA extension sequence can have a length from about 1 nucleotide to about 400 nucleotides. The tracrRNA extension sequence can have a length of more than 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, or 200 nucleotides. The tracrRNA extension sequence can have a length from about 20 to about 5000 or more nucleotides. The tracrRNA extension sequence can have a length of less than 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 nucleotides. The tracrRNA extension sequence can comprise less than 10 nucleotides in length. The tracrRNA extension sequence can be 10-30 nucleotides in length. The tracrRNA extension sequence can be 30-200 nucleotides in length.

[0127] The tracrRNA extension sequence can comprise a functional moiety (e.g., astability control sequence, ribozyme, endoribonuclease binding sequence). The functional moiety can comprise a transcriptional terminator segment (i.e., a transcription termination sequence). The functional moiety can have a total length from about 10 nucleotides (nt) to about 100 nucleotides, from about 10 nt to about 20 nt, from about 20 nt to about 30 nt, from about 30 nt to about 40 nt, from about 40 nt to about 50 nt, from about 50 nt to about 60 nt, from about 60 nt to about 70 nt, from about 70 nt to about 80 nt, from about 80 nt to about 90 nt, or from about 90 nt to about 100 nt, from about 15 nt to about 80 nt, from about 15 nt to about 50 nt, from about 15 nt to about 40 nt, from about 15 nt to about 30 nt, or from about 15 nt to about 25 nt.

[0128] In some embodiments, a tracrRNA may be a 5' tracrRNA. In some embodiments,a 5’ tracrRNA sequence can comprise a sequence with at least about 30%, about 40%, about 50%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, or 100% sequence identity to a reference tracrRNA sequence (e.g., a tracrRNA from S. pyogenes).

[0129] In some embodiments, a sgRNA may comprise a linker sequence with a lengthfrom about 3 nucleotides to about 100 nucleotides. In Jinek et al., supra, for example, a -51- 99975521.7Attorney Docket No: 124540-829360 simple 4 nucleotide "tetraloop" (-GAAA-) was used (Jinek et al., Science, 2012, 337(6096):816-821). An illustrative linker has a length from about 3 nucleotides (nt) to about 90 nt, from about 3 nt to about 80 nt, from about 3 nt to about 70 nt, from about 3 nt to about 60 nt, from about 3 nt to about 50 nt, from about 3 nt to about 40 nt, from about 3 nt to about 30 nt, from about 3 nt to about 20 nt, from about 3 nt to about 10 nt. For example, the linker can have a length from about 3 nt to about 5 nt, from about 5 nt to about 10 nt, from about 10 nt to about 15 nt, from about 15 nt to about 20 nt, from about 20 nt to about 25 nt, from about 25 nt to about 30 nt, from about 30 nt to about 35 nt, from about 35 nt to about 40 nt, from about 40 nt to about 50 nt, from about 50 nt to about 60 nt, from about 60 nt to about 70 nt, from about 70 nt to about 80 nt, from about 80 nt to about 90 nt, or from about 90 nt to about 100 nt. The linker of a single molecule guide nucleic acid can be between 4 and 40 nucleotides. The linker can be at least about 100, 500, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, or 7000 or more nucleotides. The linker can be at most about 100, 500, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, or 7000 or more nucleotides.

[0130] Linkers can comprise any of a variety of sequences, although in some examples thelinker will not comprise sequences that have extensive regions of homology with other portions of the guide RNA, which might cause intramolecular binding that could interfere with other functional regions of the guide. In Jinek et al., supra, a simple 4 nucleotide sequence -GAAA- was used (Jinek et al., Science, 2012, 337(6096):816-821), but numerous other sequences, including longer sequences can likewise be used.

[0131] The linker sequence can comprise a functional moiety. For example, the linkersequence can comprise one or more features, including an aptamer, a ribozyme, a protein- interacting hairpin, a protein binding site, a CRISPR array, an intron, or an exon. The linker sequence can comprise at least about 1, 2, 3, 4, or 5 or more functional moieties. In some examples, the linker sequence can comprise at most about 1, 2, 3, 4, or 5 or more functional moieties.

[0132] In various aspects, the scaffold sequence of the gRNAs provided herein may be 30to 200 nt in length, 40 to 200 nt in length, 50 to 200 nt in length, 60 to 200 nt in length, 70 nt to 200 nt in length, 80 nt to 200 nt in length, 90 to 200 nt in length, 100 to 200 nt in length, 110 to 200 nt in length, 120 to 200 nt in length, 130 to 200 nt in length, 140 to 200 nt in length, 150 to 200 nt in length, 160 to 200 nt in length, 180 to 200 nt in length, 30 to 190 nt in length, 40 to 190 nt in length, 30 to 190 nt in length, 40 to 190 nt in length, 50 to 190 nt in -52- 99975521.7Attorney Docket No: 124540-829360 length, 60 to 190 nt in length, 70 nt to 190 nt in length, 80 nt to 190 nt in length, 90 to 190 nt in length, 100 to 190 nt in length, 110 to 190 nt in length, 120 to 190 nt in length, 130 to 190 nt in length, 140 to 190 nt in length, 150 to 190 nt in length, 160 to 190 nt in length, or 180 to 190 nt in length. In an aspect, the scaffold sequence may be 70 nt to 200 nt in length. In an aspect, the scaffold sequence may be 140 to 190 nt in length.

[0133] In various aspects, the gRNAs comprise a scaffold sequence having at least 70%,at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to any one of SEQ ID NOs: 438-515 and SEQ ID NOs: 673-802. In various aspects, the gRNAs comprise a scaffold sequence comprising or consisting of any one of SEQ ID NOs: 438-515 and SEQ ID NOs: 673-802. In various aspects, the gRNAs comprise a scaffold sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to any one of SEQ ID NOs: 440, 512, 514, 515, which correspond to Scaffolds S003, S075, S077, and S078 as described further in the Examples below. In various aspects, the gRNAs comprise a scaffold sequence comprising or consisting of any one of SEQ ID NOs: 440, 512, 514, and 515. In various aspects, the gRNAs comprise a scaffold sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to any one of SEQ ID NOs: 673, 675, 735, 737, and 738, which correspond to Scaffolds BS-S001, BS-S003, BS-S075, BS-S077 and BS-S078 as described further in the Examples below. In various aspects, the gRNAs comprise a scaffold sequence comprising or consisting of any one of SEQ ID NOs: 673, 675, 735, 737, and 738.

[0134] Exemplary full-length gRNAs comprising both a spacer and gRNA sequence fortargeting exemplary genes or regulatory targets (HBB, HBP1, MeCP2, PGK promoter, and GC2) are provided herein as SEQ ID NOs: 552-587 as shown in the Tables 8A-8F below. Table 8A – gRNAs targeting Tet0 Target gRNA name gRNA (SCAFFOLD + spacer) SEQ ID NO: Gene Tet0 g004 GACGGTAACTTAAAGTGCCGAAGGCTG 552 AGGAGATGGATTAAATATATAAGGTTT TGACCAACTATATACCATCATCACTCGG TAAGGGTTAATCCTAACAATGTGTGAC CGTTAGGCGTTCCAAAGAAATGGAATG TTAATctccctatcagtgatagaga -53- 99975521.7Attorney Docket No: 124540-829360 Tet0 g030 GACGGTAACTTAAAGTGCCGAAGGCTG 553 AGGAGATGGATTAAATATATAAGGTTT TGACCAACTATATACCATCATCACTCGG TAAGGGTTAATCCTAACAATGTGTGAC CGTTAGGCGTTCCAAAGAAATGGAATG TTAATtctctatcactgatagggag Table 8B – gRNAs targeting HBB Target gRNA gRNA (SCAFFOLD + spacer) SEQ ID NO: Gene name HBB g051 GACGGTAACTTAAAGTGCCGAAGGCTGAGGAG 554 ATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAA CAATGTGTGACCGTTAGGCGTTCCAAAGAAATG GAATGTTAATgtgccagaagagccaaggac HBB g052 GACGGTAACTTAAAGTGCCGAAGGCTGAGGAG 555 ATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAA CAATGTGTGACCGTTAGGCGTTCCAAAGAAATG GAATGTTAATgaagagccaaggacaggta HBB g053 GACGGTAACTTAAAGTGCCGAAGGCTGAGGAG 556 ATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAA CAATGTGTGACCGTTAGGCGTTCCAAAGAAATG GAATGTTAATgggctgggcataaaagtca HBB g054 GACGGTAACTTAAAGTGCCGAAGGCTGAGGAG 557 ATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAA CAATGTGTGACCGTTAGGCGTTCCAAAGAAATG GAATGTTAATgtggagccacaccctagggt HBB g055 GACGGTAACTTAAAGTGCCGAAGGCTGAGGAG 558 ATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAA CAATGTGTGACCGTTAGGCGTTCCAAAGAAATG GAATGTTAATatctactcccaggagcaggg HBB g056 GACGGTAACTTAAAGTGCCGAAGGCTGAGGAG 559 ATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAA CAATGTGTGACCGTTAGGCGTTCCAAAGAAATG GAATGTTAATttactgccctgtggggca HBB g081 GACGGTAACTTAAAGTGCCGAAGGCTGAGGAG 560 ATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAA -54- 99975521.7Attorney Docket No: 124540-829360 CAATGTGTGACCGTTAGGCGTTCCAAAGAAATG GAATGTTAATgtccttggctcttctggcac HBB g082 GACGGTAACTTAAAGTGCCGAAGGCTGAGGAG 561 ATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAA CAATGTGTGACCGTTAGGCGTTCCAAAGAAATG GAATGTTAATtacctgtccttggctcttc HBB g083 GACGGTAACTTAAAGTGCCGAAGGCTGAGGAG 562 ATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAA CAATGTGTGACCGTTAGGCGTTCCAAAGAAATG GAATGTTAATtgacttttatgcccagccc HBB g084 GACGGTAACTTAAAGTGCCGAAGGCTGAGGAG 563 ATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAA CAATGTGTGACCGTTAGGCGTTCCAAAGAAATG GAATGTTAATaccctagggtgtggctccac HBB g085 GACGGTAACTTAAAGTGCCGAAGGCTGAGGAG 564 ATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAA CAATGTGTGACCGTTAGGCGTTCCAAAGAAATG GAATGTTAATccctgctcctgggagtagat HBB g086 GACGGTAACTTAAAGTGCCGAAGGCTGAGGAG 565 ATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAA CAATGTGTGACCGTTAGGCGTTCCAAAGAAATG GAATGTTAATtgccccacagggcagtaa Table 8C - gRNAs targeting HBP1 Target gRNA gRNA (SCAFFOLD + spacer) SEQ ID NO: Gene name HBP1 g041 GACGGTAACTTAAAGTGCCGAAGGCTGAGGAG 566 ATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAA CAATGTGTGACCGTTAGGCGTTCCAAAGAAATG GAATGTTAATggctaaactccacccatgggt HBP1 g042 GACGGTAACTTAAAGTGCCGAAGGCTGAGGAG 567 ATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAA CAATGTGTGACCGTTAGGCGTTCCAAAGAAATG GAATGTTAATatagtcttagagtatccagtg HBP1 g043 GACGGTAACTTAAAGTGCCGAAGGCTGAGGAG 568 ATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAA -55- 99975521.7Attorney Docket No: 124540-829360 CAATGTGTGACCGTTAGGCGTTCCAAAGAAATG GAATGTTAATgtgaggccaggggccggcggc HBP1 g044 GACGGTAACTTAAAGTGCCGAAGGCTGAGGAG 569 ATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAA CAATGTGTGACCGTTAGGCGTTCCAAAGAAATG GAATGTTAATtgggtcatttcacagagg HBP1 g071 GACGGTAACTTAAAGTGCCGAAGGCTGAGGAG 570 ATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAA CAATGTGTGACCGTTAGGCGTTCCAAAGAAATG GAATGTTAATacccatgggtggagtttagcc HBP1 g072 GACGGTAACTTAAAGTGCCGAAGGCTGAGGAG 571 ATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAA CAATGTGTGACCGTTAGGCGTTCCAAAGAAATG GAATGTTAATcactggatactctaagactat HBP1 g073 GACGGTAACTTAAAGTGCCGAAGGCTGAGGAG 572 ATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAA CAATGTGTGACCGTTAGGCGTTCCAAAGAAATG GAATGTTAATgccgccggcccctggcctcac HBP1 g074 GACGGTAACTTAAAGTGCCGAAGGCTGAGGAG 573 ATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAA CAATGTGTGACCGTTAGGCGTTCCAAAGAAATG GAATGTTAATcctctgtgaaatgaccca Table 8D– gRNAs targeting MeCP2 Target gRNA gRNA (SCAFFOLD + spacer) SEQ ID NO: Gene name MECP2 g091 GACGGTAACTTAAAGTGCCGAAGGCTGAGGAG 574 ATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAA CAATGTGTGACCGTTAGGCGTTCCAAAGAAATG GAATGTTAATgagctctcacccccatctct MECP2 g101 GACGGTAACTTAAAGTGCCGAAGGCTGAGGAG 575 ATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAA CAATGTGTGACCGTTAGGCGTTCCAAAGAAATG GAATGTTAATtttccctggccgaaatggac -56- 99975521.7Attorney Docket No: 124540-829360 MECP2 g113 GACGGTAACTTAAAGTGCCGAAGGCTGAGGAG 576 ATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAA CAATGTGTGACCGTTAGGCGTTCCAAAGAAATG GAATGTTAATcctacttgttcctgctagat MECP2 g114 GACGGTAACTTAAAGTGCCGAAGGCTGAGGAG 577 ATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAA CAATGTGTGACCGTTAGGCGTTCCAAAGAAATG GAATGTTAATtctagcaggaacaagtaggt MECP2 g115 GACGGTAACTTAAAGTGCCGAAGGCTGAGGAG 578 ATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAA CAATGTGTGACCGTTAGGCGTTCCAAAGAAATG GAATGTTAATcagccctctctccgagagg MECP2 g116 GACGGTAACTTAAAGTGCCGAAGGCTGAGGAG 579 ATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAA CAATGTGTGACCGTTAGGCGTTCCAAAGAAATG GAATGTTAATtcacagccaatgacgggc MECP2 g117 GACGGTAACTTAAAGTGCCGAAGGCTGAGGAG 580 ATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAA CAATGTGTGACCGTTAGGCGTTCCAAAGAAATG GAATGTTAATtttcggccagggaaaagg Table 8E - gRNAs targeting PGK promoter Target gRNA gRNA (SCAFFOLD + spacer) SEQ ID NO: Gene name PGK g171 GACGGTAACTTAAAGTGCCGAAGGCTGAGGAG 581 ATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAA CAATGTGTGACCGTTAGGCGTTCCAAAGAAATG GAATGTTAATcgttgaccgaatcaccgacc PGK g172 GACGGTAACTTAAAGTGCCGAAGGCTGAGGAG 582 ATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAA CAATGTGTGACCGTTAGGCGTTCCAAAGAAATG GAATGTTAATccggagcgcacgtcggcagt PGK g173 GACGGTAACTTAAAGTGCCGAAGGCTGAGGAG 583 ATGGATTAAATATATAAGGTTTTGACCAACTAT -57- 99975521.7Attorney Docket No: 124540-829360 ATACCATCATCACTCGGTAAGGGTTAATCCTAA CAATGTGTGACCGTTAGGCGTTCCAAAGAAATG GAATGTTAATgttcctgcccgcgcggtgtt Table 8F - gRNAs targeting GC2 ((GGGCC)n Repeats) Target gRNA gRNA (SCAFFOLD + spacer) SEQ ID NO: Gene name GC2 g026 GACGGTAACTTAAAGTGCCGAAGGCTGAGGAG 584 ATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAA CAATGTGTGACCGTTAGGCGTTCCAAAGAAATG GAATGTTAATgggccggggccggggccggg GC2 g027 GACGGTAACTTAAAGTGCCGAAGGCTGAGGAG 585 ATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAA CAATGTGTGACCGTTAGGCGTTCCAAAGAAATG GAATGTTAATcggccccggccccggccccg GC2 g028 GACGGTAACTTAAAGTGCCGAAGGCTGAGGAG 586 ATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAA CAATGTGTGACCGTTAGGCGTTCCAAAGAAATG GAATGTTAATggccccggccccggccccgg GC2 g029 GACGGTAACTTAAAGTGCCGAAGGCTGAGGAG 587 ATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAA CAATGTGTGACCGTTAGGCGTTCCAAAGAAATG GAATGTTAATgccccggccccggccccggc

[0135] The sgRNAs provided herein may be chemically modified. In some embodiments,a chemically modified gRNA is a gRNA that comprises at least one nucleotide with a chemical modification, e.g., a 2′-O-methyl sugar modification. In some embodiments, a chemically modified gRNA comprises a modified nucleic acid backbone. In some embodiments, a chemically modified gRNA comprises a 2'-O-methyl-phosphorothioate residue. In some embodiments, chemical modifications enhance stability, reduce the likelihood or degree of innate immune response, and / or enhance other attributes, as described in the art.

[0136] In some embodiments, a modified gRNA may comprise a modified backbone, forexample, phosphorothioates, phosphotriesters, morpholinos, methyl phosphonates, short chain alkyl or cycloalkyl intersugar linkages or short chain heteroatomic or heterocyclic -58- 99975521.7Attorney Docket No: 124540-829360 intersugar linkages. Cyclohexenyl nucleic acid oligonucleotide mimetics are described in Wang et al., J. Am. Chem. Soc., 2000, 122: 8595-8602.

[0137] In some embodiments, a modified gRNA may comprise one or more substitutedsugar moieties, e.g., one of the following at the 2' position: OH, SH, SCH3, F, OCN, OCH3, OCH3 O(CH2)n CH3, O(CH2)n NH2, or O(CH2)n CH3, where n is from 1 to about 10; C1 to C10 lower alkyl, alkoxy, substituted lower alkyl, alkaryl or aralkyl; Cl; Br; CN; CF3; OCF3; O-, S-, or N-alkyl; O-, S-, or N-alkenyl; SOCH3; SO2 CH3; ONO2; NO2; N3; NH2; heterocycloalkyl; heterocycloalkaryl; aminoalkylamino; polyalkylamino; substituted silyl; an RNA cleaving group; a reporter group; an intercalator; 2'-O-(2-methoxyethyl); 2'-methoxy (2'-O-CH3); 2'-propoxy (2'-OCH2 CH2CH3); and 2'-fluoro (2'-F). Similar modifications may also be made at other positions on the gRNA, particularly the 3' position of the sugar on the 3' terminal nucleotide and the 5' position of 5' terminal nucleotide. In some examples, both a sugar and an internucleoside linkage, i.e., the backbone, of the nucleotide units can be replaced with novel groups.

[0138] Guide RNAs can also include, additionally or alternatively, nucleobase (oftenreferred to in the art simply as "base") modifications or substitutions. As used herein, "unmodified" or "natural" nucleobases include adenine (A), guanine (G), thymine (T), cytosine (C), and uracil (U). Modified nucleobases include nucleobases found only infrequently or transiently in natural nucleic acids, e.g., hypoxanthine, 6-methyladenine, 5- Me pyrimidines, particularly 5-methylcytosine (also referred to as 5-methyl-2' deoxycytosine and often referred to in the art as 5-Me-C), 5-hydroxymethylcytosine (HMC), glycosyl HMC and gentobiosyl HMC, as well as synthetic nucleobases, e.g., 2-aminoadenine, 2- (methylamino)adenine, 2-(imidazolylalkyl)adenine, 2-(aminoalklyamino)adenine or other heterosubstituted alkyladenines, 2-thiouracil, 2-thiothymine, 5-bromouracil, 5- hydroxymethyluracil, 8-azaguanine, 7-deazaguanine, N6 (6-aminohexyl)adenine, and 2,6- diaminopurine. Kornberg, A., DNA Replication, W. H. Freeman & Co., San Francisco, pp75- 77, 1980; Gebeyehu et al., Nucl. Acids Res.1997, 15:4513. A "universal" base known in the art, e.g., inosine, can also be included.5-Me-C substitutions have been shown to increase nucleic acid duplex stability by 0.6-1.2 ºC. (Sanghvi, Y. S., in Crooke, S. T. and Lebleu, B., eds., Antisense Research and Applications, CRC Press, Boca Raton, 1993, pp.276-278) and are aspects of base substitutions.

[0139] Modified nucleobases can comprise other synthetic and natural nucleobases, suchas 5-methylcytosine (5-me-C), 5-hydroxymethyl cytosine, xanthine, hypoxanthine, 2- -59- 99975521.7Attorney Docket No: 124540-829360 aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2- thiocytosine, 5-halouracil and cytosine, 5-propynyl uracil and cytosine, 6-azo uracil, cytosine and thymine, 5-uracil (pseudo-uracil), 4-thiouracil, 8-halo, 8-amino, 8-thiol, 8- thioalkyl, 8- hydroxyl and other 8-substituted adenines and guanines, 5-halo particularly 5- bromo, 5- trifluoromethyl and other 5-substituted uracils and cytosines, 7-methylquanine and 7- methyladenine, 8-azaguanine and 8-azaadenine, 7-deazaguanine and 7-deazaadenine, and 3- deazaguanine and 3-deazaadenine.

[0140] Complexes of a Genome-targeting Nucleic Acid and an Endonuclease

[0141] As noted above, the Cas12f polypeptides of the present disclosure can interact withthe sgRNAs provided herein to form a complex which interacts with a target polynucleotide. Accordingly, and as described further below, methods of using the Cas12f polypeptide to alter a nucleic acid require delivery of both the Cas12f polypeptide and the sgRNA to the nucleic acid. In some aspects, the Cas12f polypeptide and the sgRNA are pre-complexed prior to delivery to the target nucleic acid. Such pre-complexed material is known as a ribonucleoprotein particle (RNP). The Cas12f polypeptide in the RNP can be flanked at the N-terminus, the C-terminus, or both the N-terminus and C-terminus by one or more nuclear localization signals (NLSs). For example, a Cas12f polypeptide can be flanked by two NLSs, one NLS located at the N-terminus and the second NLS located at the C-terminus. The NLS can be any NLS known in the art, such as a SV40 NLS. The molar ratio of sgRNA to Cas12f polypeptide in the RNP can range from about 1:1 to about 10:1. For example, the molar ratio of sgRNA to the Cas12f polypeptide in the RNP can be 3:1. Nucleic Acids Encoding System Components

[0142] The present disclosure provides a nucleic acid comprising a nucleotide sequenceencoding an sgRNA of the disclosure, a Cas12f polypeptide of the disclosure, and / or any nucleic acid or proteinaceous molecule necessary to carry out the aspects of the methods of the disclosure. The encoding nucleic acids can be RNA, DNA, or a combination thereof.

[0143] A nucleic acid or nucleotide sequence encoding a Cas12f polypeptide and / or ansgRNA may be provided in an expression construct or expression vector. The phrase "expression vector" generally refers to a nucleotide sequence that encodes the Cas12f polypeptide and / or an sgRNA, together with other regulatory components / sequences. These expression vectors typically include at least suitable promoter sequences and optionally, -60- 99975521.7Attorney Docket No: 124540-829360 transcription termination signals. An additional factor necessary or helpful in promoting or inducing expression can also be used as described herein. A nucleic acid or DNA or nucleotide sequence encoding a Cas12f polypeptide and / or an sgRNA of the present disclosure is incorporated into a DNA construct capable of introduction into and expression in an in vitro cell culture. Specifically, a DNA construct is suitable for replication in a prokaryotic host, such as bacteria, e.g., E. coli, or can be introduced into a cultured mammalian, plant, insect, (e.g., Sf9), yeast, fungi or other eukaryotic cell lines.

[0144] Expression constructs disclosed herein could be prepared using recombinanttechniques in which nucleotide sequences encoding a Cas12f polypeptide and / or an sgRNA are expressed in a suitable cell, e.g. cultured cells or cells of a multicellular organism, such as described in Ausubel et al., "Current Protocols in Molecular Biology", Greene Publishing and Wiley-Interscience, New York (1987) and in Sambrook and Russell (2001, supra); both of which are incorporated herein by reference in their entirety. Also see, Kunkel (1985) Proc. Natl. Acad. Sci.82:488 (describing site directed mutagenesis) and Roberts et al. (1987) Nature 328:731-734 or Wells, J.A., et al. (1985) Gene 34: 315 (describing cassette mutagenesis).

[0145] The expression constructs, vectors and related nucleic acids of the presentdisclosure may comprise one or more regulatory sequences. The term "regulatory sequence" is intended to include, for example, promoters, enhancers and other expression control elements (e.g., polyadenylation signals). Such regulatory sequences are well known in the art and are described, for example, in Goeddel; Gene Expression Technology: Methods in Enzymology, 1990, 185, Academic Press, San Diego, CA. Regulatory sequences include those that direct constitutive expression of a nucleotide sequence in many types of host cells, and those that direct expression of the nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences). It will be appreciated by those skilled in the art that the design of the expression vector can depend on such factors as the choice of the target cell, the level of expression desired, and the like.

[0146] A transcriptional regulatory sequence typically includes a heterologous enhancer orpromoter that is recognized by the host. The selection of an appropriate promoter sequence generally depends upon the host cell selected for the expression of a DNA segment. Examples of suitable promoter sequences include prokaryotic, and eukaryotic promoters well known in the art (see, e.g. Sambrook and Russell, 2001, supra). Non-limiting examples of suitable eukaryotic promoters (i.e., promoters functional in a eukaryotic cell) include those -61- 99975521.7Attorney Docket No: 124540-829360 from cytomegalovirus (CMV) immediate early, herpes simplex virus (HSV) thymidine kinase, early and late SV40, long terminal repeats (LTRs) from retrovirus, human elongation factor-1 α promoter (EF1α), chicken beta-actin promoter (CBA), ubiquitin C promoter (UBC), a hybrid construct comprising the cytomegalovirus enhancer fused to the chicken beta-actin promoter (CAG), a hybrid construct comprising the cytomegalovirus enhancer fused to the promoter, the first exon, and the first intron of chicken beta-actin gene (CAG or CAGGS), murine stem cell virus promoter (MSCV), phosphoglycerate kinase-1 locus promoter (PGK), and mouse metallothionein-I promoter. A promoter can be an inducible promoter (e.g., a heat shock promoter, tetracycline-regulated promoter, steroid-regulated promoter, metal-regulated promoter, estrogen receptor-regulated promoter, etc.). The promoter can be a constitutive promoter (e.g., CMV promoter, UBC promoter, CAG promoter). In some cases, the promoter can be a spatially restricted and / or temporally restricted promoter (e.g., a tissue specific promoter, a cell type specific promoter, etc.). Other promoters such as such as the trp, lac and phage promoters, tRNA promoters and glycolytic enzyme promoters are known and available (see, e.g. Sambrook and Russell, 2001, supra). In some cases, the promoter can be an inducible form (e.g. TetO promoter) and it is only expressed upon the introduction of inducing reagent, which is typically small molecules (e.g. doxycycline).

[0147] An expression vector includes the replication system and transcriptional andtranslational regulatory sequences together with the insertion site for the polypeptide encoding segment can be employed. In most cases, the replication system is only functional in the cell that is used to make the vector (bacterial cell as E. Coli). Most plasmids and vectors do not replicate in the cells infected with the vector. Examples of workable combinations of cell lines and expression vectors are described in Sambrook and Russell (2001, supra) and in Metzger et al. (1988) Nature 334: 31-36. For example, suitable expression vectors can be expressed in, yeast, e.g. S. cerevisiae, e.g., insect cells, e.g., Sf9 cells, mammalian cells, e.g., CHO cells and bacterial cells, e.g., E. coli. A cell may thus be a prokaryotic or eukaryotic host cell. A cell may be a cell that is suitable for culture in liquid or on solid media.

[0148] In some aspects, the current disclosure also encompasses vectors that facilitatetransfer of nucleic acids encoding the disclosed receptors between cells, such as, but not limited to, plasmids, transposons, cosmids, chromosomes, artificial chromosomes, viruses, virions, and the like. A vector may also be a chemical vector, such as a lipid complex or -62- 99975521.7Attorney Docket No: 124540-829360 naked DNA. In some aspects the vector may be a viral vector. A viral vector may comprise an expression construct as described herein. In some aspects, the vector is a gene therapy vector. A gene therapy vector is a vector that is suitable for gene therapy. Vectors that are suitable for gene therapy are described in Anderson 1998, Nature 392: 25-30; Walther and Stein, 2000, Drugs 60: 249-71; Kay et al., 2001, Nat. Med.7: 33-40; Russell, 2000, J. Gen. Virol.81: 2573-604; Amado and Chen, 1999, Science 285: 674-6; Federico, 1999, Curr. Opin. Biotechnol.10: 448-53; Vigna and Naldini, 2000, J. Gene Med.2: 308-16; Marin et al., 1997, Mol. Med. Today 3: 396-403; Peng and Russell, 1999, Curr. Opin. Biotechnol.10: 454-7; Sommerfelt, 1999, J. Gen. Virol.80: 3049-64; Reiser, 2000, Gene Ther.7: 910-3; and references cited therein.

[0149] One type of vector contemplated by the instant disclosure is a “plasmid” whichrefers to a circular double-stranded DNA loop into which additional nucleic acid segments can be ligated. Another type of vector is a viral vector, wherein additional nucleic acid segments can be ligated into the viral genome. Certain vectors are capable of autonomous replication in a host cell into which they are introduced (e.g., bacterial vectors having a bacterial origin of replication and episomal mammalian vectors). Other vectors (e.g., non- episomal mammalian vectors) are integrated into the genome of a host cell upon introduction into the host cell, and thereby are replicated along with the host genome.

[0150] Expression vectors contemplated include, but are not limited to, viral vectors basedon vaccinia virus, poliovirus, adenovirus, adeno-associated virus, SV40, herpes simplex virus, human immunodeficiency virus, retrovirus (e.g., Murine Leukemia Virus, spleen necrosis virus, and vectors derived from retroviruses such as Rous Sarcoma Virus, Harvey Sarcoma Virus, avian leukosis virus, a lentivirus, human immunodeficiency virus, myeloproliferative sarcoma virus, and mammary tumor virus) and other recombinant vectors. Other vectors contemplated for eukaryotic target cells include, but are not limited to, the vectors pXT1, pSG5, pSVK3, pBPV, pMSG, and pSVLSV40 (Pharmacia). Other vectors can be used so long as they are compatible with the host cell. In some aspects, the vectors contemplated herein comprise an adenoviral vector and / or a lentiviral vector.

[0151] In some examples, a vector can comprise one or more transcription and / ortranslation control elements. Depending on the host / vector system utilized, any of a number of suitable transcription and translation control elements, including constitutive and inducible promoters, transcription enhancer elements, transcription terminators, etc. can be used in the -63- 99975521.7Attorney Docket No: 124540-829360 expression vector. The vector can be a self-inactivating vector that either inactivates the viral sequences or the components of the CRISPR machinery or other elements.

[0152] An advantageous feature of the Cas12f polypeptides provided herein is theircompact size. The small size of BsCas12f or dBsCas12f makes packaging using a single AAV vector feasible, as these Cas proteins are only 1.2 kb (399 amino acids) in size. AAV vectors have a cargo limit of approximately 4.7 kb. The other epigenetic machinery components excluding the Cas protein (1.6-4.1kb) typically require around 3 kb of space (including the promoter (0.5-2kb), demethylase / methylase (1.6-2.2kb), linker (0.1kb), polyadenine tail (0.05-0.5kb) and gRNA-expressing cassette (0.15-0.4b)). This leaves 1-1.7 kb available for the Cas protein when using a single AAV vector. Most existing Cas proteins like SpCas9 (4.1 kb), SaCas9 (3.1kb) and CasMINI (1.6kb) exceed this size limit, but the compact BsCas12f or dBsCas12f proteins at 1.2 kb can be readily packaged. Their small size makes BsCas12f or dBsCas12f well-suited for delivery via a single AAV vector along with other epigenetic modulators, including but not limited to VP64 (0.15kb), VPR (1.5 kb) for CRISPR activation (CRISPRa), KRAB (0.5 kb) for CRISPR inhibition (CRISPRi), and base editors like CBE and ABE.

[0153] Accordingly, in some aspects, a vector comprising an expression constructcomprising a polynucleotide encoding a Cas12f polypeptide and an expression construct comprising a polynucleotide encoding one or more than one sgRNA is provided. The vector may, in certain aspects, be an adenoviral or a lentiviral vector.

[0154] Exemplary sequences of vectors provided by the present disclosure are discussedfurther in the Examples below and are provided herein as SEQ ID NOs 588-605. Gene Editing Systems

[0155] Consistent with the foregoing or related aspects, a gene editing system is providedherein comprising (1) any Cas12f polypeptide provided herein, or an encoding polynucleotide thereof and (2) a single molecule guide RNA (sgRNA) or an encoding polynucleotide thereof.

[0156] The gene editing system may, in some aspects, comprise a Cas12f polypeptide thatis catalytically dead so that it no longer has any nuclease activity (e.g., a Cas12f polypeptide comprising one or more substitutions selected from D216A and D390A according to SEQ ID NO: 1). In some aspects, the gene editing system may comprise a Cas12f polypeptide that is catalytically active and has nuclease activity. In all aspects herein, the gene editing system -64- 99975521.7Attorney Docket No: 124540-829360 comprises a Cas12f polypeptide comprising one or more substitutions improve its targeting and / or editing efficiency relative to a wildtype Cas12f polypeptide (e.g., relative to a Cas12f polypeptide comprising SEQ ID NO: 1).

[0157] In certain exemplary aspects, the gene editing systems provided herein maycomprise a Cas12f polypeptide having an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to any one of SEQ ID NOs: 16-391. For instance, the gene editing systems provided herein may comprise a Cas12f polypeptide having an amino acid sequence of any one of SEQ ID NOs: 16-391. In an aspect, the gene editing systems provided herein may comprise a Cas12f polypeptide having an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to any one of SEQ ID NOs 329-338. For instance, the gene editing systems provided herein may comprise a Cas12f polypeptide having an amino acid sequence of any one of SEQ ID NOs: 329-338. In some aspects, the gene editing system may comprise a fusion protein of a Cas12f polypeptide linked to a functional polypeptide (e.g., a VPR transcriptional activator, a Tet polypeptide or a DNMT3A polypeptide). In some aspects, the gene editing system comprises a fusion protein having an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to any one of SEQ ID NOs: 416, 418, 420, 422, 424, 426, 428, 430, 432, 434, and 436. In some aspects, the gene editing system comprises a fusion protein having an amino acid sequence comprising or consisting of any one of SEQ ID NOs: 416, 418, 420, 422, 424, 426, 428, 430, 432, 434, and 436.

[0158] In certain exemplary aspects, the gene editing systems provided herein maycomprise an sgRNA comprising a scaffold sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to any one of SEQ ID NOs: 438-515. For instance, in some aspects, the gene editing systems provided herein may comprise an sgRNA comprising a scaffold sequence of any one of SEQ ID NOs: 438-515. In certain exemplary aspects, the gene editing systems provided herein may comprise an sgRNA comprising a scaffold sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to any one of SEQ ID NOs: 440, 512, 514, and 515. For instance, in some aspects, the gene editing systems provided herein may -65- 99975521.7Attorney Docket No: 124540-829360 comprise an sgRNA comprising a scaffold sequence of any one of SEQ ID NOs: 440, 512, 514, and 515.

[0159] In further aspects, the gene editing system may further comprise a transgene. In anaspect, the transgene may be provided as a nucleic acid that is inserted into a site of a double stranded break (i.e., induced by a Cas12f polypeptide herein) using homologous recombination (HR). The transgene is therefore synonymous with a donor nucleic acid. The transgene may include a nucleic acid encoding a protein of interest (i.e., a therapeutic protein of interest) or may be designed to otherwise correct a defect (i.e., mismatch, mutation, deletion, insertion) in a target gene of interest. In an aspect, the transgene may be packaged or provided in an episome.

[0160] The gene editing systems provided herein may be provided in one or more nucleicacids encoding the Cas12f polypeptide and / or the sgRNA. In various aspects, as described above, the nucleic acids may be packaged into a vector (e.g., an adenoviral or lentiviral vector). In preferred embodiments, the entire gene editing system (comprising nucleic acids encoding the Cas12f polypeptide and the sgRNA, along with associated regulatory sequences to promote their expression) are packaged into a single vector (e.g., an adenoviral vector).

[0161] The gene editing system provided herein may be formulated for administrationand / or delivery to a cell (in vitro) or to a subject (in vivo). Suitable formulations for such applications are described further below. II. Formulations and Administrations

[0162] Guide RNAs, polynucleotides, e.g., polynucleotides that encode a Cas12fpolypeptide, and Cas12f polypeptides as described herein may be formulated and delivered to a cell and / or a subject in any manner known in the art.

[0163] Introduction of the complexes, polypeptides, and nucleic acids of the disclosureinto cells can occur by viral or bacteriophage infection, transfection, conjugation, protoplast fusion, lipofection, electroporation, nucleofection, calcium phosphate precipitation, polyethyleneimine (PEI)-mediated transfection, DEAE-dextran mediated transfection, liposome-mediated transfection, particle gun technology, calcium phosphate precipitation, direct micro-injection, nanoparticle-mediated nucleic acid delivery, and the like.

[0164] In addition to the foregoing, in vivo administration of the guide RNAs and / orpolynucleotides and / or polypeptides of the disclosure to a subject may comprise formulating the compositions provided above with one or more pharmaceutically acceptable excipients -66- 99975521.7Attorney Docket No: 124540-829360 such as carriers, solvents, stabilizers, adjuvants, diluents, etc., depending upon the particular mode of administration and dosage form. Guide RNAs and / or polynucleotides compositions can be formulated to achieve a physiologically compatible pH and range from a pH of about 3 to a pH of about 11, about pH 3 to about pH 7, depending on the formulation and route of administration. In some cases, the pH can be adjusted to a range from about pH 5.0 to about pH 8. In some cases, the compositions can comprise a therapeutically effective amount of at least one compound as described herein, together with one or more pharmaceutically acceptable excipients. Optionally, the compositions can comprise a combination of the compounds described herein or can include a second active ingredient useful in the treatment or prevention of bacterial growth (for example and without limitation, anti-bacterial or anti- microbial agents), or can include a combination of reagents of the present disclosure.

[0165] Suitable excipients include, for example, carrier molecules that include large,slowly metabolized macromolecules such as proteins, polysaccharides, polylactic acids, polyglycolic acids, polymeric amino acids, amino acid copolymers, and inactive virus particles. Other exemplary excipients can include antioxidants (for example and without limitation, ascorbic acid), chelating agents (for example and without limitation, EDTA), carbohydrates (for example and without limitation, dextrin, hydroxyalkylcellulose, and hydroxyalkylmethylcellulose), stearic acid, liquids (for example and without limitation, oils, water, saline, glycerol and ethanol), wetting or emulsifying agents, pH buffering substances, and the like.

[0166] Guide RNA polynucleotides (RNA or DNA) and / or polynucleotide(s) (RNA orDNA) encoding the Cas12f polypeptides of the present disclosure can be delivered by viral or non-viral delivery vehicles known in the art. Alternatively, the Cas12f polypeptide(s) can be delivered by viral or non-viral delivery vehicles known in the art, such as electroporation or lipid nanoparticles. In further alternative aspects, the Cas12f polypeptides can be delivered as one or more polypeptides, either alone or pre-complexed with one or more guide RNAs.

[0167] Polynucleotides can be delivered by non-viral delivery vehicles including, but notlimited to, nanoparticles, liposomes, ribonucleoproteins, positively charged peptides, small molecule RNA-conjugates, aptamer-RNA chimeras, and RNA-fusion protein complexes. Some exemplary non-viral delivery vehicles are described in Peer and Lieberman, Gene Therapy, 2011, 18: 1127–1133 (which focuses on non-viral delivery vehicles for siRNA that are also useful for delivery of other polynucleotides). -67- 99975521.7Attorney Docket No: 124540-829360

[0168] For polynucleotides of the disclosure, the formulation may be selected from any ofthose taught, for example, in International Application PCT / US2012 / 069610, incorporated herein by reference in its entirety.

[0169] Polynucleotides, such as guide RNA, sgRNA, and mRNA encoding a Cas12fpolypeptide provided herein, as well as vectors comprising them, may be delivered to a cell or a subject by a lipid nanoparticle (LNP).

[0170] A LNP refers to any particle having a diameter of less than 1000 nm, 500 nm, 250nm, 200 nm, 150 nm, 100 nm, 75 nm, 50 nm, or 25 nm. Alternatively, a nanoparticle may range in size from 1-1000 nm, 1-500 nm, 1-250 nm, 25-200 nm, 25-100 nm, 35-75 nm, or 25- 60 nm.

[0171] LNPs may be made from cationic, anionic, or neutral lipids. Neutral lipids, such asthe fusogenic phospholipid DOPE or the membrane component cholesterol, may be included in LNPs as 'helper lipids' to enhance transfection activity and nanoparticle stability. Limitations of cationic lipids include low efficacy owing to poor stability and rapid clearance, as well as the generation of inflammatory or anti-inflammatory responses.

[0172] LNPs may also be comprised of hydrophobic lipids, hydrophilic lipids, or bothhydrophobic and hydrophilic lipids.

[0173] Any lipid or combination of lipids that are known in the art can be used to producean LNP. Examples of lipids used to produce LNPs are: DOTMA, DOSPA, DOTAP, DMRIE, DC-cholesterol, DOTAP–cholesterol, GAP-DMORIE–DPyPE, and GL67A–DOPE–DMPE– polyethylene glycol (PEG). Examples of cationic lipids are: 98N12-5, C12-200, DLin-KC2- DMA (KC2), DLin-MC3-DMA (MC3), XTC, MD1, and 7C1. Examples of neutral lipids are: DPSC, DPPC, POPC, DOPE, and SM. Examples of PEG-modified lipids are PEG-DMG, PEG-CerC14, and PEG-CerC20.

[0174] The lipids can be combined in any number of molar ratios to produce an LNP. Inaddition, the polynucleotide(s) can be combined with lipid(s) in a wide range of molar ratios to produce an LNP.

[0175] A recombinant adeno-associated virus (AAV) vector can be used for delivery.Techniques to produce rAAV particles, in which an AAV genome to be packaged that includes the polynucleotide to be delivered, rep and cap genes, and helper virus functions are provided to a cell are standard in the art. Production of rAAV typically requires that the following components are present within a single cell (denoted herein as a packaging cell): a -68- 99975521.7Attorney Docket No: 124540-829360 rAAV genome, AAV rep and cap genes separate from (i.e., not in) the rAAV genome, and helper virus functions. The AAV rep and cap genes may be from any AAV serotype for which recombinant virus can be derived and may be from a different AAV serotype than the rAAV genome ITRs, including, but not limited to, AAV serotypes described herein. Production of pseudotyped rAAV is disclosed in, for example, international patent application publication number WO 01 / 83692. III. Methods Methods of Modifying a Nucleic Acid and / or Modifying Gene Expression

[0176] Further aspects of the present disclosure are directed to methods of using thesgRNAs and / or Cas12f polypeptides (including the gene editing systems provided herein) to modify a target nucleic acid and / or alter gene expression in a target cell. These methods may comprise, in some instances, direct genome editing which generally refers to a process of modifying the nucleotide sequence of a genome, preferably in a precise or pre-determined manner. The methods may, alternatively, not result in a modification of a nucleotide sequence, but instead result in alteration of gene expression by modifying the activity of a regulatory agent (i.e., a repressor, an activator, and the like). Also provided are methods for screening certain modifications in the Cas12f polypeptides above for improved function.

[0177] As noted above, the Cas12f polypeptides provided herein are components of theCRISPR-endonuclease system which is a naturally occurring defense mechanism in prokaryotes that has been repurposed as an RNA-guided DNA-targeting platform used for gene editing. CRISPR systems include Types I, II, III, IV, V, and VI systems. The Cas12f polypeptides provided herein function as Type V-F systems. All CRISPR systems rely on a CRISPR associated protein (e.g., Cas12f) and a gRNA (comprising a spacer sequence and a scaffold sequence) - to target a nucleic acid and effect one or more changes at the target site.

[0178] The spacer sequence drives sequence recognition and specificity of the CRISPR-endonuclease complex through Watson-Crick base pairing, typically with a ~20 nucleotide (nt) sequence in the target DNA. Changing the sequence of the 3’- 20 nt in the spacer sequence allows targeting of the Cas12f-gRNA complex to specific loci. The Cas12f-gRNA complex only binds DNA sequences that contain a sequence match to the spacer sequence of the gRNA if the target sequence is preceded by a specific short DNA motif (with the sequence CCN) referred to as a protospacer adjacent motif (PAM). The gRNA itself binds to the Cas12f protein through the scaffold sequence, thereby forming a Cas12f-gRNA complex which can then associate with the target DNA. Once associated, the Cas12f-gRNA complex -69- 99975521.7Attorney Docket No: 124540-829360 can alter gene expression by directly modifying the nucleic acid itself (i.e., by inserting, deleting, and / or mutating one or more base pairs in the nucleic acid) or by modifying the regulation of the target nucleic acid (i.e., via epigenetic means).

[0179] In some embodiments, methods are provided herein for modifying a target DNA.These methods may, in some aspects, comprise contacting the target DNA with a complex formed between a Cas12f polypeptide provided herein and an sgRNA provided herein, wherein the sgRNA comprises a spacer sequence capable of hybridizing to a target sequence in the target DNA, and wherein the complex modifies the target DNA. In some aspects, the methods of modifying a target DNA may further alter gene expression if at least one genetic modification is introduced within or near at least one gene locus that alters expression of that gene relative to an unmodified gene locus in an unmodified cell.

[0180] In various aspects, methods of modifying DNA can comprise using the Cas12fpolypeptides provided herein as site-directed nucleases to cut deoxyribonucleic acid (DNA) at precise target locations in a genome, thereby creating single-strand or double-strand DNA breaks at particular locations within the genome. Single stranded breaks can be used to precisely edit single or multiple bases in a nucleic acid (a method known as prime editing). Double stranded breaks can be and regularly are repaired by natural, endogenous cellular processes, such as homology-directed repair (HDR) and non-homologous end joining (NHEJ), as described in Cox et al., “Therapeutic genome editing: prospects and challenges,”, Nature Medicine, 2015, 21(2), 121-31. These two main DNA repair processes consist of a family of alternative pathways. NHEJ directly joins the DNA ends resulting from a double- strand break, sometimes with the loss or addition of nucleotide sequence, which may disrupt or enhance gene expression. HDR utilizes a homologous sequence, or donor sequence, as a template for inserting a defined DNA sequence at the break point. The homologous sequence can be in the endogenous genome, such as a sister chromatid. Alternatively, the donor sequence can be an exogenous polynucleotide, such as a plasmid, a single-strand oligonucleotide, a double-stranded oligonucleotide, a duplex oligonucleotide or a virus, that has regions (e.g., left and right homology arms) of high homology with the nuclease-cleaved locus, but which can also contain additional sequence or sequence changes including deletions that can be incorporated into the cleaved target locus. A third repair mechanism can be microhomology-mediated end joining (MMEJ), also referred to as "Alternative NHEJ,” in which the genetic outcome is similar to NHEJ in that small deletions and insertions can occur at the cleavage site. MMEJ can make use of homologous sequences of a few base pairs -70- 99975521.7Attorney Docket No: 124540-829360 flanking the DNA break site to drive a more favored DNA end joining repair outcome, and recent reports have further elucidated the molecular mechanism of this process; see, e.g., Cho and Greenberg, Nature, 2015, 518, 174-76; Kent et al., Nature Structural and Molecular Biology, 2015, 22(3):230-7; Mateos-Gomez et al., Nature, 2015, 518, 254-57; Ceccaldi et al., Nature, 2015, 528, 258-62. In some instances, it may be possible to predict likely repair outcomes based on analysis of potential microhomologies at the site of the DNA break.

[0181] Each of these mechanisms can be used to create desired genetic modifications. Forinstance, in one example, a method provided herein can comprise using the Cas12f polypeptides provided herein to create one or two DNA breaks, the latter as double-strand breaks or as two single-stranded breaks, in the target locus as near the site of intended mutation.

[0182] As a result of the single or double stranded breaks introduced by the Cas12fpolypeptides of the present disclosure, a mutation may occur at or near a gene locus which results in an alteration of the function / expression or activity of that gene locus. As used herein, “mutation” may include substitutions, additions, and deletions or any combination thereof in the nucleic acid at or near a gene locus. These mutations can, in some aspects, result in alterations (i.e., substitutions, insertions, or deletions) in the resulting protein product of the targeted gene. In some aspects, the mutation can convert an amino acid in the encoded protein to another amino acid (e.g., glycine, serine, threonine, cysteine, valine, leucine, isoleucine, methionine, proline, phenylalanine, tyrosine, tryptophan, aspartic acid, glutamic acid, asparagine, glutamine, histidine, lysine, or arginine). In some aspects, the mutation converts the encoded amino acid to alanine. In some aspects, the mutation converts the encoded amino acid to a non-natural amino acid (e.g., selenomethionine). In some aspects, the mutation converts the encoded amino acid to amino acid mimics (e.g., phosphomimics). The mutation can be a conservative mutation. For example, the nucleic acid mutation can convert one or more amino acids in the encoded protein to another amino acid that resemble the size, shape, charge, polarity, conformation, and / or rotamers of the original amino acid(s) (e.g., cysteine / serine substitution, lysine / asparagine substitution, histidine / phenylalanine substitution). The mutation can cause a shift in reading frame and / or the creation of a premature stop codon. Mutations can also cause changes to regulatory regions of genes or loci that affect expression of one or more genes.

[0183] Accordingly, in various aspects, the methods of modifying a target DNA cancomprise introducing one or more single or double stranded breaks at the target sequence. In -71- 99975521.7Attorney Docket No: 124540-829360 various aspects, the methods further comprise providing an exogenous nucleic acid (i.e., a transgene) to be inserted at or near the site of the double stranded break (e.g., using homologous recombination).

[0184] In further aspects, the Cas12f-gRNA complexes provided herein further comprise afusion protein of the Cas12f polypeptide and another functional polypeptide, which may directly or indirectly modify a target nucleic acid once brought into proximity by the Cas12f- gRNA complex. In these methods, the Cas12f polypeptide may comprise one or more substitutions that decrease nuclease activity relative to a Cas12f polypeptide that does not have these substitutions. In some aspects, the one or more substitutions results in a Cas12f polypeptide that does not have any nuclease catalytic activity. In some aspects, the Cas12f polypeptide comprises a D216A and / or D390A substitution according to SEQ ID NO: 1, or any other aspartic acid (D) to alanine (A) substitutions at an equivalent location that reduces or eliminates the nuclease activity of the Cas polypeptide (see e.g., Table 3 above). When the Cas12f polypeptide is a modified form that has no substantial nucleic acid-cleaving activity, it is referred to herein as "enzymatically inactive", “inactive”, “deactivated” or “dead” and is indicated by the prefix “d” (i.e., “dCas12f”).

[0185] As noted above, the Cas12f polypeptides may be provided as fusion proteins thatcomprise one or more functional polypeptides that directly or indirectly modify a target nucleic acid and / or otherwise alter gene expression at a target gene locus. In various aspects, the functional polypeptide may have methyltransferase activity, DNA repair activity, DNA damage activity, dismutase activity, alkylation activity, demethylation activity, dealkylation activity, depurination activity, oxidation activity, deoxidation activity, pyrimidine dimer forming activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, glycosylase activity, acetyl transferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, deSUMOylation activity, ribosylation activity, deribosylation activity, myristoylation activity, demyristoylation activity, glycosylation activity (e.g., from O- GlcNAc transferase) , deglycosylation activity or any combination thereof. In some aspects, the functional domain may comprise ten-eleven translocation methylcytosine dioxygenase (Tet, Tet1, Tet2 or mini Tet2 (Tet-mini)), DNA methyl transferase (e.g. DNMT3A and / or DNMT3L), VP16, VP64, VPR, ABE, CBE, a Krüppel-associated box (KRAB) or any catalytically active domain thereof. These functional polypeptides may, in certain -72- 99975521.7Attorney Docket No: 124540-829360 embodiments, increase or decrease methylation at or near the gene locus – which functionally increases or decreases protein expression.

[0186] In various aspects of the present disclosure, the methods of modifying a targetnucleic acid (DNA) and / or modifying gene expression can be performed in a cell. The cell may in certain aspects, be in vitro or in vivo. The cell may be a human cell.

[0187] In some non-limiting examples, a reporter cell is provided that further compriseone or more constructs that express a reporter protein (i.e., a florescent protein), where the constructs are operably linked to a regulatory sequence that may be modified by the methods of the present disclosure. In these aspects, the cell may act as a reporter cell line – expressing the reporter protein in proportion to the activity of the gene editing system provided herein. In some aspects, the cell may comprise one or more TurboRFP expressing constructs under control of a TRE operator comprising a TetO operator, wherein the cell expresses RFP (red florescent protein) only in the presence of a targeted transcriptional activator binding to the TetO operator of at least one TurboRFP expressing construct. In some aspects, each TetO operator in each TurboRFP expressing construct may further comprise a PAM sequence of the Cas12f polypeptide herein (i.e., TTTA, TTTC or TCCA). In some aspects, the reporter cell comprises a vector having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 593-595. Methods of Treatment

[0188] Further aspects of the present disclosure are directed to methods of diagnosing,preventing and / or treating a disease in a subject by modifying a target nucleic acid and / or modifying gene expression in at least one cell of the subject, wherein modifying the target DNA and / or modifying expression of the target gene diagnoses, prevents and / or treats the disease.

[0189] In some embodiments, a composition comprising a Cas12f polypeptide and ansgRNA of the present disclosure, or one or more polynucleotides encoding the Cas12f polypeptide and the sgRNA, may be administered to a subject, e.g., a human subject, who has, is suspected of having, or is at risk for a disease. In some embodiments, a composition may be administered to a subject who does not have, is not suspected of having or is not at risk for a disease. In some embodiments, a subject is a healthy human. In some embodiments, a subject e.g., a human subject, who has, is suspected of having, or is at risk -73- 99975521.7Attorney Docket No: 124540-829360 for a genetically inheritable disease. In some embodiments, the subject is suffering or is at risk of developing symptoms indicative of a disease.

[0190] In some embodiments, a genetically modified cell is provided, wherein thegenetically modified cell is edited using a method provided herein (i.e., using a Cas12f polypeptide and / or an sgRNA of the present disclosure). In some aspects, the genetically modified cell is obtained from a subject. In some aspects, a composition comprising one or more genetically modified cells is administered to a subject in need thereof, according to any method known in the art.

[0191] The terms "administering," "introducing", “implanting”, “engrafting” and"transplanting" are used interchangeably in the context of the placement of cells, e.g., genetically modified cells, into a subject, by a method or route that results in at least partial localization of the introduced cells at a desired site. The cells e.g., genetically modified cells, or their differentiated progeny can be administered by any appropriate route that results in delivery to a desired location in the subject where at least a portion of the implanted cells or components of the cells remain viable. The period of viability of the cells after administration to a subject can be as short as a few hours, e.g., twenty-four hours, to a few days, to as long as several years, or even the lifetime of the subject, i.e., long-term engraftment.

[0192] In some embodiments, a composition comprising cells as described herein may beadministered by a suitable route, which may include intravenous administration, e.g., as a bolus or by continuous infusion over a period of time. In some embodiments, intravenous administration may be performed by intramuscular, intraperitoneal, intracerebrospinal, subcutaneous, intra-articular, intrasynovial, or intrathecal routes (lumbar or intracisterna magna). In some embodiments, a composition may be in solid form, aqueous form, or a liquid form. In some embodiments, an aqueous or liquid form may be nebulized or lyophilized. In some embodiments, a nebulized or lyophilized form may be reconstituted with an aqueous or liquid solution.

[0193] A cell composition can also be emulsified or presented as a liposome composition,provided that the emulsification procedure does not adversely affect cell viability. The cells and any other active ingredient can be mixed with excipients that are pharmaceutically acceptable and compatible with the active ingredient, and in amounts suitable for use in the therapeutic methods described herein. -74- 99975521.7Attorney Docket No: 124540-829360

[0194] Additional agents included in a cell composition can include pharmaceuticallyacceptable salts of the components therein. Pharmaceutically acceptable salts include the acid addition salts (formed with the free amino groups of the polypeptide) that are formed with inorganic acids, such as, for example, hydrochloric or phosphoric acids, or such organic acids as acetic, tartaric, mandelic and the like. Salts formed with the free carboxyl groups can also be derived from inorganic bases, such as, for example, sodium, potassium, ammonium, calcium or ferric hydroxides, and such organic bases as isopropylamine, trimethylamine, 2- ethylamino ethanol, histidine, procaine and the like.

[0195] Physiologically tolerable carriers are well known in the art. Exemplary liquidcarriers are sterile aqueous solutions that contain no materials in addition to the active ingredients and water, or contain a buffer such as sodium phosphate at physiological pH value, physiological saline or both, such as phosphate-buffered saline. Still further, aqueous carriers can contain more than one buffer salt, as well as salts such as sodium and potassium chlorides, dextrose, polyethylene glycol and other solutes. Liquid compositions can also contain liquid phases in addition to and to the exclusion of water. Exemplary of such additional liquid phases are glycerin, vegetable oils such as cottonseed oil, and water-oil emulsions. The amount of an active compound used in the cell compositions that is effective in the treatment of a particular disorder or condition can depend on the nature of the disorder or condition and can be determined by standard clinical techniques.

[0196] In some embodiments, a composition comprising cells may be administered to asubject, e.g., a human subject, who has, is suspected of having, or is at risk for a disease. In some embodiments, a composition may be administered to a subject who does not have, is not suspected of having or is not at risk for a disease. In some embodiments, a subject is a healthy human. In some embodiments, a subject e.g., a human subject, who has, is suspected of having, or is at risk for a genetically inheritable disease. In some embodiments, the subject is suffering or is at risk of developing symptoms indicative of a disease.

[0197] In any of the foregoing aspects, the disease may comprise a genetic disease. Inother aspects, the disease may comprise a hereditary genetic disease. For example, in some aspects, the disease may comprise Rett Syndrome, X-linked severe combined immunodeficiency (SCID), Fragile X Syndrome, CDKL5 Deficiency Disease, Menkes Syndrome, Alport Syndrome, Immunodeficiency with Hyper IgM, Adrenoleukodystrophy, Hemophilia A, Progressive myoclonus epilepsy, spinocerebellar ataxias, neuronal intranuclear inclusion disease, Glutaminase Deficiency, Huntington’s Disease, Dentatorubral- -75- 99975521.7Attorney Docket No: 124540-829360 Pallidoluysian Atrophy, Spinal Bulbar muscular atrophy, oculopharyngeal muscular dystrophy, Myotonic Dystrophy, Friedreich ataxia, CANVAS, Angelman syndrome, Alzheimer’s Disease, transthyretin amyloidosis (ATTR), transthyretin amyloid cardiomyopathy (ATTR-CM), cystic fibrosis (CF), hereditary angioedema, diabetes, progressive pseudohypertrophic muscular dystrophy, Duchenne muscular dystrophy (DMD), Becker muscular dystrophy (BMD), spinal muscular atrophy (SMA), alpha-1 antitrypsin deficiency, Pompe disease, myotonic dystrophy, Huntington’s disease (HTT), fragile X syndrome, Friedreich ataxia, amyotrophic lateral sclerosis (ALS), frontotemporal dementia, hereditary chronic kidney disease, hyperlipidemia, Leber congenital amaurosis (LCA), sickle cell disease, thalassemia (e.g., β-thalassemia), Parkinson’s disease (PD), myelodysplastic syndrome (MDS), retinitis pigmentosa (RP), age-related macular degeneration (AMD), Hepatitis B, nonalcoholic fatty liver disease (NAFLD), Acquired Immune Deficiency Syndrome, corneal dystrophy (CD), hypercholesterolemia, familial hypercholesterolemia (FH), heart disease (e.g., hypertrophic cardiomyopathy (HCM)) and cancer.

[0198] For instance, in various aspects, the disease may be selected from: Rett Syndrome,X-linked severe combined immunodeficiency (SCID), Fragile X Syndrome, CDKL5 Deficiency Disease, Menkes Syndrome, Alport Syndrome, Immunodeficiency with Hyper IgM, Adrenoleukodystrophy, Hemophilia A, Progressive myoclonus epilepsy, spinocerebellar ataxias, neuronal intranuclear inclusion disease, Glutaminase Deficiency, Huntington’s Disease, dentatorubral-pallidoluysian atrophy, Spinal Bulbar muscular atrophy, oculopharyngeal muscular dystrophy, Myotonic Dystrophy, Friedreich ataxia, and CANVAS. IV. Definitions

[0199] Unless defined otherwise, all technical and scientific terms used herein have themeaning commonly understood by a person skilled in the art to which this disclosure relates. The following references provide one of skill with a general definition of many of the terms used in this disclosure: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed.1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). As used herein, the following terms have the meanings ascribed to them unless specified otherwise.

[0200] When introducing elements of the present disclosure or the preferred aspects(s)thereof, the articles "a", "an", "the" and "said" are intended to mean that there are one or more of the elements. The terms "comprising", "including" and "having" are intended to be -76- 99975521.7Attorney Docket No: 124540-829360 inclusive and mean that there may be additional elements other than the listed elements. Wherever the terms “comprising” or “including” are used, it should be understood the disclosure also expressly contemplates and encompasses additional aspects “consisting of” the disclosed elements, in which additional elements other than the listed elements are not included.

[0201] “Improved targeting and / or editing efficiency”: As used herein, the term“improved targeting and / or editing efficiency” refers to an increased ability to bind to a target nucleic acid, associate with a target nucleic acid, associate with or complex with a gRNA, increased nuclease ability, increase specificity for spacer matching sequence, increase PAM recognition ability, or have an increased ability to prevent NHEJ.

[0202] “Cas12f polypeptide”: As used herein the term “Cas12f polypeptide” comprisesany polypeptide or protein comprising at least one domain derived from a Cas12f nuclease (e.g., a “Cas12f domain”). The Cas12f polypeptides may consist of a Cas12f domain or may further comprise one or more functional domains.

[0203] “Deletion”: As used herein, the term “deletion”, which may be usedinterchangeably with the terms “genetic deletion” or “knock-out”, generally refers to a genetic modification wherein a site or region of genomic DNA is removed by any molecular biology method, e.g., methods described herein, e.g., by delivering to a site of genomic DNA an endonuclease and at least one gRNA. Any number of nucleotides can be deleted. In some embodiments, a deletion involves the removal of at least one, at least two, at least three, at least four, at least five, at least ten, at least fifteen, at least twenty, or at least 25 nucleotides. In some embodiments, a deletion involves the removal of 10-50, 25-75, 50-100, 50-200, or more than 100 nucleotides. In some embodiments, a deletion involves the removal of an entire target gene. In some embodiments, a deletion involves the removal of part of a target gene, e.g., all or part of a promoter and / or coding sequence of a gene. In some embodiments, a deletion involves the removal of a transcriptional regulator, e.g., a promoter region, of a target gene. In some embodiments, a deletion involves the removal of all or part of a coding region such that the product normally expressed by the coding region is no longer expressed, is expressed as a truncated form, or expressed at a reduced level. In some embodiments, a deletion leads to a decrease in expression of a gene relative to an unmodified cell.

[0204] “Insertion”: As used herein, the term “insertion” which may be usedinterchangeably with the terms “genetic insertion” or “knock-in”, generally refers to a genetic -77- 99975521.7Attorney Docket No: 124540-829360 modification wherein a polynucleotide is introduced or added into a site or region of genomic DNA by any molecular biological method, e.g., methods described herein, e.g., by delivering to a site of genomic DNA an endonuclease and at least one gRNA. In some embodiments, an insertion may occur within or near a site of genomic DNA that has been the site of a prior genetic modification, e.g., a deletion or insertion-deletion mutation. In some embodiments, an insertion occurs at a site of genomic DNA that partially overlaps, completely overlaps, or is contained within a site of a prior genetic modification, e.g., a deletion or insertion-deletion mutation. In some embodiments, an insertion occurs at a safe harbor locus. In some embodiments, an insertion involves the introduction of a polynucleotide that encodes a protein of interest. In general, a polynucleotide to be inserted is flanked by sequences (e.g., homology arms) having substantial sequence homology with genomic DNA at or near the site of insertion.

[0205] “Nuclease”: As used herein, the term “nuclease” generally refers to an enzymethat cleaves phosphodiester bonds within a polynucleotide. In some embodiments, a nuclease specifically cleaves phosphodiester bonds within a DNA polynucleotide. In some embodiments, a nuclease may introduce one or more single-stranded breaks (SSBs) and / or one or more double-stranded breaks (DSBs).

[0206] “Genetic modification”: As used herein, the term “genetic modification” generallyrefers to a site of genomic DNA that has been genetically edited or manipulated using any molecular biological method, e.g., methods described herein, e.g., by delivering to a site of genomic DNA an endonuclease and at least one gRNA. Example genetic modifications include insertions, deletions, duplications, inversions, and translocations, and combinations thereof. In some embodiments, a genetic modification is a deletion. In some embodiments, a genetic modification is an insertion. In other embodiments, a genetic modification is an insertion-deletion mutation (or indel), such that the reading frame of the target gene is shifted leading to an altered gene product or no gene product.

[0207] “Guide RNA (gRNA)”: As used herein, the term “guide RNA” or “gRNA”generally refers to short ribonucleic acid that can interact with, e.g., bind to, to an endonuclease and bind, or hybridize to a target genomic site or region.

[0208] The terms “polypeptide” and “protein” are used interchangeably to refer to apolymer of amino acid residues. -78- 99975521.7Attorney Docket No: 124540-829360

[0209] “Polynucleotide”: As used herein, the term “polynucleotide”, which may be usedinterchangeably with the term “nucleic acid” or “nucleic acid molecule generally refers to a biomolecule that comprises two or more nucleotides. A polynucleotide may be a DNA or RNA molecule or a hybrid DNA / RNA molecule. A polynucleotide may be single-stranded or double-stranded. In some embodiments, a polynucleotide is a site or region of genomic DNA. In some embodiments, a polynucleotide is an endogenous gene that is comprised within the genome of an unmodified cell or a genetically engineered cell. In some embodiments, a polynucleotide is an exogenous polynucleotide that is not integrated into genomic DNA. In some embodiments, a polynucleotide is an exogenous polynucleotide that is integrated into genomic DNA. In some embodiments, a polynucleotide is a plasmid or an adeno-associated viral vector. In some embodiments, a polynucleotide is a circular or linear molecule. The terms “nucleic acid encoding ...”, or “nucleic acid molecule encoding ... ”, should be understood as referring to the sequence of nucleotides which encodes a polypeptide or an RNA (e.g., a gRNA).

[0210] A “Gene” or “coding sequence” refers to a DNA or RNA region (the transcribedregion) which “encodes” a particular polypeptide such as a Cas12f polypeptide or fragments thereof. A coding sequence may also be a nucleic acid encoding an RNA as a final product (i.e., a gRNA). Generally, when the gene or coding sequence encodes a polypeptide the coding sequence is transcribed (DNA) and translated (RNA) into a polypeptide when placed under the control of an appropriate regulatory region, such as a promoter. A gene may comprise several operably linked fragments, such as a promoter, a 5’ leader sequence, an intron, a coding sequence and a 3’ nontranslated sequence, comprising a polyadenylation site or a signal sequence. A chimeric or recombinant gene (such as the one encoding novel Cas12f polypeptide as identified herein and operably linked to a promoter) is a gene not normally found in nature, such as a gene in which for example the promoter is not associated in nature with part or all of the transcribed DNA region. “Expression of a gene” refers to the process wherein a gene is transcribed into an RNA and / or translated into an active protein.

[0211] “Safe harbor locus”: As used herein, the term “safe harbor locus” generally refersto any location, site, or region of genomic DNA that may be able to accommodate a genetic insertion into said location, site, or region without adverse effects on a cell. In some embodiments, a safe harbor locus is an intragenic or extragenic region. In some embodiments, a safe harbor locus is a region of genomic DNA that is typically transcriptionally silent. In some embodiments, a safe harbor locus is a AAVS1 (PPP1 R12C), -79- 99975521.7Attorney Docket No: 124540-829360 ALB, Angptl3, ApoC3, ASGR2, CCR5, FIX (F9), G6PC, Gys2, HGD, Lp(a), Pcsk9, Serpina1, TF, or TTR locus. In some embodiments, a safe harbor locus is described in Sadelain, M. et al., “Safe harbours for the integration of new DNA in the human genome,” Nature Reviews Cancer, 2012, Vol 12, pages 51-58.

[0212] As used herein, the terms "target site", "target sequence", or “nucleic acid locus”refer to a nucleic acid sequence that defines a portion of a nucleic acid sequence to be modified or edited and to which a homologous recombination composition is engineered to target.

[0213] “Within or near a gene”: As used herein, the term “within or near a gene” refers toa site or region of genomic DNA that is an intronic or extronic component of a said gene or is located proximal to a said gene. In some embodiments, a site of genomic DNA is within a gene if it comprises at least a portion of an intron or exon of said gene. In some embodiments, a site of genomic DNA located near a gene may be at the 5’ or 3’ end of said gene (e.g., the 5’ or 3’ end of the coding region of said gene). In some embodiments, a site of genomic DNA located near a gene may be a promoter region or repressor region that modulates the expression of said gene. In some embodiments, a site of genomic DNA located near a gene may be on the same chromosome as said gene. In some embodiments, a site or region of genomic DNA is near a gene if it is within 50Kb, 40Kb, 30Kb, 20Kb, 10Kb, 5Kb, 1Kb, or closer to the 5’ or 3’ end of said gene (e.g., the 5’ or 3’ end of the coding region of said gene).

[0214] The terms "upstream" and "downstream" refer to locations in a nucleic acidsequence relative to a fixed position. Upstream refers to a position in a sequence that is 5' (i.e., nearer the 5' end of the strand) relative to the fixed position, and downstream refers to the region that is 3' (i.e., nearer the 3' end of the strand) relative to the fixed position.

[0215] As used herein, a “regulatory sequence” refers to any genetic element that isknown to the skilled person to drive or otherwise regulate expression of nucleic acids in a cell. Such sequences include without limitation promoters, transcription terminators, enhancers, repressors, silencers, kozak sequences, polyA sequences, and the like. A regulatory sequence can, for example, be inducible, non-inducible, constitutive, cell-cycle regulated, metabolically regulated, and the like. A regulatory sequence may be a promoter. As used herein, the term "promoter" refers to a nucleic acid fragment that functions to control the transcription of one or more genes (or coding sequence), located upstream with respect to -80- 99975521.7Attorney Docket No: 124540-829360 the direction of transcription of the transcription initiation site of the gene, and is structurally identified by the presence of a binding site for DNA-dependent RNA polymerase, transcription initiation sites and any other DNA sequences, including, but not limited to transcription factor binding sites, repressor and activator protein binding sites, and any other sequences of nucleotides known to one of skill in the art to act directly or indirectly to regulate the amount of transcription from the promoter. A "constitutive" promoter is a promoter that is active under most physiological and developmental conditions. An "inducible" promoter is a promoter that is regulated depending on physiological or developmental conditions. A "tissue specific" promoter is preferentially active in specific types of differentiated cells / tissues, such as preferably a T cell. A preferred promoter is the MSCV promoter. Non-limiting examples of suitable promoters include EF1α, MSCV, EF1 alpha-HTLV-1 hybrid promoter, Moloney murine leukemia virus (MoMuLV or MMLV), Gibbon Ape Leukemia virus (GALV), murine mammary tumor virus (MuMTV or MMTV), Rous sarcoma virus (RSV), MHC class II, clotting Factor IX, insulin promoter, PDX1 promoter, CD11, CD4, CD2, gp47 promoter, PGK, Beta-globin, UbC, MND, and derivatives (i.e. variants) thereof. Examples of these promoters are further described in Poletti and Mavilio (2021), Viruses 13:8;1526, Kuroda et al. (2008), J Gene Med 10(11):1163-1175, Milone et al. (2009), Mol Ther 17:8;1453-1464, and Klein et al. (2008), J Biomed Biotechnol 683505, all of which are incorporated herein by reference in their entireties.

[0216] “Operably linked” is defined herein as a configuration in which a control sequencesuch as a promoter sequence or regulating sequence is appropriately placed at a position relative to the nucleotide sequence of interest (i.e., encoding for a Cas12f polypeptide) such that the promoter or control or regulating sequence directs or affects the transcription and / or production or expression of the nucleotide sequence of interest, in a cell and / or in a subject. For instance, a promoter is operably linked to a coding sequence if the promoter is able to initiate or regulate the transcription or expression of a coding sequence, in which case the coding sequence should be understood as being “under the control of” the promoter.

[0217] A “DNA construct” or “nucleic acid construct” prepared for introduction into aparticular host may include a replication system recognized by the host, an intended DNA segment (e.g. plasmid) encoding a desired polypeptide, and transcriptional and translational initiation and termination regulatory sequences operably linked to the polypeptide-encoding segment. The term “operably linked” has already been defined herein. For example, a promoter or enhancer is operably linked to a coding sequence if it stimulates the transcription -81- 99975521.7Attorney Docket No: 124540-829360 of the sequence. DNA for a signal sequence is operably linked to DNA encoding a polypeptide if it is expressed as a preprotein that participates in the secretion of a polypeptide. Generally, a DNA sequence that is operably linked are contiguous, and, in the case of a signal sequence, both contiguous and in reading frame. However, enhancers need not be contiguous with a coding sequence whose transcription they control. Linking is accomplished by ligation at convenient restriction sites or at adapters or linkers inserted in lieu thereof, or by gene synthesis.

[0218] Vector: A vector may comprise a nucleic acid construct or an expression constructas earlier defined herein. A vector as described herein may be selected from any genetic element known in the art which can facilitate transfer of nucleic acids between cells, such as, but not limited to, plasmids, transposons, cosmids, chromosomes, artificial chromosomes, viruses, virions, and the like. A vector may also be a chemical vector, such as a lipid complex or naked DNA. ‘’Naked DNA’’ or ‘’naked nucleic acid’’ refers to a nucleic acid molecule that is not contained in encapsulating means that facilitates delivery of a nucleic acid into the cytoplasm of a target host cell. Naked DNA may be circular or linear (linearized DNA sequence). Optionally, a naked nucleic acid can be associated with standard means used in the art for facilitating its delivery of the nucleic acid to the target host cell, for example to facilitate the transport of the nucleic acid through the cell membrane. In some examples, vectors can be capable of directing the expression of nucleic acids to which they are operatively linked. Such vectors are referred to herein as "recombinant expression vectors", or more simply "expression vectors", which serve equivalent functions. In addition, such vectors are also referred to herein as comprising an “expression construct” which is generally a nucleic acid encoding a protein (or an RNA such as an sgRNA) operably linked to a promoter or other regulatory sequence that directs its expression.

[0219] A "transgene" is herein defined as a gene or a nucleic acid molecule (i.e., amolecule encoding a Cas12f polypeptide and / or an sgRNA as provided herein) that has been newly introduced into a cell, i.e., a gene that may be present but may normally not be expressed or expressed at an insufficient level in a cell. The transgene may comprise sequences that are native to the cell, sequences that naturally do not occur in the cell and it may comprise combinations of both. A transgene may contain sequences coding for Cas12f polypeptide and / or an sgRNA and comprising the polypeptide as identified and / or additional proteins as earlier identified herein that may be operably linked to appropriate regulatory sequences for expression of the sequences coding for a Cas12f polypeptide and / or an sgRNA -82- 99975521.7Attorney Docket No: 124540-829360 A transgene may also include a nucleic acid encoding a protein of interest or a donor nucleic acid for insertion at a targeted nucleic acid site using homologous recombination. Preferably, the transgene is not integrated into the host cell’s genome.

[0220] A “wild-type” protein amino acid sequence can refer to a sequence that is naturallyoccurring and encoded by a germline genome. A species can have one wild-type sequence, or two or more wild-type sequences (for example, with one canonical wild-type sequence and one or more non-canonical wild-type sequences). A wild-type protein amino acid sequence can be a mature form of a protein that has been processed to remove N-terminal and / or C- terminal residues, for example, to remove a signal peptide. An amino acid sequence that is “derived from” a wild-type sequence or other amino acid sequence disclosed herein can refer to an amino acid sequence that differs by one or more amino acids compared to the reference amino acid sequence, for example, containing one or more amino acid insertions, deletions, or substitutions as disclosed herein. The terms “derivative,” “variant,” “variations” and “fragment,” when used herein with reference to a polypeptide, refers to a polypeptide related to a wild-type polypeptide, for example either by amino acid sequence, structure (e.g., secondary and / or tertiary), activity (e.g., enzymatic activity) and / or function. Derivatives, variants, “variations” and fragments of a polypeptide can comprise one or more amino acid variations (e.g., mutations, insertions, and deletions), truncations, modifications, or combinations thereof compared to a wild-type polypeptide. A part or fragment of a polypeptide may correspond to at least 1%, at least 2%, at least 3 %, at least 4%, at least 5%, at least 10%, at least 20%, at least 30%, at least 40% of the length of a polypeptide, such as a polypeptide having an amino acid sequence identified by a specific SEQ ID NO., or having at least 50%, or at least 60%, or at least 70%, or at least 80%, or at least 90% of the length (in amino acids) of the polypeptide.

[0221] Within the context of the present application, a protein is represented by an aminoacid sequence, and correspondingly a nucleic acid molecule or a polynucleotide is represented by a nucleic acid sequence. It should be understood that for each reference to a specific amino acid sequence using a unique sequence identifier (SEQ ID NO.), the sequence may be replaced by a polypeptide represented by an amino acid sequence comprising a sequence that has at least 60% sequence identity or similarity with the reference amino acid sequence. Another preferred level of sequence identity or similarity is 65%. Another preferred level of sequence identity or similarity is 70%. Another preferred level of sequence identity or similarity is 75%. Another preferred level of sequence identity or similarity is -83- 99975521.7Attorney Docket No: 124540-829360 80%. Another preferred level of sequence identity or similarity is 85%. Another preferred level of sequence identity or similarity is 90%. Another preferred level of sequence identity or similarity is 95%. Another preferred level of sequence identity or similarity is 98%. Another preferred level of sequence identity or similarity is 99%.

[0222] Each amino acid sequence described herein by virtue of its identity or similaritypercentage with a given amino acid sequence respectively has in a further preferred aspect an identity or a similarity of at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% with the given nucleotide or amino acid sequence, respectively. The terms “homology”, “sequence identity” and the like are used interchangeably herein. Sequence identity is described herein as a relationship between two or more amino acid (polypeptide or protein) sequences or two or more nucleic acid (polynucleotide) sequences, as determined by comparing the sequences. In a preferred aspect, sequence identity is calculated based on the full length (in amino acids or nucleotides) of two given SEQ ID NOS or based on a portion thereof. A portion of a full-length sequence may be referred to as a fragment, and preferably means at least 50%, 60%, 70%, 80%, 90%, or 100% of the length (in amino acids or nucleotides) of a reference sequence. "Identity" also refers to the degree of sequence relatedness between two amino acid sequences, or between two nucleic acid sequences, as the case may be, as determined by the match between strings of such sequences. The degree of sequence identity between two sequences can be determined, for example, by comparing the two sequences using computer programs commonly employed for this purpose, such as global or local alignment algorithms. Non-limiting examples include BLASTp, BLASTn, Clustal W, MAFFT, Clustal Omega, AlignMe, Praline, GAP, BESTFIT, or another suitable method or algorithm. A Needleman and Wunsch global alignment algorithm can be used to align two sequences over their entire length or part thereof (part thereof may mean at least 50%, 60%, 70%, 80%, 90% of the length of the sequence), maximizing the number of matches and minimizes the number of gaps. Default settings can be used and preferred program is Needle for pairwise alignment (in an aspect, EMBOSS Needle 6.6.0.0, gap open penalty 10, gap extent penalty: 0.5, end gap penalty: false, end gap open penalty: 10 , end gap -84- 99975521.7Attorney Docket No: 124540-829360 extent penalty: 0.5 is used) and MAFFT for multiple sequence alignment ( in an aspect, MAFFT v7Default value is: BLOSUM62 [bl62], Gap Open: 1.53, Gap extension: 0.123, Order: aligned , Tree rebuilding number: 2, Guide tree output: ON [true], Max iterate: 2 , Perform FFTS: none is used).

[0223] "Similarity" between two amino acid sequences is determined, for example, bycomparing the amino acid sequence and its conserved amino acid substitutes of one polypeptide to the sequence of a second polypeptide. Similar algorithms used for determination of sequence identity may be used for determination of sequence similarity. Optionally, in determining the degree of amino acid similarity, the skilled person may also take into account so-called conservative amino acid substitutions. As used herein, “conservative” amino acid substitutions refer to the interchangeability of residues having similar side chains. Examples of classes of amino acid residues for conservative substitutions are given in the Tables below: Acidic Residues Asp (D) and Glu (E)Basic Residues Lys (K), Arg (R), and His (H)Hydrophilic Uncharged Residues Ser (S), Thr (T), Asn (N), and Gln (Q) Aliphatic Uncharged Residues Gly (G), Ala (A), Val (V), Leu (L), and Ile (I) Non-polar Uncharged Residues Cys (C), Met (M), and Pro (P)Aromatic Residues Phe (F), Tyr (Y), and Trp (W)Alternative conservative amino acid residue substitution classes : 1A S T2 D E3 N Q4 R K5 I L M6 F Y WAlternative physical and functional classifications of amino acid residues: Alcohol group-containing residues S and TAliphatic residues I, L, V, and MCycloalkenyl-associated residues F, H, W, and YHydrophobic residues A, C, F, G, H, I, L, M, R, T, V, W, and Y Negatively charged residues D and EPolar residues C, D, E, H, K, N, Q, R, S, and T-85- 99975521.7Attorney Docket No: 124540-829360 Positively charged residues H, K, and RSmall residues A, C, D, G, N, P, S, T, and VVery small residues A, G, and SResidues involved in turn formation A, C, D, E, G, H, K, N, Q, R, S, P and TFlexible residues Q, T, K, S, G, P, D, E, and R

[0224] For example, a group of amino acids having aliphatic side chains is glycine,alanine, valine, leucine, and isoleucine; a group of amino acids having aliphatic-hydroxyl side chains is serine and threonine; a group of amino acids having amide-containing side chains is asparagine and glutamine; a group of amino acids having aromatic side chains is phenylalanine, tyrosine, and tryptophan; a group of amino acids having basic side chains is lysine, arginine, and histidine; and a group of amino acids having sulfur-containing side chains is cysteine and methionine. Preferred conservative amino acids substitution groups are valine-leucine-isoleucine, phenylalanine-tyrosine, lysine-arginine, alanine-valine, and asparagine-glutamine. Substitutional variants of the amino acid sequence disclosed herein are those in which at least one residue in the disclosed sequences has been removed and a different residue inserted in its place. Preferably, the amino acid change is conservative. Preferred conservative substitutions for each of the naturally occurring amino acids are as follows: Ala to Ser; Arg to Lys; Asn to Gln or His; Asp to Glu; Cys to Ser or Ala; Gln to Asn; Glu to Asp; Gly to Pro; His to Asn or Gln; Ile to Leu or Val; Leu to Ile or Val; Lys to Arg; Gln or Glu; Met to Leu or Ile; Phe to Met, Leu or Tyr; Ser to Thr; Thr to Ser; Trp to Tyr; Tyr to Trp or Phe; and, Val to Ile or Leu.

[0225] A polynucleotide described herein may comprise one or more nucleic acids eachencoding a polypeptide, all operably linked to (i.e., in a functional relationship with) one or more regulatory sequences, such as a promoter. Such a polynucleotide may alternatively be referred to herein as a ‘’nucleic acid construct’’ or ‘’construct’’. An “expression construct” or “nucleic acid construct” carries a genome that is able to stabilize and remain episomal in a cell. Within the context of the present disclosure, a cell may mean to encompass a cell used to make the construct or a cell wherein the construct will be administered. Alternatively, a construct is capable of integrating into a cell's genome, e.g. through homologous recombination or otherwise. An exemplary expression construct is one wherein a nucleotide sequence encoding a Cas12f polypeptide or an sgRNA is operably linked to a promoter as defined herein wherein said promoter is capable of directing expression of said nucleotide sequence (i.e. coding sequence) in a cell. Such a preferred expression construct is said to comprise an expression cassette. A viral expression construct is an expression construct -86- 99975521.7Attorney Docket No: 124540-829360 which is intended to be used in gene therapy. It is designed to comprise part of a viral genome as later defined herein.

[0226] “Transduction” refers to the delivery of a Cas12f polypeptide and / or sgRNA into arecipient host cell by a viral vector. For example, transduction of a cell by a retroviral or lentiviral vector as described herein leads to transfer of the genome contained in that vector into the transduced cell. In an aspect, the vector is a lentiviral vector. In an aspect, the vector is an adenoviral vector. A “Host cell” refers to the cell into which the DNA delivery (transduction) takes place.

[0227] As used herein, a “cell” refers to any mammalian cell, preferably a human cell. Inan aspect, the cell is an induced pluripotent stem cell (iPSC) or any cell derived or differentiated from an iPSC. In some aspects, the cell is from a cell line such as, but not limited to, HeLa cells, HEK293 cells, CHO cells, NIH / 3T3 cells, Vero cells, MDCK cells, COS-7 cells, HepG2 cells, MCF-7 cells, A549 cells, Jurkat cells, U2OS cells, PC12 cells, BHK-21 cells, L929 cells, SH-SY5Y cells, RAW 264.7 cells, Caco-2 cells, MDA-MB-231 cells, and Neuro-2a cells, YAC-1, THP-1, and U937.

[0228] A “genetically modified” or “modified cell” refers to a cell in which the nuclear,organellar or extrachromosomal nucleic acid sequences of a cell has been transformed, modified or transduced using recombinant DNA technology to comprise a heterologous nucleic acid molecule, and is used interchangeably with “engineered cell,” “transformed cell,” and “transduced cell.” In one aspect, a genetically modified cell as disclosed herein expresses a protein encoded by a nucleic acid molecule engineered in such manner to contain an insertion of at least one nucleotide, a deletion of at least one nucleotide, and / or a substitution of at least one nucleotide in a sequence encoding at least one heterologous protein. "Engineered cells" refers herein to cells having been engineered, e.g., by the introduction of an exogenous nucleic acid sequence as defined herein. Such a cell has been genetically modified for example by the introduction of for example one or more mutations, insertions and / or deletions in the endogenous gene and / or insertion of a genetic construct in the genome. The modification may have been introduced using recombinant DNA technology. An engineered cell may refer to a cell in isolation or in culture. Engineered cells may be "transduced cells" wherein the cells have been infected with e.g., a modified virus, for example, a retrovirus may be used but other suitable viruses may also be contemplated such as lentiviruses. Non-viral methods may also be used, such as transfections. Engineered cells may thus also be "stably transfected cells" or "transiently transfected cells". Transfection -87- 99975521.7Attorney Docket No: 124540-829360 refers to non-viral methods to transfer DNA (or RNA) to cells such that a gene is expressed. Transfection methods are widely known in the art, such as calcium phosphate transfection, PEG transfection, and liposomal or lipoplex transfection of nucleic acids. Such a transfection may be transient but may also be a stable transfection wherein cells can be selected that have the gene construct integrated in their genome. In various cases, a genetically engineered cell is edited via a CRISPR system using the Cas12f polypeptides and / or sgRNAs provided herein.

[0229] “Pharmaceutical composition” means a mixture of substances suitable foradministering to an individual that includes a pharmaceutical agent. As used herein a pharmaceutical composition comprises one or more of receptors, vectors, cells disclosed herein compounded with suitable pharmaceuticals carriers or excipients.

[0230] “Treatment” or “therapy” of a subject refers to any type of intervention or processperformed on, or the administration of an active agent to, the subject with the objective of reversing, alleviating, ameliorating, inhibiting, slowing down or preventing the onset, progression, development, severity or recurrence of a symptom, complication, condition or biochemical indicia associated with a disease.

[0231] Subject: As used herein, the term “patient”, “subject”, or “test subject” refersinterchangeably to any organism to which provided compound or compounds described herein are administered in accordance with the present disclosure, e.g., for experimental, diagnostic, prophylactic, and / or therapeutic purposes. In some embodiments, the subject is a mammal (e.g., mice, rats, rabbits, non-human primates, humans, insects, worms etc.). In some embodiments, a subject is non-human primate or rodent. In some embodiments, a subject is a human. In some embodiments, a subject has, is suspected of having, or is at risk for, a disease or disorder. In some embodiments, a subject has one or more symptoms of a disease or disorder. As used herein, a “patient population” or “population of subjects” refers to a plurality of patients or subjects.

[0232] The term "effective amount" as used herein is defined as the amount of themolecules of the present disclosure that are necessary to result in the desired physiological change in the cell or tissue to which it is administered. The term "therapeutically effective amount" as used herein is defined as the amount of the molecules of the present disclosure that achieves a desired effect with respect to cancer. In this context, a “desired effect” is synonymous with “an anti-tumor activity” as earlier defined herein. A skilled artisan readily -88- 99975521.7Attorney Docket No: 124540-829360 recognizes that in many cases the molecules may not provide a cure but may provide a partial benefit, such as alleviation or improvement of at least one symptom or parameter. In some embodiments, a physiological change having some benefit is also considered therapeutically beneficial. Thus, in some embodiments, an amount of molecules that provides a physiological change is considered an "effective amount" or a "therapeutically effective amount."

[0233] As used herein, the disclosure of numerical ranges by numerical endpoints includesall numbers encompassed by that range (e.g., “1 to 5” includes but is not limited to 1, 1.25, 1.5, 1.75, 2, 2.3, 2.5, 2.8, 3, 3.1,3.3, 3.8, 3.9, 4, 4.25, 4.5, 4.75 and 5). Unless otherwise indicated, all numbers used herein to express quantities, amounts, dimensions, measurements, and the like should be understood as encompassing the specific quantities, amounts, dimensions, measurements and so on, and also as encompassing such instances modified by the term “about.” Accordingly, unless indicated to the contrary, the numerical descriptions set forth herein may vary while remaining well within the teachings of the present disclosure. At the very least, each numerical value should be construed in view of the number of significant digits and by applying routine rounding techniques. As various changes could be made in the above-described cells and methods without departing from the scope of the disclosure, it is intended that all matter contained in the above description and in the examples given below, shall be interpreted as illustrative and not in a limiting sense. As various changes could be made in the above-described cells and methods without departing from the scope of the disclosure, it is intended that all matter contained in the above description and in the examples given below, shall be interpreted as illustrative and not in a limiting sense. V. Examples

[0234] The examples below describe generation and characterization of specific compactCas nucleases according to the present disclosure. Example 1: Cas12f protein design

[0235] Eight wildtype Cas12f sequences including BsCas12f1 (from Blautia species, SEQID NO: 1), UnCas12f1 (from Uncultured bacterium, SEQ ID NO: 5), CnCas12f1 (from Clostridium novyi, SEQ ID NO: 2), Pt1Cas12f1 (from Parageobacillus thermoglucosidasius, SEQ ID NO: 3), Ob3Cas12f1 (from Oscillospiraceae bacterium, SEQ ID NO: 7), Un2Cas12f1 (from Uncultured bacterium, SEQ ID NO: 6), Pt2Cas12f1 (from Parageobacillus thermoglucosidasius, SEQ ID NO: 12) and SpCas12f1 (from Syntrophomonas palmitatica, SEQ ID NO: 4) were obtained. These eight Cas12f1s and their gRNAs were tested for their ability to target DNA in reporter cells. For the testing, each of -89- 99975521.7Attorney Docket No: 124540-829360 the Cas12fs were mutated into nuclease activity-deficient Cas (dCas) form by introducing a D216A and D390A mutation into the wildtype BsCas12f nuclease (SEQ ID NO: 1) to generate dBsCas12f (SEQ ID NO: 393) or equivalent mutations into the remaining Cas12f nucleases (provided herein in Table 3 above), according to the sequence and structure of each nuclease. The creation of dCas enabled the screening using a reporter cell platform, the CRISPRa-based reporter - TetO-TurboRFP (further details in Example 3).

[0236] A nucleic acid encoding dBsCas12f1 above was synthesized using Integrated DNATechnologies’ gBlock service. The synthesized DNA fragment was subsequently cloned into a pCS1-VPR vector (SEQ ID NO: 588) to create a fusion protein with an open reading frame containing a nucleic acid encoding NLS-dCas-XTEN linker-VPR-NLS under the control of a UBC promoter. This fusion protein (SEQ ID NO: 418, encoded by a nucleic acid of SEQ ID NO: 419) contains the dBsCas12f1 protein linked to a VPR transcriptional activator (VP64- p65-Rta) via a flexible XTEN linker. The pCS1-VPR vector also carries an EGFP-Puro- expressing cassette, under the control of EF1a promoter. FIG.1 presents a schematic of an empty pCS1-VPR vector and FIG.2 provides a schematic of the pCS1-dBsCas12-VPR vector encoding the NLS-dCas-XTEN linker-VPR fusion protein. Example 2 – sgRNA cloning and vector generation

[0237] A wildtype gRNA scaffold (174 nt, SEQ ID NO: 438) predicted by XiangfengKong, et al., 2023 (Nat. Communications 2023 April 11; 14(1):2046, incorporated herein by reference in its entirety) together with a 20-nt long spacer sequence (SEQ ID NO: 516) targeting a Tet operator (TetO) in the CRISPRa reporter system (see Example 3, below) were synthesized using Integrated DNA Technologies’ service cloned into a pMINI vector (see FIG.3, SEQ ID NO: 591). The pMINI vector contains a U6 promoter to drive the expression of gRNA. The wildtype scaffold was further engineered to a shorter version (140-nt) with stronger efficiency (SEQ ID NO: 440). Wildtype scaffolds for each of the remaining Cas12f polypeptides tested in Example 1 were also obtained (see e.g., SEQ ID NOs: 666-672). FIG. 4 depicts the pMINI-BsCas12f gRNA (174-nt)-TetO vector (SEQ ID NO: 591). FIG.5 depicts the pMINI-BsCas12f gRNA (140-nt)-TetO vector (SEQ ID NO: 592). For ease of reference, Table 9, below provides the gRNA scaffolds used in this Example. Table 9 gRNA scaffold Sequence SEQ ID NO: -90- 99975521.7Attorney Docket No: 124540-829360 Wildtype BsCas12f gRNA ACCGCTTCACTTAGAGTGAAGGTGGGCTGCTT 438 scaffold (174 nt) GCATCAGCCTAATGTCGAGAAGTGCTTTCTTCG GAAAGTAACCCTCGAAACAAAGAAAGGAATG CAAC Shortened gRNA scaffold ACAGGGCGATTTAACGTCCTAAGGCTGAGAGA 440 (140 nt) (also referred to as AGTTCCTTCTACTCGGCAAGGGTTAATCTCGAT S003, below) TGTTGTGTTACCGATCGAGCGTTTCACAGAAAT GTGAAATGTAAAT WT dUnCas12f1 scaffold AGTCGAGAAGTGCCGTAATAAGCATCTAAAAA 666 TGCCTAACGGTAACACTCGATAAGGTAGTCCT GCTAGGCAGGCTGAAACCCTAGCCACAAAATC CGGCTAGGCATCATACGAAAATGTATGATGTG A WT dCnCas12f1 scaffold AAGGGACGACTTCCCGTCCCAAAATCGAGATA 667 GTGGTCCTGATTCTTTGATTTCAAAGCGGACAA TACACTCGATAAGGTTAAGATGCACATAGGAA TCCGTGCATGGGTCACAGAAATGTGACTTGAA GG WT dPt1Cas12f1 scaffold AATGTTATTCCATAATAACATTTGATGCACACG 668 ATTCCTCCCTACAGTAGTTAGGTATAGCCGAA AGGTAGAGACTAAATCTGTAGTTGGAGTGGGC CGCTTGCATCGGCCTAAAGTTGAGAAGTGTCA GACTCTGATAACCCTCAACGACGATATTCTTTA TTTCGGAAACGAATGAAGGAATGCAAC WT dOb3Cas12f1 scaffold AGGGTGAGGGTATAGATAAAACGCATAAGGTA 669 GTATGCCAAATATGTGCTATAACCACTCGCTA AGCCGAAAAAACCTTAGTTTATGATGGCAACT AAGCACACTATGAAAGTAGTGTGTAAAC WT dUn2Cas2f1 scaffold CTCTGTTTCGCGCGCCAGGGCAGTTAGGTGCCC 670 TAAAAGAGCGAAGTGGCCGAAAGGAAAGGCT AACGCTTCTCTAACGCTACGGCGACCTTGGCG AAATGCCATCAATACCACGC GAAAAACGCGTGGATTGAAAC WT dPt2Cas12f1 scaffold ACCGCTTCACTTAGAGTGAAGGTGGGCTGCTT 671 GCATCAGCCTAATGTCGAGAAGTGCTTTCTTCG GAAAGTAACCCTCGAAACAAAGAAAGGAATG CAAC WT dSpCas12f1 scaffold ACAGGGCGATTTAACGTCCTAAGGCTGAGAGA 672 AGTTCCTTCTACTCGGCAAGGGTTAATCTCGAT TGTTGTGTTACCGATCGAGCGTTTCACAGAAAT GTGAAATGTAAAT Example 3- Testing Cas12f for DNA binding ability

[0238] The Cas12f polypeptides described in Example 1 were tested for activity using aTurboRFP system. The rationale of this system is intended to test the target DNA binding ability of the selected dCas. As shown in FIG.6, The reporter is designed to entail a TurboRFP-expressing cassette under the control of TRE, which contains 7 tandem repeats of -91- 99975521.7Attorney Docket No: 124540-829360 Tet operators. Without proper transcriptional modulators, the TurboRFP gene remains silenced and no red fluorescence will be generated. If a specific and efficient dCas-VPR fusion protein exists, which acts as a targeted transcriptional activator, its binding to TetO operator in TRE will activate the downstream TurboRFP gene and thus generate red fluorescence.

[0239] To generate reporter cells, three lentivector plasmid containing the TRE-TurboRFP-expressing cassettes were generated by VectorBuilder. All three reporter plasmids contain identical backbone sequences with TRE containing seven tandem repeats of Tet operator sequences. However, the 4 nucleotides immediately upstream of each of the Tet operator were modified to contain either TTTA, TTTC or TCCA DNA sequences. This difference enables the testing of a variety of dCas species, which were predicted to have different protospacer adjacent motif (PAM) preferences. The three reporter plasmids were packaged into lentiviral vectors by VectorBuilder. The lentiviruses were transduced into HEK293T cells and purified by clonal selection to ensure genomic integration and consistency of the reporter gene. FIG.7 depicts a representative plasmid (specifically, the TCCA plasmid (SEQ ID NO: 593). The reporter plasmids containing the TTTA and TTTC DNA sequences are provided herein as SEQ ID NO: 594 and 595, respectively.

[0240] Outcome measurement: The dCas expressing plasmids (e.g., pCS1-dBsCas12f-VPR, SEQ ID NO: 589) and gRNA plasmids (e.g., pMINI-BsCas12f gRNA (174-nt)-TetO, SEQ ID NO: 591) were transfected into the reporter cells using Lipofectamine 2000. For the transfection, the reporter cells were seeded into 24-well plates at 5 x 10^4 cells / well. Each of the wells were transfected with one kind of dCas expressing plasmids and one kind of gRNA plasmids with the ratio at 900ng:300ng. For each well, 3.6 µl of Lipofectamine was used. The fluorescence was determined after 72 hrs of transfection. The fluorescence of the reporter reflects the efficiency of the dCas-VPR fusion protein (the “editor”). The stronger the TurboRFP fluorescence, the stronger the editor is. The red fluorescence was determined using NovoCyte flow cytometer. FIG.8A and 8B show that dBsCas12f (SEQ ID NO: 393), dUnCas12f1 (SEQ ID NO: 397), and dCnCas12f (SEQ ID NO: 394) all showed activity in this system.

[0241] In another experiment, dAsCas12f (SEQ ID NO: 407), dUnCas12f (SEQ ID NO:397), dPt1Cas12f (SEQ ID NO; 395) and dUn2Cas12f (SEQ ID NO: 398) were tested using the TTTA reporter system, dOb3Cas12f (SEQ ID NO: 399) and dSpCas12f (SEQ ID NO: 396) were tested with the TTTC reporter system and dPt2Cas12f (SEQ ID NO:404), -92- 99975521.7Attorney Docket No: 124540-829360 dCnCas12f (SEQ ID NO: 394) and dBsCas12f (SEQ ID NO: 393) were tested using the TCCA reporter system. Results from these experiments are summarized in FIG.8C. In further experiments below, dBsCas12f was tested and optimized using the TCCA reporter system. Example 4 – Additional Cas12f polypeptides

[0242] In further experiments, the dCas12 protein (e.g., SEQ ID NO: 393) were furtherengineered using the same pCS1-VPR vector to introduce different mutations into the wildtype dBsCas12f. The mutations include: Y48F, Y48R, A50R, A50Y, T66Q, S98H, S98R, T99R, E109M, E109R, N114Q, N114A, F119Y, F119H, M122R, S131R, S131H, G176R, G176H, D183K, K184R, N195Q, N195R, A223R, G266R, A276R, Y306F, Y306M, K318R, K318G and D353R. The geometric mean fluorescence ratio was set as the outcome measurement to evaluate the improvement and results are shown in FIG.9. Based on these results, it was concluded that T66, E109, S131, D183, K184, N195, K220, A276, Y306, K318 and D353 are important amino acids and they can be further engineered to improve the editing efficiency. Example 5 – Systematic Optimization of bsCas9 using Saturated Mutagenesis

[0243] In further experiments, the BsdCas12 protein was again engineered using the samepCS1-VPR vector to introduce 1 to 7 simultaneous mutations into the wildtype dBsCas12f. The selection of mutations for testing was based on sequence alignment and artificial intelligence as described below.

[0244] First, using sequence alignment between other Cas12fs, 21 different mutationswere predicted as important for activity and then mutated and tested using the system described above. These substitutions included: Y48(F / R), A50(R / Y), T66Q, S98(H / R), T99R, E109(M / R), N114(Q / A), F119(Y / H), M122R, S131(H / R), G176(R / H), D183K, K184R, N195Q, N195R, K220R, A223R, G266R, A276R, Y306F, Y306M, K318R, K318G, and DS53R. Table 10, below shows the percentage of TurboRFP+ cells and calculated efficiency relative to WT for each mutation tested. From this data, 10 amino acid residues were then selected for saturated mutagenesis and analysis: E109, S131, E181, D183, K184, N195, K220, Y306, K318, and D353. Table 10: Selective Mutagenesis of Residues of Interest TurboRFP+ Population / WT SEQ ID TurboRFP+ Population / WT SEQ ID Mutation cell (%) (%) NO: Mutation cell (%) (%) NO: -93- 99975521.7Attorney Docket No: 124540-829360 WT* 28 100 393 G176R 10.3 36.71 34 Y48F 14.2 50.85 16 G176H 15.3 54.59 35 Y48R 6.8 24.35 17 D183K 27.5 98.23 36 A50R 18 64.29 18 K184R 30.1 107.41 37 A50Y 21.3 76.09 19 N195Q 34.5 123.26 38 T66Q 22.2 79.07 20 N195R 49.4 176.4 39 S98H 7.4 26.38 21 K220R 26 92.92 40 S98R 6.1 21.63 23 A223R 4.3 15.18 41 T99R 5.5 19.65 24 G266R 7.4 26.28 42 E109M 34.4 122.86 25 A276R 22.1 78.88 43 E109R 22.7 81.04 26 Y306F 24.8 88.57 44 N114Q 5.7 20.24 27 Y306M 25.5 91.18 45 N114A 7.5 26.73 28 K318R 28.3 101.06 46 F119Y 10.5 37.61 29 K318G 26.4 94.12 47 F119H 26.5 94.74 30 D353R 46.7 166.67 48 M122R 5.6 20.06 31 S131H 29.5 105.18 32 S131R 25.9 92.34 33 * “WT” in all examples herein refers to dBsCas12 polypeptide comprising D216A D390A relative to SEQ ID NO: 1, unless otherwise stated.

[0245] In another method, ESM (please define) was used for AI-based prediction ofpossible substitutions that might improve activity. The results of the AI based prediction is shown in FIG.11. From this method, positions T307, N312 and W143 were flagged as important for further analysis.

[0246] 10 residues (E109, S131, E181, D183, K184, N195, K220, Y306, K318, andD353) were systematically mutated to each of the 20 canonical amino acids and tested in the TetO reporter system described above. In addition, two additional residues (T66 and A276) were also systematically mutated and will be tested in the Tet0 reporter system. As shown in Table 11, below, certain residues showed improved efficiency relative to wildtype (right columns). These results were then used to select the best variant at each position to generate mutants with 1-7 mutations. Table 11: Saturated Mutagenesis at Residues of Interest -94- 99975521.7Attorney Docket No: 124540-829360 TurboRFP+ Efficiency / WT SEQ TurboRFP+ Efficiency / WT SEQ ID E109 S131 cell (%) (%) ID NO cell (%) (%) NO E109M 22.12 132.04 25 S131H 23.34 120.53 32E109R 10.4 121.87 26 S131R 30.96 112.08 33E109A 16.52 120.01 49 S131A 33.14 140.01 66E109C 26.88 164.45 50 S131C 32.64 138.08 67E109D 6.55 90.73 51 S131D 28.05 116.01 68E109F 20.08 126.36 52 S131E 28.38 112.41 69E109G 2.03 80.79 53 S131F 30.11 112.01 70 E109H 34.9 164.4 54 S131G 28.9 113.05 71E109I 27.88 154.78 55 S131I 30.07 123.59 72 E109K 15.07 115.41 56 S131K 31.25 119.58 73E109L 28.09 141.73 57 S131L 31.16 123.49 74E109N 19.34 134.16 58 S131M 32.55 127.59 75E109P 1.35 58.95 59 S131N 32.01 120.89 76E109Q 22.69 131.02 60 S131P 9.08 62.88 77 E109S 20.22 144.85 61 S131Q 30.68 154.87 78E109T 33.79 147.92 62 S131T 25.72 110.23 79E109V 19.03 126.09 63 S131V 33.01 136.23 80E109W 31.4 153.41 64 S131W 27.08 117.12 81E109Y 23.98 128.64 65 S131Y 14.7 82.96 82TurboRFP+ Efficiency / WT SEQ TurboRFP+ Efficiency / WT SEQ ID E181 D183 cell (%) (%) ID NO cell (%) (%) NO E181A 32.28 109.51 83 D183K 30.7 103.19 36E181C 24.82 84.51 84 D183A 33.31 118.25 102E181D 11.59 58.91 85 D183C 25.78 94.93 103E181F 13.8 64.94 86 D183E 14.12 71.87 104E181G 32.01 95.91 87 D183F 14.39 67.94 105E181H 15.3 60.25 88 D183G 22.88 77.56 106E181I 19.04 72.7 89 D183H 21.02 68.54 107E181K 18.36 63.98 90 D183I 24.63 85.56 108 -95- 99975521.7Attorney Docket No: 124540-829360 E181L 28.42 93.89 91 D183L 18.93 69.35 109E181M 34.09 102.82 92 D183M 22.8 88.18 110E181N 32.56 114.25 93 D183N 35.85 124.82 111E181P 6.88 41.83 94 D183P 12.75 62.8 112E181Q 38.25 145.74 95 D183Q 28.59 84.37 113E181R 8.02 35.45 96 D183R 21.1 84 114E181S 27.06 90.75 97 D183S 24.49 97.29 115E181T 16.26 65.82 98 D183T 35.95 129.95 116E181V 19.37 75.32 99 D183V 21.43 95.2 117E181W 9.55 64.72 100 D183W 14.62 64.5 118E181Y 6.47 59.16 101 D183Y 10.62 57.44 119SEQ SEQ ID TurboRFP+ Efficiency / WT TurboRFP+ Efficiency / WT K184 ID N195 NO: cell (%) (%) cell (%) (%) NO: K184R 25.81 102.79 37 N195Q 19.62 119.64 38K184A 8.14 56.19 120 N195R 29.78 126.5 39K184C 7.04 56.15 121 N195A 30.89 144.34 138K184D 4.39 53.46 122 N195C 17.69 110.85 139K184E 7.34 76.55 123 N195D 2.91 74.65 140K184F 7.47 49.58 124 N195E 9.76 86.36 141K184G 7.23 43.77 125 N195F 16.4 107.05 142K184H 11.7 69.3 126 N195G 12.47 92.32 143K184I 8.6 53.36 127 N195H 57.22 239.4 144K184L 9.01 71.51 128 N195I 16.28 101.43 145 K184M 7.2 57.79 129 N195K 54.68 232.71 146K184N 7.02 51.17 130 N195L 16.29 97.68 147K184P 7.36 51.43 131 N195M 9.62 76.53 148K184Q 8.91 69.75 132 N195P 1.41 68.13 149K184S 8.87 91.51 133 N195S 20.32 119.46 150K184T 8.61 77.28 134 N195T 15.72 101.61 151K184V 6.73 68.26 135 N195V 13.05 92.57 152K184W 7.96 62.7 136 N195W 47.66 219.19 153-96- 99975521.7Attorney Docket No: 124540-829360 K184Y 7.58 63.87 137 N195Y 29.79 139.59 154SEQ SEQ ID TurboRFP+ Efficiency / WT TurboRFP+ Efficiency / WT K220 ID Y306 NO: cell (%) (%) cell (%) (%) NO: K220R 23.01 89.39 40 Y306F 16.04 79.21 44K220A 12.61 58.27 155 Y306M 21.47 90.41 45K220C 12.67 63.89 156 Y306A 18.03 100.1 173K220D 8.67 59.43 157 Y306C 24.82 94.18 174K220E 9.55 55.74 158 Y306D 7.56 61.34 175K220F 13.38 62.2 159 Y306E 23.45 98.33 176K220G 11.85 57.12 160 Y306G 23.98 93.89 177K220H 15.06 62.64 161 Y306H 27.86 118.8 178K220I 12.65 60 162 Y306I 12.83 64.57 179 K220L 14.39 58.75 163 Y306K 49.95 235.36 180K220M 14.02 60.73 164 Y306L 16.33 70.48 181K220N 13.57 64.17 165 Y306N 29.12 114.85 182K220P 22.97 86.17 166 Y306P 7.46 56.61 183K220Q 13.73 63.85 167 Y306Q 31.85 127.12 184K220S 10.28 53.33 168 Y306R 33.81 168.29 185K220T 9.61 55.61 169 Y306S 19.24 97.1 186K220V 11.14 64.35 170 Y306T 21.41 95.47 187K220W 11.41 60.93 171 Y306V 15.6 72.6 188K220Y 6.64 37.95 172 Y306W 19.87 78.56 189SEQ SEQ ID TurboRFP+ Efficiency / WT TurboRFP+ Efficiency / WT K318 ID D353 NO: cell (%) (%) cell (%) (%) NO: K318R 22.61 90.56 46 D353R 46.56 166.02 48K318G 18.48 91.06 47 D353A 44.52 170.38 207K318A 19.22 79.81 190 D353C 42.8 150.93 208K318C 12.99 77.54 191 D353E 34.56 128.45 209K318D 8.9 56.3 192 D353F 41.24 150.83 210K318E 8.23 55.34 193 D353G 41 146.37 211-97- 99975521.7Attorney Docket No: 124540-829360 K318F 12.72 57.57 194 D353H 42.35 146.49 212K318H 19.16 92.06 195 D353I 43.91 161.78 213 K318I 16.85 69.3 196 D353K 50.23 189.54 214K318L 15.2 69.44 197 D353L 38.74 151.04 215K318M 13.62 70.12 198 D353M 42.55 160.3 216K318N 15.34 66.55 199 D353N 36.45 133.02 217K318P 14.99 61.51 200 D353P 34.67 117.19 218K318Q 18.02 92.45 201 D353Q 45.25 156.94 219K318S 22.15 107.98 202 D353S 36.49 135.33 220K318T 10.26 58.88 203 D353T 44.91 178.8 221K318V 15.85 72.46 204 D353V 43.39 168.5 222K318W 16.15 70.43 205 D353W 41.77 176.02 223K318Y 18.67 81.43 206 D353Y 40.13 153.21 224SEQ SEQ ID TurboRFP+ Efficiency / WT TurboRFP+ Efficiency / WT T66 ID A276 NO: cell (%) (%) cell (%) (%) NO: T66F to be225 A276F TBD TBD 243determined (TBD) TBD T66L TBD TBD 226 A276L TBD TBD 244T66I TBD TBD 227 A276I TBD TBD 245 T66M TBD TBD 228 A276M TBD TBD 246T66V TBD TBD 229 A276V TBD TBD 247T66S TBD TBD 230 A276S TBD TBD 248T66P TBD TBD 231 A276P TBD TBD 249T66A TBD TBD 232 A276T TBD TBD 250T66Y TBD TBD 233 A276Y TBD TBD 251T66H TBD TBD 234 A276H TBD TBD 252T66N TBD TBD 235 A276Q TBD TBD 253T66K TBD TBD 236 A276N TBD TBD 254T66D TBD TBD 237 A276K TBD TBD 255T66E TBD TBD 238 A276D TBD TBD 256-98- 99975521.7Attorney Docket No: 124540-829360 T66C TBD TBD 239 A276E TBD TBD 257T66W TBD TBD 240 A276C TBD TBD 258T66R TBD TBD 241 A276W TBD TBD 259T66G TBD TBD 242 A276G TBD TBD 260* “WT” in all examples herein refers to dBsCas12 polypeptide comprising D216A D390A relative to SEQ ID NO: 1, unless otherwise stated.

[0247] In further experiments, targeted substitutions at residues T307, N312, W143 weretested. These substitutions and results are summarized in the Table 12 below. Table 12 – Additional targeted substitutions Mutation TurboRFP+ cell (%) Population / WT (%) SEQ ID NO:T307G 14.75 102.48 261 T307K 26.93 139.81 262 T307N 25.88 138.45 263 N312V 18.11 110.47 264 N312Y 15.62 97.88 265 W143T 3.16 76.46 266 * “WT” in all examples herein refers to dBsCas12 polypeptide comprising D216A D390A relative to SEQ ID NO: 1, unless otherwise stated.

[0248] It was then tested whether combining substitutions could further improve editingefficiency. To this end, double, triple, quadruple, quintuple, sextuple substitutions were tested. In addition, variants comprising seven substitutions were also generated. Some were tested and data is provided below, relative to the variant with six substitutions (i.e., SEQ ID NO: 338). The rest will be tested as described herein. A list of these mutation combinations and their performances are shown in Tables 13A-13H below. Table 13A – Double Substitution Combinations TurboRFP+ cell Efficiency / WT Substitution (N195K based) SEQ ID NO: (%) (%) N195K 69.6 320.24 146 N195K, E109H 75.94 425.8 267 N195K, S131Q 74.49 382.45 268 N195K, S131C 75.62 407.14 269 -99- 99975521.7Attorney Docket No: 124540-829360 N195K, E181Q 81.6 555.48 270 N195K, D183T 80.7 486.61 271 N195K, Y306K 87.06 825.94 272 N195K, T307K 79.69 495.84 273 N195K, D353K 81.29 640.17 274 TurboRFP+ cell Efficiency / WT Substitution (N195H based) SEQ ID NO: (%) (%) N195H 60.39 305.9 144 N195H, E109H 72.51 434.45 275 N195H, S131C 71.88 385.93 276 N195H, E181Q 79.06 465.43 277 N195H, D183T 75.52 433.13 278 N195H, Y306K 83.49 769.61 279 N195H, T307K 76.61 506.05 280 N195H, D353K 76.07 544.99 281 * “WT” in all examples herein refers to dBsCas12 polypeptide comprising D216A D390A relative to SEQ ID NO: 1, unless otherwise stated. Table 13B – Additional Double Substitution Combinations TurboRFP+ cell Efficiency / WT Substitutions SEQ ID NO: (%) (%) N195H, S131Q TBD TBD 282 T66F, A112V TBD TBD 283 E109I, S162G TBD TBD 284 E181N, L175H TBD TBD 285 E181S, Q249K TBD TBD 286 E181Y, R247H TBD TBD 287 K184G, A355T TBD TBD 288 K184L, D322Y TBD TBD 289 K184P, V219A TBD TBD 290 K184P, Y258H TBD TBD 291 K184V, A112V TBD TBD 292 K318M, S263Y TBD TBD 293 -100- 99975521.7Attorney Docket No: 124540-829360 D353A, R124M TBD TBD 294 S131G, A389V TBD TBD 295 S131G, R110W TBD TBD 296 D183Y, D368Y TBD TBD 297 Table 13C – Triple Substitution Combinations Substitution (N195K, TurboRFP+ Efficiency / WT SEQ ID NO: Y306K-based) cell (%) (%) N195K,Y306K,E109H 72.56 423.34 298 N195K,Y306K,S131C 80.48 639.06 299 N195K,Y306K,S131Q 81.17 589.53 300 N195K,Y306K,E181Q 83.71 674.61 301 N195K,Y306K,D183T 83.02 663.48 302 N195K,Y306K,T307K 83.04 756.87 303 N195K,Y306K,D353K 77.63 561.07 304 Substitution (N195H, TurboRFP+ Efficiency / WT SEQ ID NO: Y306K-based) cell (%) (%) N195H, Y306K E109H, 76.79 503.93 305 N195H, Y306K S131C, 80.31 616.36 306 N195H, Y306K S131Q, 77.64 577.24 307 N195H, Y306K E181Q, 84.19 688.09 308 N195H, Y306K D183T, 80.68 606.94 309 N195H, Y306K,T307K 79.93 623.43 310 N195H, Y306K,D353K 75.15 520.18 311 * “WT” in all examples herein refers to dBsCas12 polypeptide comprising D216A D390A relative to SEQ ID NO: 1, unless otherwise stated. -101- 99975521.7Attorney Docket No: 124540-829360 Table 13D – Additional Triple Substitution Combinations Substitution TurboRFP+ Efficiency / WT cell (%) (%) SEQ ID NO: E181G, N195R, D353R TBD TBD 312 D76G, S131Q, N195H TBD TBD 313 N195K, S131C, N317D TBD TBD 314 E109H,N195K,D353K TBD TBD 315 S131C,N195K,D353K TBD TBD 316 S131Q,N195K,D353K TBD TBD 317 E181Q,N195K,D353K TBD TBD 318 D183T,N195K,D353K TBD TBD 319N195K,T307K,D353K TBD TBD 320 E109H,N195H,D353K TBD TBD 321 S131C,N195H,D353K TBD TBD 322 S131Q,N195H,D353K TBD TBD 323 E181Q,N195H,D353K TBD TBD 324 D183T,N195H,D353K TBD TBD 325 N195H,T307K,D353K TBD TBD 326 E109H,N195H,D353K TBD TBD 327 S131C,N195H,D353K TBD TBD 312 S131Q,N195H,D353K TBD TBD 313 E181Q,N195H,D353K TBD TBD 314 D183T,N195H,D353K TBD TBD 315 N195H,T307K,D353K TBD TBD 316 N195H, S131Q, A342E TBD TBD 317 Table 13E – Quadruple Substitution Combinations Substitution TurboRFP+ Efficiency / WT (N195K,Y306K, T307K SEQ ID NO: cell (%) (%) based) E109M, E181G, N195R, TBD TBD 328 D353R -102- 99975521.7Attorney Docket No: 124540-829360 N195K, Y306K, T307K, 329 75.19 486.28 E109HN195K, Y306K, T307K S131C, 82.81 689.38 330N195K, Y306K, T307K, 331 79.59 713.21 S131Q N195K, Y306K, T307K, 332 82.71 727.37 E181Q N195K, Y306K, T307K, 333 81.65 746.08 D183T N195K, Y306K, T307K, 334 79.92 667.73 D353K * “WT” in all examples herein refers to dBsCas12 polypeptide comprising D216A D390A relative to SEQ ID NO: 1, unless otherwise stated. Table 13F – Quintuple Substitution Combinations Substitution TurboRFP+ cell SEQ ID Efficiency / WT (%) (N195K,Y306K, T307K based (%) NO:N195K,Y306K,T307K, S131C, E181Q 83.2 657.03 335N195K,Y306K,T307K, S131C, E183T 83.46 739.54 336N195K,Y306K,T307K, E181Q, E183T 82.76 816.14 337* “WT” in all examples herein refers to dBsCas12 polypeptide comprising D216A D390A relative to SEQ ID NO: 1, unless otherwise stated. Table 13G – Sextuple Substitution Combinations TurboRFP+ cell Substitution (%) Efficiency / WT (%) SEQ ID NO: S131C, E181Q, E183T, 338 83.66 835.35 N195K,Y306K,T307K * “WT” in all examples herein refers to dBsCas12 polypeptide comprising D216A D390A relative to SEQ ID NO: 1, unless otherwise stated. -103- 99975521.7Attorney Docket No: 124540-829360 Table 13H = Septuple Substitution Combinations Substitution (N195K, Y306K, TurboRFP+ cell Efficiency / SEQ T307K, S131C E181Q, D183T SEQ ID NO: (%) ID NO: 338 (%) based) N195K, Y306K, T307K, S131C, 71.8 152.5 339 E181Q, D183T, I10R N195K, Y306K, T307K, S131C, 71.8 124.4 340 E181Q, D183T, I10H N195K, Y306K, T307K, S131C, 67.1 101.0 341 E181Q, D183T, D76R N195K, Y306K, T307K, S131C, 65.9 105.2 342 E181Q, D183T, D76G N195K, Y306K, T307K, S131C, 47.3 55 343 E181Q, D183T, N103R N195K, Y306K, T307K, S131C, 23.7 28.1 344 E181Q, D183T, G104R N195K, Y306K, T307K, S131C, 11.7 14.4 345 E181Q, D183T, G104M N195K, Y306K, T307K, S131C, 63.2 100.3 346 E181Q, D183T, G111R N195K, Y306K, T307K, S131C, 48.1 67 347 E181Q, D183T, G111H N195K, Y306K, T307K, S131C, TBD TBD 348 E181Q, D183T, I201A N195K, Y306K, T307K, S131C, TBD TBD 349 E181Q, D183T, I201C N195K, Y306K, T307K, S131C, TBD TBD 350 E181Q, D183T, I201D N195K, Y306K, T307K, S131C, TBD TBD 351 E181Q, D183T, I201E N195K, Y306K, T307K, S131C, TBD TBD 352 E181Q, D183T, I201F -104- 99975521.7Attorney Docket No: 124540-829360 N195K, Y306K, T307K, S131C, TBD TBD 353 E181Q, D183T, I201G N195K, Y306K, T307K, S131C, TBD TBD 354 E181Q, D183T, I201H N195K, Y306K, T307K, S131C, TBD TBD 355 E181Q, D183T, I201K N195K, Y306K, T307K, S131C, TBD TBD 356 E181Q, D183T, I201L N195K, Y306K, T307K, S131C, TBD TBD 357 E181Q, D183T, I201M N195K, Y306K, T307K, S131C, TBD TBD 358 E181Q, D183T, I201N N195K, Y306K, T307K, S131C, TBD TBD 359 E181Q, D183T, I201P N195K, Y306K, T307K, S131C, TBD TBD 360 E181Q, D183T, I201Q N195K, Y306K, T307K, S131C, TBD TBD 361 E181Q, D183T, I201R N195K, Y306K, T307K, S131C, TBD TBD 362 E181Q, D183T, I201S N195K, Y306K, T307K, S131C, TBD TBD 363 E181Q, D183T, I201T N195K, Y306K, T307K, S131C, TBD TBD 364 E181Q, D183T, I201V N195K, Y306K, T307K, S131C, TBD TBD 365 E181Q, D183T, I201W N195K, Y306K, T307K, S131C, TBD TBD 366 E181Q, D183T, I201Y N195K, Y306K, T307K, S131C, TBD TBD 367 E181Q, D183T, P209A N195K, Y306K, T307K, S131C, TBD TBD 368 E181Q, D183T, P209C -105- 99975521.7Attorney Docket No: 124540-829360 N195K, Y306K, T307K, S131C, TBD TBD 369 E181Q, D183T, P209D N195K, Y306K, T307K, S131C, TBD TBD 370 E181Q, D183T, P209E N195K, Y306K, T307K, S131C, TBD TBD 371 E181Q, D183T, P209F N195K, Y306K, T307K, S131C, TBD TBD 372 E181Q, D183T, P209G N195K, Y306K, T307K, S131C, TBD TBD 373 E181Q, D183T, P209H N195K, Y306K, T307K, S131C, TBD TBD 374 E181Q, D183T, P209I N195K, Y306K, T307K, S131C, TBD TBD 375 E181Q, D183T, P209K N195K, Y306K, T307K, S131C, TBD TBD 376 E181Q, D183T, P209L N195K, Y306K, T307K, S131C, TBD TBD 377 E181Q, D183T, P209M N195K, Y306K, T307K, S131C, TBD TBD 378 E181Q, D183T, P209N N195K, Y306K, T307K, S131C, TBD TBD 379 E181Q, D183T, P209Q N195K, Y306K, T307K, S131C, TBD TBD 380 E181Q, D183T, P209R N195K, Y306K, T307K, S131C, TBD TBD 381 E181Q, D183T, P209S N195K, Y306K, T307K, S131C, TBD TBD 382 E181Q, D183T, P209T N195K, Y306K, T307K, S131C, TBD TBD 383 E181Q, D183T, P209V N195K, Y306K, T307K, S131C, TBD TBD 384 E181Q, D183T, P209W -106- 99975521.7Attorney Docket No: 124540-829360 N195K, Y306K, T307K, S131C, TBD TBD 385 E181Q, D183T, P209Y N195K, Y306K, T307K, S131C, TBD TBD 386 E181Q, D183T, I159A N195K, Y306K, T307K, S131C, TBD TBD 387 E181Q, D183T, I159R N195K, Y306K, T307K, S131C, TBD TBD 388 E181Q, D183T, I196A N195K, Y306K, T307K, S131C, TBD TBD 389 E181Q, D183T, I196R N195K, Y306K, T307K, S131C, TBD TBD 390 E181Q, D183T, A355K N195K, Y306K, T307K, S131C, TBD TBD 391 E181Q, D183T, A355R

[0249] FIG. 12 shows a plot demonstrating that increasing the number of substitutionscontinuously improved the performance of the Cas12f. After these experiments, it was determined that variants having at least four substitutions (i.e.., SEQ ID NOs: 329-334), five substitutions (i.e., SEQ ID NOs: 335-337) and six substitutions (SEQ ID NO: 338 performed maximally (with SEQ ID NO: 338 performing at 835% efficiency relative to WT). The highest performer (SEQ ID NO: 338) as “CasNano” and was used in further experiments described below.

[0250] As shown in the Tables above, additional variants with alternative combinations ofsubstitutions were generated and provided herein. These will be tested for improved targeting and / or editing efficacy using the methods described herein. Example 6 – Optimizing sgRNA scaffolds

[0251] In another set of experiments, the bsCas12f gRNA scaffold was modified toimprove its efficiency. Specifically, 77 different scaffolds were prepared (i.e., SEQ ID NOs: 439-515, named S002-S078) and were tested for improved editing efficiency relative to a WT bsCas12f scaffold (SEQ ID NO: 438, S001). In these experiments, each scaffold was cloned into the pMINI vector with the TetO targeting spacer sequence as described above and tested for efficiency in combination with the wildtype pSC1-dBsCas12f1-VPR using the same transfection conditions described above in Example 3. FIG.10 depicts data from the first set -107- 99975521.7Attorney Docket No: 124540-829360 of scaffolds (SEQ ID NOs:439-470 , S002-S033) and shows that scaffolds corresponding to S003 (SEQ ID NO: 440), S021 (SEQ ID NO: 458), S002 (SEQ ID NO: 439), S030 (SEQ ID NO: 467) and S031 (SEQ ID NO: 468) all were shown to have the potential to achieve higher editing efficiency compared to WT (S001, SEQ ID NO: 438). Tables 14A-14C, below provides the editing efficiency relative to WT across all 77 scaffolds. From this data, the following scaffolds were tagged as useful to increase editing efficiency in the CRISPR editing system herein: S003 (SEQ ID NO: 440), S075 (SEQ ID NO: 512), S077 (SEQ ID NO: 514) and S078 (SEQ ID NO: 515). Table 14A: Scaffolds 1-33 Scaffold TurboRFP Efficiency / SEQ ID Length version + cell (%) WT (%) NO:S001 (wt) 29.77 100 174 438S002 47.01 148.59 166 439 S003 73.17 349.92 140 440 S004 7.59 71.19 151 441 S005 7.33 70.51 166 442 S006 7.05 56.73 150 443 S007 7.83 63.98 112 444 S008 7.3 66.62 160 445 S009 8.37 66.86 166 446 S010 6.78 65.84 154 447 S011 9.26 69.28 146 448 S012 6.77 71.66 134 449 S013 5.83 55.49 121 450 S014 41.67 129.22 178 451 S015 8.96 58.74 178 452 S016 8.66 72.69 178 453 S017 8.62 65.85 178 454 S018 7 68.44 178 455 S019 7.92 69.98 178 456 S020 8.02 71.44 178 457 -108- 99975521.7Attorney Docket No: 124540-829360 Scaffold TurboRFP Efficiency / SEQ ID Length version + cell (%) WT (%) NO: S021 37.18 122.76 178 458 S022 34.39 111.7 175 459 S023 43.4 135.27 174 460 S024 8.87 60.34 173 461 S025 10.06 65.12 173 462 S026 7.48 62.53 171 463 S027 7.22 62.61 173 464 S028 8.12 67.83 172 465 S029 8.27 67.51 175 466 S030 24.4 78.43 175 467 S031 31.45 92.35 172 468 S032 8.18 67.8 148 469 S033 8.95 68.68 174 470 Table 14B: Scaffolds 34-67 Scaffold TurboRFP Efficiency / Length SEQ ID NO: version + cell (%) WT (%) S001 21.91 100 174 438 S034 21.18 97.77 175 471 S035 9.11 84.02 176 472 S036 24.11 103.39 175 473 S037 6.51 78.83 176 474 S038 43.29 105.19 175 475 S039 4.74 90.61 134 476 S040 4.92 93.79 133 477 S041 4.92 86.59 141 478 S042 5.15 90.12 141 479 S043 3.99 84 120 480 S044 3.73 100.1 126 481 -109- 99975521.7Attorney Docket No: 124540-829360 Scaffold TurboRFP Efficiency / Length SEQ ID NO: version + cell (%) WT (%) S045 5.2 82.85 108 482 S046 4.33 89.66 100 483 S047 5.94 80.99 130 484 S048 4.67 119.93 120 485 S049 4.87 95.66 104 486 S050 4.91 91.69 116 487 S051 4.57 99.97 120 488 S052 4.25 98.55 135 489 S053 115.63 181.36 140 490 S054 201.43 278.53 140 491 S055 17.69 91.56 140 492 S056* 3.84 82.44 178* 493 and 494* S057* 59.88 139.09 178* 493 and 494* S058 5.8 84.43 98 495 S059 4.38 86.16 92 496 S060 5.69 81.57 99 497 S061 5.5 86.33 104 498 S062 5.74 80.39 85 499 S063 4.63 90.17 80 500 S064 4.66 96.55 81 501 S065 4.15 78.41 81 502 S066 9.47 29.96 94 503 S067 11.25 25.82 106 504 *S056 and S057 were a split scaffold system where the gRNA spacer was positioned either between (S056) or after (S057) two sequences (SEQ ID NO: 493 and 494). SEQ ID NO: 493 is a true scaffold sequence paired in both instances with a prequeosine1-1 riboswitch aptamer (evopreQ1, SEQ ID NO: 494) as described in Nelson et al., Nat Biotechnol.2022 Mar;40(3):402-410. doi: 10.1038 / s41587-021-01039-7, which is incorporated herein by reference in its entirety). The evopreQ1 (SEQ ID NO: 494) was either put at the end of gRNA (S056, after scaffold (SEQ ID NO: 493) and spacer) or at the beginning of gRNA (S057, before the scaffold (SEQ ID NO: 493). Therefore, in the “S056” scaffold, the order of the tested scaffold+spacer was: SEQ ID NO: 493, spacer, SEQ ID NO: 494. In S056 the order of the tested scaffold+spacer was SEQ ID NO: 494, SEQ ID NO: 493, spacer. -110- 99975521.7Attorney Docket No: 124540-829360 Table 14C: Scaffolds 68-78 Scaffold TurboRFP Efficiency / WT Length SEQ ID NO: version + cell (%) (%) S001 16.09 100 174 438 S068 66.33 264.01 112 505 S069 3.58 87.19 111 506 S070 3.84 86.95 109 507 S071 4.2 101.34 107 508 S072 3.16 93.01 105 509 S073 2.93 113.25 120 510 S074 7.13 43.14 141 511 S075 56.81 221.37 141 512 S076 53.38 183.25 141 513 S077 72.6 360.81 141 514 S078 67.87 318.82 141 515 Example 7 - Summary of CasNano Optimization

[0252] Gene therapy and gene editing hold great potential for treating genetic disorders,but delivery challenges remain. Adeno-associated viruses (AAVs) are promising delivery vehicles; however, their limited packaging capacity of ~4.5 kb hinders the delivery of many CRISPR-based gene editing tools, particularly those involving large epigenetic effectors like TET enzymes, DNMTs or VPR. Epigenetic editing enables precise modulation of gene expression without altering the DNA sequence, but the size of effector proteins combined with conventional CRISPR / Cas systems exceeds AAV packaging limits. To address this, the development of ultracompact CRISPR / Cas systems accommodating epigenetic effectors within AAV capacity is crucial. These compact systems would enable efficient delivery of epigenetic editing tools and expand possibilities for multiplex editing and tissue-specific targeting. By overcoming delivery bottlenecks, we identified and engineered an ultracompact CRISPR / Cas system originated from a commensal bacteria species which we named CasNano (FIG.13). CasNano has the potential to revolutionize gene therapy and gene editing, enabling novel therapeutic strategies for a wide range of genetic disorders. -111- 99975521.7Attorney Docket No: 124540-829360

[0253] As shown in FIG 13, CasNano has an ultracompact size that enables packaging ofdiverse cargoes into a single AAV vector. FIG.14 also shows the PAM sequence of CasNano compared to the PAM sequence of SpCas9. This figure shows that the PAM sequence of CasNano advantageously matches that of Cas9, which means that any potential target of Cas9 can also be targeted by CasNano.

[0254] FIG. 15A-15C summarizes the engineering of CasNano that was detailed in earlierExamples. Specifically, CasNano optimization was achieved through direct evolution of CasNano (e.g., in FIG.15A and FIG.15B, see Example 5 above) and its associated gRNA scaffold (e.g., see FIG.15C and Examples 2 and 6 above). For Cas engineering, critical amino acid residues, identified by alignment-based analysis and ESM-guided protein modeling, were subjected to saturated mutagenesis. Combinations of optimized residues synergistically enhanced targeting efficiency. This optimization resulted in a new CasNano that had improved editing efficiency compared to wildltype with up to 835.35% compared by MFI or 337.07% compared by RFP+ cell ratio (FIG.15B). Additionally, the tracrRNA (scaffold) component of the associated gRNAs was re-engineered to improve stability and interaction with the CasNano protein, further contributing to the system's enhanced targeting efficiency (FIG.15C).

[0255] FIG. 15B shows specifically how increasing the number of substitutions in theunderlying Cas protein improved efficiency. Highly optimized sequences having quadruple, quintuple or sextuple substitutions were selected for further experimentation as described below and referred to as ‘CasNano’ – i.e., Cas Nano V4.1 (SEQ ID NO: 329) , CasNano V4.2 (SEQ ID NO:330 CasNano V4.3 (SEQ ID NO: 331), CasNano V4.4 (SEQ ID NO: 332), CasNano V4.5 (SEQ ID NO: 333), CasNano V4.6 (SEQ ID NO:334), CasNano V5.1 (SEQ ID NO: 335), CasNano V5.2 (SEQ ID NO: 336), CasNano V5.3 (SEQ ID NO:337) and CasNano V6 (SEQ ID NO:338). In select experiments, the V6 version of CasNano (SEQ ID NO: 338) is simply referred to as “CasNano”.

[0256] FIG. 15C summarizes data detailed in Example 6, above, where the tracrRNAcomponent was re-engineered to improve stability and interaction with the CasNano protein (SEQ ID NO: 338), further contributing to the system's enhanced targeting efficiency. Certain scaffolds showed improved stability, specifically, S075 (SEQ ID NO: 512), S003 (SEQ ID NO: 440) and S077 (SEQ ID NO: 514). FIG.15D shows results from an experiment testing the off-target sensitivity of CasNano. CasNano was used with different spacer sequences with a single mismatch relative to the target sequence (indicated in red in the figure) and the -112- 99975521.7Attorney Docket No: 124540-829360 resulting activity was compared to CasNano activity with a spacer sequence that completely matched the target sequence (bottom). Table 15 shows the mismatched spacer sequences, alone, or in combination with the spacer sequence. It was surprisingly shown that optimized CasNano V6 (SEQ ID NO: 338) has superior target specificity and limited mismatch tolerance particularly when the mismatch occurs near the PAM sequence (in blue on the left). This suggests that CasNano has the potential for very minimal off-target editing, in contrast with Cas9. Table 15 - gRNAs tested in FIG.15D. gRNA gRNA Spacer SEQ ID Spacer + Scaffold SEQ ID Name NO: NO: g004 ctccctatcagtgatagaga 516 GACGGTAACTTAAAGTGCCGAA 552 GGCTGAGGAGATGGATTAAATA TATAAGGTTTTGACCAACTATAT ACCATCATCACTCGGTAAGGGT TAATCCTAACAATGTGTGACCGT TAGGCGTTCCAAAGAAATGGAA TGTTAATctccctatcagtgatagaga g181 ctccctatcagtgatagagC 612 GACGGTAACTTAAAGTGCCGAA 632 GGCTGAGGAGATGGATTAAATA TATAAGGTTTTGACCAACTATAT ACCATCATCACTCGGTAAGGGT TAATCCTAACAATGTGTGACCGT TAGGCGTTCCAAAGAAATGGAA TGTTAATctccctatcagtgatagagC g182 ctccctatcagtgatagaTa 613 GACGGTAACTTAAAGTGCCGAA 633 GGCTGAGGAGATGGATTAAATA TATAAGGTTTTGACCAACTATAT ACCATCATCACTCGGTAAGGGT TAATCCTAACAATGTGTGACCGT TAGGCGTTCCAAAGAAATGGAA TGTTAATctccctatcagtgatagaTa g183 ctccctatcagtgatagCga 614 GACGGTAACTTAAAGTGCCGAA 634 GGCTGAGGAGATGGATTAAATA TATAAGGTTTTGACCAACTATAT ACCATCATCACTCGGTAAGGGT TAATCCTAACAATGTGTGACCGT TAGGCGTTCCAAAGAAATGGAA TGTTAATctccctatcagtgatagCga g184 ctccctatcagtgataTaga 615 GACGGTAACTTAAAGTGCCGAA 635 GGCTGAGGAGATGGATTAAATA TATAAGGTTTTGACCAACTATAT ACCATCATCACTCGGTAAGGGT TAATCCTAACAATGTGTGACCGT -113- 99975521.7Attorney Docket No: 124540-829360 TAGGCGTTCCAAAGAAATGGAA TGTTAATctccctatcagtgataTaga g185 ctccctatcagtgatCgaga 616 GACGGTAACTTAAAGTGCCGAA 636 GGCTGAGGAGATGGATTAAATA TATAAGGTTTTGACCAACTATAT ACCATCATCACTCGGTAAGGGT TAATCCTAACAATGTGTGACCGT TAGGCGTTCCAAAGAAATGGAA TGTTAATctccctatcagtgatCgaga g186 ctccctatcagtgaGagaga 617 GACGGTAACTTAAAGTGCCGAA 637 GGCTGAGGAGATGGATTAAATA TATAAGGTTTTGACCAACTATAT ACCATCATCACTCGGTAAGGGT TAATCCTAACAATGTGTGACCGT TAGGCGTTCCAAAGAAATGGAA TGTTAATctccctatcagtgaGagaga g187 ctccctatcagtgCtagaga 618 GACGGTAACTTAAAGTGCCGAA 638 GGCTGAGGAGATGGATTAAATA TATAAGGTTTTGACCAACTATAT ACCATCATCACTCGGTAAGGGT TAATCCTAACAATGTGTGACCGT TAGGCGTTCCAAAGAAATGGAA TGTTAATctccctatcagtgCtagaga g188 ctccctatcagtTatagaga 619 GACGGTAACTTAAAGTGCCGAA 639 GGCTGAGGAGATGGATTAAATA TATAAGGTTTTGACCAACTATAT ACCATCATCACTCGGTAAGGGT TAATCCTAACAATGTGTGACCGT TAGGCGTTCCAAAGAAATGGAA TGTTAATctccctatcagtTatagaga g189 ctccctatcagGgatagaga 620 GACGGTAACTTAAAGTGCCGAA 640 GGCTGAGGAGATGGATTAAATA TATAAGGTTTTGACCAACTATAT ACCATCATCACTCGGTAAGGGT TAATCCTAACAATGTGTGACCGT TAGGCGTTCCAAAGAAATGGAA TGTTAATctccctatcagGgatagaga g190 ctccctatcaTtgatagaga 621 GACGGTAACTTAAAGTGCCGAA 641 GGCTGAGGAGATGGATTAAATA TATAAGGTTTTGACCAACTATAT ACCATCATCACTCGGTAAGGGT TAATCCTAACAATGTGTGACCGT TAGGCGTTCCAAAGAAATGGAA TGTTAATctccctatcaTtgatagaga g191 ctccctatcCgtgatagaga 622 GACGGTAACTTAAAGTGCCGAA 642 GGCTGAGGAGATGGATTAAATA TATAAGGTTTTGACCAACTATAT -114- 99975521.7Attorney Docket No: 124540-829360 ACCATCATCACTCGGTAAGGGT TAATCCTAACAATGTGTGACCGT TAGGCGTTCCAAAGAAATGGAA TGTTAATctccctatcCgtgatagaga g192 ctccctatAagtgatagaga 623 GACGGTAACTTAAAGTGCCGAA 643 GGCTGAGGAGATGGATTAAATA TATAAGGTTTTGACCAACTATAT ACCATCATCACTCGGTAAGGGT TAATCCTAACAATGTGTGACCGT TAGGCGTTCCAAAGAAATGGAA TGTTAATctccctatAagtgatagaga g193 ctccctaGcagtgatagaga 624 GACGGTAACTTAAAGTGCCGAA 644 GGCTGAGGAGATGGATTAAATA TATAAGGTTTTGACCAACTATAT ACCATCATCACTCGGTAAGGGT TAATCCTAACAATGTGTGACCGT TAGGCGTTCCAAAGAAATGGAA TGTTAATctccctaGcagtgatagaga g194 ctccctCtcagtgatagaga 625 GACGGTAACTTAAAGTGCCGAA 645 GGCTGAGGAGATGGATTAAATA TATAAGGTTTTGACCAACTATAT ACCATCATCACTCGGTAAGGGT TAATCCTAACAATGTGTGACCGT TAGGCGTTCCAAAGAAATGGAA TGTTAATctccctCtcagtgatagaga g195 ctcccGatcagtgatagaga 626 GACGGTAACTTAAAGTGCCGAA 646 GGCTGAGGAGATGGATTAAATA TATAAGGTTTTGACCAACTATAT ACCATCATCACTCGGTAAGGGT TAATCCTAACAATGTGTGACCGT TAGGCGTTCCAAAGAAATGGAA TGTTAATctcccGatcagtgatagaga g196 ctccAtatcagtgatagaga 627 GACGGTAACTTAAAGTGCCGAA 647 GGCTGAGGAGATGGATTAAATA TATAAGGTTTTGACCAACTATAT ACCATCATCACTCGGTAAGGGT TAATCCTAACAATGTGTGACCGT TAGGCGTTCCAAAGAAATGGAA TGTTAATctccAtatcagtgatagaga g197 ctcActatcagtgatagaga 628 GACGGTAACTTAAAGTGCCGAA 648 GGCTGAGGAGATGGATTAAATA TATAAGGTTTTGACCAACTATAT ACCATCATCACTCGGTAAGGGT TAATCCTAACAATGTGTGACCGT TAGGCGTTCCAAAGAAATGGAA TGTTAATctcActatcagtgatagaga -115- 99975521.7Attorney Docket No: 124540-829360 g198 ctAcctatcagtgatagaga 629 GACGGTAACTTAAAGTGCCGAA 649 GGCTGAGGAGATGGATTAAATA TATAAGGTTTTGACCAACTATAT ACCATCATCACTCGGTAAGGGT TAATCCTAACAATGTGTGACCGT TAGGCGTTCCAAAGAAATGGAA TGTTAATctAcctatcagtgatagaga g199 cGccctatcagtgatagaga 630 GACGGTAACTTAAAGTGCCGAA 650 GGCTGAGGAGATGGATTAAATA TATAAGGTTTTGACCAACTATAT ACCATCATCACTCGGTAAGGGT TAATCCTAACAATGTGTGACCGT TAGGCGTTCCAAAGAAATGGAA TGTTAATcGccctatcagtgatagaga g200 Atccctatcagtgatagaga 631 GACGGTAACTTAAAGTGCCGAA 651 GGCTGAGGAGATGGATTAAATA TATAAGGTTTTGACCAACTATAT ACCATCATCACTCGGTAAGGGT TAATCCTAACAATGTGTGACCGT TAGGCGTTCCAAAGAAATGGAA TGTTAATAtccctatcagtgatagaga Example 8 - Use of CasNano for Gene Activation – VPR mediated

[0257] In further experiments, CRISPRa editors with different versions of the Cas12fpolypeptides herein (e.g., CasNano V4 (SEQ ID NO: 332, CasNano V5 (SEQ ID NO: 336), or CasNanoV6 (SEQ ID NO: 338)) were compared with SpCas9 for transcriptional activation activity using an HEK293 CRISPRa reporter as described in Xu et al., 2021 (Mol Cell.2021 Oct 21; 81(20);4333-4345.e4, which is incorporated herein by reference in its entirety). Two different gRNA spacers were tested (g004 (SEQ ID NO: 516) and g030 (SEQ ID NO: 517)) and paired with scaffolds designed for Cas12f (e.g., SEQ ID NO: 493) or for spCas9 (e.g., SEQ ID NO: 654) to form four total gRNAs tested: SEQ ID NOs: 552 and 553 for CasNano and SEQ ID NOs: 655 and 656 for SpCas9. Results are shown in FIG.16A.

[0258] In a second experiment, specific endogenous genes (HBG1 and HBB) weretargeted for activation in these cells using a CasNano-VPR construct. Different sets of gRNAs (Set A and Set B) for each gene target were tested. Set A gRNA spacer sequence for HBB corresponded to SEQ ID NOs: 518-523, with the full-length gRNA (spacer + scaffold) corresponding to SEQ ID NOs: 554-559, respectively. Set B gRNA spacer sequence for HBB corresponded to SEQ ID NOs: 524-529, with the full-length gRNA (spacer + scaffold) corresponding to SEQ ID NOs: 560-565, respectively. Set A gRNA spacer sequence for -116- 99975521.7Attorney Docket No: 124540-829360 HBG1 corresponded to SEQ ID NOs: 530-533, with the full-length gRNA (spacer + scaffold) corresponding to SEQ ID NOs: 566-569, respectively. Set B gRNA spacer sequence for HBG1 corresponded to SEQ ID NOs: 534-537, with the full-length gRNA (spacer + scaffold) corresponding to SEQ ID NOs: 570-573, respectively. The full sequences for SEQ ID NOs: 518-537 and 554-573 are provided in Tables 7B-7C and 8B-8C above.

[0259] FIG. 16B provides data plots showing endogenous gene activation triggered bydCasNano-VPR in HEK293 cells, with superior results shown for the Set A gRNAs for both gene targets. Example 9 - Use of CasNano for DNA Demethylation

[0260] In another experiment, different epigenetic editors were tested to compare theirdemethylation efficiency using different Tet proteins. As shown in FIG.17A, combining Tet proteins with Cas editors can be difficult because their large sizes prohibit including both in a single vector (like AAV). In this example, different truncated versions of Tet proteins were generated that retained their catalytic activity but were much smaller than full length Tet. These two truncated Tet proteins were called TET1-CD (derived from the catalytic domain of Tet1) and TETmini (derived from the catalytic domain of Tet2). The full sequences of Tet1- CD (SEQ ID NO: 409) and TetMini (SEQ ID NO: 408) are provided in Table 4 above.

[0261] DNA fragments encoding CasNano V6 (protein: SEQ ID NO: 338) or dSpCas9(protein: SEQ ID NO; 653, nucleic acid: SEQ ID NO: 654) and either Tet1-CD (SEQ ID NO: 409) or TetMini (SEQ ID NO:408) were cloned into a pCS1- vector (SEQ ID NO: 588) to create a fusion protein with an open reading frame containing a nucleic acid encoding NLS- dCas9-XTEN -TETmini-NLS (SEQ ID NO: 426), NLS-dCasNano-XTEN-Tet1CD-NLS (SEQ ID NO: 422), NLS-dCasNano-XTEN -TETmini-NLS (SEQ ID NO:420) or NLS- TETmini-XTEN-CasNano (SEQ ID NO: 424) under the control of a UBC promoter as shown in plasmid schematics in FIGs.19A -19D, respectively.

[0262] After individually testing these different combinations, a fusion protein comprisingdCasNano-TETmini (i.e., third from the top in FIG.17A having an amino acid sequence of SEQ ID NO: 420) was found to have strongest demethylation ability. This was further tested by measuring demethylation of MeCP2 promoter region in HEK293 cells after incubation with dCasNano-TETmini. After 7 days of incubation, the average methylated CpG dropped from 24% to 17% (FIG.17B). Example 10 - Use of CasNano for DNA Methylation -117- 99975521.7Attorney Docket No: 124540-829360

[0263] In another experiment, different epigenetic editors were tested to compare theirgene silencing efficiency. Specifically, methylation efficiency was tested for CasNano V6 (SEQ ID NO: 338) in constructs with different human or mice DNMT3A polypeptides. Specifically, hDNMT3A-L (SEQ ID NO: 410), hDNMT3A (SEQ ID NO: 411), hDNMT3L (SEQ ID NO: 412), H3-hDNMT3L (SEQ ID NO: 413), and mDNMT3A-L (SEQ ID NO:414) were all tested on a C9Orf72 G4C2 repeat reporter cell line.

[0264] FIG. 20A-20E depict schematics of plasmids used in these experiments.Specifically, DNA fragments encoding CasNano V6 (SEQ ID NO: 338) and one of the above referenced DNMT3A polypeptides were cloned into a pCS1 vector to create a fusion protein with an open reading frame containing a nucleic acid encoding NLS-dCasNano-XTEN - hDNMT3A-3L (SEQ ID NO: 428), NLS--dCasNano -XTEN-hDNMT3A (SEQ ID NO: 430), NLS- dCasNano-XTEN – hDNMT3L (SEQ ID NO:432), NLS- dCasNano-XTEN – hH3- hNMT3L (SEQ ID NO: 434 ) or NLS- dCasNano-XTEN – mDNMT3A-3L (SEQ ID NO: 436) under the control of a UBC promoter as shown in plasmid schematics in FIGs.20A- 20E, respectively.

[0265] FIG. 18A illustrates how the reporter line contains a TurboRFP reporter geneoperably linked to a PGK promoter sandwiched between (GGGGCC)n repeat (G4C2 repeat) sections. To suppress expression of the reporter (RFP), either the PGK promoter or the G4G2 repeats may be targeted for methylation by the CasNano-DNMT3A protein. To target the PGK promoter, gRNA spacer sequences corresponding to SEQ ID NOs: 545-547 were used, with the full gRNAs (scaffold + spacer) corresponding to SEQ ID NOs: 581-583. To target the G4C2 repeats, gRNA spacer sequences corresponding to SEQ ID NOs: 548-551 were used, with the full gRNAs (scaffold + spacer) corresponding to SEQ ID NOs: 584-587. For ease of reference, SEQ ID NOs 545-551 and 581-587 are provided earlier in Tables 7E-7F and 8E-8F. FIGs.18B-18C shows relative TurboRFP expression in each experiment following targeting of the PGK promoter region (FIG.18B) or the G4C2 repeats (FIG.18C). Example 11 – Optimizing Spacer length for Cas12f

[0266] In another example, whether the length of a spacer paired with a Cas12fpolypeptide could impact activity was tested. Specifically, 9 different gRNAs comprising the scaffold sequence of S003 (SEQ ID NO: 440) and a spacer sequence having 16-24 nucleotides in length were tested in combination with CasNano V6 (SEQ ID NO: 338) using the Tet0 reporter system described previously. Each of the gRNA spacers targeted Tet0 (like -118- 99975521.7Attorney Docket No: 124540-829360 SEQ ID NOs 516 and 517 shown previously). Table 16 below provides the spacer sequences used. Table 16 Spacer Length Sequence SEQ ID NO:24-mer ctccctatcagtgatagagaacgt 65723-mer ctccctatcagtgatagagaacg 65822-mer ctccctatcagtgatagagaac 65921-mer ctccctatcagtgatagagaa 66020-mer ctccctatcagtgatagaga 66119-mer ctccctatcagtgatagag 66218-mer ctccctatcagtgataga 66317-mer ctccctatcagtgatag 66416-mer ctccctatcagtgata 665

[0267] Results were analyzed using the TurboRFP cell line and reported as a percentageof positive cells (FIG.22A) or using Geo MFI (FIG.22B). It was found surprisingly that activity improved as the spacer sequence was shortened, particularly when analyzed using Geo MFI. Example 12 – Additional Mutagenesis of dCasNano

[0268] Mutagenesis was conducted on several amino acids of dCasNano including E109,S131, E181, D183, K184, N195, K220, Y306, K318, on D353 to assess the impact of these residues on dCasNano activity and function. Targeting activity was assessed by linking dCasNano to VPR (CRISPRa) to activate RFP expression. Higher efficiency (fluorescence intensity) or higher RFP+ cell population were the desired effect. Data was represented as percentage compared to WT levels. Thus, a value of 100% corresponds to no improvement following mutagenesis. Data was collected following 3 days of incubation on TCCA-NGG reporter cells. As shown in FIG.23A and FIG.23B, several amino acid substitutions resulted in striking improvements in dCasNano activity. Additionally, select CasNano variants including N195K, Y306K, T307K, S131C, E181Q, D183T, I10R, D76K, G111N, I159R, and I196V displayed significantly increased activity in TCCA-NGG reporter cells (FIGs.24A, 24B).

[0269] To further maximize CasNano activity, select CasNano variants (D76K, G111N,I159R, and I196V) were paired with different gRNA scaffold sequences (S001, S003, S075, S077, and S078) and targeting activity was assessed by linking dCasNano to VPR (CRISPRa) to activate RFP expression. Results were obtained following 3 days of incubation -119- 99975521.7Attorney Docket No: 124540-829360 on TCCA-NGG reporter cells. Higher readings of RFP fluorescence correspond to improved CasNano efficiency. As shown in FIG.24C, the CasNano I196V variant when paired with the gRNA scaffold sequence S077 resulted in the highest level of targeting efficiency.

[0270] CasNano targeting activity was next compared to the targeting activity of dSpCas9.The TCCA-NGG reporter was designed to be targetable by both dCasNano and dSpCas9 simultaneously using exact the same sequence as the gRNA spacer sequence to activate RFP expression. A scrambled spacer sequence was utilized as a negative control. FIG.25A shows a representative schematic depicting the targetable sequences in the sense and antisense strands of CasNano and SpCas9. Both dCasNano-VPR and dSpCas9-VPR displayed similar targeting efficiency when using the sense strand as the spacer sequence, however dCasNano- VPR displayed a striking improvement in targeting efficiency when using the antisense strand as the spacer sequence when compared to dSpCas9-VPR (FIG.25B). Example 14 – CasNano PAM preference and catalytic site analysis

[0271] We built different reporter constructs, which contain the same spacer anddownstream RFP expression cassettes. However, the four nucleotides preceding the spacer (in the 5’ sequence) were modified from the base version (TCCA) to test the effect of the PAM sequence on CasNano activity. In total, 13 PAM sequence variations were tested: TCCA, ACCA, GCCA, CCCA, TACA, TTCA, TGCA, TCAA, TCTA, TCGA, TCCT, TCCG, TCCC. The position -1 to -4 indicates the position next to the spacer. As shown in FIG.26, dCasNano displayed a significant preference for a cytosine residing at positions -2 and -3.

[0272] CasNano variants including D216A, D390A, and D216A + D390A were expressedin an RFP expressing reporter cells to assess double strand break formation. A decrease of RFP fluorescence in the RFP expressing reporter cells indicated the formation of a double strand break and resulting frame shift at the target site. When D216, D390 or both were mutated (substituted with alanine), CasNano became a deactivated Cas (i.e. losing DNA cutting activity but preserved DNA-binding activity) (FIG.27). Example 15 – Epigenetic editing with CasNano-TETmini and CasNano-DMNT

[0273] The activity of CasNano, when linked to a catalytically active truncated Tetpolypeptide (TETmini) was next assessed. TETmini was linked to CasNano in two constructs. A first construct comprised dCasNano, a nuclear localization sequence, TETmini, and the regulatory element sequence CW3SL. A second construct comprised dCasNano, a -120- 99975521.7Attorney Docket No: 124540-829360 nuclear localization sequence, TETmini, and the regulatory element sequence sPA. The dCasNano-TETmini epigenetic editors were delivered to HEK293 cells via an AAV vector. As shown in FIGs.28A and 28B, both dCasNano-TETmini constructs decreased methylation levels at the MeCP2 locus in HEK293 cells when compared to control cells.

[0274] CasNano was also linked to the DNA methyl transferases DNMT3A andDNMT3L in a variety of constructs (FIGs.29A and 29B). These constructs were delivered to HEK293 cells via AAV vector and methylation levels of the MeCP2 locus were quantifiedusing bisulfite sequencing. As shown in FIGs. 29A and 29B, the dCasNano-DMNTconstructs increased methylation level at the MeCP2 locus in HEK293 cells.

[0275] Thus, CasNano can be functionalized with epigenetic regulators to controlmethylation levels at targeted genetic loci. Example 16 – Additional scaffold sequences

[0276] FIG. 30 depicts an improved annotated CasNano scaffold structure diagram.Additional scaffold sequences are provided in Table 17. Table 17: Scaffold Sequences Scaffold # Scaffold sequence Length SEQ ID NO: TTGTATTGATGTTATATATAAATATATAGCAGTTACGGTAACTTAAAGT GCCGAAGGCTGAGGAGATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAACAATGTGTGACCGTTA BS-S001GGCGTTCCAAAGAAATGGAATGTTAAT 174 673ATGTTATATATAAATATATAGCAGTTACGGTAACTTAAAGTGCCGAAGG CTGAGGAGATGGATTAAATATATAAGGTTTTGACCAACTATATACCATC ATCACTCGGTAAGGGTTAATCCTAACAATGTGTGACCGTTAGGCGTTCC BS-S002AAAGAAATGGAATGTTAAT 166 674ACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAG GTTTTGACCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAAC BS-S003AATGTGTGACCGTTAGGCGTTCCAAAGAAATGGAATGTTAAT 140 675TTGTATTGATGTTATATATAAATATATAGCAGGGCTGAGGAGATGGATT AAATATATAAGGTTTTGACCAACTATATACCATCATCACTCGGTAAGGG TTAATCCTAACAATGTGTGACCGTTAGGCGTTCCAAAGAAATGGAATGT BS-S004TAAT 151 676TTGTATTGATGTTATATATAAATATATAGCAGTTACGGTAACTTAAAGT GCCGAAGGCTGAGGAGATGGATTAAATATATAAGCAACTATATACCATC ATCACTCGGTAAGGGTTAATCCTAACAATGTGTGACCGTTAGGCGTTCC BS-S005AAAGAAATGGAATGTTAAT 166 677TTGTATTGATGTTATATATAAATATATAGCAGTTACGGTAACTTAAAGT GCCGAAGGCTGAGGAGATGGATTAAATACCATCATCACTCGGTAAGGGT TAATCCTAACAATGTGTGACCGTTAGGCGTTCCAAAGAAATGGAATGTT BS-S006AAT 150 678TTGTATTGATGTTATATATAAATATATAGCAGTTACGGTAACTTAAAGT GCCGAAAGGGTTAATCCTAACAATGTGTGACCGTTAGGCGTTCCAAAGA BS-S007AATGGAATGTTAAT 112 679TTGTATTGATGTTATATATAAATATATAGCAGTTACGGTAACTTAAAGT BS-S008GCCGAAGGCTGAGGAGATGGATTAAATATATAAGGTTTTGACCAACTAT 160 680-121- 99975521.7Attorney Docket No: 124540-829360 ATACCATCATCACTCGGTACAATGTGTGACCGTTAGGCGTTCCAAAGAA ATGGAATGTTAAT TTGTATTGATGTTATATATAAATATATAGCAGTTACGGTAACTTAAAGT GCCGAAGGCTGAGGAGATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAACAATGTGTGACGTTCC BS-S009AAAGAAATGGAATGTTAAT 166 681TTGTATTGATGTTATATATAAATATATAGCAGTTACGGTAACTTAAAGT GCCGAAGGCTGAGGAGATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAACAATGTGTGACCGTTA BS-S010GGTTAAT 154 682TTGTATTGATGTTATATATAAATATATAGCAGTTACGGTAACTTAAAGT GCCGAAGGCTGAGGAGATGGATTAAATATATAAGGTTTTGACCAACTAT BS-S011ATACCATCATCACTCGGTAAGGGTTAATCCTAACAATGTGTGATTAAT 146 683TTGTATTGATGTTATATATAAATATATAGCAGTTACGGTAACTTAAAGT GCCGAAGGCTGAGGAGATGGATTAAATATATAAGGTTTTGACCAACTAT BS-S012ATACCATCATCACTCGGTAAGGGTTAATCCTTTAAT 134 684TTGTATTGATGTTATATATAAATATATAGCAGTTACGGTAACTTAAAGT GCCGAAGGCTGAGGAGATGGATTAAATATATAAGGTTTTGACCAACTAT BS-S013ATACCATCATCACTCGGTTTAAT 121 685TTGTATTGATGTTATGGATATAAATATCCATAGCAGTTACGGTAACTTA AAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAGGTTTTGACCAA CTATATACCATCATCACTCGGTAAGGGTTAATCCTAACAATGTGTGACC BS-S014GTTAGGCGTTCCAAAGAAATGGAATGTTAAT 178 686TTGTATTGATGTTATATATAAATATATAGCAGTTACGGTAACTTAAAGT GCCGAAGGCTGAGGAGATGGATTAAATATATAAGGGGTTTTGACCCCAA CTATATACCATCATCACTCGGTAAGGGTTAATCCTAACAATGTGTGACC BS-S015GTTAGGCGTTCCAAAGAAATGGAATGTTAAT 178 687TTGTATTGATGTTATATATAAATATATAGCAGTTACGGTAACTTAAAGT GCCGAAGGCTGAGGAGATGGATTAAATAGGTATAAGGTTTTGACCAACT CCATATACCATCATCACTCGGTAAGGGTTAATCCTAACAATGTGTGACC BS-S016GTTAGGCGTTCCAAAGAAATGGAATGTTAAT 178 688TTGTATTGATGTTATATATAAATATATAGCAGTTACGGTAACTTAAAGT GCCGAAGGCTGAGGAGATGGGGATTAAATATATAAGGTTTTGACCAACT ATATACCCCATCATCACTCGGTAAGGGTTAATCCTAACAATGTGTGACC BS-S017GTTAGGCGTTCCAAAGAAATGGAATGTTAAT 178 689TTGTATTGATGTTATATATAAATATATAGCAGTTACGGTAACTTAAAGT GCCGAAGGCTGGGAGGAGATGGATTAAATATATAAGGTTTTGACCAACT ATATACCATCATCACCCTCGGTAAGGGTTAATCCTAACAATGTGTGACC BS-S018GTTAGGCGTTCCAAAGAAATGGAATGTTAAT 178 690TTGTATTGATGTTATATATAAATATATAGCAGTTACGGTAACTTAAAGT GCCGAAGGCTGAGGAGATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGGGTTAATCCCCTAACAATGTGTGACC BS-S019GTTAGGCGTTCCAAAGAAATGGAATGTTAAT 178 691TTGTATTGATGTTATATATAAATATATAGCAGTTACGGTAACTTAAAGT GCCGAAGGCTGAGGAGATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAACAATGTGTGACCCCGT BS-S020TAGGGGCGTTCCAAAGAAATGGAATGTTAAT 178 692TTGTATTGATGTTATATATAAATATATAGCAGTTACGGTAACTTAAAGT GCCGAAGGCTGAGGAGATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAACAATGTGTGACCGTTA BS-S021GGCGTTCCCCAAAGAAATGGGGAATGTTAAT 178 693TTGTATTGATGTTATATATAAATATATAGCAGTTACGGTAACTTAAAGT GCCGAAGGCTGAGGAGATGGATTAAATATATAAGGTTATTGACCAACTA TATACCATCATCACTCGGTAAGGGTTAATCCTAACAATGTGTGACCGTT BS-S022AGGCGTTCCAAAGAAATGGAATGTTAAT 175 694TTGTATTGATGTTATATATAAATATATAGCAGTTACGGTAACTTAAAGT GCCGAAGGCTGAGGAGATGGATTAAATATATAAGGGTTTGCCCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAACAATGTGTGACCGTTA BS-S023GGCGTTCCAAAGAAATGGAATGTTAAT 174 695-122- 99975521.7Attorney Docket No: 124540-829360 TTGTATTGATGTTATATATAAATATATAGCAGTTACGGTAACTTAAAGT GCCGAAGGCTGAGGAGATGGATTAAATATATAGGTTTTGACCAACTATA TACCATCATCACTCGGTAAGGGTTAATCCTAACAATGTGTGACCGTTAG BS-S024GCGTTCCAAAGAAATGGAATGTTAAT 173 696TTGTATTGATGTTATATATAAATATATAGCAGTTACGGTAACTTAAAGT GCCGAAGGCTGAGGAGATGGATTAAATATATAAGGTTTTGACCACTATA TACCATCATCACTCGGTAAGGGTTAATCCTAACAATGTGTGACCGTTAG BS-S025GCGTTCCAAAGAAATGGAATGTTAAT 173 697TTGTATTGATGTTATATATAAATATATAGCAGTTACGGTAACTTAAAGT GCCGAAGGCTGAGGAGATGGAAATATATAAGGTTTTGACCAACTATATA CCATCATCACTCGGTAAGGGTTAATCCTAACAATGTGTGACCGTTAGGC BS-S026GTTCCAAAGAAATGGAATGTTAAT 171 698TTGTATTGATGTTATATATAAATATATAGCAGTTACGGTACTTAAAGTG CCGAAGGCTGAGGAGATGGATTAAATATATAAGGTTTTGACCAACTATA TACCATCATCACTCGGTAAGGGTTAATCCTAACAATGTGTGACCGTTAG BS-S027GCGTTCCAAAGAAATGGAATGTTAAT 173 699TTGTATTGATGTTATATATAAATATATAGCAGTTACGGTAACTTAAAGT GCCGAAGGCTGAGGAGATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTATCCTAACAATGTGTGACCGTTAGG BS-S028CGTTCCAAAGAAATGGAATGTTAAT 172 700TTGTATTGATGTTATATATAAATATATAGCAGTTACGGTAACTTAAAGT GCCGAAGGCTGAGGAGATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAACAATGTGTGACCGTTT BS-S029AGGCGTTCCAAAGAAATGGAATGTTAAT 175 701TTGTATTGATGTTATATATAAATATATAGCAGTTACGGTAACTTAAAGT GCCGAAGGCTGAGGAGATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAACAATGTGTGACCGTTA BS-S030GGCGTTCCAAAGAAAATGGAATGTTAAT 175 702TTGTATTGATGTTATATATAAATATATAGCAGTTACGGTAACTTAAAGT GCCGAAGGCTGAGGAGATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAACAATGTGTGACCGTTA BS-S031GGCGTTCCAAAAATGGAATGTTAAT 172 703TTGTATTGATGTTATATATAAATATATAGCAGTTACGGTAACTTAAAGT GCCGAAGGCTGAGGAGATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAACAATGTGTGACCGTAA BS-S032T 148 704TTGTATTGATGTTATATATAAATATATAGCAGTTACGGTAACTTAAAGT GCCGAAGGCTGAGGAGATGGATTAAATATATAAGGTTTTGACCAACTAT ATACCATCATCACTCGGTAAGGGTTAATCCTAACAATGTGTGACCGTTA BS-S033GGCGTTCCAAAGAAATGGAATGTGAAC 174 705TTGTATTGATGTTATATATAAATATATAGCAGTTACGGTAACTTAAAGT GCCGAAGGCTGAGGAGATGGATTAAATATATAAGGGTATTGCCCAACTA TATACCATCATCACTCGGTAAGGGTTAATCCTAACAATGTGTGACCGTT BS-S034AGGCGTTCCAAAGAAATGGAATGTTAAT 175 706TTGTATTGATGTTATATATAAATATATAGCAGTTACGGTAACTTAAAGT GCCGAAGGCTGAGGAGATGGATTAAATATATAAGGGTAATTGCCCAACT ATATACCATCATCACTCGGTAAGGGTTAATCCTAACAATGTGTGACCGT BS-S035TAGGCGTTCCAAAGAAATGGAATGTTAAT 176 707TTGTATTGATGTTATATATAAATATATAGCAGTTACGGTAACTTAAAGT GCCGAAGGCTGAGGAGATGGATTAAATATATAAGGGTGTTGCCCAACTA TATACCATCATCACTCGGTAAGGGTTAATCCTAACAATGTGTGACCGTT BS-S036AGGCGTTCCAAAGAAATGGAATGTTAAT 175 708TTGTATTGATGTTATATATAAATATATAGCAGTTACGGTAACTTAAAGT GCCGAAGGCTGAGGAGATGGATTAAATATATAAGGGTGGTTGCCCAACT ATATACCATCATCACTCGGTAAGGGTTAATCCTAACAATGTGTGACCGT BS-S037TAGGCGTTCCAAAGAAATGGAATGTTAAT 176 709ACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAG GTTTTGACCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAAC BS-S038AATGTGTGACCGTTAGGCCCAAAAATGGGTTAAT 175 710-123- 99975521.7Attorney Docket No: 124540-829360 ACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAG GTTTTGACCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAAG BS-S039TGACCGTTAGGCGTTCCAAAGAAATGGAATGTTAAT 134 711ACACTTAAAGTGAAGGCTGAGGAGATGGATTAAATATATAAGGTTTTGA CCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAACAATGTGT BS-S040GACCGTTAGGCGTTCCAAAGAAATGGAATGTTAAT 133 712ACGGTAACTTAAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAA GGTTTTGACCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAA BS-S041CAATGTGTGACCGTTAGGCGTTCCAAAGAAATGGAATGTTAAT 141 713ACGGTAACTTGAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAA GGTTTTGACCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAA BS-S042CAATGTGTGACCGTTAGGCGTTCCAAAGAAATGGAATGTTAAT 141 714AGGCTGAGGAGATGGATTAAATATATAAGGTTTTGACCAACTATATACC ATCATCACTCGGTAAGGGTTAATCCTAACAATGTGTGACCGTTAGGCGT BS-S043TCCAAAGAAATGGAATGTTAAT 120 715ACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAAGGTTTTGA CCTACCATCATCACTCGGTAAGGGTTAATCCTAACAATGTGTGACCGTT BS-S044AGGCGTTCCAAAGAAATGGAATGTTAAT 126 716ACGGTAACTTAAAGTGCCGAAGGCTGAGGAGGTTTTGACCATCACTCGG TAAGGGTTAATCCTAACAATGTGTGACCGTTAGGCGTTCCAAAGAAATG BS-S045GAATGTTAAT 108 717ACGGTAACTTAAAGTGCCGAAGGCTGAGGTTTTGACCTCGGTAAGGGTT AATCCTAACAATGTGTGACCGTTAGGCGTTCCAAAGAAATGGAATGTTA BS-S046AT 100 718ACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATGTTTT GACATATACCATCATCACTCGGTAAGGGTTAATCCTAACAATGTGTGAC BS-S047CGTTAGGCGTTCCAAAGAAATGGAATGTTAAT 130 719ACGGTAACTTAAAGTGCCGAAGGCTGAGGAATATAAGGTTTTGACCAAC TATATCACTCGGTAAGGGTTAATCCTAACAATGTGTGACCGTTAGGCGT BS-S048TCCAAAGAAATGGAATGTTAAT 120 720ACGGTAACTTAAAGTGCCGAAGGATATAAGGTTTTGACCAACTATTAAG GGTTAATCCTAACAATGTGTGACCGTTAGGCGTTCCAAAGAAATGGAAT BS-S049GTTAAT 104 721ACGGTAACTTAAAGTGCCGAAGGGATTAAATATATAAGGTTTTGACCAA CTATATACTAAGGGTTAATCCTAACAATGTGTGACCGTTAGGCGTTCCA BS-S050AAGAAATGGAATGTTAAT 116 722ACGGTAACTTAAAGTGCCGAAGGTGGATTAAATATATAAGGTTTTGACC AACTATATACCATAAGGGTTAATCCTAACAATGTGTGACCGTTAGGCGT BS-S051TCCAAAGAAATGGAATGTTAAT 120 723ACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAG GTTTTGACCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAAC BS-S052AATGTGTGACCGTTAGGCGTTCCAAAGAAATGGAATG 135 724ACGGGAACTTAAAGTCCCGAAGGCTGAGGAGATGGATTAAATATATAAG GTTTTGACCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAAC BS-S053AATGTGTGACCGTTAGGCGTTCCAAAGAAATGGAATGTTAAT 140 725ACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAG GTTTTGACCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAAC BS-S054AATGTGTGACCGTTAGGCGTTCCAAAGAAATGGAACGTTAAT 140 726ACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATGTATAAG GTTTTGACCAACTACATACCATCATCACTCGGTAAGGGTTAATCCTAAC BS-S055AATGTGTGACCGTTAGGCGTTCCAAAGAAATGGAATGTTAAT 140 727ACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAG GTTTTGACCAACTATATACCATCATCACTCGGTACCGGCGTTCCAAAGA BS-S068AATGGAATGTTAAT 112 728ACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAG GTTTTGACCAACTATATACCATCATCACTCGGTACGGCGTTCCAAAGAA BS-S069ATGGAATGTTAAT 111 729ACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAG GTTTTGACCAACTATATACCATCATCACTCGGTACCGTTCCAAAGAAAT BS-S070GGAATGTTAAT 109 730-124- 99975521.7Attorney Docket No: 124540-829360 ACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAG GTTTTGACCAACTATATACCATCATCACTCGGTACTTCCAAAGAAATGG BS-S071AATGTTAAT 107 731ACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAG GTTTTGACCAACTATATACCATCATCACTCGGTACCCAAAGAAATGGAA BS-S072TGTTAAT 105 732GACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAA GGTTTTGACCAACTATATACCATCATCACTCGGTAACAATGTTAGGCGT BS-S073TCCAAAGAAATGGAATGTTAAT 120 733GACGGTAACTTAAAGTGCCGAAGGCGGAGGAGATGGATTAAATATATAA GGTTTTGACCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAA BS-S074CAATGTGTGACCGTTAGGCGTTCCAAAGAAATGGAATGTTAAT 141 734GACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAA GGGTTTGCCCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAA BS-S075CAATGTGTGACCGTTAGGCGTTCCAAAGAAATGGAACGTTAAT 141 735GACGGTAACTTAAAGTGCCGAAGGCCGAGGAGATGGATTAAATATATAA GGGTTTGCCCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAA BS-S076CAATGTGTGACCGTTAGGCGTTCCAAAGAAATGGAACGTTAAT 141 736GACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAA GGGTTTGCCCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAA BS-S077CAATGTGTGACCGTTAGGCGTTCCAAAGAAATGGAATGTTAAT 141 737GACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAA GGTTTTGACCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAA BS-S078CAATGTGTGACCGTTAGGCGTTCCAAAGAAATGGAACGTTAAT 141 738ACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAG GTTTTGACCAACTATATACCATCATCACTCGGTAAGTAACAATGTGTGA BS-S079CCGTTAGGCGTTCCAAAGAAATGGAATGTTAAT 131 739ACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAG GTTTTGACCAACTATATACCATCATCACTCGGTAAGTAACAATTTAGGC BS-S080GTTCCAAAGAAATGGAATGTTAAT 122 740ACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAG GTTTTGACCAACTATATACCATCATCACTCGGTAAGTAACGTTCCAAAG BS-S081AAATGGAATGTTAAT 113 741ACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAG GTTTTGACCAACTATATACCATCATCACTCGGTAAGTCCAAAGAAATGG BS-S082AAT 101 742ACGGTAACTTAAAGTGCCGAAGAGGAGATGGATTAAATATATAAGGTTT BS-S083TGACCAACTATATACCATCATCACTAAGTCCAAAGAAATGGAAT 93 743ACGGTAACTTAAAGTGCCGAAGATGGATTAAATATATAAGGTTTTGACC BS-S084AACTATATACCATCATCACTAAGTCCAAAGAAATGGAAT 88 744ACGGTAACTTAAAGTGCCGAAGGGATTAAATATATAAGGTTTTGACCAA BS-S085CTATATACCCATCACTAAGTCCAAAGAAATGGAAT 84 745ACGGTAACTTAAAGTGCCGAAGGGATTAAATATATAAGGTTTTGACCAA BS-S086CTATATACCCATCAGTCCAAAGAAATGGAAT 80 746ACGGTAACTTAAAGTGCCGAAGGGATTAAATATATAAGGTTTTGACCAA BS-S087CTATATACCCATCACAAAGAAATGGAAT 77 747ACGGTAACTTAAAGTGCCGAAGGGATTAAATATATAAGGTTTTGACCAA BS-S088CTATATACCCATCAGTCCAATGGAATGGAAT 80 748ACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAG GTTTTGACCAACTATATACCATCATCACTCGGTAAGTCCAATGGAATGG BS-S089AAT 101 749ACGGTAACTTAAAGTGCCGAAGAGGAGATGGATTAAATATATAAGGTTT BS-S090TGACCAACTATATACCATCATCACTCGTAAGTCCAATGGAATGGAAT 96 750ACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAG GTTTTGACCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAAC BS-S091AATGTGTGACCGTTAGGCGTTCGAATGTTAAT 130 751ACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAG GTTTTGACCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAAC BS-S092AATGTGTGACCGTTAGGCGTTCCATGGAATGTTAAT 134 752-125- 99975521.7Attorney Docket No: 124540-829360 ACGGCGCCGAAGGCTGAGGAGATGGATTAAATATATAAGGTTTTGACCA ACTATATACCATCATCACTCGGTAAGGGTTAATCCTAACAATGTGTGAC BS-S093CGTTAGGCGTTCCAAAGAAATGGAATGTTAAT 130 753ACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAG GTTTTGACCAACTATATACCATCATCACTCGGTAAGGCCTAACAATGTG BS-S094TGACCGTTAGGCGTTCCAAAGAAATGGAATGTTAAT 134 754ACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAG GTTTTGACCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAAC BS-S095AATGTGTGACCGGCGTTCCAAAGAAATGGAATGTTAAT 136 755ACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAG GTTTTGACCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAAC BS-S096AATTTAGGCGTTCCAAAGAAATGGAATGTTAAT 131 756ACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAG GTTTTGACCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAAC BS-S097AATCGTTCCAAAGAAATGGAATGTTAAT 126 757AGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAGGTTTTGACCAAC TATATACCATCATCACTCGGTAAGGGTTAATCCTAACAATGTGTGACCG BS-S098TTAGGCGTTCCAAAGAAATGGAATGTTAAT 128 758AAGGCTGAGGAGATGGATTAAATATATAAGGTTTTGACCAACTATATAC CATCATCACTCGGTAAGGGTTAATCCTAACAATGTGTGACCGTTAGGCG BS-S099TTCCAAAGAAATGGAATGTTAAT 121 759ACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAG ACCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAACAATGTG BS-S100TGACCGTTAGGCGTTCCAAAGAAATGGAATGTTAAT 134 760ACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAG GTTTTGACCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTCGT BS-S101TCCAAAGAAATGGAATGTTAAT 120 761ACGGTAACTTAAAGTGCCGAAGGCTGAGGAAACTATATACCATCATCAC TCGGTAAGGGTTAATCCTAACAATGTGTGACCGTTAGGCGTTCCAAAGA BS-S102AATGGAATGTTAAT 112 762ACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAG GTTTTGACCAACTATATACCACTCGGTAAGGGTTAATCCTAACAATGTG BS-S103TGACCGTTAGGCGTTCCAAAGAAATGGAATGTTAAT 134 763CAGGATAAATTTGAATGCCAAAAGTGTGGCTTTACTCTTAATGCAGATC ATAATGCAGCAATAAATATTGCACGTAAGTAAACATTGATAAAACTAAC CTTTATTTACAAATTTAAAATAATGCAATATAATTGTATTGATGTTATA TATAAATATATAGCAGTTACGGTAACTTAAAGTGCCGAAGGCTGAGGAG ATGGATTAAATATATAAGGTTTTGACCAACTATATACCATCATCACTCG GTAAGGGTTAATCCTAACAATGTGTGACCGTTAGGCGTTCCAAAATTGA ATAATTATATTAGTTTATTGTCTATCATAGATAGTGTAAAAAACACAGG BS-S104GGGTAGACAAAATAAGTAATATAAGTATATGGTAGCTA 381 764CAGGATAAATTTGAATGCCAAAAGTGTGGCTTTACTCTTAATGCAGATC ATAATGCAGCAATAAATATTGCACGTAAGTAAACATTGATAAAACTAAC CTTTATTTACAAATTTAAAATAATGCAATATAATTGTATTGATGTTATA TATAAATATATAGCAGTTACGGTAACTTAAAGTGCCGAAGGCTGAGGAG ATGGATTAAATATATAAGGTTTTGACCAACTATATACCATCATCACTCG BS-S105GTAAGGGTTAATCCTAACAATGTGTGACCGTTAGGCGTTCCAAAATT 292 765CAGGATAAATTTGAATGCCAAAAGTGTGGCTTTACTCTTAATGCAGATC ATAATGCAGCAATAAATATTGCACGTAAGTAAACATTGATAAAACTAAC CTTTATTTACAAATTTAAAATAATGCAATATAATTGTATTGATGTTATA TATAAATATATAGCAGTTACGGTAACTTAAAGTGCCGAAGGCTGAGGAG ATGGATTAAATATATAAGGTTTTGACCAACTATATACCATCATCACTCG BS-S106GTAAGGGTTAATCCTAACAATGTGTGACCGTT 277 766CAGGATAAATTTGAATGCCAAAAGTGTGGCTTTACTCTTAATGCAGATC ATAATGCAGCAATAAATATTGCACGTAAGTAAACATTGATAAAACTAAC CTTTATTTACAAATTTAAAATAATGCAATATAATTGTATTGATGTTATA TATAAATATATAGCAGTTACGGTAACTTAAAGTGCCGAAGGCTGAGGAG ATGGATTAAATATATAAGGTTTTGACCAACTATATACCATCATCACTCG BS-S107GTAAGGGTTAATCCTAACAATGTGTGACCGTTAGGCGTTCCAAA 289 767-126- 99975521.7Attorney Docket No: 124540-829360 CAGGATAAATTTGAATGCCAAAAGTGTGGCTTTACTCTTAATGCAGATC ATAATGCAGCAATAAATATTGCACGTAAGTAAACATTGATAAAACTAAC CTTTATTTACAAATTTAAAATAATGCAATATAATTGTATTGATGTTATA TATAAATATATAGCAGTTACGGTAACTTAAAGTGCCGAAGGCTGAGGAG ATGGATTAAATATATAAGGTTTTGACCAACTATATACCATCATCACTCG BS-S108GTAAGGGTTAATCCTAACAATGTGTGACCGTTAGGCGTTCCAAA 289 768CGTAAGTAAACATTGATAAAACTAACCTTTATTTACAAATTTAAAATAA TGCAATATAATTGTATTGATGTTATATATAAATATATAGCAGTTACGGT AACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAGGTTTT GACCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAACAATGT GTGACCGTTAGGCGTTCCAAAATTGAATAATTATATTAGTTTATTGTCT ATCATAGATAGTGTAAAAAACACAGGGGGTAGACAAAATAAGTAATATA BS-S109AGTATATGGTAGCTA 309 769cGTAAGTAAACATTGATAAAACTAACCTTTATTTACAAATTTAAAATAA TGCAATATAATTGTATTGATGTTATATATAAATATATAGCAGTTACGGT AACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAGGTTTT GACCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAACAATGT BS-S110GTGACCGTTAGGCGTTCCAAA 217 770CGTAAGTAAACATTGATAAAACTAACCTTTATTTACAAATTTAAAATAA TGCAATATAATTGTATTGATGTTATATATAAATATATAGCAGTTACGGT AACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAGGTTTT GACCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAACAATGT BS-S111GTGACCGTT 205 771cGTAAGTAAACATTGATAAAACTAACCTTTATTTACAAATTTAAAATAA TGCAATATAATTGTATTGATGTTATATATAAATATATAGCAGTTACGGT AACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAGGTTTT GACCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAACAATGT BS-S112GTGACCGTTAGGCGTTCCAAAGAAATGGAATGTTAAT 217 772CGTAAGTAAACATTGATAAAACTAACCTTTATTTACAAATTTAAAATAA TGCAATATAATTGTATTGATGTTATATATAAATATATAGCAGTTACGGT AACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAGGTTTT GACCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAACAATGT BS-S113GTGACCGTTAGGCGTTCCAAAGAAATGGAATGTT 230 773TGTTATATATAAATATATAGCAGTTACGGTAACTTAAAGTGCCGAAGGC TGAGGAGATGGATTAAATATATAAGGTTTTGACCAACTATATACCATCA TCACTCGGTAAGGGTTAATCCTAACAATGTGTGACCGTTAGGCGTTCCA AAATTGAATAATTATATTAGTTTATTGTCTATCATAGATAGTGTAAAAA BS-S114ACACAGGGGGTAGACAAAATAAGTAATATAAGTATATGGTAGCTA 241 774TGTTATATATAAATATATAGCAGTTACGGTAACTTAAAGTGCCGAAGGC TGAGGAGATGGATTAAATATATAAGGTTTTGACCAACTATATACCATCA TCACTCGGTAAGGGTTAATCCTAACAATGTGTGACCGTTAGGCGTTCCA BS-S115AAGAAATGGAATGCC 162 775TGTTATATATAAATATATAGCAGTTACGGTAACTTAAAGTGCCGAAGGC TGAGGAGATGGATTAAATATATAAGGTTTTGACCAACTATATACCATCA BS-S116TCACTCGGTAAGGGTTAATCCTAACAATGTGTGACCGTT 137 776TGTTATATATAAATATATAGCAGTTACGGTAACTTAAAGTGCCGAAGGC TGAGGAGATGGATTAAATATATAAGGTTTTGACCAACTATATACCATCA TCACTCGGTAAGGGTTAATCCTAACAATGTGTGACCGTTAGGCGTTCCA BS-S117AAGAAATGGAATGTTAAT 165 777TGTTATATATAAATATATAGCAGTTACGGTAACTTAAAGTGCCGAAGGC TGAGGAGATGGATTAAATATATAAGGTTTTGACCAACTATATACCATCA TCACTCGGTAAGGGTTAATCCTAACAATGTGTGACCGTTAGGCGTTCCA BS-S118AAGAAATGGAATGTT 162 778ACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAG GTTTTGACCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAAC AATGTGTGACCGTTAGGCGTTCCAAAATTGAATAATTATATTAGTTTAT TGTCTATCATAGATAGTGTAAAAAACACAGGGGGTAGACAAAATAAGTA BS-S119ATATAAGTATATGGTAGCTA 216 779ACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAG GTTTTGACCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAAC BS-S120AATGTGTGACCGTTAGGCGTTCCAAAGAAATGGAATGCC 137 780-127- 99975521.7Attorney Docket No: 124540-829360 ACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAG GTTTTGACCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAAC BS-S121AATGTGTGACCGTT 112 781ACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAG GTTTTGACCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAAC BS-S122AATGTGTGACCGTTAGGCGTTCCAAAGAAATGGAATGTTAAT 140 782ACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAG GTTTTGACCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAAC BS-S123AATGTGTGACCGTTAGGCGTTCCAAAGAAATGGAATGTT 137 783GCTGAGGAGATGGATTAAATATATAAGGTTTTGACCAACTATATACCAT CATCACTCGGTAAGGGTTAATCCTAACAATGTGTGACCGTTAGGCGTTC CAAAATTGAATAATTATATTAGTTTATTGTCTATCATAGATAGTGTAAA BS-S124AAACACAGGGGGTAGACAAAATAAGTAATATAAGTATATGGTAGCTA 194 784GCTGAGGAGATGGATTAAATATATAAGGTTTTGACCAACTATATACCAT CATCACTCGGTAAGGGTTAATCCTAACAATGTGTGACCGTTAGGCGTTC BS-S125CAAAGAAATGGAATGCC 115 785GCTGAGGAGATGGATTAAATATATAAGGTTTTGACCAACTATATACCAT BS-S126CATCACTCGGTAAGGGTTAATCCTAACAATGTGTGACCGTT 90 786GCTGAGGAGATGGATTAAATATATAAGGTTTTGACCAACTATATACCAT CATCACTCGGTAAGGGTTAATCCTAACAATGTGTGACCGTTAGGCGTTC BS-S127CAAAGAAATGGAATGTTAAT 118 787GCTGAGGAGATGGATTAAATATATAAGGTTTTGACCAACTATATACCAT CATCACTCGGTAAGGGTTAATCCTAACAATGTGTGACCGTTAGGCGTTC BS-S128CAAAGAAATGGAATGTT 115 788CGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAGG TTTTGACCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAACA BS-S129ATGTGTGACCGTTAGGCGTTCCAAAGAAATGGAATGTTAAT 139 789ACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAG GTTTTGACCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAAC BS-S130AATGTGTGACCGTTAGGCGTTCCAAAGAAATGGAATGTT 137 790ACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAG GTTTTGACCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAAC BS-S131AATGTGTGACCGTTAGGCGTTCCAAAGAAATGGAATGCCAAT 140 791ACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAG GTTTTGACCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAAC BS-S132AATGTGTGACCGTTAGGCGTTCCAAAGAAATGGAATGCC 137 792GACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAA GGTTTTGACCAACTATATACCATCATCACTCGGTGGCGTTCCAAAGAAA BS-S133TGGAATGTTAAT 110 793TAATGCAATATAATTGTATTGATGTTATATATAAATATATAGCAGTTAC GGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAGGT TTTGACCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAACAA BS-S134TGTGTGACCGTTAGGCGTTCCAAAGAAATGGAATGTTAAT 187 794GACGGTAAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAGGTT TTGACCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAACAAT BS-S135GTGTGACCGTTAGGCGTTCCAAAGAAATGGAATGTTAAT 137 795GACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAA GGTTTTGACCAACTATATACCATCATCACTCGGTAAGGGTTAGAAATGA BS-S136CCGTTAGGCGTTCCAAAGAAATGGAATGTTAAT 131 796GACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAA GGTTTTGACCAACTATATACCATCATCACTCGGTAAGGGTTATTCGTGA BS-S137CCGTTAGGCGTTCCAAAGAAATGGAATGTTAAT 131 797GACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAA GGTTTTGACCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAA BS-S138CAATGTGTGACCGTTAGGCGTTCCGAAAGGAATGTTAAT 137 798GACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAA GGTTTTGACCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAA BS-S139CAATGTGTGACCGTTAGGCGTTCCTTCGGGAATGTTAAT 137 799-128- 99975521.7Attorney Docket No: 124540-829360 GACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAA GGTTTCGACCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAA BS-S140CAATGTGTGACCGTTAGGCGTTCCAAAGAAATGGAATGTTAAT 141 800GACGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAA GGTGAAAACCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAA BS-S141CAATGTGTGACCGTTAGGCGTTCCAAAGAAATGGAATGTTAAT 141 801GCGGTAACTTAAAGTGCCGAAGGCTGAGGAGATGGATTAAATATATAAG GTTTCGACCAACTATATACCATCATCACTCGGTAAGGGTTAATCCTAAC BS-S142AATGTGTGACCGTTAGGCGTTCCGAAAGGAACGTTAAT 136 802

[0277] A list of Sequences disclosed herein is provided in the Table below.SEQUENCES Substitutions to WT (for SEQ Cas12f variants). All ID Name substitutions are relative to Prt / N Species NO: SEQ ID NO: 1 unless otherwise t stated. 1 BsCas12f WT N / A Prt Blautia species 2CnCas12f1 N / A Prt Clostridium novyi3 Pt1Cas12f1 N / A PrtParageobacillus thermoglucosidasius 4SpCas12f1 N / A Prt Syntrophomonas palmitatica5 UnCas12f1 N / A Prt Uncultured bacterium 6 Un2Cas12f1 N / A Prt Uncultured bacterium 7 Ob3Cas12f1 N / A Prt Oscillospiraceae bacterium 8 Cb1Cas12f1 N / A Prt Clostridia bacterium 9 EsCas12f1 N / A Prt Eubacterium siraeum 10 RhgCas12f1 N / A Prt Ruminiclostridium hungatei 11 Cb3Cas12f1 N / A Prt Clostridium botulinum 12 Pt2Cas12f1 N / A Prt Parageobacillus thermoglucosidasius 13 CrCas12f1 N / A Prt Cellulosilyticum ruminicola 14 ChCas12f1 N / A Prt Clostridium hiranonis strain DSM 13275 15 AsCas12f1 N / A Prt Acidibacillus sulfuroxidans 16 dBsCas12f1 variant Y48F Prt synthetic construct 17 dBsCas12f1 variant Y48R Prt synthetic construct 18 dBsCas12f1 variant A50R Prt synthetic construct 19 dBsCas12f1 variant A50Y Prt synthetic construct 20 dBsCas12f1 variant T66F Prt synthetic construct 21 dBsCas12f1 variant T66Q Prt synthetic construct 22 dBsCas12f1 variant S98H Prt synthetic construct 23 dBsCas12f1 variant S98R Prt synthetic construct 24 dBsCas12f1 variant T99R Prt synthetic construct 25 dBsCas12f1 variant E109M Prt synthetic construct 26 dBsCas12f1 variant E109R Prt synthetic construct -129- 99975521.7Attorney Docket No: 124540-829360 27 dBsCas12f1 variant N114Q Prt synthetic construct 28 dBsCas12f1 variant N114A Prt synthetic construct 29 dBsCas12f1 variant F119Y Prt synthetic construct 30 dBsCas12f1 variant F119H Prt synthetic construct 31 dBsCas12f1 variant M122R Prt synthetic construct 32 dBsCas12f1 variant S131H Prt synthetic construct 33 dBsCas12f1 variant S131R Prt synthetic construct 34 dBsCas12f1 variant G176R Prt synthetic construct 35 dBsCas12f1 variant G176H Prt synthetic construct 36 dBsCas12f1 variant D183K Prt synthetic construct 37 dBsCas12f1 variant K184R Prt synthetic construct 38 dBsCas12f1 variant N195Q Prt synthetic construct 39 dBsCas12f1 variant N195R Prt synthetic construct 40 dBsCas12f1 variant K220R Prt synthetic construct 41 dBsCas12f1 variant A223R Prt synthetic construct 42 dBsCas12f1 variant G266R Prt synthetic construct 43 dBsCas12f1 variant A276R Prt synthetic construct 44 dBsCas12f1 variant Y306F Prt synthetic construct 45 dBsCas12f1 variant Y306M Prt synthetic construct 46 dBsCas12f1 variant K318R Prt synthetic construct 47 dBsCas12f1 variant K318G Prt synthetic construct 48 dBsCas12f1 variant D353R Prt synthetic construct 49 dBsCas12f1 variant E109A Prt synthetic construct 50 dBsCas12f1 variant E109C Prt synthetic construct 51 dBsCas12f1 variant E109D Prt synthetic construct 52 dBsCas12f1 variant E109F Prt synthetic construct 53 dBsCas12f1 variant E109G Prt synthetic construct 54 dBsCas12f1 variant E109H Prt synthetic construct 55 dBsCas12f1 variant E109I Prt synthetic construct 56 dBsCas12f1 variant E109K Prt synthetic construct 57 dBsCas12f1 variant E109L Prt synthetic construct 58 dBsCas12f1 variant E109N Prt synthetic construct 59 dBsCas12f1 variant E109P Prt synthetic construct 60 dBsCas12f1 variant E109Q Prt synthetic construct 61 dBsCas12f1 variant E109S Prt synthetic construct 62 dBsCas12f1 variant E109T Prt synthetic construct 63 dBsCas12f1 variant E109V Prt synthetic construct 64 dBsCas12f1 variant E109W Prt synthetic construct 65 dBsCas12f1 variant E109Y Prt synthetic construct 66 dBsCas12f1 variant S131A Prt synthetic construct 67 dBsCas12f1 variant S131C Prt synthetic construct 68 dBsCas12f1 variant S131D Prt synthetic construct -130- 99975521.7Attorney Docket No: 124540-829360 69 dBsCas12f1 variant S131E Prt synthetic construct 70 dBsCas12f1 variant S131F Prt synthetic construct 71 dBsCas12f1 variant S131G Prt synthetic construct 72 dBsCas12f1 variant S131I Prt synthetic construct 73 dBsCas12f1 variant S131K Prt synthetic construct 74 dBsCas12f1 variant S131L Prt synthetic construct 75 dBsCas12f1 variant S131M Prt synthetic construct 76 dBsCas12f1 variant S131N Prt synthetic construct 77 dBsCas12f1 variant S131P Prt synthetic construct 78 dBsCas12f1 variant S131Q Prt synthetic construct 79 dBsCas12f1 variant S131T Prt synthetic construct 80 dBsCas12f1 variant S131V Prt synthetic construct 81 dBsCas12f1 variant S131W Prt synthetic construct 82 dBsCas12f1 variant S131Y Prt synthetic construct 83 dBsCas12f1 variant E181A Prt synthetic construct 84 dBsCas12f1 variant E181C Prt synthetic construct 85 dBsCas12f1 variant E181D Prt synthetic construct 86 dBsCas12f1 variant E181F Prt synthetic construct 87 dBsCas12f1 variant E181G Prt synthetic construct 88 dBsCas12f1 variant E181H Prt synthetic construct 89 dBsCas12f1 variant E181I Prt synthetic construct 90 dBsCas12f1 variant E181K Prt synthetic construct 91 dBsCas12f1 variant E181L Prt synthetic construct 92 dBsCas12f1 variant E181M Prt synthetic construct 93 dBsCas12f1 variant E181N Prt synthetic construct 94 dBsCas12f1 variant E181P Prt synthetic construct 95 dBsCas12f1 variant E181Q Prt synthetic construct 96 dBsCas12f1 variant E181R Prt synthetic construct 97 dBsCas12f1 variant E181S Prt synthetic construct 98 dBsCas12f1 variant E181T Prt synthetic construct 99 dBsCas12f1 variant E181V Prt synthetic construct 100 dBsCas12f1 variant E181W Prt synthetic construct 101 dBsCas12f1 variant E181Y Prt synthetic construct 102 dBsCas12f1 variant D183A Prt synthetic construct 103 dBsCas12f1 variant D183C Prt synthetic construct 104 dBsCas12f1 variant D183E Prt synthetic construct 105 dBsCas12f1 variant D183F Prt synthetic construct 106 dBsCas12f1 variant D183G Prt synthetic construct 107 dBsCas12f1 variant D183H Prt synthetic construct 108 dBsCas12f1 variant D183I Prt synthetic construct 109 dBsCas12f1 variant D183L Prt synthetic construct 110 dBsCas12f1 variant D183M Prt synthetic construct 111 dBsCas12f1 variant D183N Prt synthetic construct -131- 99975521.7Attorney Docket No: 124540-829360 112 dBsCas12f1 variant D183P Prt synthetic construct 113 dBsCas12f1 variant D183Q Prt synthetic construct 114 dBsCas12f1 variant D183R Prt synthetic construct 115 dBsCas12f1 variant D183S Prt synthetic construct 116 dBsCas12f1 variant D183T Prt synthetic construct 117 dBsCas12f1 variant D183V Prt synthetic construct 118 dBsCas12f1 variant D183W Prt synthetic construct 119 dBsCas12f1 variant D183Y Prt synthetic construct 120 dBsCas12f1 variant K184A Prt synthetic construct 121 dBsCas12f1 variant K184C Prt synthetic construct 122 dBsCas12f1 variant K184D Prt synthetic construct 123 dBsCas12f1 variant K184E Prt synthetic construct 124 dBsCas12f1 variant K184F Prt synthetic construct 125 dBsCas12f1 variant K184G Prt synthetic construct 126 dBsCas12f1 variant K184H Prt synthetic construct 127 dBsCas12f1 variant K184I Prt synthetic construct 128 dBsCas12f1 variant K184L Prt synthetic construct 129 dBsCas12f1 variant K184M Prt synthetic construct 130 dBsCas12f1 variant K184N Prt synthetic construct 131 dBsCas12f1 variant K184P Prt synthetic construct 132 dBsCas12f1 variant K184Q Prt synthetic construct 133 dBsCas12f1 variant K184S Prt synthetic construct 134 dBsCas12f1 variant K184T Prt synthetic construct 135 dBsCas12f1 variant K184V Prt synthetic construct 136 dBsCas12f1 variant K184W Prt synthetic construct 137 dBsCas12f1 variant K184Y Prt synthetic construct 138 dBsCas12f1 variant N195A Prt synthetic construct 139 dBsCas12f1 variant N195C Prt synthetic construct 140 dBsCas12f1 variant N195D Prt synthetic construct 141 dBsCas12f1 variant N195E Prt synthetic construct 142 dBsCas12f1 variant N195F Prt synthetic construct 143 dBsCas12f1 variant N195G Prt synthetic construct 144 dBsCas12f1 variant N195H Prt synthetic construct 145 dBsCas12f1 variant N195I Prt synthetic construct 146 dBsCas12f1 variant N195K Prt synthetic construct 147 dBsCas12f1 variant N195L Prt synthetic construct 148 dBsCas12f1 variant N195M Prt synthetic construct 149 dBsCas12f1 variant N195P Prt synthetic construct 150 dBsCas12f1 variant N195S Prt synthetic construct 151 dBsCas12f1 variant N195T Prt synthetic construct 152 dBsCas12f1 variant N195V Prt synthetic construct 153 dBsCas12f1 variant N195W Prt synthetic construct -132- 99975521.7Attorney Docket No: 124540-829360 154 dBsCas12f1 variant N195Y Prt synthetic construct 155 dBsCas12f1 variant K220A Prt synthetic construct 156 dBsCas12f1 variant K220C Prt synthetic construct 157 dBsCas12f1 variant K220D Prt synthetic construct 158 dBsCas12f1 variant K220E Prt synthetic construct 159 dBsCas12f1 variant K220F Prt synthetic construct 160 dBsCas12f1 variant K220G Prt synthetic construct 161 dBsCas12f1 variant K220H Prt synthetic construct 162 dBsCas12f1 variant K220I Prt synthetic construct 163 dBsCas12f1 variant K220L Prt synthetic construct 164 dBsCas12f1 variant K220M Prt synthetic construct 165 dBsCas12f1 variant K220N Prt synthetic construct 166 dBsCas12f1 variant K220P Prt synthetic construct 167 dBsCas12f1 variant K220Q Prt synthetic construct 168 dBsCas12f1 variant K220S Prt synthetic construct 169 dBsCas12f1 variant K220T Prt synthetic construct 170 dBsCas12f1 variant K220V Prt synthetic construct 171 dBsCas12f1 variant K220W Prt synthetic construct 172 dBsCas12f1 variant K220Y Prt synthetic construct 173 dBsCas12f1 variant Y306A Prt synthetic construct 174 dBsCas12f1 variant Y306C Prt synthetic construct 175 dBsCas12f1 variant Y306D Prt synthetic construct 176 dBsCas12f1 variant Y306E Prt synthetic construct 177 dBsCas12f1 variant Y306G Prt synthetic construct 178 dBsCas12f1 variant Y306H Prt synthetic construct 179 dBsCas12f1 variant Y306I Prt synthetic construct 180 dBsCas12f1 variant Y306K Prt synthetic construct 181 dBsCas12f1 variant Y306L Prt synthetic construct 182 dBsCas12f1 variant Y306N Prt synthetic construct 183 dBsCas12f1 variant Y306P Prt synthetic construct 184 dBsCas12f1 variant Y306Q Prt synthetic construct 185 dBsCas12f1 variant Y306R Prt synthetic construct 186 dBsCas12f1 variant Y306S Prt synthetic construct 187 dBsCas12f1 variant Y306T Prt synthetic construct 188 dBsCas12f1 variant Y306V Prt synthetic construct 189 dBsCas12f1 variant Y306W Prt synthetic construct 190 dBsCas12f1 variant K318A Prt synthetic construct 191 dBsCas12f1 variant K318C Prt synthetic construct 192 dBsCas12f1 variant K318D Prt synthetic construct 193 dBsCas12f1 variant K318E Prt synthetic construct 194 dBsCas12f1 variant K318F Prt synthetic construct 195 dBsCas12f1 variant K318H Prt synthetic construct -133- 99975521.7Attorney Docket No: 124540-829360 196 dBsCas12f1 variant K318I Prt synthetic construct 197 dBsCas12f1 variant K318L Prt synthetic construct 198 dBsCas12f1 variant K318M Prt synthetic construct 199 dBsCas12f1 variant K318N Prt synthetic construct 200 dBsCas12f1 variant K318P Prt synthetic construct 201 dBsCas12f1 variant K318Q Prt synthetic construct 202 dBsCas12f1 variant K318S Prt synthetic construct 203 dBsCas12f1 variant K318T Prt synthetic construct 204 dBsCas12f1 variant K318V Prt synthetic construct 205 dBsCas12f1 variant K318W Prt synthetic construct 206 dBsCas12f1 variant K318Y Prt synthetic construct 207 dBsCas12f1 variant D353A Prt synthetic construct 208 dBsCas12f1 variant D353C Prt synthetic construct 209 dBsCas12f1 variant D353E Prt synthetic construct 210 dBsCas12f1 variant D353F Prt synthetic construct 211 dBsCas12f1 variant D353G Prt synthetic construct 212 dBsCas12f1 variant D353H Prt synthetic construct 213 dBsCas12f1 variant D353I Prt synthetic construct 214 dBsCas12f1 variant D353K Prt synthetic construct 215 dBsCas12f1 varian...

Claims

Attorney Docket No: 124540-829360 CLAIMS1. A Cas12f polypeptide having one or more amino acid substitutions relative to an aminoacid sequence of a wildtype Cas12f nuclease obtained from a Blautia species, wherein the Cas12f polypeptide has improved targeting and / or editing efficiency relative to the wildtype bsCas12f (bsCas12f).

2. The Cas12f polypeptide of claim 1, wherein the wildtype Cas12f nuclease has an aminoacid sequence of SEQ ID NO:1, and the one or more amino acid substitutions are selected from a substitution at any one or more of I10, Y48, A50, T66, D76, S98, T99, E109, I110, G111, N114, F119, N103, M122R, S131, W143, I159, G176, E181, D183, K184, N195, I196, I201, P209, K220, A223, G266, A276, Y306, T307, N312, K318, A355, and D353 in SEQ ID NO: 1.

3. The Cas 12f polypeptide of claim 2, wherein the one or more amino acid substitutions areat one or more residues of N195, Y306, T307, E109, S131, E181, D183, D353, D183, I10, D76, G111, I159, I196 in SEQ ID NO: 1.

4. The Cas12f polypeptide of any one of claims 1 to 3, wherein the one or more amino acidsubstitutions are selected from I10R, Y48F, Y48R, A50R, A50Y, T66Q, D76K, S98H, S98R, T99R, E109M, E109R, G111N, N114Q, N114A, F119Y, F119H, M122R, S131R, S131H, S131C, W143T, I159R, G176R, G176H, E181Q, D183K, D183T, K184R, N195Q, N195R, N195K, N195H, I196V, K220R, A223R, G266R, A276R, Y306F, Y306M, Y306K, T307K, N312V, N312Y, K318R, K318G and D353R.

5. The Cas12f polypeptide of any one of claims 1 to 4, wherein Cas12f polypeptidecomprises two or more amino acid substitutions selected from: N195K, N195H, E109H, S131W, S131C, E181Q, D183T, Y306K, T307K, I10R, D76K, G111N, I159R, I196V, and D353K.

6. The Cas12f polypeptide of any one of claims 1 to 5, wherein the Cas12f polypeptidecomprises three or more amino acid substitutions selected from: N195K, N195H, Y306K, D353K, S131Q, E181Q, D183T, I10R, D76K, G111N, I159R, I196V, and T307K,7. The Cas12f polypeptide of any one of claims 1 to 5, wherein the Cas12f polypeptidecomprises four or more amino acid substitutions selected from N195K,Y306K, T307K, E109H, S131C, S131Q, E181Q, D183T, I10R, D76K, G111N, I159R, I196V, and D353K. -154- 99975521.7Attorney Docket No: 124540-8293608. The Cas12f polypeptide of any one of claims 1 to 5, wherein the Cas12f polypeptidecomprises five or more amino acid substitutions selected from N195K,Y306K, T307K, S131C, E181Q I10R, D76K, G111N, I159R, I196V, and D183T,9. The Cas12f polypeptide of any one of claims 1 to 5, wherein the Cas12f polypeptidecomprises six or more amino acid substitutions selected from N195K, Y306K, T307K, S131C, E181Q, I10R, D76K, G111N, I159R, I196V, and D183T.

10. The Cas12f polypeptide of any one of claims 1 to 5, wherein the Cas12f polypeptidecomprises seven or more amino acid substitutions selected from N195K, Y306K, T307K, S131C, E181Q, D183T, and I10R.

11. The Cas12f polypeptide of any one of claims 1 to 5, wherein the Cas12f polypeptidecomprises eight or more amino acid substitutions selected from N195K, Y306K, T307K, S131C, E181Q, D183T, I10R, and D76K.

12. The Cas12f polypeptide of any one of claims 1 to 5, wherein the Cas12f polypeptidecomprises nine or more amino acid substitutions selected from N195K, Y306K, T307K, S131C, E181Q, D183T, I10R, D76K, and G111N.

13. The Cas12f polypeptide of any one of claims 1 to 5, wherein the Cas12f polypeptidecomprises ten or more amino acid substitutions selected from N195K, Y306K, T307K, S131C, E181Q, D183T, I10R, D76K, G111N, and I159R.

14. The Cas12f polypeptide of any one of claims 1 to 5, wherein the Cas12f polypeptidecomprises eleven or more amino acid substitutions selected from N195K, Y306K, T307K, S131C, E181Q, D183T, I10R, D76K, G111N, I159R, and I196V.

15. The Cas12f polypeptide of any one of claims 1 to 14, wherein the Cas12f polypeptide hasincreased editing activity relative to the wildtype Cas12f nuclease.

16. The Cas12f polypeptide of any one of claims 1 to 14 wherein the Cas12f polypeptide hasreduced nuclease activity relative to the wildtype Cas12f nuclease.

17. The Cas12f polypeptide of claim 16, further comprising one or more amino acidsubstitutions selected from D216A and D390A.

18. The Cas12f polypeptide of any one of claims 1 to 17, comprising or consisting of anamino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least -155- 99975521.7Attorney Docket No: 124540-829360 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs: 16-391 and SEQ ID NOs: 803-807.

19. The Cas12f polypeptide of claim 18, comprising or consisting of an amino acid sequenceof SEQ ID NO: 329-338 and SEQ ID NOs: 803-807.

20. The Cas12f polypeptide of claim 19, comprising or consisting of any amino acidsequence of SEQ ID NO: 338.

21. The Cas12f polypeptide of claim 19, comprising or consisting of any amino acidsequence of SEQ ID NO: 807.

22. A fusion protein comprising a Cas12f polypeptide of any one of claims 1 to 21 and afunctional polypeptide.

23. The fusion protein of claim 22, wherein the functional polypeptide comprises a baseediting domain, for example, a deaminase or a catalytic domain thereof, a base excising domain, an uracil glycosylase inhibitor (UGI) or a catalytic domain thereof, an uracil glycosylase (UNG) or a catalytic domain thereof, a methylpurine glycosylase (MPG) or a catalytic domain thereof, a methylase or a catalytic domain thereof, a demethylase or a catalytic domain thereof, an transcription activating domain (e.g., VP64 or VPR) , an transcription inhibiting domain (e.g., KRAB moiety or SID moiety) , a reverse transcriptase or a catalytic domain thereof, an exonuclease (e.g., T5E) or a catalytic domain thereof, a destabilized domain (e.g., destabilized domains (DD) of E. coli dihydrofolate reductase (ecDHFR), a histone residue modification domain, a nuclease catalytic domain (e.g., FokI) , a transcription modification factor, a light gating factor, a chemical inducible factor, a chromatin visualization factor, a targeting polypeptide for providing binding to a cell surface portion on a target cell or a target cell type, a reporter (e.g., fluorescent) polypeptide or a detection label (e.g., GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, BFP) , a localization signal, a polypeptide targeting moiety, a DNA binding domain (e.g., MBP, Lex A DBD, Gal4 DBD) , an epitope tag (e.g., His, myc, V5, FLAG, HA, VSV-G, Trx, etc.) , a transcription release factor, an HDAC, a moiety having ssRNA cleavage activity, a moiety having dsRNA cleavage activity, a moiety having ssDNA cleavage activity, a moiety having dsDNA cleavage activity, a DNA or RNA ligase, a functional domain exhibiting activity to modify a target DNA, selected from the group consisting of: methyltransferase activity, DNA repair activity, DNA damage activity, dismutase activity, alkylation activity, dealkylation activity, depurination -156- 99975521.7Attorney Docket No: 124540-829360 activity, oxidation activity, deoxidation activity, pyrimidine dimer forming activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, glycosylase activity, acetyl transferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, deSUMOylation activity, ribosylation activity, deribosylation activity, myristoylation activity, demyristoylation activity, glycosylation activity (e.g., from O-GlcNAc transferase) , deglycosylation activity, reverse transcriptase activity, and a catalytic domain thereof, and a functional fragment (e.g., a functional truncation) thereof, or any combination thereof.

24. The fusion protein of claim 22 or 23, wherein the functional polypeptide comprises aTen-eleven translocation methylcytosine dioxygenase (Tet, Tet2 or Tet3, or a functional or catalytic domain thereof), DNA methyl transferase (e.g. DNMT3A and / or DNMT3L, or a functional or catalytic domain thereof), VP16, VP64, VPR, ABE, CBE, or a Krüppel- associated box (KRAB).

25. The fusion protein of claim 24, wherein the functional polypeptide comprises acatalytically active truncated Tet polypeptide comprising an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 408 or 409.

26. The fusion protein of claim 24, wherein the functional polypeptide comprises a DNMT3Aprotein.

27. The Cas12f polypeptide of any one of claims 1 to 21 or the fusion protein of any one ofclaims 22 to 26, further comprising a nuclear localization signal (NLS).

28. A nucleic acid encoding the Cas12f polypeptide of any one of claims 1 to 21 or the fusionprotein of any one of claims 22 to 26.

29. A single molecule guide RNA (sgRNA), the sgRNA comprising a scaffold sequencecapable of forming a complex with the Cas12f polypeptide of any one of claims 1 to 21 and / or the fusion protein of any one of claims 22 to 26 and a spacer sequence capable of hybridizing to a target sequence of a target DNA, thereby guiding the complex to the target DNA. -157- 99975521.7Attorney Docket No: 124540-82936030. The sgRNA of claim 29, wherein the scaffold sequence comprises a nucleic acidsequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to any one of SEQ ID NOs: 438-515 and SEQ ID NOs: 673-802.

31. The sgRNA of claim 29 or 30, wherein the single-molecule guide RNA is chemicallymodified.

32. The sgRNA of any one of claims 29 to 31, wherein the sgRNA is pre-complexed with theCas12f polypeptide of any one of claims 1 to 21 or a fusion protein of any one of claims 22 to 26.

33. A polynucleotide encoding the sgRNA of any one of claims 29 to 32.

34. A vector comprising one or more expression constructs comprising the polynucleotide ofclaim 28 and / or the polynucleotide of claim 33.

35. The vector of claim 34, comprising the polynucleotide of claim 28 and the polynucleotideof claim 33.

36. The vector of claim 34 or 35, wherein the vector is an adenoviral, adeno-associated viral(AAV) or lentiviral vector.

37. The vector of claim 36, wherein the vector comprises a nucleic acid sequence having atleast 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 558-605.

38. A cell comprising a polynucleotide of claim 28, the polynucleotide of claim 33, the vectorof any one of claims 34 to 37, or any combination thereof.

39. The cell of claim 38, further comprising one or more TurboRFP expressing constructsunder control of a TRE operator comprising seven TetO operators, wherein the cell expresses RFP only in the presence of a targeted transcriptional activator binding to the TetO operator of at least one TurboRFP expressing construct.

40. The cell of claim 39, wherein each TetO operator in each TurboRFP expressing constructcomprises a PAM sequence selected from TTTA, TTTC or TCCA preceding each TetO operator sequence. -158- 99975521.7Attorney Docket No: 124540-82936041. The cell of any one of claims 38 to 40, comprising a vector having at least 70%, at least75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 593-595.

42. The cell of any one of claims 38 to 41, wherein the cell is mammalian cell.

43. The cell of claim 42, wherein the cell is a human cell.

44. A gene editing system comprising:(1) the Cas12f polypeptide of any one of claims 1-21, a fusion protein of any one of claims 22-26 or the polynucleotide of claim 28 and (2) a single molecule guide RNA (sgRNA) of any one of claims 29 to 32 or the polynucleotide of claim 33, and optionally, (3) a transgene encoding a therapeutic protein of interest, which is optionally comprised in an episome.

45. The gene editing system of claim 44, wherein the system is packaged into an adenoviral,adeno-associated viral (AAV) or lentiviral vector.

46. The gene editing system of claim 44 or 45, wherein the system is further packaged into aliposome or lipid nanoparticle.

47. A method for modifying a target DNA, comprising contacting the target DNA with acomplex formed between a Cas12f polypeptide of any one of claims 1 to 21 or a fusion protein of any one of claims 22 to 26 and a sgRNA of any one of claims 29 to 32, wherein the sgRNA comprises a spacer sequence capable of hybridizing to a target sequence in the target DNA, and wherein the complex modifies the target DNA.

48. The method of claim 47, wherein the complex modifies the target DNA by introducingone or more single or double stranded breaks at the target sequence.

49. The method of claim 48, further comprising providing an exogenous nucleic acid,wherein the exogenous nucleic acid is inserted at the site of the double stranded break in the target sequence.

50. The method of claim 49, wherein the complex modifies the target DNA by editing a baseat or near the target sequence. -159- 99975521.7Attorney Docket No: 124540-82936051. A method for modifying expression of a target gene, the method comprising: contacting anucleic acid comprising the target gene with a complex formed between a Cas12f polypeptide of any one of claims 1 to 21 or a fusion protein of any one of claims 22 to 26 and a sgRNA of any one of claims 29 to 32, wherein the sgRNA comprises a spacer sequence capable of hybridizing to a target sequence in the target gene, and wherein the complex modifies expression of the target gene.

52. The method of claim 51, wherein the complex comprises fusion protein comprising afunctional polypeptide that modifies expression of the target gene.

53. The method of claim 52, wherein the functional polypeptide has methyltransferaseactivity, DNA repair activity, DNA damage activity, dismutase activity, alkylation activity, demethylation activity, dealkylation activity, depurination activity, oxidation activity, deoxidation activity, pyrimidine dimer forming activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, glycosylase activity, acetyl transferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, deSUMOylation activity, ribosylation activity, deribosylation activity, myristoylation activity, demyristoylation activity, glycosylation activity (e.g., from O-GlcNAc transferase) , deglycosylation activity or any combination thereof.

54. The method of claim 51 or 52, wherein the functional polypeptide comprises a Ten-eleven translocation methylcytosine dioxygenase (Tet, Tet2 or Tet3, or a functional or catalytic domain thereof (e.g., TetMini)), DNA methyl transferase (e.g. DNMT3A and / or DNMT3L or a functional or catalytic domain thereof), VP16, VP64, VPR, ABE, CBE, or a Krüppel-associated box (KRAB).

55. The method of any one of claims 51 to 54, wherein the target DNA and / or the target geneare in a cell, and the method further comprises delivering to the cell the Cas12f polypeptide or a polynucleotide encoding the Cas12f polypeptide and the sgRNA or a polynucleotide encoding the sgRNA.

56. The method of claim 55, wherein the cell is a human cell.

57. The method of claim 55 or 56, wherein the cell is in vitro or in vivo.-160- 99975521.7Attorney Docket No: 124540-82936058. A method for diagnosing, preventing or treating a disease in a subject in need thereof, themethod comprising: modifying a target DNA and / or modifying expression of a target gene in at least one cell of the subject according to the method of any one of claims 51 to 57, wherein modifying the target DNA and / or modifying expression of the target gene diagnoses, prevents and / or treats the disease.

59. The method of claim 58, wherein the disease comprises a genetic disease.

60. The method of claim 58 or 59, wherein the subject is a human.-161- 99975521.7

Citation Information

Patent Citations

  • Novel crispr-CAS12f systems and uses thereof

    WO2023208000A1