Programmable DNA binding proteins and methods of use thereof

Programmable DNA binding proteins, like Cas12 enzymes, enhance CRISPR-Cas systems' efficiency and adaptability, addressing limitations in existing technologies for precise nucleic acid modifications and gene editing.

WO2025162433A1PCT designated stage Publication Date: 2025-08-07ACCUREDIT THERAPEUTICS (SUZHOU) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/075403
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-05
Filing Date
2025-01-27
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

Existing CRISPR-Cas systems have limitations in terms of effectiveness, specificity, PAM adaptability, and multiplexing capabilities, hindering their versatility in practical applications.

Method used

Development of programmable DNA binding proteins, such as Cas12 enzymes, with specific amino acid sequences that enhance gene editing efficiency and adaptability, combined with engineered guide RNAs and CRISPR-Cas systems, allowing for precise nucleic acid modifications and gene editing.

Benefits of technology

The programmable DNA binding proteins exhibit increased gene editing efficiency and specificity, overcoming limitations of existing systems, enabling more versatile applications in fields like therapy and agriculture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025075403_07082025_PF_FP_ABST
    Figure CN2025075403_07082025_PF_FP_ABST
Patent Text Reader

Abstract

Provided herein are programmable DNA binding proteins, and nucleic acids that encode the programmable DNA binding proteins, wherein each of the programmable DNA binding proteins comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identical to any one of SEQ ID NOs: 1-92 or 241-253. Also provided herein are products (e.g., an engineered, non-naturally occurring CRISPR-Cas system, a recombinant expression system, a cell, a pharmaceutical composition or a kit), methods of use of the programmable DNA binding proteins.Further provided herein is an engineered guide RNA.
Need to check novelty before this filing date? Find Prior Art

Description

PROGRAMMABLE DNA BINDING PROTEINS AND METHODS OF USE THEREOF

[0001] The present application claims the benefit of the priority applications PCT / CN2024 / 075387 filed on February 2, 2024 and PCT / CN2024 / 103769 filed on July 5, 2024, both of which are incorporated by reference in their entirety.TECHNICAL FIELD

[0002] This disclosure relates to programmable DNA binding proteins for use with, e.g., nucleic acid-guided nucleases for making modifications in target nucleic acid sequences.BACKGROUND

[0003] Crispr-Cas systems have found extensive applications in fields such as therapy, agriculture, and various industries. Nevertheless, the range of characterized systems remains relatively narrow, and existing systems possess inherent limitations in terms of their effectiveness, specificity, PAM (Protospacer Adjacent Motif) adaptability, multiplexing capabilities, and accommodating larger genes. These constraints hinder their versatility in practical applications.SUMMARY

[0004] Provided herein are programmable DNA binding proteins comprising an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identical to any one of SEQ ID NOs: 1-92 or 241-253. In some embodiments, the programmable DNA binding protein comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identical to any one of SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 4, SEQ ID NO:7, SEQ ID NO: 48, SEQ ID NO: 50, SEQ ID NO: 53, SEQ ID NO: 54, SEQ ID NO: 55, or SEQ ID NO: 81. In some embodiments, the programmable DNA binding protein comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identical to any one of SEQ ID NO: 1, SEQ ID NO: 7, or SEQ ID NO: 50, SEQ ID NO: 53. In some embodiments, the programmable DNA binding protein exhibits a 5%, 10%, 20%, 30%, 40%, or 50%increased gene editing efficiency relative to an LbCpf1 endonuclease, wherein the LbCpf1 endonuclease comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identical to SEQ ID NO: 94.

[0005] Also provided herein are nucleic acids comprising a sequence encoding a programmable DNA binding protein, wherein the programmable DNA binding protein comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identical to any one of SEQ ID NOs: 1-92 or 241-253. In some embodiments, the nucleic acid comprises a nucleic acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identical to any one of SEQ ID NOs: 96-105, 237 or 254-267. In some embodiments, the nucleic acid comprises a nucleic acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identical to any one of SEQ ID NO: 237, SEQ ID NO: 254, SEQ ID NO: 100, SEQ ID NO: 99, SEQ ID NO: 97, SEQ ID NO: 98, SEQ ID NO: 96, SEQ ID NOs: 101-104, or SEQ ID NO: 267. In some embodiments, the nucleic acid comprises a nucleic acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identical to any one of SEQ ID NO:96, SEQ ID NO: 99, SEQ ID NO: 100, SEQ ID NO: 104, or SEQ ID NO: 105. In some embodiments, the programmable DNA binding protein exhibits a 5%, 10%, 20%, 30%, 40%, or 50%increased gene editing efficiency relative to an LbCpf1 endonuclease, wherein the LbCpf1 endonuclease comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identical to SEQ ID NO: 94.

[0006] Also provided herein are engineered, non-naturally occurring Clustered Regularly Interspaced Short Palindromic Repeat (CRISPR) -associated (Cas) systems comprising (a) a guide RNA (gRNA) or a nucleic acid encoding the guide RNA, wherein the gRNA comprises a direct repeat sequence and a spacer sequence; and (b) a programmable DNA binding protein or a nucleic acid encoding the programmable DNA binding protein, wherein the programmable DNA binding protein comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identical to any one of SEQ ID NOs: 1-92 or 241-253, wherein the programmable DNA binding protein binds to the gRNA, and wherein the spacer sequence is at least partially complementary to a target nucleic acid. In some embodiments, the programmable DNA binding protein comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identical to any one of SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO:4, SEQ ID NO: 7, SEQ ID NO: 48, SEQ ID NO: 50, SEQ ID NO: 53, SEQ ID NO: 54, SEQ ID NO: 55, or SEQ ID NO: 81. In some embodiments, the programmable DNA binding protein comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identical to any one of SEQ ID NO: 1, SEQ ID NO: 7, SEQ ID NO: 50, or SEQ ID NO: 53.

[0007] In some embodiments, the guide RNA comprises any one of SEQ ID NOs: 106-199 or 229-236. In some embodiments, wherein the guide RNA comprises one or more modifications selected from 2' O-methyl, 2' fluoro, or phosphorothioate modification, or combination thereof. In some embodiments, the spacer sequence is complementary to at least 15 nucleotides of the target nucleic acid. In some embodiments, wherein the spacer sequence comprises one or more modifications selected from 2' O-methyl, 2' fluoro, or phosphorothioate modification, or combination thereof. In some embodiments, the guide RNA comprising one or more modifications is referred to as “engineered guide RNA” .

[0008] In some embodiments, the engineered guide RNA comprises:

[0009] a) a DNA-targeting segment comprising a nucleotide sequence that is complementary to a target sequence in a target DNA molecule; and

[0010] b) a protein-binding segment configured to bind to a programmable DNA binding protein comprising an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identical to any one of SEQ ID NOs: 1-92 or 241-253;

[0011] wherein the engineered guide RNA comprises a nucleotide modification pattern selected from 2' O-methyl, 2' fluoro, phosphorothioate modification, or combination thereof.

[0012] In some embodiments, the modifications in the engineered guide RNA are present in scaffold sequence and / or spacer sequence, optionally, wherein the scaffold sequence comprises any one of SEQ ID NOs: 268-287.

[0013] In some embodiments, the engineered guide RNA comprises any one of SEQ ID NOs: 297-326.

[0014] In some embodiments, the engineered, non-naturally occurring CRISPR-Cas system further comprises a nuclear localization signal (NLS) operably linked to the endonuclease, wherein the gene-editing complex is thereby localized to the nucleus of a cell. In some embodiments, the NLS is an N-terminal NLS and / or a C-terminal NLS. In some embodiments, wherein the NLS comprises an amino acid sequence of any one of SEQ ID NOs: 238-240.

[0015] Also provided herein are recombinant expression systems for an engineered, non-naturally occurring CRISPR-Cas system comprising (i) a nucleic acid sequence encoding a programmable DNA binding protein; and (ii) a nucleic acid sequence encoding a DNA targeting sequence, wherein the programmable DNA binding protein comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identical to any one of SEQ ID NOs: 1-92 or 241-253.

[0016] In some embodiments, the DNA targeting sequence comprises a guide RNA (gRNA) . In some embodiments, the guide RNA comprises any one of SEQ ID NOs: 106-199 or 229-236 or 297-326. In some embodiments, (i) and (ii) are comprised within a same vector or comprised within different vectors.

[0017] Also provided herein are cells comprising any one of the programmable DNA binding proteins, any one of the nucleic acids, any one of the engineered, non-naturally occurring CRISPR-Cas systems, or any one of the recombinant expression systems described herein.

[0018] Also provided herein are pharmaceutical compositions comprising any one of the recombinant expression systems or any one of the cells described herein and a pharmaceutically acceptable carrier.

[0019] Also provided herein is a kit comprising: (a) a programmable DNA binding protein or a nucleic acid encoding thereof, an engineered guide RNA, an engineered, non-naturally occurring CRISPR-Cas system, any one of the recombinant expression systems or any one of the cells described herein, and (b) an instruction.

[0020] Also provided herein are methods of editing a nucleobase of a target nucleic acid sequence, the method comprising contacting the target nucleic acid sequence with any one of the recombinant expression systems or any one of the engineered, non-naturally occurring CRISPR-Cas systems described herein, wherein the DNA targeting sequence binds the target nucleic acid sequence, thereby editing the nucleobase of the target nucleic acid sequence.

[0021] In some embodiments, the present disclosure relates to use of a programmable DNA binding protein or a nucleic acid encoding thereof, an engineered guide RNA, an engineered, non-naturally occurring CRISPR-Cas system, any one of the recombinant expression systems or any one of the engineered, non-naturally occurring CRISPR-Cas systems, or a cell described herein in the manufacture of a kit for editing a nucleobase of a target nucleic acid sequence.

[0022] In some embodiments, the editing comprises a mutation of a single nucleobase of the target nucleic acid sequence, wherein the mutation comprises an insertion, a deletion, or a substitution. In some embodiments, the editing comprises a gene sequence insertion, a gene sequence deletion, or a gene sequence replacement. In some embodiments, the editing comprises an epigenetic modification of the target nucleic acid sequence, wherein the epigenetic modification comprises methylation or de-methylation of the target nucleic acid sequence. In some embodiments, the editing comprises an epigenetic modification of the target nucleic acid sequence, wherein the epigenetic modification modulates a gene expression by activation or repression of the target nucleic acid sequence.

[0023] Also provided herein are methods of cleaving a target nucleic acid sequence, the method comprising contacting the target nucleic acid sequence with any one of the recombinant expression systems or any one of the engineered, non-naturally occurring CRISPR-Cas systems described herein, wherein (i) the DNA targeting sequence binds the target nucleic acid sequence and (ii) the programmable DNA binding protein cleaves the target nucleic acid sequence. In some embodiments, the target nucleic acid sequence is a single-stranded DNA or a double-stranded DNA. In some embodiments, the cleaving of the target nucleic acid sequence generates an insertion of a deletion of a nucleobase of the target nucleic acid sequence. In some embodiments, the target nucleic acid sequence comprises a protospacer adjacent motif (PAM) sequence.

[0024] In some embodiments, the present disclosure relates to use of a programmable DNA binding protein or a nucleic acid encoding thereof, an engineered guide RNA, an engineered, non-naturally occurring CRISPR-Cas system, any one of the recombinant expression systems or any one of the engineered, non-naturally occurring CRISPR-Cas systems, or a cell described herein in the manufacture of a kit for cleaving a target nucleic acid sequence.

[0025] Also provided herein are methods of modulating gene expression of a target nucleic acid sequence, the method comprising contacting the target nucleic acid sequence with any one of the recombinant expression systems or any one of the engineered, non-naturally occurring CRISPR-Cas systems, wherein the DNA targeting sequence binds the target nucleic acid sequence, thereby modulating gene expression of the target nucleic acid sequence. In some embodiments, the gene expression of the target nucleic acid sequence is upregulated. In some embodiments, the gene expression of the target nucleic acid sequence is downregulated.

[0026] In some embodiments, the present disclosure relates to use of a programmable DNA binding protein or a nucleic acid encoding thereof, an engineered guide RNA, an engineered, non-naturally occurring CRISPR-Cas system, any one of the recombinant expression systems or any one of the engineered, non-naturally occurring CRISPR-Cas systems, or a cell described herein in the manufacture of a kit for modulating gene expression of a target nucleic acid sequence.

[0027] Also provided herein are methods of modifying histone at a target nucleic acid sequence, the method comprising contacting the target nucleic acid sequence with any one of the recombinant expression systems or any one of the engineered, non-naturally occurring CRISPR-Cas systems, wherein the DNA targeting sequence binds the target nucleic acid sequence, thereby modifying histone at the target nucleic acid sequence. In some embodiments, this histone modification comprises histone acetylation, phosphorylation, or methylation.

[0028] In some embodiments, the present disclosure relates to use of a programmable DNA binding protein or a nucleic acid encoding thereof, an engineered guide RNA, an engineered, non-naturally occurring CRISPR-Cas system, any one of the recombinant expression systems or any one of the engineered, non-naturally occurring CRISPR-Cas systems, or a cell described herein in the manufacture of a kit for modifying histone at a target nucleic acid sequence.

[0029] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Methods and materials are described herein for use in the present invention; other, suitable methods and materials known in the art can also be used. The materials, methods, and examples are illustrative only and not intended to be limiting. All publications, patent applications, patents, sequences, database entries, and other references mentioned herein are incorporated by reference in their entirety. In case of conflict, the present specification, including definitions, will control.

[0030] Other features and advantages of the invention will be apparent from the following detailed description and figures, and from the claims.BRIEF DESCRIPTION OF DRAWINGS

[0031] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing (s) will be provided by the Office upon request and payment of the necessary fee.

[0032] FIG. 1 shows results from in vitro PAM cleavage assays wherein eight enzymes displayed in vitro cleavage activity and a preference for different 5’ PAM sequences.

[0033] FIG. 2 shows results depicting reporter editing activities of the candidate Cas enzymes, wherein cells exhibit a green fluorescent signal upon successful editing by the Cas enzymes.

[0034] FIG. 3 shows results from assessments of enzymatic activities of the candidate Cas enzymes at specific endogenous sites in human cells.

[0035] FIG. 4 shows results from assessment of enzymatic activities of the candidate Cas-CBE enzymes (Cas-CBE-ART_HT_43) at specific inserted sites in human cells.

[0036] FIG. 5 shows results from Cas-CBE PAM assays wherein one enzyme displayed C->T mutation activity in cell and a preference for 5’ PAM sequences.

[0037] FIG. 6 shows results from assessment of enzymatic activities of the candidate Cas-CBE enzymes (Cas-CBE-ART_HT_41) at specific inserted sites in human cells.

[0038] FIG. 7 shows results from Cas-CBE PAM assays wherein one enzyme displayed C->T mutation activity in cell and a preference for 5’ PAM sequences.

[0039] FIG. 8 shows results from assessment of enzymatic activities of the candidate Cas-CBE enzymes (Cas-CBE-ART_HT_41 and Cas-CBE-ART_HT_43) at multiple endogenous sites in human cells.

[0040] FIG. 9 shows results from assessment of enzymatic activities of the mutated Cas-CBE enzymes (D900A for Cas-CBE-ART_S72, E1048A for Cas-CBE-ART_HT_46, and D968A for Cas-CBE-ART_HT_47) at multiple endogenous sites in human cells.

[0041] FIG. 10 shows results from assessment of epigenetic editing capabilities of the Cas-CRISPRoff-V2 systems in human cells.DETAILED DESCRIPTION

[0042] Disclosed herein are programmable DNA binding proteins (e.g., Cas12 enzymes) that, when present in a cell and in conjunction with binding to a guide polynucleotide (e.g., gRNA) , can specifically bind to a target polynucleotide sequence (i.e., via complementary base pairing between bases of the bound guide nucleic acid and bases of the target polynucleotide sequence) and catalyze (e.g., cleave) the target polynucleotide sequence.

[0043] Programmable DNA binding protein

[0044] As used herein, the term “programmable DNA binding protein” can refer to a protein or polypeptide that is capable of catalyzing (e.g., cleaving) internal regions in a nucleic acid (e.g., DNA or RNA) . In some embodiments, a programmable DNA binding protein can include an endonuclease.

[0045] A programmable DNA binding protein can itself comprise one or more domains. For example, a programmable DNA binding protein can comprise one or more nuclease domains. In some embodiments, a nuclease domain of a programmable DNA binding protein can comprise an endonuclease. In some embodiments, an endonuclease can cleave a single-stranded nucleic acid molecule. In some embodiments, an endonuclease can cleave a single strand of a double-stranded nucleic acid. In some embodiments, an endonuclease can cleave both strands of a double-stranded nucleic acid molecule. In some embodiments, a programmable DNA binding protein can comprise a deoxyribonuclease. In some embodiments, a programmable DNA binding protein can comprise a ribonuclease. In some embodiments, a nuclease domain of a programmable DNA binding protein can cut zero, one, or two strands of a target polynucleotide sequence.

[0046] In some embodiments, a nuclease domain of a programmable DNA binding protein is capable of cleaving only one strand of the two strands in a duplexed nucleic acid molecule (e.g., DNA) . In some embodiments, a nuclease domain of a programmable DNA binding protein can be derived from a fully catalytically active (e.g., natural) form of a programmable DNA binding protein. In some embodiments, a nuclease domain of a programmable DNA binding protein can be derived from a fully catalytically active (e.g., wild-type) form of a programmable DNA binding protein by introducing one or more mutations into the active the active, wild-type programmable DNA binding protein.

[0047] In some embodiments, a programmable DNA binding protein can include a Cas12 protein. In some embodiments, a programmable DNA binding protein comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identical to any one of SEQ ID NOs: 1-95 (Table 1) . In some embodiments, a programmable DNA binding protein comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identical to any one of SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 4, SEQ ID NO:7, SEQ ID NO: 48, SEQ ID NO: 50, SEQ ID NO: 53, SEQ ID NO: 54, SEQ ID NO: 55, or SEQ ID NO: 81. In some embodiments, a programmable DNA binding protein comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identical to any one of SEQ ID NO: 1, SEQ ID NO: 7, SEQ ID NO: 50 or SEQ ID NO: 53.

[0048] In some embodiments, a programmable DNA binding protein (e.g., a Cas12 enzyme) can specifically bind to a target polynucleotide sequence that is targeted for gene editing, and catalyze (e.g., cleave) the target polynucleotide sequence. In some embodiments, a programmable DNA binding protein can exhibit increased gene editing efficiency relative to a wild-type Cas12 protein (e.g., Cas12a, full sequence listed under the UniProtKB accession number U2UMQ6) . In some embodiments, a programmable DNA binding protein can exhibit increased gene editing efficiency relative to an LbCpf1 endonuclease (SEQ ID NO: 94) . In some embodiments, a programmable DNA binding protein exhibits a 5%, 10%, 20%, 30%, 40%, or 50%increased gene editing efficiency relative to an LbCpf1 endonuclease, wherein the LbCpf1 endonuclease comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identical to SEQ ID NO: 94.

[0049] Table 1 -Amino Acid Sequences of Programmable DNA Binding Proteins

[0050] Also described herein are nucleic acids comprising a sequence encoding a programmable DNA binding protein (e.g., endonuclease) , wherein the programmable DNA binding protein comprises a nucleic acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identical to any one of SEQ ID NOs: 96-105 (Table 2) . In some embodiments, the programmable DNA binding protein comprises a nucleic acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identical to any one of SEQ ID NO: 96, SEQ ID NO: 97, SEQ ID NO: 98, SEQ ID NO: 99, SEQ ID NO: 100, SEQ ID NO: 101, SEQ ID NO: 102, SEQ ID NO: 103, SEQ ID NO: 104, or SEQ ID NO: 105. In some embodiments, the programmable DNA binding protein comprises a nucleic acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identical to any one of SEQ ID NO: 96, SEQ ID NO: 99, SEQ ID NO: 100, SEQ ID NO: 104, or SEQ ID NO: 105. In some embodiments, the programmable DNA binding protein exhibits a 5%, 10%, 20%, 30%, 40%, or 50%increased gene editing efficiency relative to an LbCpf1 endonuclease, wherein the LbCpf1 endonuclease comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identical to SEQ ID NO: 94.

[0051] Table 2 -Nucleic Acid Sequences of Endonucleases

[0052] Gene-Editing Complex

[0053] Disclosed herein are gene-editing complexes that are capable of editing, modifying, or altering a target nucleotide of a polynucleotide. As used herein, the term “gene-editing complex” can refer to a complex of components required for recognizing and editing, modifying, or altering a nucleobase of a target polynucleotide sequence. In some embodiments, a gene-editing complex can include (i) a programmable DNA binding protein (e.g., a nuclease, e.g., an endonuclease) , and (ii) a DNA targeting sequence (e.g., guide polynucleotide, guide RNA or gRNA) . In some embodiments, components of the gene-editing complex can be associated with each other via covalent bonds, noncovalent interactions, or any combination of associations and interactions thereof. In some embodiments, a programmable DNA binding protein can be targeted to a target polynucleotide sequence by a DNA targeting sequence (e.g., guide polynucleotide, guide RNA or gRNA) . In some embodiments, the programmable DNA binding protein can be covalently bound to the DNA targeting sequence.

[0054] In some embodiments, a gene-editing complex can include CRISPR-Cas components. As used herein, the term “CRISPR” refers to a technique of sequence specific genetic manipulation relying on the clustered regularly interspaced short palindromic repeats pathway, which, unlike RNA interference, regulates gene expression at a transcriptional level. In some embodiments, a gene-editing complex can include an engineered, non-naturally occurring Clustered Regularly Interspaced Short Palindromic Repeat (CRISPR) -associated (Cas) system comprising (a) a guide RNA (gRNA) or a nucleic acid encoding the guide RNA, wherein the gRNA comprises a direct repeat sequence and a spacer sequence, and (b) a programmable DNA binding protein or a nucleic acid encoding the programmable DNA binding protein, wherein the programmable DNA binding protein binds to the gRNA, and wherein the spacer sequence is at least partially complementary to a target nucleic acid. In some embodiments, the spacer sequence is complementary to at least 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides of the target nucleic acid. In some embodiments, the spacer sequence can be between 10 to 30 (e.g., 15 to 30, 20 to 30, 25 to 30, 10 to 25, 15 to 25, 20 to 25, 10 to 20, 15 to 20, or 10 to 15) nucleotides in length.

[0055] Programmable DNA binding protein

[0056] In some embodiments, the programmable DNA binding protein comprises a Cas effector. In some embodiments, a programmable DNA binding protein can include an enzyme or protein that uses CRISPR sequences as a guide to recognize and cleave specific nucleic acid strands that are complementary to the CRISPR sequence. A programmable DNA binding protein can associate with a CRISPR sequence to bind to, and alter, DNA or RNA target sequences. In some embodiments, a programmable DNA binding protein can be a Cas12a nuclease that makes a double-stranded break in a target DNA sequence. In some embodiments, a programmable DNA binding protein is guided to a target nucleotide sequence by a DNA targeting sequence (e.g., guide polynucleotide, guide RNA or gRNA) , where the programmable DNA binding protein generates a single-strand DNA break at the specific target polynucleotide sequence (e.g., determined by the complementary sequence of a bound guide nucleic acid) . In some embodiments, the strand of a target polynucleotide sequence that is cleaved by a gene-editing complex comprising a programmable DNA binding protein (e.g., a nuclease, e.g., an endonuclease) is the strand that is not edited by the gene-editing complex (i.e., the strand that is cleaved by the programmable DNA binding protein is opposite to a strand comprising a base to be edited) . In other embodiments, a gene-editing complex comprising a programmable DNA binding protein (e.g., a nuclease, e.g., an endonuclease) can cleave the strand of a DNA molecule which is being targeted for editing. In such embodiments, the non-targeted strand is not cleaved.

[0057] In some embodiments, a gene-editing complex can include a programmable DNA binding protein which is catalytically dead (i.e., incapable of cleaving a target polynucleotide sequence) . As used herein, the terms "catalytically dead" and "nuclease dead" are used interchangeably to refer to a programmable DNA binding protein which has one or more mutations and / or deletions resulting in its inability to cleave a strand of a nucleic acid while retaining its ability, and specificity, to bind to a target polynucleotide. In some embodiments, a catalytically dead programmable DNA binding protein can lack nuclease activity as a result of specific point mutations in the one or more nuclease domains of the programmable DNA binding protein.

[0058] In some embodiments, a programmable DNA binding protein can comprise a type V CRISPR / Cas effector protein. Type V CRISPR / Cas effector proteins are a subtype of Class 2 CRISPR / Cas effector proteins. For examples of type V CRISPR / Cas systems and their effector proteins (e.g., Cas12 family proteins such as Cas12a) , see, e.g., Shmakov et al., Nat Rev Microbial. 2017 March; 15 (3) : 169-182: "Diversity and evolution of class 2 CRISPR-Cas systems. " Examples can include, but are not limited to, Cas12 family (Cas12a, Cas12b, Cas12c) , C2c4, C2c8, C2c5, C2c10, and C2c9; as well as CasX (Cas12e) and CasY (Cas12d) . Also see, e.g., Koonin et al., Curr Opin Microbial. 2017 June; 37: 67-78: 25 "Diversity, classification and evolution of CRISPR-Cas systems. "

[0059] For non-limiting examples of a programmable DNA binding protein, see, e.g., Makarova et al. "Classification and Nomenclature of CRISPR-Cas Systems: Where from Here? " CRISPR J. 2018 Oct; 1: 325-336. doi: 10.1089 / crispr. 2018.0033; Yan et al., "Functionally diverse type V CRISPR-Cas systems" Science. 2019 Jan 4; 363 (6422) : 88-91. doi: 10. l 126 / science. aav7271, the entire contents of each are hereby incorporated by reference. In some embodiments, examples of a programmable DNA binding protein can include, but are not limited to, Casl2a / Cpfl (UniProtKB: A0A7C9H0Z9) , Casl2b / C2cl (UniProtKB: T0D7A2) , Casl2c / C2c3 (UniProtKB: A0A9E2NRP4) , Casl2d / CasY (UniProtKB: A0A9E2QZL7) , Casl2e / CasX, Casl2g, Casl2h, and Casl2i. Non-limiting examples of Cas enzymes include Casl2a / Cpfl, Casl2b / C2cl, Casl2c / C2c3, Casl2d / CasY, Casl2e / CasX, Cas12f, Casl2g, Casl2h, Casl2i, Cas12j, Cas12l, Cas12m, Cas12n, Type V Cas effector proteins, an IscB protein, a TnpB protein, a Fanzor protein, or modified or engineered versions thereof. In some embodiments, a programmable DNA binding protein can comprise a Cas12a protein, wherein the programmable DNA binding protein has DNA binding ability and nuclease activity. In some embodiments, a programmable DNA binding protein can comprise a Cas12m protein, wherein the programmable DNA binding protein has DNA binding ability and RNA endonuclease activity. Other nucleic acid programmable DNA binding proteins are also within the scope of this disclosure, although they may not be specifically listed in this disclosure.

[0060] DNA Targeting Sequences

[0061] In some embodiments, a DNA targeting sequence can include a guide nucleic acid. In some embodiments, a DNA targeting sequence can include a guide RNA (gRNA) . As used herein, the term “gRNA” or “guide RNA” refers to a RNA sequence used to target specific genes for correction (e.g., gene editing) employing the CRISPR technique. Techniques of designing gRNAs and donor therapeutic polynucleotides for target specificity are well known in the art. For example, see, e.g., Doench, J., et al. Nature biotechnology 2014; 32 (12) : 1262-7 and Graham, D., et al. Genome Biol. 2015; 16: 260.

[0062] In some embodiments, a DNA targeting sequence includes at least one single guide RNA ( "sgRNA" or "gNRA" ) . In some embodiments, a DNA targeting sequence includes at least one tracrRNA. In some embodiments, a DNA targeting sequence requires a protospacer adjacent motif (PAM) sequence to guide a programmable DNA binding protein (e.g., a nuclease, e.g., an endonuclease) to the target polynucleotide sequence. In some embodiments, a DNA targeting sequence does not require protospacer adjacent motif (PAM) sequence to guide a programmable DNA binding protein (e.g., a nuclease, e.g., an endonuclease) to the target polynucleotide sequence. The programmable DNA binding protein (e.g., a nuclease, e.g., an endonuclease) of the gene-editing complexes disclosed herein can recognize a target polynucleotide sequence by associating with a DNA targeting sequence. A DNA targeting sequence (e.g., guide polynucleotide, guide RNA or gRNA) is typically single-stranded and can be programmed to site-specifically bind (i.e., via complementary base pairing) to a target sequence of a polynucleotide, thereby directing a gene-editing complex delivered in conjunction with a DNA targeting sequence to the target sequence. A DNA targeting sequence can be DNA. A DNA targeting sequence can be RNA. In some embodiments, a DNA targeting sequence comprises natural nucleotides (e.g., adenosine) . In some embodiments, a DNA targeting sequence comprises non-natural (or unnatural) nucleotides (e.g., peptide nucleic acid or nucleotide analogs) . In some embodiments, the targeting region of a DNA targeting sequence can be at least 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides in length. In some embodiments, a targeting region of a DNA targeting sequence can be between 10 to 30 (e.g., 15 to 30, 20 to 30, 25 to 30, 10 to 25, 15 to 25, 20 to 25, 10 to 20, 15 to 20, or 10 to 15) nucleotides in length. In some embodiments, a DNA targeting sequence (e.g., guide polynucleotide, guide RNA or gRNA) can include a sequence from SEQ ID NOs: 106-199 (Table 3) or SEQ ID NOs: 229-236 (Table 12-2) .

[0063] Table 3 -gRNA Scaffold Sequence (5’-3’)

[0064] Modifications of gRNAs

[0065] In some embodiments, the gRNA is chemically modified. A gRNA having one or more modified nucleosides or nucleotides is called a “modified” gRNA or “chemically modified” gRNA, to describe the presence of one or more non-naturally and / or naturally occurring components or configurations that are used instead of or in addition to the canonical A, G, C, and U residues. In some embodiments, a gRNA comprises a hybrid DNA-RNA guide, in which one or more DNA nucleotides replaces one or more RNA nucleotides in the polynucleotide sequence of the gRNA. In some embodiments, a modified gRNA is synthesized with a non-canonical nucleoside or nucleotide, is here called “modified. ” Modified nucleosides and nucleotides can include one or more of: (i) alteration, e.g., replacement, of one or both of the non-linking phosphate oxygens and / or of one or more of the linking phosphate oxygens in the phosphodiester backbone linkage (an exemplary backbone modification) ; (ii) alteration, e.g., replacement, of a constituent of the ribose sugar, e.g., of the 2' hydroxyl on the ribose sugar (an exemplary sugar modification) ; (iii) wholesale replacement of the phosphate moiety with “dephospho” linkers (an exemplary backbone modification) ; (iv) modification or replacement of a naturally occurring nucleobase, including with a non-canonical nucleobase (an exemplary base modification) ; (v) replacement or modification of the ribose-phosphate backbone (an exemplary backbone modification) ; (vi) modification of the 3' end or 5' end of the oligonucleotide, e.g., removal, modification or replacement of a terminal phosphate group or conjugation of a moiety, cap or linker (such 3' or 5' cap modifications may comprise a sugar and / or backbone modification) ; and (vii) modification or replacement of the sugar (an exemplary sugar modification) .

[0066] Chemical modifications such as those listed above can be combined to provide modified gRNAs having nucleosides and nucleotides (collectively “residues” ) that can have two, three, four, or more modifications. For example, a modified residue can have a modified sugar and / or a modified nucleobase. In some embodiments, every base of a gRNA is modified, e.g., all bases have a modified phosphate group, such as a phosphorothioate group. In certain embodiments, all, or substantially all, of the phosphate groups of an gRNA molecule are replaced with phosphorothioate groups. In some embodiments, modified gRNAs comprise at least one modified residue at or near the 5' end of the RNA. In some embodiments, modified gRNAs comprise at least one modified residue at or near the 3' end of the RNA.

[0067] In some embodiments, the gRNA comprises one, two, three or more modified residues. In some embodiments, at least 5% (e.g., at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100%) of the positions in a modified gRNA are modified nucleosides or nucleotides.

[0068] Unmodified nucleic acids can be prone to degradation by, e.g., intracellular nucleases or those found in serum. For example, nucleases can hydrolyze nucleic acid phosphodiester bonds. Accordingly, in one aspect the gRNAs described herein can contain one or more modified nucleosides or nucleotides, e.g., to introduce stability toward intracellular or serum-based nucleases. In some embodiments, the modified gRNA molecules described herein can exhibit a reduced innate immune response when introduced into a population of cells (e.g., in vivo and ex vivo) . The term “innate immune response” includes a cellular response to exogenous nucleic acids, including single stranded nucleic acids, which involves the induction of cytokine expression and release, particularly the interferons, and cell death.

[0069] In some embodiments of a backbone modification, the phosphate group of a modified residue can be modified by replacing one or more of the oxygens with a different substituent. Further, the modified residue, e.g., modified residue present in a modified nucleic acid, can include the wholesale replacement of an unmodified phosphate moiety with a modified phosphate group as described herein. In some embodiments, the backbone modification of the phosphate backbone can include alterations that result in either an uncharged linker or a charged linker with unsymmetrical charge distribution.

[0070] Examples of modified phosphate groups include phosphorothioate, phosphoroselenates, borano phosphates, borano phosphate esters, hydrogen phosphonates, phosphoroami dates, alkyl or aryl phosphonates and phosphotriesters. The phosphorous atom in an unmodified phosphate group is achiral. However, replacement of one of the non-bridging oxygens with one of the above atoms or groups of atoms can render the phosphorous atom chiral. The stereogenic phosphorous atom can possess either the “R” configuration (herein Rp) or the “S” configuration (herein Sp) . The backbone can also be modified by replacement of a bridging oxygen (i.e., the oxygen that links the phosphate to the nucleoside) with nitrogen (bridged phosphoroamidates) , sulfur (bridged phosphorothioates) and / or carbon (bridged methylenephosphonates) . The replacement can occur at either linking oxygen or at both of the linking oxygens.

[0071] The phosphate group can be replaced by non-phosphorus containing connectors in certain backbone modifications. In some embodiments, the charged phosphate group can be replaced by a neutral moiety. Examples of moieties which can replace the phosphate group can include, without limitation, e.g., methyl phosphonate, hydroxylamino, siloxane, carbonate, carboxy methyl, carbamate, amide, thioether, ethylene oxide linker, sulfonate, sulfonamide, thioformacetal, formacetal, oxime, methyleneimino, methylenemethylimino, methylenehydrazo, methylenedimethylhydrazo and methyleneoxymethylimino.

[0072] Scaffolds that can mimic nucleic acids can also be constructed wherein the phosphate linker and ribose sugar are replaced by nuclease resistant nucleoside or nucleotide surrogates. Such modifications may comprise backbone and sugar modifications. In some embodiments, the nucleobases can be tethered by a surrogate backbone. Examples include, without limitation, the morpholino, cyclobutyl, pyrrolidine and peptide nucleic acid (PNA) nucleoside surrogates.

[0073] The modified nucleosides and modified nucleotides can include one or more modifications to the sugar group, i.e., at sugar modification. For example, the 2' hydroxyl group (OH) can be modified, e.g., replaced with a number of different “oxy” or “deoxy” substituents. In some embodiments, modifications to the 2' hydroxyl group can enhance the stability of the nucleic acid since the hydroxyl can no longer be deprotonated to form a 2'-alkoxide ion.

[0074] Examples of 2' hydroxyl group modifications can include alkoxy or aryloxy (OR, wherein “R” can be, e.g., alkyl, cycloalkyl, aryl, aralkyl, heteroaryl or a sugar) ; polyethyleneglycols (PEG) , O (CH2CH2O) n CH2CH2OR wherein R can be, e.g., H or optionally substituted alkyl, and n can be an integer from 0 to 20 (e.g., from 0 to 4, from 0 to 8, from 0 to 10, from 0 to 16, from 1 to 4, from 1 to 8, from 1 to 10, from 1 to 16, from 1 to 20, from 2 to 4, from 2 to 8, from 2 to 10, from 2 to 16, from 2 to 20, from 4 to 8, from 4 to 10, from 4 to 16, and from 4 to 20) . In some embodiments, the 2' hydroxyl group modification can be 2'-O-Me. In some embodiments, the 2' hydroxyl group modification can be a 2'-fluoro modification, which replaces the 2' hydroxyl group with a fluoride. In some embodiments, the 2' hydroxyl group modification can include “locked” nucleic acids (LNA) in which the 2' hydroxyl can be connected, e.g., by a C1-6 alkylene or C1-6 heteroalkylene bridge, to the 4' carbon of the same ribose sugar, where exemplary bridges can include methylene, propylene, ether, or amino bridges; O-amino (wherein amino can be, e.g., NH2; alkylamino, dialkylamino, heterocyclyl, arylamino, diarylamino, heteroarylamino, or diheteroarylamino, ethylenediamine, or polyamino) and aminoalkoxy, O- (CH2) n-amino, (wherein amino can be, e.g., NH2; alkylamino, dialkylamino, heterocyclyl, arylamino, diarylamino, heteroarylamino, or diheteroarylamino, ethylenediamine, or polyamino) . In some embodiments, the 2' hydroxyl group modification can include “unlocked” nucleic acids (UNA) in which the ribose ring lacks the C2'-C3' bond. In some embodiments, the 2' hydroxyl group modification can include the methoxy ethyl group (MOE) , (OCH2CH2OCH3, e.g., a PEG derivative) .

[0075] “Deoxy” 2' modifications can include hydrogen (i.e. deoxyribose sugars, e.g., at the overhang portions of partially dsRNA) ; halo (e.g., bromo, chloro, fluoro, or iodo) ; amino (wherein amino can be, e.g., NEE; alkylamino, dialkylamino, heterocyclyl, arylamino, diarylamino, heteroarylamino, diheteroarylamino, or amino acid) ; NH (CH2CH2NH) nCH2CH2-amino (wherein amino can be, e.g., as described herein) , -NHC (0) R (wherein R can be, e.g., alkyl, cycloalkyl, aryl, aralkyl, heteroaryl or sugar) , cyano; mercapto; alkyl-thio-alkyl; thioalkoxy; and alkyl, cycloalkyl, aryl, alkenyl and alkynyl, which may be optionally substituted with, e.g., an amino as described herein.

[0076] The sugar modification can comprise a sugar group which may also contain one or more carbons that possess the opposite stereochemical configuration than that of the corresponding carbon in ribose. Thus, a modified nucleic acid can include nucleotides containing, e.g, arabinose, as the sugar. The modified nucleic acids can also include abasic sugars. These abasic sugars can also be further modified at one or more of the constituent sugar atoms. The modified nucleic acids can also include one or more sugars that are in the L form, e.g., L-nucleosides.

[0077] The modified nucleosides and modified nucleotides described herein, which can be incorporated into a modified nucleic acid, can include a modified base, also called a nucleobase. Examples of nucleobases include, but are not limited to, adenine (A) , guanine (G) , cytosine (C) , and uracil (U) . These nucleobases can be modified or wholly replaced to provide modified residues that can be incorporated into modified nucleic acids. The nucleobase of the nucleotide can be independently selected from a purine, a pyrimidine, a purine analog, or pyrimidine analog. In some embodiments, the nucleobase can include, for example, naturally-occurring and synthetic derivatives of a base.

[0078] In embodiments employing a dual guide RNA, each of the crRNA and the tracr RNA can contain modifications. Such modifications may be at one or both ends of the crRNA and / or tracr RNA. In embodiments having an sgRNA, one or more residues at one or both ends of the sgRNA may be chemically modified, or the entire sgRNA may be chemically modified. Certain embodiments comprise a 5' end modification. Certain embodiments comprise a 3' end modification. In certain embodiments, one or more or all of the nucleotides in single stranded overhang of a guide RNA molecule are deoxynucleotides.

[0079] In some embodiments, a gRNA can have one or more modifications. In some embodiments, the modification includes a 2'-O-methyl (2'-O-Me) modified nucleotide. In some embodiments, the modification includes a phosphorothioate (PS) bond between nucleotides.

[0080] The terms “mA, ” “mC, ” “mU, ” or “mG” may be used to denote a nucleotide that has been modified with 2’ -O-Me.

[0081] In some embodiments, the guide RNA comprises a gRNA scaffold sequence selected from any one of SEQ ID NOs: 106-199 shown in Table 3, or any one of SEQ ID NOs: 229-236 shown in Table 12-2, or any one of SEQ ID NOs: 268-287 shown in Table 17-1, Table 17-2 or Table 17-3. In some embodiments, the guide RNA comprises a spacer sequence selected from any one of SEQ ID NOs: 288-296 shown in Table 17-1, Table 17-2 or Table 17-3. In some embodiments, the guide RNA comprises one or more modifications selected from 2' O-methyl, 2' fluoro, phosphorothioate modification, or combination thereof. In some embodiments, the guide RNA comprises any one of SEQ ID NOs: 297-326.

[0082] Nuclear Localization Signal (NLS)

[0083] In some embodiments, a gene-editing complex can further include a nuclear localization signal (NLS) operably linked to the programmable DNA binding protein to localize the gene-editing complex to the nucleus of a cell. As used herein, a “nuclear localization signal” refers to a short stretch of amino acids that mediates the transport of cargo proteins into the nucleus. In some embodiments, a nuclear localization signal includes a high proportion of positively charged lysines or arginines exposed on the protein surface. In some embodiments, a nuclear localization signal includes about 7 to 20 (e.g., about 8 to 20, about 9 to 20, about 10 to 20, about 12 to 20, about 14 to 20, about 16 to 20, about 18 to 20, about 7 to 18, about 8 to 18, about 9 to 18, about 10 to 18, about 12 to 18, about 14 to 18, about 16 to 18, about 7 to 16, about 8 to 16, about 9 to 16, about 10 to 16, about 12 to 16, about 14 to 16, about 7 to 14, about 8 to 14, about 9 to 14, about 10 to 14, about 12 to 14, about 7 to 12, about 8 to 12, about 9 to 12, about 10 to 12, about 7 to 10, about 8 to 10, about 9 to 10, about 7 to 9, about 8 to 9, or about 7 to 8) amino acids. In some embodiments, the NLS is located at the N-terminus of a cargo protein. In some embodiments, the NLS is located at the C-terminus of a cargo protein.

[0084] Recombinant Expression Systems

[0085] Disclosed herein are recombinant expression systems for a gene-editing complex, wherein the gene-editing complex is capable of editing, modifying, or altering a target nucleotide of a polynucleotide. As used herein, a “recombinant expression system” refers to a system wherein a recombinant DNA is cloned into a vector introduced in a specific expression system (e.g., mammalian, bacteria, yeast, or insect cells) to support the expression of a gene of interest and the production of a recombinant protein. In some embodiments, a recombinant expression system for a gene-editing complex includes (i) a nucleic acid sequence encoding a programmable DNA binding protein, and (ii) a nucleic acid sequence encoding a DNA targeting sequence.

[0086] In some embodiments, a recombinant expression system for a gene-editing complex includes a nucleic acid sequence encoding a programmable DNA binding protein, wherein the programmable DNA binding protein comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identical to any one of SEQ ID NOs: 1-95. In some embodiments, the programmable DNA binding protein comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identical to any one of SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 4, SEQ ID NO: 7, SEQ ID NO: 48, SEQ ID NO: 50, SEQ ID NO: 53, SEQ ID NO: 54, SEQ ID NO: 55, or SEQ ID NO: 81. In some embodiments, the programmable DNA binding protein comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identical to any one of SEQ ID NO:1, SEQ ID NO: 7, SEQ ID NO: 50, or SEQ ID NO: 53.

[0087] In some embodiments, a recombinant expression system for a gene-editing complex includes a nucleic acid sequence encoding a DNA targeting sequence, wherein the DNA targeting sequence comprises a guide RNA (gRNA) . In some embodiments, the guide RNA comprises any one of SEQ ID NOs: 106-199 or 229-236. In some embodiments, wherein the guide RNA comprises one or more modifications selected from 2' O-methyl, 2' fluoro, or phosphorothioate modification, or combination thereof.

[0088] In some embodiments, a recombinant expression system for a gene-editing complex includes (i) a nucleic acid sequence encoding a programmable DNA binding protein, and (ii) a nucleic acid sequence encoding a DNA targeting sequence, wherein (i) and (ii) are comprised within a same vector. In some embodiments, a recombinant expression system for a gene-editing complex includes (i) a nucleic acid sequence encoding a programmable DNA binding protein, and (ii) a nucleic acid sequence encoding a DNA targeting sequence, wherein (i) and (ii) are comprised within different vectors. Non-limiting examples of vectors can include plasmids, transposons, cosmids, and viral vectors (e.g., any adenoviral vectors (e.g., pSV or pCMV vectors) , adeno-associated virus (AAV) vectors, lentivirus vectors, and retroviral vectors) , and any vectors. In some embodiments, a vector is a viral vector. In some embodiments, the viral vector is a lentiviral vector. In some embodiments, a vector can, e.g., include sufficient cis-acting elements for expression; other elements for expression can be supplied by the host mammalian cell or in an in vitro expression system. Skilled practitioners will be capable of selecting suitable vectors and mammalian cells for making any of the recombinant expression systems described herein.

[0089] Cells / Pharmaceutical Compositions

[0090] Also disclosed herein are cells that include any one of the programmable DNA binding proteins, any one of the vectors, any one of the gene-editing complexes, or any one of the recombinant expression systems described herein. In some embodiments, any one of the programmable DNA binding proteins, any one of the vectors, any one of the gene-editing complexes, or any one of the recombinant expression systems described herein can be introduced into a cell, e.g., a mammalian cell. Non-limiting examples of a mammalian cell can include, but are not limited to, a human cell, a rodent cell (e.g., a rat cell or a mouse cell) , a rabbit cell, a dog cell, a cat cell, a porcine cell, or a non-human primate cell. In some embodiments, any one of the programmable DNA binding proteins, any one of the vectors, any one of the gene-editing complexes, or any one of the recombinant expression systems described herein can be delivered into the cytoplasm of a cell. In some embodiments, any one of the programmable DNA binding proteins, any one of the vectors, any one of the gene-editing complexes, or any one of the recombinant expression systems described herein can be delivered into the cell by chemical transfection, non-chemical transfection, particle-based transfection, or viral transfection.

[0091] Also disclosed herein are pharmaceutical compositions comprising any one of the recombinant expression systems or any one of the cells described herein and a pharmaceutically acceptable carrier. The pharmaceutical compositions can be formulated in any way and can be administered in a variety of unit dosage forms depending upon the condition or disease and the degree of illness, the general medical condition of each patient, the resulting preferred method of administration and the like. Details on techniques for formulation and administration of pharmaceuticals are well described in the scientific and patent literature, see, e.g., Remington: The Science and Practice of Pharmacy, 21st ed., 2005.

[0092] Methods of Editing a Nucleobase

[0093] Disclosed herein are methods of editing a nucleobase of a target nucleic acid sequence, wherein a method includes contacting the target nucleic acid sequence with any one of the recombinant expression systems or any one of the gene-editing complexes described herein, wherein the DNA targeting sequence binds the target nucleic acid sequence. In some embodiments, the editing of the nucleobase of the target nucleic acid sequence comprises a DNA modification, wherein the DNA modification can include insertion, deletion, substitution, deletion-insertion, duplication, inversion, frameshift, repeat expansion, translocation, replacement, or combinations thereof, of the DNA. In some embodiments, the method can edit a single nucleobase of a target nucleic acid sequence. In some embodiments, the method can edit one or more nucleobases of a target nucleic acid sequence.

[0094] In some embodiments, the editing comprises a mutation of a single nucleobase of the target nucleic acid sequence, wherein the mutation comprises an insertion, a deletion, or a substitution. In some embodiments, the mutation is selected from the group consisting of an insertion, a deletion, a substitution, a deletion-insertion, a duplication, an inversion, a frameshift, a repeat expansion, a translocation, and combinations thereof. In some embodiments, the editing comprises a gene sequence insertion, a gene sequence deletion, or a gene sequence replacement.

[0095] In some embodiments, the editing comprises an epigenetic modification of the target nucleic acid sequence, wherein the epigenetic modification comprises methylation or de-methylation of the target nucleic acid sequence. In some embodiments, the epigenetic modification can comprise hydroxymethylation, histone modification, microRNA regulation, acetylation, phosphorylation, ubiquitination, or sumoylation. In some embodiments, the editing comprises an epigenetic modification of the target nucleic acid sequence, wherein the epigenetic modification comprises modulating a gene expression by activation or repression of the target nucleic acid sequence.

[0096] Also disclosed herein are methods of cleaving a target nucleic acid sequence that include contacting the target nucleic acid sequence with any one of the recombinant expression systems or any one of the engineered, non-naturally occurring CRISPR-Cas systems disclosed herein, wherein the DNA targeting sequence binds the target nucleic acid sequence, thereby cleaving the target nucleic acid sequence. In some embodiments, the target nucleic acid sequence is a single-stranded DNA or a double-stranded DNA. In some embodiments, the cleaving of the target nucleic acid sequence generates an insertion of a deletion of a nucleobase of the target nucleic acid sequence.

[0097] In some embodiments, the target nucleic acid sequence comprises a protospacer adjacent motif (PAM) sequence. In some embodiments, the target nucleic acid sequence comprises a PAM sequence selected from the group consisting of TTTS, TTTN, NTTN, NYTN, TTTV, RTYN, NNNN, TTTA, TTTC, NTTY and GTTC.

[0098] Also disclosed herein are methods of modulating gene expression of a target nucleic acid sequence, wherein a method includes contacting the target nucleic acid sequence with any one of the recombinant expression systems or any one of the gene-editing complexes described herein, wherein the DNA targeting sequence binds the target nucleic acid sequence. As used herein, “modulating” can refer to modifying, regulating, or altering gene expression of the target nucleic acid in a cell. In some embodiments, modulation of gene expression can include increasing translation of a target nucleic acid. In some embodiments, modulation of gene expression can include suppressing translation of a target nucleic acid. In some embodiments, gene expression of a target nucleic acid is upregulated. In some embodiments, gene expression of a target nucleic acid is downregulated. In some embodiments, the programmable DNA binding protein can be fused to a transcriptional activator protein, wherein the transcriptional activator protein upregulates gene expression of a target nucleic acid. As used herein, a “transcriptional activator protein” refers to a protein (e.g., a transcription factor) that increases transcription of one or more genes. In some embodiments, the programmable DNA binding protein can be fused to a transcriptional repressor protein, wherein the transcriptional repressor protein downregulates gene expression of a target nucleic acid. As used herein, a “transcription repressor protein” refers to a sequence-specific DNA binding protein that inhibits expression of one or more genes.

[0099] Also disclosed herein are methods of modifying histone at a target nucleic acid sequence, wherein a method includes contacting the target nucleic acid sequence with any one of the recombinant expression systems or any one of the gene-editing complexes described herein, wherein the DNA targeting sequence binds the target nucleic acid sequence. In some embodiments, the programmable DNA binding protein can be fused to a histone modifying protein, wherein the histone modifying protein modifies histone substrates at the target nucleic acid. In some embodiments, histone modification can include histone acetylation, phosphorylation or methylation. In some embodiments, a method of modifying histone at a target nucleic acid sequence can lead to modulating gene expression of the target nucleic acid.

[0100] In some embodiments, any one of the recombinant expression systems or any one of the gene-editing complexes described herein can be used for DNA / RNA detection, tracking and labeling of nucleic acids, enrichments assays, detecting circulating tumor DNA, preparing next generation library, drug screening, disease diagnosis and prognosis, and treating various genetic disorders. In some embodiments, any one of the recombinant expression systems or any one of the gene-editing complexes described herein can comprise or be associated with one or more functional domains (e.g., fusion protein, linker peptides) . In some embodiments, a functional domain can comprise methylase activity, demethylase activity, transcription activation activity, transcription repression activity, transcription release factor activity, histone modification activity, RNA cleavage activity, DNA cleavage activity, nucleic acid binding activity, and / or switch activity (e.g., light inducible) .

[0101] EXAMPLES

[0102] The compositions and methods disclosed herein are further described in the following examples, which do not limit the scope of the compositions and methods described in the claims. Example 1: Mining of Cas enzymes from genomic and metagenomic sources

[0103] A computational pipeline was used to expand a database of class 2 CRISPR-Cas systems from both genomic and metagenomic sources, following established principles (Koonin and Makarova, 2022; Makarova et al., 2020) . 92 promising Cas12 enzymes and their potential gRNA scaffold sequences were identified (Tables 1 and 3) , optimized for human expression (few of them shown in Table 2) , and cloned into mammalian expression vectors. Two widely used Cas proteins, type V LbCpf1 and type II SpCas9, were included as positive controls. Then, the PAM preferences and editing activities of these Cas enzymes were observed.

[0104] Example 2: PAM Determination for Cas enzymes

[0105] Targeted plasmid libraries (Table 5) were designed with a single unique spacer sequence (Table 4) , flanked on the 5' side by 7 base pairs of randomized sequences (5'-NNNNNNN-3') for PAM analysis. The Ribonucleoprotein (RNP) complexes for each Cas enzyme were derived from HEK293T cell lysates transfected with Cas enzyme (80 ng) and gRNA (40 ng) expression plasmids. This transfection process was conducted in a 96-well plate format (1.5E4 per well) using Fugene HD (commercially available from Promega) , following the manufacturer's instructions. The in vitro PAM cleavage assays were performed using prepared targeted plasmid libraries and cell lysates following procedures previously described (Russell et al., 2021) . Out of the 92 candidate enzymes, eight displayed clear in vitro cleavage activity and a preference for different 5’ PAM sequences (FIG. 1) .

[0106] Table 4 -Spacer Sequence

[0107] Example 3: Reporter editing activity characterization for Cas enzymes

[0108] To assess the editing activity of Cas enzymes in mammalian cells, co-transfections were conducted in HEK293T cells (1.5E4) by introducing Cas enzyme plasmids (70 ng) , gRNA plasmids (40 ng) , and reporter plasmids (35 ng) . The spacer sequence of gRNA was listed in Table 4. For each Cas enzyme, a reporter plasmid with corresponding PAM sequence was used (Table 6) . The transfection was carried out in a 96-well plate format (1.5E4 cells per well) using Fugene HD, following the manufacturer's instructions. After 48 hours, the cells were analyzed using FACS (Fluorescence-Activated Cell Sorting) . The reporter plasmid was designed to contain a disrupted GFP expression cassette, and a gRNA targeted site with preferred PAM sequence was inserted in the cassette. Once cleaved inside cells, there is a percentage of reporter plasmid that will get repaired and GFP expression will be restored. Therefore, it was anticipated that cells would exhibit a green fluorescent signal upon successful editing (Zhang et al., 2022) . The results depicting the reporter editing activities of the Cas enzymes described herein can be found in FIG. 2 and Table 7.

[0109] Table 5 -Plasmid Sequence for PAM Assay

[0110] Table 6 -Plasmid Sequence for Reporter activities

[0111] Table 7 -Reporter Editing Activity

[0112] Example 4: Endogenous editing activity characterization for Cas enzymes

[0113] Further assessments of the enzymatic activities of the Cas enzymes were conducted at specific endogenous sites in human cells. To do this, a set of gRNAs were designed targeting various human endogenous loci for each enzyme and their individual activities were evaluated (Target sequence of those sgRNAs shown in Table 8) . In each experimental condition, HEK293T cells (1.5E4) were transfected with 80 ng of the Cas enzyme plasmid and 40 ng of the gRNA plasmid using Fugene HD, following the manufacturer's instructions. After a 48-hour incubation period, genomic DNA was harvested from these cells and NGS-based amplicon sequencing was performed to analyze DNA editing efficiencies, as illustrated in FIG. 3 and Table 9.

[0114] Table 8 -Target Sequence of sgRNAs

[0115] Table 9 -Percentage of Editing at Endogenous Loci in Human Cell Lines

[0116] Note: NA: Not applicable.

[0117] Example 5: Characterizing DNA binding &unwinding activities of Cas enzymes by base editing based assays

[0118] To assess the potential DNA binding and unwinding activity of the Cas12 enzymes without detectable DNA cleavage activity, the Cas12 enzymes were put in cytosine base editor architecture (Cas-CBE) and their base editing activity was tested. Firstly, a stable HEK293T cell pool was established with a single unique spacer sequence (Table 4) inserted in the genome flanked on the 5' side by potential PAM sequence (TTTC) . The cells were then transfected with Cas-CBE enzyme (80 ng) and corresponding gRNA (40 ng) expression plasmids, following the provided instructions mentioned above. After a 72-hour incubation period, genomic DNA was harvested from these cells for NGS-based amplicon sequencing to analyze DNA editing efficiencies. Among all tested enzymes, Cas-CBE using ART_HT_43 as the Cas enzyme (Cas-CBE-ART_HT_43 DNA and Protein sequences listed in Table 1 and 2, separately) was shown to be active in introducing C->T mutations and results were depicted in FIG. 4.

[0119] Subsequently, a stable HEK293T cell pool was established with a single unique spacer sequence (Table 4) inserted in the genome flanked on the 5' side by seven base pairs of randomized sequences (5'-NNNNNNN-3') for Cas-CBE PAM analysis. Following transfection with Cas-CBE enzyme (80 ng) and corresponding gRNA (40 ng) expression plasmids, according to the provided instructions mentioned above, NGS-based amplicon sequencing was performed after a 72-hour incubation period. Reads exhibiting C>T mutation were extracted for PAM analysis, and the resultant PAM sequence was illustrated in FIG. 5.

[0120] Example 6: Characterizing DNA binding &unwinding activities of Cas enzymes by base editing based assays

[0121] Another enzyme, ART_HT_41, was assessed in the same way as described in Example 5 and was also shown to be effective in forms of Cas-CBE (Table 10) . It was shown to be active in introducing C->T mutations in transgenic HEK293T cell pools as depicted in FIG. 6, and its PAM sequence was illustrated in FIG. 7.

[0122] Example 7: Base editing activity characterization for Cas enzymes

[0123] The inventors further fused naturally occurring or mutated dead Cas enzymes to DNA deaminases to test their potential as base editors.

[0124] For ART_HT_41 and ART_HT_43, the inventors used their Cas-CBE forms and designed 5 gRNAs (Table 8) targeting various endogenous loci in human cells to test their activities. In each experimental condition, the inventors transfected HEK293T cells (1.5E4) with 80 ng of the Cas-CBE enzyme plasmid and 40 ng of the gRNA plasmid using Fugene HD, following the manufacturer's instructions. After a 72-hour incubation period, the inventors harvested genomic DNA from these cells and performed NGS-based amplicon sequencing to analyze DNA base editing efficiencies, as illustrated in FIG. 8.

[0125] For ART_S72, ART_HT_46, ART_HT_47, the inventors first mutated them to generate their corresponding dead Cas enzymes (D900A for ART_S72, E1048A for ART_HT_46 and D968A for ART_HT_47) , then used their dead versions to generate Cas-CBEs. Base editing activities of those Cas-CBEs were characterized as described above. For all Cas-CBEs, clear base editing activities were observed, as illustrated in FIG. 9.

[0126] Example 8: Epigenetic editing activity characterization for Cas enzymes

[0127] To assess the epigenetic editing capabilities of these novel Cas enzymes (ART_S72, ART_HT_46, ART_HT47) , the inventors engineered their inactive forms to fuse with DNMT3A, DNMT3L, and the KRAB domain, creating the new Cas-CRISPRoff-V2 systems (Table 10) . These systems were evaluated in a stable GFP reporter cell line where GFP expression is regulated by the methylation status of CpG islands in the snrpn promoter. The inventors used a specific snrpn promoter-targeting gRNA (spacer sequence: CAACGCAATGGAGCGAGGAA) with each Cas-CRISPRoff-V2 to assess their ability to suppress GFP expression. In each experiment, HEK293T cells (1.5E4) were transfected with 80 ng of the Cas-CRISPRoff-V2 enzyme plasmid and 40 ng of the gRNA plasmid using Fugene HD, according to the manufacturer's instructions. GFP expression was analyzed at various time points (day 3, day 7, and day 14) , as depicted in FIG. 10. Results clearly demonstrated that all three tested Cas enzymes could induce sustained gene repression through epigenetic editing.

[0128] Table 10 -Sequences of Cas-CBE and Cas-CRISPRoff-V2

[0129] Example 9: Rational Engineering of Cas enzyme ART_HT_46 and ART_S72

[0130] To enhance the activity of the native ART_HT_46 (SEQ ID NO: 53) and ART_S72 (SEQ ID NO: 7) Cas enzymes, the inventors employed structure-guided protein engineering. Initially, the inventors focused on the negatively charged residues (aspartate and glutamate) in both ART_HT_46 and ART_S72 Cas enzymes. A subset of these residues was systematically mutated to positively charged arginine (Table 11-1, Table 11-2) . It was hypothesized that this modification would increase the binding affinity of ART_HT_46 and ART_S72 for the negatively charged target DNA.

[0131] Next, the inventors targeted non-negatively charged residues and mutated them to arginine to enhance interactions between the Cas nucleases and the PAM duplex, as well as between the catalytic pocket and the ssDNA substrate (Table 11-3) .

[0132] To validate the targeting efficiency of the rationally designed ART_HT_46 and ART_S72 variants, further assessments were performed in human cells. For each ART_HT_46 variant, three gRNAs (gRNA-SSA, gRNA-TTR, gRNA-RUNX1, SEQ ID NO: 232 to SEQ ID NO: 234) , each containing a direct repeat (DR) (Table 12-2) , were designed to target specific human endogenous loci. HEK293T cells (96-well plate, 2 × 104 cells) were transfected with 88 ng of Cas enzyme plasmid and 24 ng of each gRNA plasmid using Fugene HD. After 72 hours of incubation, genomic DNA was harvested, and NGS-based amplicon sequencing was performed to assess DNA editing efficiencies (Table 11-1, Table 11-2, Table 11-3) .

[0133] Table 12-1 shows the representative sequence of plasmids used for directed evolution in Example 10. Table 12-2 shows the gRNA sequences, and Table 12-3 shows the encoding nucleotide sequence and NLS sequences.

[0134] Table 11-1 -Editing efficiency of ART_HT_46 variants in HEK293T cells

[0135] Table 11-2 -Editing efficiency of ART_S72 variants in HEK293T cells

[0136] Table 11-3 -Editing efficiency of ART_HT_46 variants in HEK293T cells

[0137] Table 12-1 -Sequence of plasmid

[0138] Table 12-2 -gRNA sequences

[0139] Table 12-3 -NLS sequences

[0140] Example 10: Directed Evolution of Cas enzyme ART_HT_46

[0141] The inventors used directed evolution in bacteria to isolate mutant proteins with improved targeting efficiency. A positive bacterial selection system was constructed to enrich improved variants. The system exploits Cas enzyme-mediated cleavage of a plasmid encoding the toxin-producing ccdB gene (SEQ ID NO: 228, see Table 12-1) to enrich for surviving colonies. The expression of ccdB is regulated by an inducible arabinose promoter (pBAD) and can be conditionally turned on. The expression of ART_HT_46 variants is controlled by a Tet-on promoter to control Cas enzyme expression period (SEQ ID NO: 227, see Table 12-1) . Two different strategies were employed to improve the selection pressures, including: (1) Using targeting gRNA with less efficient TCC or CCA PAM sequences; (2) Using shortened targeting gRNAs (17bp spacer sequence) (SEQ ID NO: 229 to SEQ ID NO: 231, see Table 12-2) . Both strategies ensured that only ART_HT_46 variants with higher targeting efficiency than the wild-type could be enriched in the positive bacterial selection system.

[0142] The inventors constructed a site saturation mutagenesis library of ART_HT_46 as directed evolution input. Approximately 300 ng of the ART_HT_46 library plasmids was electroporated into 50 μL of competent Top10 cells carrying a selection plasmid encoding both the arabinose-inducible ccdB toxin gene and the gRNA expression element (for example, SEQ ID NO: 228, see Table 12-2) . After a 1-hour recovery in 10 mL LB medium at 37℃, the bacterial culture was induced with 100 ng / μL anhydrotetracycline (aTc) for 1 hour. The culture was then plated on agar plates containing both arabinose and ampicillin. Positive colonies from the arabinose-containing plate were selected, and plasmids were prepared for sequencing after three rounds of selection (Table 13-1, Table 13-3 and Table 13-3) .

[0143] Top enriched variants were selected for validation in HEK293T cells. A stable HEK293T library was constructed, containing approximately thousands of gRNA expression elements and their corresponding target sites with five distinct PAM sequences (TTTN, TTN, TCN, CTN, CCN) . About 500ng of ART_HT_46 variants were transfected into HEK293T library (24-well plate, 1 × 105 cells per well) in form of plasmid (using Fugene HD) or mRNA (using Lipofectamine MessengerMAX) . Part of the lead variants were also evaluated in Huh-7 cells with two individual gRNAs (Table 12-2: gRNA-TTR, SEQ ID NO: 233, gRNA-DNMT1, SEQ ID NO: 235) in forms of mRNA at mRNA: gRNA=1: 1 weight ratio. These variants were transfected into Huh-7 (96-well plate, 8 × 103 cells per well) using Lipofectamine MessengerMAX. Genomic DNA was harvested from those samples after 72 hours followed by NGS analysis.

[0144] ART_HT_46 variants enriched from the pressure of a shortened gRNA were validated in HEK293T library in form of DNA plasmid. The results were shown in Table14-1.

[0145] ART_HT_46 variants selected from the pressure of gRNAs with TCC or CCA PAM sequences were evaluated in HEK293T library (the results were shown in Table14-2) and Huh-7 in form of mRNA (the results were shown in Table14-3) .

[0146] Two amino acids, D315 and D582 of ART_HT_46, were further mutated to other amino acids, including D315W, D315F, D315Y, D315A, D315G, D582W, D582F, D582Y, D582A, and these variants were evaluated in HEK293T library (the results were shown in Table14-4) and Huh-7 in form of mRNA (the results were shown in Table14-5) .

[0147] Table 13-1 -Mutation frequency fold-change under screening pressure of gRNA-361

[0148] Table 13-2 -Mutation frequency fold-change under screening pressure of gRNA-362

[0149] Table 13-3 -Mutation frequency fold-change under screening pressure of gRNA-17bp

[0150] Table 14-1 -Editing efficiency of ART_HT_46 variants in HEK293T gRNA library cells

[0151] Table 14-2 -Editing efficiency of ART_HT_46 variants in HEK293T gRNA library cells

[0152] Table 14-3 -Editing efficiency of ART_HT_46 variants in Huh-7 cells

[0153] Table 14-4 -Editing efficiency of ART_HT_46 variants in HEK293T gRNA library cells

[0154] Table 14-5 -Editing efficiency of ART_HT_46 variants in Huh-7 cells

[0155] Example 11: Investigating the Combined Effects of ART_HT_46 Enhancing Mutations

[0156] The inventors combined mutations that improved ART_HT_46 targeting efficiency in mammalian cells and tested them in the HEK293T library following procedures described above. The results were shown in Table 15-1. We further evaluated part of these combinatorial variants in Primary Human Hepatocytes (PHH) with gRNA-DNMT1 (SEQ ID NO: 235) . ART_HT_46 variants and gRNAs were mixed at 1: 1 weight ratio and transfected into PHH (48-well plate, 1.4 × 106 cells per well, two different doses: 62.5 ng mRNA / gRNA or 31.25 ng mRNA / gRNA) using Lipofectamine MessengerMAX. After 72 hours of incubation, genomic DNA was harvested, and NGS-based amplicon sequencing was performed to evaluate DNA editing efficiencies. The results were shown in Table 15-2, with the mean editing efficiencies for all gRNAs in the library used as the parameter to represent the enzyme's relative activity.

[0157] Table 15-1 -Editing efficiency of ART_HT_46 variants in HEK293T gRNA library cells

[0158] Table 15-2 -Editing efficiency of ART_HT_46 variants in PHH

[0159] Example 12: ART_HT_46-M4 Coding sequence and Nuclear Localization Signals (NLS) optimization

[0160] To optimize the expression or cellular localization for ART_HT_46_M4 (DNA sequence, SEQ ID NO: 254) , the inventors performed codon optimization (SEQ ID NO: 237) , or replaced the bpNLS (SEQ ID NO: 238) with either 2xNLS-cMyc-cMyc (SEQ ID NO: 239) or 3xNLS-NLP-cMyc-cMyc (SEQ ID NO: 240) at the C-terminus of the ART_HT_46_M4 protein. These variants were evaluated in PHH using the previously described methodology. The results were shown in Table 16, indicating that the NLS replacement and the codon optimization can significantly improve the editing efficiency of the Cas enzymes.

[0161] Table 16 -Editing efficiency of ART_HT_46-M4 codon optimization or NLS optimization

[0162] Example 13: Chemical Modification of gRNA to Improve Targeting Efficiency of ART_HT_46 and ART_S72

[0163] The inventors chemically modified the gRNA sequence of ART_HT_46 and ART_S72 to further improve its targeting efficiency. The inventors selected a gRNA (SEQ ID NO: 235, Table 12-2) and performed methoxy (2’ -Ome) scanning modifications on the ribonucleotides of the gRNA direct repeat (DR) region, as well as introduced 2′-fluoro-ribonucleotide modification on the gRNA spacer region and evaluate their effects on editing activity (SEQ ID NO: 297–SEQ ID NO: 312, see Table 17-1 and Table 17-2) .

[0164] The experiment used Lipofectamine MessengerMAX transfection reagent. We divided the mRNA: gRNA transfection doses into two gradients: 100 ng: 100 ng and 12.5 ng: 12.5 ng, and transfected them into Huh-7 cells (96-well plate, 8 × 103 cells per well) . After 72 hours of culture, genomic DNA was extracted from the lysed cells. PCR library construction and NGS sequencing were performed to verify the activity of ART_HT_46 and ART_S72 with different chemically modified gRNAs in cells (Table 17-1 and Table 17-2) .

[0165] Next, the inventors combined the most active modified types from the above experiment and added a chemically modified auxiliary sequence to the 5’ end of the gRNA (SEQ ID NO: 313–SEQ ID NO: 326, Table 17-3) to further improve the gRNA activity. The cell transfection and gRNA activity validation methods were as described above. The experimental results are shown in Table 17-3. These results showed that with chemical modifications (e.g., 2' O-methyl, 2' fluoro, phosphorothioate modification, or combination thereof) , the gRNA stability and the editing efficiency of Cas nuclease could be further improved.

[0166] As shown in Table 17-1, Table 17-2 and Table 17-3, the engineered gRNA sequence is constructed by directly linking a scaffold sequence (at 5’ -end) and a spacer sequence (at 3’ -end) . For example, the engineered gRNA of SEQ ID NO: 297 (i.e., 5’ -mU*A*AUUUCUACUAAGUGUAGAUCUGAUGGUCCAUGUCUGU*mU*mA*mC-3’ ) is constructed by directly linking the scaffold sequence of SEQ ID NO: 268 (at 5’ -end) to the spacer sequence of SEQ ID NO: 288 (at 3’ -end) .

[0167] Example 14: In Vivo Targeting Efficiency Evaluation of ART_HT_46 variants

[0168] To evaluate the targeting activity of ART_HT_46-M4 in the mouse liver, the inventors used gRNA (SEQ ID NO: 326, Table 17-3) targeting the mouse Ttr gene, with an mRNA-to-gRNA ratio of 1: 1 (w / w) , and formulated them into LNP as described in a previous patent WO / 2023 / 185697A2. LNPs were used for intravenous injections at doses of 2, 1, 0.75, 0.5, and 0.25mg / kg. Seven days later, genomic DNA was extracted from the mouse liver, followed by targeted amplicon library preparation and NGS sequencing. The experimental results were shown in Table 18, indicating that these variants can improve the editing efficiency in a dose-dependent manner in vivo, and can be used as desired Cas enzymes.

[0169] Table 18. In vivo evaluation of ART_HT_46 variants for editing efficiency

[0170] Those skilled in the art will further appreciate that the present invention may be embodied in other specific forms without departing from the spirit or central attributes thereof. In that the foregoing description of the present invention discloses only exemplary embodiments thereof, it is to be understood that other variations are contemplated as being within the scope of the present invention. Accordingly, the present invention is not limited to the particular embodiments that have been described in detail herein. Rather, reference should be made to the appended claims as indicative of the scope and content of the invention.

Claims

1.A programmable DNA binding protein comprising an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identical to any one of SEQ ID NOs: 1-92 or 241-253.2.The programmable DNA binding protein of claim 1, wherein the programmable DNA binding protein comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identical to any one of SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 4, SEQ ID NO: 7, SEQ ID NO: 48, SEQ ID NO: 50, SEQ ID NO: 53, SEQ ID NO: 54,SEQ ID NO: 55, or SEQ ID NO: 81.3.The programmable DNA binding protein of claim 2, wherein the programmable DNA binding protein comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identical to any one of SEQ ID NO: 1, SEQ ID NO: 7, or SEQ ID NO: 50, SEQ ID NO: 53.4.The programmable DNA binding protein of any one of claims 1-3, wherein the programmable DNA binding protein exhibits a 5%, 10%, 20%, 30%, 40%, or 50%increased gene editing efficiency relative to an LbCpf1 endonuclease, wherein the LbCpf1 endonuclease comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identical to SEQ ID NO: 94.5.A nucleic acid comprising a sequence encoding a programmable DNA binding protein, wherein the programmable DNA binding protein comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identical to any one of SEQ ID NOs: 1-92 or 241-253.6.The nucleic acid of claim 5, wherein the nucleic acid comprises a nucleic acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identical to any one of SEQ ID NOs: 96-105, 237 or 254-267.7.The nucleic acid of claim 6, wherein the nucleic acid comprises a nucleic acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identical to any one of SEQ ID NO: 237, SEQ ID NO: 254, SEQ ID NO: 100, SEQ ID NO: 99, SEQ ID NO: 97, SEQ ID NO: 98, SEQ ID NO: 96, SEQ ID NOs: 101-104, or SEQ ID NO: 267.8.The nucleic acid of any one of claims 4-7, wherein the programmable DNA binding protein exhibits a 5%, 10%, 20%, 30%, 40%, or 50%increased gene editing efficiency relative to an LbCpf1 endonuclease, wherein the LbCpf1 endonuclease comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identical to SEQ ID NO: 94.9.An engineered, non-naturally occurring Clustered Regularly Interspaced Short Palindromic Repeat (CRISPR) -associated (Cas) system comprising:(a) a guide RNA (gRNA) or a nucleic acid encoding the guide RNA, wherein the gRNA comprises a direct repeat sequence and a spacer sequence; and(b) a programmable DNA binding protein or a nucleic acid encoding the programmable DNA binding protein, wherein the programmable DNA binding protein comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identical to any one of SEQ ID NOs: 1-92 or 241-253,wherein the programmable DNA binding protein binds to the gRNA, and wherein the spacer sequence is at least partially complementary to a target nucleic acid.10.The engineered, non-naturally occurring CRISPR-Cas system of claim 9, wherein the programmable DNA binding protein comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identical to any one of SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 4, SEQ ID NO: 7, SEQ ID NO: 48, SEQ ID NO: 50, SEQ ID NO: 53, SEQ ID NO: 54, SEQ ID NO: 55, or SEQ ID NO: 81.11.The engineered, non-naturally occurring CRISPR-Cas system of claim 10, wherein the programmable DNA binding protein comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identical to any one of SEQ ID NO: 1, SEQ ID NO: 7, SEQ ID NO: 50, or SEQ ID NO: 53.12.The engineered, non-naturally occurring CRISPR-Cas system of claim 9, wherein the guide RNA comprises any one of SEQ ID NOs: 106-199 or 229-236, optionally, wherein the guide RNA comprises one or more modifications selected from 2' O-methyl, 2' fluoro, phosphorothioate modification, or combination thereof.13.The engineered, non-naturally occurring CRISPR-Cas system of any one of claims 9-12, wherein the spacer sequence is complementary to at least 15 nucleotides of the target nucleic acid, optionally, wherein the spacer sequence comprises one or more modifications selected from 2' O-methyl, 2' fluoro, phosphorothioate modification, or combination thereof.14.The engineered, non-naturally occurring CRISPR-Cas system of any one of claims 9-13, further comprising a nuclear localization signal (NLS) operably linked to the endonuclease, wherein the gene-editing complex is thereby localized to the nucleus of a cell.15.The engineered, non-naturally occurring CRISPR-Cas system of claim 14, wherein the NLS is an N-terminal NLS and / or a C-terminal NLS, preferably, wherein the NLS comprises an amino acid sequence of any one of SEQ ID NOs: 238-240.16.A recombinant expression system for an engineered, non-naturally occurring CRISPR-Cas system comprising:(i) a nucleic acid sequence encoding a programmable DNA binding protein, and(ii) a nucleic acid sequence encoding a DNA targeting sequence,wherein the programmable DNA binding protein comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identical to any one of SEQ ID NOs: 1-92 or 241-253.17.The recombinant expression system of claim 16, wherein the DNA targeting sequence comprises a guide RNA (gRNA) .18.The recombinant expression system of claim 17, wherein the guide RNA comprises any one of SEQ ID NOs: 106-199 or 229-236 or 297-326.19.The recombinant expression system of any one of claims 16-18, wherein (i) and (ii) are comprised within a same vector or comprised within different vectors.20.A cell comprising a programmable DNA binding protein of any one of claims 1-4, a nucleic acid of any one of claims 5-8, an engineered, non-naturally occurring CRISPR-Cas system of any one of claims 9-15, or a recombinant expression system of any one of claims 16-19.21.A pharmaceutical composition comprising a recombinant expression system of any one of claims 16-19 or the cell of claim 20 and a pharmaceutically acceptable carrier.22.A method of editing a nucleobase of a target nucleic acid sequence, the method comprising contacting the target nucleic acid sequence with a recombinant expression system of any one of claims 16-19 or an engineered, non-naturally occurring CRISPR-Cas system of any one of claims 9-15, wherein the DNA targeting sequence binds the target nucleic acid sequence, thereby editing the nucleobase of the target nucleic acid sequence.23.The method of claim 22, wherein the editing comprises a mutation of a single nucleobase of the target nucleic acid sequence, wherein the mutation comprises an insertion, a deletion, or a substitution.24.The method of claim 22, wherein the editing comprises a gene sequence insertion, a gene sequence deletion, or a gene sequence replacement.25.The method of claim 22, wherein the editing comprises an epigenetic modification of the target nucleic acid sequence, wherein the epigenetic modification comprises methylation or de-methylation of the target nucleic acid sequence.26.The method of claim 22, wherein the editing comprises an epigenetic modification of the target nucleic acid sequence, wherein the epigenetic modification modulates a gene expression by activation or repression of the target nucleic acid sequence.27.A method of cleaving a target nucleic acid sequence, the method comprising contacting the target nucleic acid sequence with a recombinant expression system of any one of claims 16-19 or an engineered, non-naturally occurring CRISPR-Cas system of any one of claims 9-15, wherein (i) the DNA targeting sequence binds the target nucleic acid sequence and (ii) the programmable DNA binding protein cleaves the target nucleic acid sequence.28.The method of claim 27, wherein the target nucleic acid sequence is a single-stranded DNA or a double-stranded DNA.29.The method of claim 27, wherein the cleaving of the target nucleic acid sequence generates an insertion of a deletion of a nucleobase of the target nucleic acid sequence.30.The method of any one of claims 22-29, wherein the target nucleic acid sequence comprises a protospacer adjacent motif (PAM) sequence.31.A method of modulating gene expression of a target nucleic acid sequence, the method comprising contacting the target nucleic acid sequence with a recombinant expression system of any one of claims 16-19 or an engineered, non-naturally occurring CRISPR-Cas system of any one of claims 9-15, wherein the DNA targeting sequence binds the target nucleic acid sequence, thereby modulating gene expression of the target nucleic acid sequence.32.The method of claim 31, wherein the gene expression of the target nucleic acid sequence is upregulated or downregulated.33.A method of modifying histone at a target nucleic acid sequence, the method comprising contacting the target nucleic acid sequence with a recombinant expression system of any one of claims 16-19 or an engineered, non-naturally occurring CRISPR-Cas system of any one of claims 9-15, wherein the DNA targeting sequence binds the target nucleic acid sequence, thereby modifying histone at the target nucleic acid sequence.34.The method of claim 33, wherein the histone modification comprises histone acetylation, phosphorylation, or methylation.35.An engineered guide RNA comprising:a) a DNA-targeting segment comprising a nucleotide sequence that is complementary to a target sequence in a target DNA molecule; andb) a protein-binding segment configured to bind to a programmable DNA binding protein comprising an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identical to any one of SEQ ID NOs: 1-92 or 241-253;wherein the engineered guide RNA comprises a nucleotide modification pattern selected from 2' O-methyl, 2' fluoro, phosphorothioate modification, or combination thereof.36.The engineered guide RNA of claim 35, wherein the modifications are present in scaffold sequence and / or spacer sequence, optionally, wherein the scaffold sequence comprises any one of SEQ ID NOs: 268-287.37.The engineered guide RNA of claim 35, wherein the engineered guide RNA comprises any one of SEQ ID NOs: 297-326.38.A kit, comprising:(a) a programmable DNA binding protein of any one of claims 1-4, or a nucleic acid of any one of claims 5-8, or an engineered, non-naturally occurring CRISPR-Cas system of any one of claims 9-15, or a recombinant expression system of any one of claims 16-19, or the cell of claim 20, or an engineered guide RNA of any one of claims 35-37; and(b) an instruction.

Citation Information

Patent Citations

  • Novel crispr-cas systems for genome editing

    CN113166744A

  • Novel CRISPR-Cas12i system

    CN114015674A

  • Novel, non-naturally occurring crispr-cas nucleases for genome editing

    EP3943600A1

  • Cas12a mutant genes and polypeptides encoded by same

    WO2020146297A1