Highly efficient RNA-aptamer recruitment-mediated DNA base editors and their use for targeted genome modification

The CasRCure system addresses the limitations of conventional gene editing by using RNA-aptamer-mediated base editing to achieve precise and efficient genome modifications in various cells without DNA breaks, offering a safer and more effective therapeutic approach.

JP7814052B2Active Publication Date: 2026-02-16RUTGERS THE STATE UNIV

Patent Information

Application Number
JP2022517222
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-09-17
Filing Date
2020-09-16
Publication Date
2026-02-16
Estimated Expiration
2040-09-16

AI Technical Summary

Technical Problem

Conventional gene editing technologies like ZFNs, TALENs, and CRISPR systems face limitations such as off-target effects, DNA double-strand breaks, and low efficiency in somatic cells, especially for precise modifications like point mutations, due to reliance on DSBs and error-prone NHEJ pathways.

Method used

A novel RNA-aptamer-mediated base editing system, known as CasRCure (CRC), uses a nuclease-deficient Cas9 protein fused with a recruiting RNA motif to recruit cytidine or adenine deaminases, enabling precise base modifications without DSBs or HDR, with a modular design allowing flexible effector recruitment.

Benefits of technology

The CRC system achieves high precision and specificity in genome editing, reducing off-target effects and enhancing efficacy in both prokaryotic and eukaryotic cells, including mammalian cells, with improved activity windows and reduced collateral damage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007814052000056
    Figure 0007814052000056
  • Figure 0007814052000057
    Figure 0007814052000057
  • Figure 0007814052000058
    Figure 0007814052000058
Patent Text Reader

Abstract

The present invention discloses a system for targeted gene editing and related uses. Also disclosed are related cells.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 62 / 901,584, filed September 17, 2019, the disclosure of which is incorporated herein by reference.

[0002] Technical Field The present invention relates to a system for targeted genome modification and uses thereof. [Background technology]

[0003] Gene editing technologies, such as zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), or clustered regularly interspaced short palindromic repeats (CRISPR) systems, provide powerful tools for biotechnology and biomedical research in general. These technologies have also raised hopes for the systematic development of targeted therapies for genetic diseases, cancer, and viral infections. However, gene editing technologies have significant limitations that must be addressed before they can be widely used in clinical settings. First, conventional gene editing systems rely on generating DNA double-strand breaks (DSBs) at the target site, which can potentially have deleterious consequences, especially if unintended off-target activity is high (1, 2). The development of strategies such as paired nickases (3), catalytically inactive Cas9 fused to dimeric nucleases (4, 5), or high-fidelity CRISPR systems (6, 7) is believed to mitigate such adverse effects; however, the actual deleterious effects of gene editing interventions may be underestimated due to limitations in detection methods for accurately assessing on-target and off-target mutagenesis. Recently, DSBs generated by CRISPR systems have been shown to induce previously unnoticed deletions and rearrangements spanning several kilobases at on-target sites (8). Similarly, insertional mutagenesis has been observed in experiments using purified Cas9 / sgRNA ribonucleoprotein complexes (RNPs), a method believed to enhance target specificity (9). Second, the introduction of precise modifications, such as point mutations, often requires the target cells to undergo homology-dependent DNA double-strand break repair (HDR) (10, 11). However, somatic cells, especially terminally differentiated somatic cells, do not have high HDR activity and instead use the error-prone non-homologous end joining (NHEJ) pathway.(12) These findings highlight the need for new gene editing systems to develop safe and effective therapies. Summary of the Invention

[0004] The present invention, in no small aspect, addresses the above-referenced needs.

[0005] In one aspect, the present invention provides a system comprising: (i) a sequence-targeting component or a polynucleotide encoding the same; (ii) an RNA scaffold or a polynucleotide encoding the same (e.g., DNA); and (iii) a first effector fusion protein or a polynucleotide encoding the same. The sequence-targeting component comprises (a) a sequence-targeting protein and (b) a targeting fusion protein having a first uracil-DNA glycosylase (UNG) inhibitor peptide (UGI). The RNA scaffold comprises (a) a nucleic acid-targeting motif comprising a guide RNA sequence complementary to a target nucleic acid sequence, (b) an RNA motif capable of binding to the sequence-targeting protein (e.g., a CRISPR motif described herein), and (c) a first recruiting RNA motif. The first effector fusion protein comprises (a) a first RNA-binding domain capable of binding to the first recruiting RNA motif, (b) a linker, and (c) an effector domain. The first effector fusion protein or effector domain has an enzymatic activity, such as a cytosine deaminating activity or an adenosine deaminating activity. In one embodiment, the exemplary system is referred to as a Cas-RNA aptamer-mediated C to U reversion (CasRCure or CRC) system. Additional exemplary systems include the second-generation CRC system CRC_AID ( A CRCnu, A CRCnu.2) and CRC_APOBEC1( A1 CRCnu., A1 CRCnu.2) (where u indicates the presence of UGI in the system).

[0006] In the system of the present invention, the target fusion protein may contain one, two, or more UGIs. The RNA scaffold may contain one, two, or more recruitment RNA motifs. Thus, the target fusion protein may further contain two or more UGIs (e.g., a second UGI). The RNA scaffold may further contain two or more recruitment RNA motifs (e.g., a second recruitment RNA motif). Preferably, one, more, or all coding sequences are codon-optimized. For example, one or more of the polynucleotides encoding the sequence-targeting protein, including the first UGI, the second UGI, the RNA-binding domain, and the effector domain, are optimized for expression in eukaryotic cells (e.g., plant cells, insect cells, or mammalian cells). Each of the sequence-targeting component and the second effector fusion protein may have a nuclear localization signal (NLS). For example, the sequence-targeting component or the first effector fusion protein contains one or more NLSs. In one embodiment, the sequence targeting component comprises two NLSs, which may be located at the N-terminus and C-terminus of the sequence targeting component, respectively, as shown in Figure 9C.

[0007] In the above system, the sequence-targeting protein may be a CRISPR protein. The sequence-targeting protein does not have nuclease activity. Examples of sequence-targeting proteins include dCas9 or nCas9 sequences from a species selected from the group consisting of Streptococcus pyogenes, Streptococcus agalactiae, Staphylococcus aureus, Streptococcus thermophilus, Streptococcus thermophilus, Neisseria meningitidis, and Treponema denticola.

[0008] In the above-mentioned RNA scaffolds, the first recruitment RNA motif and the first RNA-binding domain may be a pair selected from the group consisting of: (1) a telomerase Ku-binding motif and a Ku protein or its RNA-binding segment, (2) a telomerase Sm7-binding motif and a Sm7 protein or its RNA-binding segment, (3) an MS2 phage operator stem-loop and an MS2 coat protein (MCP) or its RNA-binding segment, (4) a PP7 phage operator stem-loop and a PP7 coat protein (PCP) or its RNA-binding segment, (5) an SfMu phage Com stem-loop and a Com RNA-binding protein or its RNA-binding segment, (6) chemically modified versions of the above-mentioned aptamers and their corresponding aptamer ligands or their RNA-binding segment, and (7) a non-natural RNA aptamer and a corresponding aptamer ligand or its RNA-binding segment.

[0009] The effector fusion protein may have a variety of suitable enzymatic activities. In one embodiment, the effector may have cytidine deamination activity, such as wild-type or genetically modified AID, CDA, APOBEC1, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D, APOBEC3F, or other APOBEC family enzymes from a species selected from the group consisting of humans, rats, mice, bats, naked mole rats, elephants, chickens, lizards, giant turtles, coelacanths, and other vertebrate species. In another embodiment, the effector may have adenine deamination activity, such as wild-type or genetically modified ADA, ADAR family enzymes, or tRNA adenosine deaminase from a species selected from the group consisting of bacteria, yeast, humans, rats, mice, bats, naked mole rats, elephants, chickens, lizards, giant turtles, coelacanths, and other vertebrate species. The linker sequence may be 0 to 100 (eg, 1 to 100, 5 to 80, 10 to 50, and 20 to 30) amino acid residues in length.

[0010] Also provided is an isolated nucleic acid encoding one or more of components (i)-(iii) of the system described above, an expression vector comprising the nucleic acid, or a host cell comprising the nucleic acid.

[0011] In a second aspect, the present invention provides a method for site-specifically modifying target DNA. The method comprises contacting a target nucleic acid with components (i) to (iii) of the system described above. The target nucleic acid may be present in a cell. The target nucleic acid may be RNA, extrachromosomal DNA, or genomic DNA located in a chromosome. The cell may be selected from the group consisting of an archaeal cell, a bacterial cell, a eukaryotic cell, a eukaryotic unicellular organism, a somatic cell, a germ cell, a stem cell, a plant cell, an algae cell, an animal cell, an invertebrate cell, a vertebrate cell, a fish cell, a frog cell, an avian cell, a mammalian cell, a porcine cell, a bovine cell, a caprine cell, a sheep cell, a rodent cell, a rat cell, a mouse cell, a horse cell, a non-human primate cell, and a human cell. The cell may be present in or derived from a human or non-human subject. The human or non-human subject may have a genetic mutation in a gene. In some embodiments, the subject has or is at risk of having a disorder caused by the genetic mutation. In that case, the site-specific modification corrects a genetic mutation or inactivates expression of a gene. In other embodiments, the subject has or is at risk of being exposed to a pathogen, and the site-specific modification inactivates a gene of the pathogen.

[0012] The present invention also provides genetically modified cells obtained by the method described above. The cells can be selected from the group consisting of stem cells, immune cells, and lymphocytes. Examples of stem cells include embryonic stem cells, ES-like stem cells, fetal stem cells, adult stem cells, pluripotent stem cells, induced pluripotent stem cells, multipotent stem cells, oligopotent stem cells, unipotent stem cells, and others described herein. Examples of immune cells include T cells, B cells, NK cells, macrophages, mixtures thereof, and others described herein. Also provided is a pharmaceutical composition comprising an effective amount of cells and a pharmaceutically acceptable carrier.

[0013] The present invention further provides a kit comprising the above-described system or one or more components thereof, which may further comprise one or more components selected from the group consisting of reagents for renaturation and / or dilution, and reagents for introducing nucleic acids or polypeptides into host cells.

[0014] The details of one or more embodiments of the invention are set forth herein below. Other features, objects, and advantages of the invention will become apparent from the specification and claims. [Brief explanation of the drawings]

[0015] [Figure 1-1] A series of diagrams show the CRC system and proof-of-principle demonstration in prokaryotic cells. A. Components of the CRC platform, from left to right: 1. sequence-targeting component dCas9 or nCas9D10A; 2. chimeric RNA scaffold containing a guide RNA motif (for sequence targeting; 2.1), a CRISPR motif (for Cas9 binding; 2.2), and a recruiting RNA aptamer motif (for recruiting the effector-RNA-binding protein fusion; 2.3); and 3. a fusion protein consisting of an effector cytidine deaminase (3.1) and an RNA aptamer protein ligand (3.2). B. Schematic of the CRC complex at the target sequence: Cas9 binds to the CRISPR RNA, and the recruiting RNA aptamer recruits the effector module, forming an active CRC complex capable of editing the target C residue (shaded) in the unpaired DNA within the CRISPR R-loop. The PAM sequence is underlined. C. RRDR cluster I region of the rpoB gene of E. coli (SEQ ID NO:2 (nucleic acid sequence) and SEQ ID NO:3 (corresponding amino acid sequence)). The PAM sequence is underlined and the critical cytosine is shaded in a gray box. The arrow represents the gRNA targeting site. The gray shading is the RRDR protein sequence. [Figure 1-2]A series of diagrams showing the CRC system and proof-of-principle in prokaryotic cells. D. Representative photographs showing surviving bacterial colonies after treatment with CRCs targeted with the indicated gRNAs expressing one MS2 copy (1xMS2). E. Quantification of the percentage of surviving cells in a similar experiment shown in D. Bars indicate the standard deviation of the mean of three independent experiments. F. Representative sequencing results of untreated cells (top row, SEQ ID NO: 4) and ACRCd-treated cells with rpoB_TS4_1xMS2 gRNA (bottom row, SEQ ID NO: 5). The target position is indicated by a black star. This C1592>T mutation results in a S531F change in the protein sequence, a mutation known to induce rifampicin resistance (23, 24). [Figure 2] This set of diagrams shows the genetic engineering of CRC modules to enhance base editing efficiency in bacterial cells. A. The effect of replacing Cas9 nickase (nCas9H840A or nCas9D10A) with dCas9 and increasing the number of recruitment motifs from 1xMS2 to 2xMS2. B. The effect of varying the linker length of the effector module. L4, L5, L10, L12, and L25 are linker peptides consisting of 4, 5, 10, 12, and 25 amino acids, respectively. C. Comparison of AID (ACRCD10A), APOEC3G (A3GCRCD10A), and APOBEC1 (A1CRCD10A) as effectors. The figure shows a representative result of three independent experiments. [Figure 3]A set of diagrams and photographs show the effect of CRC on correction of targeted mutations and global mutagenesis in human cells. A. Non-fluorescent EGFP (nfEGFP) target region (SEQ ID NO: 6). A loss-of-function A-to-G mutation in the chromophore sequence (underlined in black). One gRNA targeting the non-template strand (NT1) is indicated by an arrow, the PAM sequence is underlined, and the target cytosine is shaded in gray. The corresponding protein sequence (SEQ ID NO: 7) is shaded in gray. B. Effect on editing of extrachromosomal genes. HEK 293 cells were transiently transfected with target DNA containing nfEGFP mutants together with ACRCnu, BE4max, or BE3 components and nfEGFP_NT1 gRNA. Panels show representative plate sections under a fluorescent microscope after the indicated treatments. C. Flow cytometry analysis of cells expressing an extrachromosomal nfEGFP gene targeted by nfEGFP_NT1 and treated with ACRCnu, BE4max, and BE3. D. Flow cytometry analysis of HEK293 cells (nf2.16 cells) stably expressing a non-fluorescent EGFP mutant gene guided by nfEGFP_NT1 gRNA and treated with ACRCnu, BE4max, and BE3. E. Sequencing of sorted fluorescent cells. *G to A conversion of the targeted nucleotide (top row: SEQ ID NO: 8; bottom row: SEQ ID NO: 9). Note that base editing occurs on the complementary strand. F. Whole-exome sequencing and SNP comparison of nf2.16 cells treated with ACRCnu / nfEGFP_NT-1, ACRCnu / scramble, or untreated. Genomic DNA was isolated and subjected to whole-exome sequencing. This figure shows the global distribution of single nucleotide polymorphisms in the three treatments compared to the human reference genome (hg38), including the AID signature mutations C->T / G->A. Statistical analysis showed no significant differences across all SNP categories. G. Comparison of the occurrence of C>T and G>A events in "AID motif" sequences (WRCH / DGYW; dark gray bars) compared to "non-motif" sequences (NNCN / NGNN; light gray bars). Mutations at CpG sites were not counted to avoid overestimation due to the high mutation rate at these sites.p-values ​​were calculated using a chi-squared test. NT1: nfEGFP_NT1 gRNA (NT = non-template strand targeted). Error bars represent the standard deviation of the mean of three independent experiments. All gRNAs used for CRC treatment express two MS2 aptamers for effector recruitment. [Figure 4] A set of diagrams showing that the CRC system efficiently edits endogenous sites in the human genome (SEQ ID NOS: 10-15). HEK293 cells were treated with ACRCnu or A1CRCnu and the indicated gRNAs. A-C. Quantification of single-base mutations induced by ACRCnu at the indicated loci. D-F. Quantification of single-base mutations induced by A1CRCnu at the indicated loci. Treatments were analyzed by high-throughput sequencing to quantify the frequency of mutations induced by the systems tested in this set of experiments. The gRNA target sequence is shaded gray. All gRNAs used in this experiment express two MS2 aptamers for effector recruitment. [Figure 5] A set of diagrams showing that optimizing CRC constructs leads to enhanced base editing efficiency. Cells were treated with the indicated base editing systems targeting site 2 (SEQ ID NOs: 16, 18, and 20). High-throughput sequencing analysis reveals that efficiency was enhanced after targeting site 2 with ACRCnu.2 (A) and A1CRCnu.2 (C), reaching an efficiency comparable to that of BE4max (E). Cells were treated with the corresponding systems with scrambled gRNAs (B, D, F; SEQ ID NOs: 17, 19, and 21). Target sequences are shaded gray. All gRNAs used for CRC treatment express two MS2 aptamers for effector recruitment. [Figure 6-1]Figure 1 shows a set of diagrams and photographs demonstrating that CRC mediates efficient knockout of the GFP reporter and endogenous locus in human cells. A. Schematic of the EGFP region (SEQ ID NO: 22) targeted in these experiments. One gRNA (arrow) was designed to induce a stop codon at residue Q157 (EGFP_TS1). The PAM sequence is underlined. The corresponding protein sequence (SEQ ID NO: 23) is shown in gray shading. B. HEK293 cells expressing an EGFP transgene were treated with ACRCnu.2 and EGFP_TS1. The panels show representative plate sections under a fluorescent microscope. C. Cells from a similar experiment shown in B were subjected to flow cytometry analysis to quantify GFP loss. Error bars represent the standard deviation of the mean of at least three independent experiments. [Figure 6-2] A set of diagrams and photographs showing that CRC mediates efficient knockout of the GFP reporter and endogenous loci in human cells. D-E. High-throughput sequencing analysis of EGFP reporter cells treated with ACRCnu.2 and EGFP_TS1 (D) (SEQ ID NO: 24) or left untreated (E) (SEQ ID NO: 25). F. Schematic of the endogenous PDCD1 locus region (SEQ ID NO: 26) targeted in these experiments. One gRNA (arrow) was designed to induce a stop codon at residue Q133 (PDCD1_TS1); PAM sequence. The corresponding protein sequence (SEQ ID NO: 27) is shown in gray shading. G-H. High-throughput sequencing analysis of the endogenous PDCD1 locus treated with ACRCnu.2 and PDCD1_TS1 gRNA (G) (SEQ ID NO: 28) or left untreated (H) (SEQ ID NO: 29). TS: targeting the template strand. All gRNAs used in this experiment express two MS2 aptamers for effector recruitment. [Figure 7]A set of diagrams showing bacterial expression constructs. A-C. Schematic diagrams of the constructs used in bacterial experiments, including a DNA targeting module encoding the Cas9 variants dCas9, nCas9D10A, or nCas9H840A (A; element (1) in Figure 1A); a gRNA / recruitment module containing one or two RNA aptamer motifs (B, top and bottom rows, respectively; element (2) in Figure 1A); and an effector module encoding the fusion proteins AID_MCP, APOBEC1_MCP, or APOBEC3G_MCP (C; element (3) in Figure 1A). [Figure 8] Figure 1 shows a set of diagrams depicting mutation distribution in E. coli cells targeted to the rpoB gene sequence (SEQ ID NO: 30). Mutation distribution of clones selected on rifampicin plates after treatment. All experiments use TS4 gRNA for comparison. A. Side-by-side comparison of editing outcomes after treatment with CRC systems with different Cas9 variants (i.e., ACRCd with dCas9, ACRCH840A with nCas9H840A, and ACRCD10A with nCas9D10A). B. Side-by-side comparison of editing outcomes after treatment with CRC systems with different effector proteins (i.e., A1CRCD10A with APOBEC1 and A3GCRCD10A with APOBEC3G). RpoB genes from individual clones were PCR-amplified and sequenced for genotyping. Numbers represent the percentage of clones with a given genotype. [Figure 9]A set of diagrams showing mammalian expression constructs. A. Schematic of the first-generation ACRCnu polycistronic construct expressing the AID_L25_MCP fusion protein and nCas9D10A-UGI. The two modules are separated by a self-cleaving 2A, and their expression is driven by a CMV promoter. B. The gRNA_2×MS2 construct is expressed from a U6 promoter. C. The second-generation ACRCnu.2 system follows a similar architecture to the first generation but exhibits key differences: codon optimization, enhanced nuclear localization of the Cas9-UGI module, and increased UGI copy number. NLS: nuclear localization signal; effector: AID, APOBEC1; L25: 25-amino acid flexible linker; 2A: self-cleaving 2A peptide. [Figure 10] A set of diagrams showing the frequency of indel formation after treatment with ACRCnu and A1CRCnu targeting sites 2, 3, and 4. Histograms showing the indel analysis of the experiment are shown in Figure 4, with the CRC systems and targeting gRNAs indicated. A-C show indels induced by ACRCnu targeting sites 2, 3, and 4. D-F show indels induced by A1CRCnu targeting the same sites. The gRNA target sites are indicated by black lines. Note that indels tend to accumulate at high frequency at the gRNA target sites. [Figure 11-1] A set of diagrams showing high-throughput sequencing analysis of selected off-target sites (homologous sites) after ACRCnu and ACRCnu.2 treatment targeting site 2, site 3, or site 4. Analysis of known S. pyogenes Cas9 off-target sites (31, 32) for site 2: S2O2; site 3: S3O1, S3O2, and S3O3; and site 4: S4O1, S4O2, and S4O4 (SEQ ID NOs: 31-36). Off-target sequences are summarized in Table S5. [Figure 11-2]A set of diagrams showing high-throughput sequencing analysis of selected off-target sites (homologous sites) after ACRCnu and ACRCnu.2 treatment targeting site 2, site 3, or site 4. Analysis of known S. pyogenes Cas9 off-target sites (31, 32) for site 2: S2O2; site 3: S3O1, S3O2, and S3O3; and site 4: S4O1, S4O2, and S4O4 (SEQ ID NOs: 31-36). Off-target sequences are summarized in Table S5. [Figure 12] FIG. 1 is a set of diagrams showing the frequency and distribution of indel formation after treatment with ACRCnu.2, A1CRCnu.2, or BE4max targeting site 2. The histograms quantitate the indel frequency for the experiment shown in FIG. 5. Cells were treated with the indicated systems and gRNAs and subjected to high-throughput sequencing. The gRNA target sequence is indicated by a black line. [Figure 13] A set of diagrams showing high-throughput sequencing analysis of ACRCnu.2 targeting site 3 and site 4. HEK293 cells were treated with ACRCnu.2 and the indicated gRNAs targeting site 3 (A) (SEQ ID NO: 37) and site 4 (C) (SEQ ID NO: 38). Untreated counterparts are shown in B for site 3 and in D for site 4. Samples were then analyzed by high-throughput sequencing to quantify the frequency of mutations induced by the system. Target sequences are shaded gray. All gRNAs used in this experiment express two MS2 aptamers for effector recruitment. [Figure 14]

[0039] Figure 13A is a set of diagrams of the frequency and distribution of indel formation after treatment with ACRCnu.2 at sites 3 and 4. The histograms quantify the indel frequency for the experiments shown in Figures 13A-13D. Cells were treated with the indicated systems and gRNAs and subjected to high-throughput sequencing. The gRNA target sites are shown as black lines. [Figure 15]Figure 5 shows a set of diagrams depicting the frequency of indel formation after treatment with ACRCnu.2 targeting the EGFP transgene. The histograms represent the indel analysis of the experiment shown in Figures 5A-5F. In the experiment, ACRCnu.2 targeted EGFP using gRNA TS1 (A). The untreated counterpart is shown in B. The gRNA target sequence is indicated by a black line. [Figure 16] (A) Single nucleotide polymorphisms (SNPs) spanning the region of Site 2 (SEQ ID NO: 39) targeted by Site 2 gRNA and second-generation rat A1CRCnu.2; (B) SNPs spanning the region of Site 2 targeted by Site 2 gRNA and second-generation lizard (Anolis carolinensis) LizardA1CRCnu.2; (C) SNPs spanning the region of Site 2 targeted by Site 2 gRNA and second-generation bat (Myotis lucifugus) BatA1CRCnu.2; and (D) a set of diagrams showing SNPs spanning the region of Site 2 in untreated cells. [Figure 17] 1 is a diagram showing a comparison of C to T conversion rates at the human fetal hemoglobin promoter locus (HBF) (SEQ ID NO: 40) in K562 cells by the LizardA1CRCnu.2 (lizard Apobec1), RatA1CRCnu.2 (rat Apobec1), BE4max (BE4), and LizardA1CRCnu.2 (lizard AID) systems. The PAM motif is AGG at the 3' end. [Figure 18] 1 is a diagram showing a comparison of C to T conversion rates at the site 2 locus (SEQ ID NO: 41) in HEK293 cells by the LizardA1CRCnu.2 (labeled lizard Apobec1) and Rat A1CRCnu.2 (labeled rat Apobec1) systems. The PAM motif is GGG at the 3' end. [Figure 19]1 is a diagram showing a comparison of C to T conversion rates at the site 3 locus (SEQ ID NO: 42) in HEK293 cells by the LizardACRCnu.2 (labeled Lizard AID) and human ACRCnu.2 (labeled Human AID) systems. The PAM motif is TGG at the 3' end. [Figure 20] 1 is a diagram showing a comparison of C to T conversion rates at the site 3 locus (SEQ ID NO: 43) in HEK293 cells by the BatACRCnu.2 (labeled bat AID) and human ACRCnu.2 (labeled human AID) systems. [Figure 21] This diagram shows the conversion of a C to a T using a catalytically dead Cas9 (dCas9) version of the ACRCnu.2 construct at site 2, which contains two target Cs (C1 and C2) within the editing window. All experiments were performed using the ACRCnu.2 version of the base editing system, including both the original nCas9 version (ACRCnu.2) and the derived dCas9 version (ACRCnu.2_dCas9). As a control, the experiment included an sgRNA lacking the aptamer component of the system (ACRCnu.2_dCas9_MS2less), as the absence of the MS2 element in the system would result in a loss of editing due to a failure to recruit the deaminase upon fusion with the MCP. A scrambled non-targeting sgRNA (ACRCnu.2_dCas9_scrambled) was also included as a negative control. Data are presented as the percentage of sequenced Ts at the indicated target C residues as determined by Sanger sequencing. Error bars represent the standard deviation of the mean of three replicate experiments. DETAILED DESCRIPTION OF THE INVENTION

[0016] The present invention relates to a novel system for targeted genome modification and uses thereof. The present invention is based, at least in part, on a novel RNA-aptamer-mediated base editing system.

[0017] Conventional nuclease-dependent precise genome editing typically requires the introduction of DNA double-strand breaks (DSBs) and the activation of the homology-dependent repair (HDR) pathway. However, DSBs often have the disadvantage of being oncogenic, and HDR activity is low in somatic cells. Recently, a base editing (BE) system has been developed in which a cytidine (or adenine) deaminase effector is recruited to the target DNA sequence by direct fusion with a nuclease-deficient Cas9 protein. BE alters the target base pair without the need for DSBs or HDR.

[0018] An alternative base editing system with a modular design has also been developed. This system recruits an effector deaminase via the RNA component of the CRISPR complex. This system, named CasRCure (CRC), includes a modified gRNA with a reprogrammable RNA aptamer at its 3' end that recruits a cognate aptamer ligand fused to an effector (e.g., a deaminase effector). Using this system, targeted nucleotide modifications have been achieved with high precision in prokaryotic and eukaryotic cells, including mammalian cells. See International Publication Nos. 2018129129 and 2017011721. As disclosed herein, a new, second-generation CRC base editor CRC system with increased efficacy has been tested and further improved in mammalian cells. The second-generation CRC base editor includes one or all of the following features: First, the Cas9 protein contains one, two, or more than two UGIs. Second, the Cas9-UGI protein has at least two nuclear localization signal peptides (NLSs). Third, both the Cas9-UGI and effector proteins are codon-optimized for expression in target host cells (e.g., mammalian cells). This second-generation system / platform exhibits higher efficacy and specificity than previously disclosed first-generation CRC systems. Importantly, various effector orthologs from different species were constructed in second-generation CRC configurations. Surprisingly, some second-generation CRCs with certain orthologs, such as the lizard ortholog, exhibit unique features that differ from all previously documented base editors. For example, they have a broader activity window that allows modification of nucleotides closer to the PAM motif than the standard activity window of positions 3–9. Due to its modular design, which completely separates the nucleic acid modification module from the nucleic acid recognition module, as well as other advantages disclosed herein, the CRC base editing platform offers an alternative to recruiting effectors by fusion or direct interaction with sequence-targeting proteins, which could not effectively separate the sequence-targeting function from the nucleic acid-modifying function.Lacking the requirement for DNA DSBs and HDR, the new CRC system provides a powerful tool for genetic engineering and therapy development.

[0019] Gene Editing Platform One aspect of the present invention provides a gene editing platform that overcomes the above-mentioned limitations of conventional nuclease- and DSB-dependent genome engineering and gene editing technologies. This platform has three functional components: (1) a nuclease-deficient CRISPR / Cas-based module engineered for sequence targeting; (2) an RNA scaffold-based module for guiding the platform to the target sequence and recruiting the correction module; and (3) a non-nuclease DNA / RNA-modifying enzyme as an effector correction module, such as cytidine deaminase (e.g., activation-induced cytidine deaminase, AID). At the same time, the CasRcure system enables the anchoring of specific DNA / RNA sequences, the flexible and modular recruitment of effector DNA / RNA-modifying enzymes to specific sequences, and the triggering of somatically active cellular pathways to correct genetic information, especially point mutations.

[0020] 1A and 1B show a schematic diagram of an exemplary CasRcure system. This system includes three structural and functional components: (1) a sequence-targeting module (e.g., dCas9 protein); (2) an RNA scaffold for sequence recognition and effector recruitment (a chimeric RNA molecule containing a guide RNA (gRNA) motif, a CRISPR RNA motif, and a recruitment RNA motif); and (3) an effector (a non-nuclease DNA-modifying enzyme such as AID fused to a small protein that binds to the recruitment RNA motif). More specifically, as shown in Figure 1A, the components of the CRC platform include a sequence-targeting component 1 (dCas9 or nCas9D); 10AThe CRC complex comprises: a chimeric RNA scaffold 2 containing guide RNA motif 2.1 (for sequence targeting), CRISPR motif 2.2 (for Cas9 binding), and recruiting RNA aptamer motif 2.3 (for effector-RNA binding protein fusion); and a fusion protein 3 containing effector 3.1 (e.g., cytidine deaminase) fused to RNA aptamer ligand 3.2. Figure 1B shows a schematic diagram of the CRC complex at a target sequence, in which Cas9 binds to the CRISPR RNA and the recruiting RNA aptamer recruits the effector module, forming an active CRC complex capable of editing the target C residue in the unpaired DNA within the CRISPR R-loop, also known as the protospacer. The three components can be assembled into a single expression vector or multiple separate expression vectors. The overall and combination of these three specific components constitutes the realization of this technology platform. Although the three components of the RNA scaffold are shown in a specific 3' to 5' order in Figure 1B, the components may be arranged in different orders as needed, such as for optimization of different Cas protein variants.

[0021] As disclosed herein, there are significant differences between the recruitment mechanisms of the RNA scaffold-mediated recruitment system (CRC system) and the direct fusion system (BE system) between Cas9 and effector proteins. The modular design of the CRC system allows for flexible system engineering. Modules are interchangeable, and numerous combinations of different modules can be achieved simply by swapping the nucleotide sequences of the recruiting RNA aptamer and cognate ligand. On the other hand, recruitment of effectors via direct fusion or direct interaction of sequence-targeting units with protein components always requires re-engineering of new fusion proteins, which is technically more challenging and the outcome less predictable. Furthermore, RNA scaffold-mediated recruitment likely promotes oligomerization of effector proteins, whereas direct fusion precludes oligomer formation due to steric hindrance.

[0022] CRISPR / Cas-based gene systems are becoming dominant in the therapeutic landscape due to their relative ease of use and scalability, making them an attractive gene editing technology for developing novel applications with therapeutic value. As disclosed herein, a second-generation CRISPR base editor system utilizes certain aspects of the CRISPR / Cas system. To overcome limitations associated with the need for DSBs and HDR in conventional CRISPR / Cas gene editing systems, a sophisticated gene editing method called base editing (BE) has been developed that utilizes the DNA targeting ability of Cas9, which lacks nuclease activity, combined with the DNA editing ability of APOBEC-1, a member of the APOBEC family of enzymes in DNA / RNA cytidine deaminases (13). By directly fusing a deaminase effector to the nuclease-deficient Cas9 protein, these tools, called base editors, can introduce targeted point mutations into genomic DNA (13) or RNA (14) without generating DSBs or requiring HDR activity. Essentially, the BE system uses a nuclease-deficient CRISPR / Cas9 complex as the DNA targeting mechanism, and mutant Cas9 serves as an anchor to recruit cytidine or adenine deaminase by direct protein-protein fusion.

[0023] In contrast, the CRC system takes a different approach. More specifically, in the CRC system, the RNA component of the CRISPR / Cas9 complex serves as an anchor for effector recruitment because the RNA aptamer is contained in the RNA molecule. The RNA aptamer then recruits an effector fused to an RNA aptamer ligand. Compared with recruitment via direct protein fusion or other recruitment methods using protein components, the RNA aptamer-mediated effector recruitment mechanism has several unique features that are potentially advantageous for both system genetic engineering and achieving better functionality. For example, the RNA aptamer-mediated effector recruitment mechanism has a modular design in which the nucleic acid sequence targeting function and the effector function reside on different molecules, allowing for independent reprogramming of functional modules and multiplexing of the system. Reprogramming the CRC system simply requires changing the RNA aptamer sequence of the gRNA and replacing the cognate RNA aptamer-ligand-fused effector. Re-engineering of individual functional Cas9 fusion proteins is not required. Additionally, the smaller size of the fusion effectors potentially allows for more efficient oligomerization of functional effectors. Furthermore, because CRC does not require the generation of Cas9 fusion proteins, which would further increase the Cas9 gene / transcript size, the CRC system may be engineered to allow for more efficient packaging and delivery by viral vectors.

[0024] As disclosed herein, the present invention provides further engineering of a second-generation CRC system for high-precision base editing. As shown herein, the second-generation CRC system exhibits several important distinct features compared to the previous CRC systems (first generation) described in International Publication Nos. 2018129129 and 2017011721. The second-generation CRC system exhibits significantly increased on-target efficacy compared to first-generation CRCs. The inventors selected and optimized the configuration of second-generation CRCs with higher efficacy, fewer or no off-target effects, and higher purity (more C to T conversions rather than C to other nucleotides). Importantly, second-generation CRC systems using a wide variety of cytidine deaminases from different species and different deaminase families have been tested, and many of them exhibit distinctly different activity windows and preferred positions, as well as higher activity, than any previously described base editing systems, including BE systems. See, for example, Figures 16-20.

[0025] a. Sequence targeting module The sequence-targeting component of the system is based on the CRISPR / Cas system of bacterial species. The original functional bacterial CRISPR-Cas system requires three components: a Cas protein that provides nuclease activity, and two short non-coding RNA species called CRISPR RNA (crRNA) and trans-acting RNA (tracrRNA). These two RNA species form the so-called guide RNA (gRNA). Type II CRISPR is one of the best-characterized systems and performs targeted DNA double-strand breaks in four sequential steps. First, two non-coding RNAs, pre-crRNA and tracrRNA, are transcribed from the CRISPR locus. Second, tracrRNA hybridizes to the repeat region of the pre-crRNA molecule and mediates its processing, resulting in a mature crRNA molecule containing individual spacer sequences. Third, the mature crRNA:tracrRNA complex (i.e., the so-called guide RNA) guides a Cas nuclease (e.g., Cas9) to the target DNA through Watson-Crick base pairing between the spacer sequence in the crRNA and the complement of the protospacer sequence in the target DNA, which contains a three-nucleotide (nt) protospacer adjacent motif (PAM). The PAM sequence is essential for Cas9 targeting. Finally, the Cas nuclease mediates cleavage of the target DNA to create a double-strand break within the target site. In its natural context, the CRISPR / Cas system functions as an adaptive immune system that protects bacteria from repeated viral infections, the PAM sequence serves as a self / nonself recognition signal, and the Cas9 protein possesses nuclease activity. The CRISPR / Cas system has shown great potential for gene editing both in vitro and in vivo.

[0026] In the invention disclosed herein, a sequence recognition mechanism can be achieved in a similar manner: mutant Cas proteins, such as dCas9 proteins containing mutations in their nuclease catalytic domains and thus lacking nuclease activity, or nCas9 proteins with one catalytic domain partially mutated and thus lacking nuclease activity to generate DSBs, specifically recognize non-coding RNA scaffold molecules containing short, typically 20-nucleotide-long, spacer sequences that guide the Cas protein to its target DNA or RNA sequence. The latter are flanked by 3' PAMs.

[0027] Cas proteins Various Cas proteins can be used in the present invention. Cas protein, CRISPR-associated protein, or CRISPR protein are used interchangeably and refer to proteins of or derived from CRISPR-Cas type I, type II, or type III systems with RNA-guided DNA binding. Non-limiting examples of suitable CRISPR / Cas proteins include: Cas3, Cas4, Cas5, Cas5e (or CasD), Cas6, Cas6e, Cas6f, Cas7, Cas8a1, Cas8a2, Cas8b, Cas8c, Cas9, Cas10, Cas10d, CasF, CasG, CasH, Csy1, Csy2, Csy3, Cse1 (or CasA), Cse2 (or Ca sB), Cse3 (or CasE), Cse4 (or CasC), Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csz1, Csx15, Csf1, Csf2, Csf3, Csf4, and Cu1966. See, e.g., WO 2014144761, WO 2014144592, WO 2013176772, U.S. Patent Application Publication No. 20140273226, and U.S. Patent Application Publication No. 20140273233, the contents of which are incorporated herein by reference in their entireties.

[0028] In one embodiment, the Cas protein is derived from a type II CRISPR-Cas system. In an exemplary embodiment, the Cas protein is or is derived from a Cas9 protein. The Cas9 protein may be derived from the following: Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus species, Nocardiopsis dassonvillei, Streptomyces pristinaespiralis, Streptomyces viridochromogenes, Streptomyces viridochromogenes, Streptosporangium roseum, Streptosporangium roseum, Alicyclobacillus acidocaldarius, Bacillus pseudomycoides, Bacillus selenitidis, Bacillus selenitireducens, Exiguobacterium sibiricum, Lactobacillus delbrueckii, Lactobacillus salivarius, Microscilla marina, Burkholderiales bacterium, Polaromonas naphthalenivorans, Polaromonas sp., Crocosphaera watsonii, Cyanothece sp., Microcystis aeruginosa, Synechococcus sp.), Acetohalobium arabaticum, Ammonifex degensii, Caldicelulosiruptor becscii, Candidatus Desulforudis, Clostridium botulinum, Clostridium difficile, Finegoldia magna, Natranaerobius thermophilus, Pelotomaculum thermopropionicum, Acidithiobacillus caldas caldus, Acidithiobacillus ferrooxidans, Allochromatium vinosum, Marinobacter sp., Nitrosococcus halophilus, Nitrosococcus watsoni, Pseudoalteromonas haloplanktis, Ktedonobacter racemifer, Methanohalobium evestigatum, Anabaena variabilis, Nodularia spumigena, Nostoc sp. sp.), Arthrospira maxima, Arthrospira platensis, Arthrospira sp., Lyngbya sp.), Microcoleus chthonoplastes, Oscillatoria species, Petrotoga mobilis, Thermosipho africanus, or Acaryochloris marina.

[0029] Generally, Cas proteins contain at least one RNA-binding domain. The RNA-binding domain interacts with a guide RNA. Cas proteins may be wild-type Cas proteins or modified versions lacking nuclease activity. Cas proteins may be modified to increase nucleic acid binding affinity and / or specificity, alter enzymatic activity, and / or change other properties of the protein. For example, the nuclease (i.e., DNase, RNase) domain of the protein may be modified, deleted, or inactivated. Alternatively, proteins may be truncated to remove domains not essential for protein function. Proteins may also be truncated or modified to optimize the activity of effector domains.

[0030] In some embodiments, the Cas protein may be a mutant or fragment of a wild-type Cas protein (e.g., Cas9). In other embodiments, the Cas protein may be derived from a mutant Cas protein. For example, the amino acid sequence of the Cas9 protein may be modified to alter one or more properties of the protein (e.g., nuclease activity, affinity, stability, etc.). Alternatively, domains of the Cas9 protein that are not involved in RNA targeting may be omitted from the protein, such that the modified Cas9 protein is smaller than the wild-type Cas9 protein. In some embodiments, the system utilizes a Cas9 protein from S. pyogenes, either bacterially encoded or codon-optimized for expression in mammalian cells.

[0031] A mutant Cas protein refers to a polypeptide derivative of a wild-type protein, such as a protein having one or more point mutations, insertions, deletions, truncations, fusion proteins, or a combination thereof. The mutant has at least one of RNA-guided DNA binding activity or RNA-guided nuclease activity, or both. Generally, the modified version is at least 50% identical (e.g., any number between 50% and 100%, inclusive, e.g., 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, and 99%) to a wild-type protein, such as SEQ ID NO: 1 (from GenBank: AKE81011.1) below.

[0032] [ka]

[0033] Cas proteins (as well as other protein components described in the present invention) can be obtained as recombinant polypeptides. To prepare recombinant polypeptides, a nucleic acid encoding the recombinant polypeptide can be ligated to a fusion partner, such as glutathione-S-transferase (GST), a 6x-His epitope tag, or another nucleic acid encoding an M13 gene 3 protein. The resulting fusion nucleic acid expresses a fusion protein in a suitable host cell, which can be isolated by methods known in the art. The isolated fusion protein can be further treated, for example, by enzymatic digestion, to remove the fusion partner and obtain a recombinant polypeptide of the present invention. Alternatively, proteins can be chemically synthesized (see, e.g., Creighton, "Proteins: Structures and Molecular Principles," W.H. Freeman & Co., New York, 1983) or produced by recombinant DNA techniques as described herein. For additional guidance, those skilled in the art can consult Frederick M. Ausubel et al., Current Protocols in Molecular Biology, John Wiley & Sons, 2003; and Sambrook et al., Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Press, Cold Spring Harbor, NY, 2001.

[0034] The Cas proteins described in the present invention may be provided in purified or isolated form, or may be part of a composition. Preferably, in the case of a composition, the protein is first purified to a certain degree, more preferably to a high level of purity (e.g., about 80%, 90%, 95%, or 99%, or higher). The composition according to the present invention may be any type of desired composition, but is typically an aqueous composition suitable for use as or inclusion in a composition for RNA-guided targeting. Those skilled in the art are familiar with the various substances that can be included in such nuclease reaction compositions.

[0035] To carry out the methods disclosed herein for modifying a target nucleic acid, proteins can be produced in target cells using mRNA, protein-RNA complexes (RNPs), or any suitable expression vector. Examples of expression vectors include chromosomal, non-chromosomal, and synthetic DNA sequences, bacterial plasmids, minicircles, phage DNA, baculovirus, yeast plasmids, vectors derived from combinations of plasmids and phage DNA, and viral DNA such as vaccinia, adenovirus, fowlpox, and pseudorabies. More details are provided in the Expression Systems and Methods section below.

[0036] As disclosed herein, nuclease-inactive Cas9 (e.g., dCas9 derived from S. pyogenes D10A, H840A mutant proteins) or nuclease-deficient nickase Cas9 (e.g., nCas9 derived from S. pyogenes D10A mutant proteins) can be used. dCas9 or nCas9 can also be derived from various bacterial species. Table 1 provides a non-exhaustive list of example dCas9s and their corresponding PAM requirements. Synthetic Cas surrogates, such as those described in Rauch et al., "Programmable RNA-Guided RNA Effector Proteins Built from Human Parts." Cell, Vol. 178, No. 1, June 27, 2019, pp. 122-134. e12, can also be used.

[0037] [Table 1]

[0038] UGI In some aspects of the present disclosure, the sequence-targeting component described above comprises a targeting fusion protein having (a) a sequence-targeting protein and (b) a first uracil-DNA glycosylase (UNG) inhibitor peptide (UGI). For example, the fusion protein may comprise a Cas9 protein fused to a UGI. Such a fusion protein may exhibit increased nucleic acid editing efficiency compared to a fusion protein that does not contain a UGI domain. In some embodiments, the UGI comprises a wild-type UGI sequence or one having the following amino acid sequence: sp|P14739|UNGI_BPPB2:uracil-DNA glycosylase inhibitor (UGI)MTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML (SEQ ID NO: 44).

[0039] In some embodiments, the UGI proteins provided herein include fragments of UGI and proteins homologous to UGI or UGI fragments. For example, in some embodiments, UGI comprises a fragment of the amino acid sequence set forth above. In some embodiments, UGI comprises an amino acid sequence homologous to the amino acid sequence set forth above, or an amino acid sequence homologous to a fragment of the amino acid sequence set forth in the UGI sequence above. In some embodiments, proteins comprising UGI, or a fragment of UGI, or a homolog of UGI or a UGI fragment, are referred to as "UGI variants." UGI variants share homology with UGI or a fragment thereof. For example, UGI variants are at least about 70% (e.g., at least about 80%, 90%, 95%, 96%, 97%, 98%, 99%) to wild-type UGI or a UGI sequence as set forth above.

[0040] Suitable UGI protein and nucleotide sequences are provided herein; additional suitable UGI sequences will be known to those of skill in the art, including, for example, those published in the following references: Wang et al., "Uracil-DNA glycosylase inhibitor gene of bacteriophage PBS2 encodes a binding protein specific for uracil-DNA glycosylase." J. Biol. Chem. 264:1163-1171 (1989); Lundquist et al., "Site-directed mutagenesis and characterization of uracil-DNA glycosylase inhibitor protein." Role of specific carboxylic amino acids in complex formation with Escherichia coli uracil-DNA glycosylase. J. Biol. Chem. 272:21408-21419 (1997); Ravishankar et al., "X-ray analysis of a complex of Escherichia coli uracil DNA glycosylase (EcUDG) with a proteinaceous inhibitor." The structure elucidation of a prokaryotic UDG. Nucleic Acids Res. 26:4880-4887 (1998); and Putnam et al., Protein mimicry of DNA from crystal structures of the uracil-DNA glycosylase inhibitor protein and its complex with Escherichia coli uracil-DNA glycosylase. J Mol. Biol. 287:331-346 (1999). The contents of each of these references are incorporated herein by reference.

[0041] b. RNA scaffolds for sequence recognition and effector recruitment: The second component of the platform disclosed herein is an RNA scaffold with three subcomponents: a programmable guide RNA motif, a CRISPR RNA motif, and a recruiting RNA motif. This scaffold can be a single RNA molecule or a complex of multiple RNA molecules. As disclosed herein, the programmable guide RNA, the CRISPR RNA, and the Cas protein together form a CRISPR / Cas-based module for sequence targeting and recognition, and the recruiting RNA motif recruits a protein effector that performs gene correction through an RNA-protein binding pair. Thus, this second component connects the correction module and the sequence recognition module.

[0042] Programmable guide RNA One important subcomponent is the programmable guide RNA. Due to its simplicity and efficiency, the CRISPR-Cas system has been used to perform genome editing in the cells of various organisms. The specificity of this system is determined by the base pairing between the target DNA and the custom-designed guide RNA. By genetically engineering and adjusting the base pairing properties of the guide RNA, it can target any desired sequence as long as the target sequence contains a PAM sequence.

[0043] Among the subcomponents of the RNA scaffolds disclosed herein, the guide sequence provides target specificity. The guide sequence comprises a region complementary to and capable of hybridizing with a preselected target site of interest. In various embodiments, the guide sequence can comprise from about 10 nucleotides to more than about 25 nucleotides. For example, the region of base-pairing between the guide sequence and the corresponding target site sequence can be about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 22, 23, 24, 25, or more than 25 nucleotides in length. In an exemplary embodiment, the guide sequence is about 17-20 nucleotides in length, such as 20 nucleotides in length.

[0044] One requirement for selecting a suitable target nucleic acid is that it have a 3' PAM site / sequence. Each target sequence and its corresponding PAM site / sequence are referred to herein as a Cas target site. The type II CRISPR system, one of the best-characterized systems, requires only the Cas9 protein and a guide RNA complementary to the target sequence to affect target cleavage. The S. pyogenes type II CRISPR system uses a target site with N12-20NGG, where NGG represents the PAM site derived from S. pyogenes and N12-20 represents the 12-20 nucleotides directly 5' of the PAM site. Additional PAM site sequences from other bacterial species include NGGNG, NNNNGATT, NNAGAA, NNAGAAW, and NAAAAC. For example, see U.S. Patent Application Publication No. 20140273233, International Publication No. 2013176772, Cong et al. (2012), Science 339 (6121): 819-823; Jinek et al. (2012), Science 337 (6096): 816-821; Mali et al. (2013), Science 6121: 823-826; Gasiunas et al. (2012), Proc Natl Acad Sci USA 109 (39): E2579-E2586; Cho et al. (2013), Nature Biotechnology 31, 230-232; Hou et al., Proc Natl Acad Sci USA. 2013 Sep. 24; 110(39):15644-9; Mojica et al., Microbiology. 2009 Mar; 155(Pt3):733-40, www.addgene.org / CRISPR / . The contents of these documents are incorporated herein by reference in their entirety.

[0045] The target nucleic acid strand can be either of the double strands of the genomic DNA of the host cell.Examples of such genomic dsDNA include, but are not limited to, host cell chromosomes, mitochondrial DNA, and stable maintained plasmids.However, it should be understood that the method of the present invention can be carried out on other dsDNA present in the host cell, such as non-stable plasmid DNA, virus DNA, and phagemid DNA, regardless of the nature of the host cell dsDNA, as long as there is a Cas target site.The method can also be carried out on RNA.

[0046] CRISPR motifs In addition to the guide sequence described above, the RNA scaffolds of the present invention contain additional active or inactive subcomponents. In one example, the scaffold has a CRISPR motif with tracrRNA activity. For example, the scaffold may be a hybrid RNA molecule in which the programmable guide RNA described above is fused to a tracrRNA to mimic the natural crRNA:tracrRNA duplex. An exemplary hybrid crRNA:tracrRNA, gRNA sequence is shown below: 5'-(20nt guide)-GUUUAAGAGCUAUGCUGGAAACAGCAUAGCAAGUUUAAAUAAGGC UAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUUU-3' (SEQ ID NO: 45; Chen et al., Cell. 2013 Dec. 19; 155(7):1479-91). Various tracrRNA sequences are known in the art, including, for example, the following tracrRNAs and their active portions: As used herein, an active portion of a tracrRNA retains the ability to form a complex with a Cas protein, such as Cas9 or dCas9. See, e.g., International Publication No. 2014144592. Methods for generating crRNA-tracrRNA hybrid RNAs are known in the art. See, e.g., International Publication No. 2014099750, U.S. Patent Application Publication No. 20140179006, and U.S. Patent Application Publication No. 20140273226. The contents of these documents are incorporated herein by reference in their entirety.

[0047] [ka]

[0048] In some embodiments, the tracrRNA activity and the guide sequence are two separate RNA molecules that together form the guide RNA and associated scaffold, in which case the molecule with the tracrRNA activity must be able to interact (usually by base pairing) with the molecule with the guide sequence.

[0049] Recruitment RNA motif The third subcomponent of the RNA scaffold is the recruitment RNA motif that connects the editing module and the sequence recognition module, and this connection is critical to the platform disclosed herein.

[0050] One method for recruiting effector / DNA-editing enzymes to target sequences is by fusing the effector protein directly to dCas9. While successful sequence-specific transcriptional activation or repression has been achieved by fusing the effector enzyme ("editing module") directly to the protein required for sequence recognition (e.g., dCas9), protein-protein fusion designs can introduce spatial obstacles and are not ideal for enzymes that require the formation of multimeric complexes for their activity. Indeed, most nucleotide editing enzymes (e.g., AID or APOBEC3G) require the formation of dimers, tetramers, or higher-order oligomers for DNA-editing catalytic activity.

[0051] In contrast, the platform disclosed herein is based on RNA scaffold-mediated effector protein recruitment. More specifically, the platform utilizes various RNA motifs / RNA binding protein binding pairs. For this purpose, the RNA scaffold is designed so that the RNA motif (e.g., MS2 operator motif) that specifically binds to RNA binding protein (e.g., MS2 coat protein, MCP) is linked to the gRNA-CRISPR scaffold. The recruiting RNA motif can be fused to the 3'-end or 5'-end of the gRNA-CRISPR scaffold, or can replace the loop in the gRNA-CRISPR scaffold, specifically the tetraloop and / or stem loop 2.

[0052] As a result, the RNA scaffold component of the platform disclosed herein is a designed RNA molecule that contains not only a gRNA motif for specific DNA / RNA sequence recognition and a CRISPR RNA motif for dCas9 binding, but also a recruitment RNA motif for effector recruitment (Figure 1B). In this way, recruited effector protein fusions can be recruited to target sites via their ability to bind to the recruitment RNA motif. The flexibility of RNA scaffold-mediated recruitment allows for the relatively easy formation of functional monomers, as well as dimers, tetramers, or oligomers, near target DNA or RNA sequences. These RNA recruitment motif / binding protein pairs can be derived from naturally occurring sources (e.g., RNA phages or yeast telomerase) or artificially designed (e.g., RNA aptamers and their corresponding binding protein ligands). A non-exhaustive list of examples of recruitment RNA motif / RNA-binding protein pairs that can be used in the CasRcure system is summarized in Table 2.

[0053] [Table 2]

[0054] The sequences of the above binding pairs are listed below.

[0055] 1. Telomerase Ku-binding motif / Ku heterodimer a. Ku-binding hairpin

[0056] [ka]

[0057] b. Ku heterodimer

[0058] [ka]

[0059] 2. Telomerase Sm7 binding motif / Sm7 homoheptamer a. Sm consensus site (single-stranded)

[0060] [ka]

[0061] b. Monomeric Sm-like protein (archaea)

[0062] [ka]

[0063] 3. MS2 phage operator stem-loop / MS2 coat protein a. MS2 phage operator stem loop

[0064] [ka]

[0065] b. MS2 coat protein

[0066] [ka]

[0067] 4. PP7 Phage Operator Stem-Loop / PP7 Coat Protein a. PP7 phage operator stem loop

[0068] [ka]

[0069] b. PP7 coat protein (PCP)

[0070] [ka]

[0071] 5. SfMu Com stem-loop / SfMu Com binding protein a. SfMu Com stem loop

[0072] [ka]

[0073] b. SfMu Com-binding protein

[0074] [ka]

[0075] RNA scaffold can be either a single RNA molecule or a complex of multiple RNA molecules.For example, guide RNA, CRISPR motif and recruiting RNA motif can be three segments of one long single RNA molecule.Alternatively, one, two or three of them can be present in separate molecules.In the latter case, the three components can be linked together to form scaffold by covalent or non-covalent linkage or bond, for example, including Watson-Crick type base pairing.

[0076] In one example, an RNA scaffold may comprise two separate RNA molecules. The first RNA molecule may comprise a programmable guide RNA and a region capable of forming a stem duplex structure with a complementary region. The second RNA molecule may comprise a complementary region in addition to the CRISPR motif and the recruitment DNA motif. This stem duplex structure allows the first and second RNA molecules to form an RNA scaffold of the present invention. In one embodiment, the first and second RNA molecules each comprise a sequence (of about 6 to about 20 nucleotides) that base-pairs with the other sequence. Similarly, the CRISPR motif and the recruitment DNA motif may be present in different RNA molecules and may be combined into separate stem duplex structures.

[0077] The RNAs and related scaffolds of the present invention can be produced by various methods known in the art, including cell-based expression, in vitro transcription, and chemical synthesis. Because relatively long RNAs (200-mer or longer) can be chemically synthesized using TC-RNA chemistry (see, for example, U.S. Pat. No. 8,202,983), it is possible to produce RNAs with specialized characteristics that are superior to those possible with the four basic ribonucleotides (A, C, G, and U).

[0078] Cas protein-guide RNA scaffold complexes can be produced by recombinant techniques using host cell systems or in vitro translation-transcription systems known in the art. Details of such systems and techniques can be found, for example, in International Publication Nos. WO2014144761, WO2014144592, WO2013176772, US Patent Application Publication Nos. 20140273226, and 20140273233, the contents of which are incorporated herein by reference in their entireties. The complexes can be isolated or purified, at least to some extent, from the cellular material of the cells or in vitro translation-transcription systems in which they are produced.

[0079] qualification The RNA scaffold may contain one or more modifications. Such modifications may include the inclusion of at least one non-naturally occurring nucleotide, or a modified nucleotide or analog thereof. The modified nucleotide may have a modified ribose moiety, phosphate moiety, and / or base moiety. Modified nucleotides may include 2'-O-methyl analogs, 2'-deoxy analogs, or 2'-fluoro analogs. The nucleic acid backbone may be modified, for example, a phosphorothioate backbone may be used. Locked nucleic acids (LNA) or bridged nucleic acids (BNA) may also be used. Further examples of modified bases include, but are not limited to, 2-aminopurine, 5-bromo-uridine, pseudouridine, inosine, and 7-methylguanosine. Such modifications can be applied to any component of the CRISPR system. In a preferred embodiment, such modifications are made to the RNA component, for example, the guide RNA sequence.

[0080] In some embodiments, the above-described RNA scaffolds or subsections thereof may contain one or more modifications, e.g., base modifications, backbone modifications, etc., to provide new or enhanced characteristics to the nucleic acid (e.g., increased stability).

[0081] Modified backbones and modified internucleoside linkages Examples of suitable nucleic acids containing modifications include nucleic acids containing modified backbones or non-natural internucleoside linkages. Nucleic acids with modified backbones include those that retain a phosphorus atom in the backbone and those that do not have a phosphorus atom in the backbone.

[0082] Suitable modified oligonucleotide backbones that contain a phosphorus atom therein include, for example, phosphorothioates, chiral phosphorothioates, phosphorodithioates, phosphotriesters, aminoalkylphosphotriesters, methyl and other alkyl phosphonates, including 3'-alkylene phosphonates, 5'-alkylene phosphonates, and chiral phosphonates, phosphinates, phosphoramidates, including 3'-amino phosphoramidates and aminoalkyl phosphoramidates, phosphorodiamidates, thionophosphoramidates, thionoalkylphosphonates, thionoalkylphosphotriesters, selenophosphates, and boranophosphates having normal 3'-5' linkages, their 2'-5' linked analogs, and those having reverse polarity, where one or more internucleotide linkages are 3'-3', 5'-5', or 2'-2' linkages. Preferred oligonucleotides with reverse polarity contain a single 3'-3' linkage at the 3'-most internucleotide linkage, i.e., a single reverse nucleoside residue that may be basic (either lacking a nucleobase or having a hydroxyl group instead), and include various salts (e.g., potassium or sodium), mixed salts, and free acid forms.

[0083] In some embodiments, the subject nucleic acids contain one or more phosphorothioate and / or heteroatom internucleoside linkages, particularly -CH-NH-O-CH-, -CH-N(CH)-O-CH- (known as a methylene(methylimino) or MMI backbone), -CH-ON(CH)-CH-, -CH-N(CH)-N(CH)-CH-, and -ON(CH)-CH-CH- (natural phosphodiester internucleotide linkages are represented as -OP(=O)(OH)-O-CH-). MMI-type internucleoside linkages are disclosed in the above-referenced U.S. Pat. No. 5,489,677. Suitable amide internucleoside linkages are disclosed in U.S. Pat. No. 5,602,240.

[0084] Nucleic acids having morpholino backbone structures, such as those described in U.S. Patent No. 5,034,506, are also suitable. For example, in some embodiments, the subject nucleic acids contain a six-membered morpholino ring instead of a ribose ring. In some of these embodiments, phosphorodiamidates or other non-phosphodiester internucleoside linkages are used instead of phosphodiester linkages.

[0085] Suitable modified polynucleotide backbones that do not contain phosphorus atoms have backbones formed by short-chain alkyl or cycloalkyl internucleoside linkages, mixed heteroatom and alkyl or cycloalkyl internucleoside linkages, or one or more short-chain heteroatom or heterocyclic internucleoside linkages. These include morpholino linkages (formed in part from the sugar portion of the nucleoside), siloxane backbones, sulfide, sulfoxide, and sulfone backbones, formacetyl and thioformacetyl backbones, methyleneformacetyl and thioformacetyl backbones, riboacetyl backbones, alkene-containing backbones, sulfamate backbones, methyleneimino and methylenehydrazino backbones, sulfonate and sulfonamide backbones, those with amide backbones, and others with mixed N, O, S, and CH moieties.

[0086] mimetic The subject nucleic acids may be nucleic acid mimetics. When applied to polynucleotides, the term "mimetic" is intended to include polynucleotides in which only the furanose ring or both the furanose ring and the internucleotide linkage are replaced with non-furanose groups; replacement of only the furanose ring is also referred to in the art as a sugar surrogate. The heterocyclic base moiety or modified heterocyclic base moiety is maintained for hybridization with an appropriate target nucleic acid. One such nucleic acid, a polynucleotide mimetic that has been shown to have excellent hybridization properties, is called a peptide nucleic acid (PNA). In PNAs, the sugar backbone of a polynucleotide is replaced with an amide-containing backbone, particularly an aminoethylglycine backbone. The nucleotides are retained and are bound directly or indirectly to the aza nitrogen atoms of the amide portion of the backbone.

[0087] One polynucleotide mimetic that has been reported to have excellent hybridization properties is peptide nucleic acid (PNA). The backbone of PNA compounds is two or more linked aminoethylglycine units, giving PNA an amide-containing backbone. The heterocyclic base moiety is directly or indirectly bound to the aza nitrogen atom of the amide portion of the backbone. Representative U.S. patents that describe the preparation of PNA compounds include, but are not limited to, U.S. Pat. Nos. 5,539,082; 5,714,331; and 5,719,262.

[0088] Another class of polynucleotide mimetics under investigation is based on linked morpholino units (morpholino nucleic acids) in which a heterocyclic base is attached to a morpholino ring. Numerous linking groups have been reported to link the morpholino monomer units of morpholino nucleic acids. One class of linking groups has been selected to obtain nonionic oligomeric compounds. Nonionic morpholino-based oligomeric compounds are less likely to exhibit undesirable interactions with cellular proteins. Morpholino-based polynucleotides are nonionic mimics of oligonucleotides that are less likely to form undesirable interactions with cellular proteins (Dwaine A. Braasch and David R. Corey, Biochemistry, 2002, Vol. 41(14), pp. 4503-4510). Morpholino-based polynucleotides are disclosed in U.S. Pat. No. 5,034,506. A variety of compounds within the morpholino class of polynucleotides have been prepared with a variety of different linking groups joining the monomeric subunits.

[0089] A further class of polynucleotide mimetics is called cyclohexenyl nucleic acids (CeNA). The furanose ring normally present in DNA / RNA molecules is replaced with a cyclohexenyl ring. CeNA DMT-protected phosphoramidite monomers have been prepared and used to synthesize oligomeric compounds using classical phosphoramidite chemistry. Fully modified CeNA oligomeric compounds and oligonucleotides modified at specific positions with CeNA have been prepared and studied (see Wang et al., J. Am. Chem. Soc., 2000, vol. 122, pp. 8595-8602). In general, the incorporation of CeNA monomers into DNA strands increases the stability of DNA / RNA hybrids. CeNA oligoadenylates form complexes with RNA and DNA complements with stabilities similar to those of native complexes. Studies incorporating the CeNA structure into native nucleic acid structures have demonstrated facile conformational adaptation by NMR and circular dichroism.

[0090] Further modifications include locked nucleic acids (LNAs), in which a 2'-hydroxyl group is linked to the 4' carbon atom of the sugar ring, forming a 2'-C,4'-C-oxymethylene linkage, thereby forming a bicyclic sugar moiety. The linkage is a methylene (-CH2-) group bridging the 2' oxygen atom and the 4' carbon atom, where n is 1 or 2 (Singh et al., Chem. Commun., 1998, vol. 4, pp. 455-456). LNAs and LNA analogs exhibit extremely high duplex thermal stability with complementary DNA and RNA (Tm = +3 to +10°C), stability against 3'-exonuclease degradation, and good solubility properties. Potent, non-toxic antisense oligonucleotides containing LNAs have been described (Wahlestedt et al., Proc. Natl. Acad. Sci. USA, 2000, vol. 97, pp. 5633-5638).

[0091] The synthesis and preparation of LNA monomers adenine, cytosine, guanine, 5-methyl-cytosine, thymine, and uracil have been described, along with their oligomerization and nucleic acid recognition properties (Koshkin et al., Tetrahedron, 1998, Vol. 54, pp. 3607-3630). LNAs and their preparation are also described in WO 98 / 39352 and WO 99 / 14226.

[0092] modified sugar moiety The subject nucleic acids may also contain one or more substituted sugar moieties. Suitable polynucleotides contain sugar substituents selected from OH; F; O-, S-, or N-alkyl; O-, S-, or N-alkenyl; O-, S-, or N-alkynyl; or O-alkyl-Co-alkyl, where alkyl, alkenyl, and alkynyl are substituted or unsubstituted C1-C6. 10 Alkyl or C2-C 10 Alkenyl and alkynyl are also suitable. Particularly preferred are O((CH) n O) m CH3, O(CH2) n OCH3, O(CH2) nNH2, O(CH2) n CH3, O(CH2) n ONH2, and O(CH2) n ON((CH2) n CH3)2, where n and m are from 1 to about 10. Other suitable polynucleotides include C1 to C 10 The sugar substituents include lower alkyl, substituted lower alkyl, alkenyl, alkynyl, alkaryl, aralkyl, O-alkaryl or O-aralkyl, SH, SCH, OCN, Cl, Br, CN, CF, OCF, SOCH, SOCH, ONO, NO, N, NH, heterocycloalkyl, heterocycloalkaryl, aminoalkylamino, polyalkylamino, substituted silyl, RNA cleaving groups, reporter groups, intercalators, groups for improving the pharmacokinetic properties of oligonucleotides, or groups for improving the pharmacodynamic properties of oligonucleotides, and other substituents with similar properties. Suitable modifications include 2'-methoxyethoxy (2'-O-CHCHOCH, also known as 2'-O-(2-methoxyethyl) or 2'-MOE) (Martin et al., Helv. Chim. Acta, 1995, 78, 486-504), i.e., alkoxyalkoxy groups. Further suitable modifications include the 2'-dimethylaminooxyethoxy, also known as 2'-DMAOE, i.e., O(CH2)2ON(CH3)2 group, and 2'-dimethylaminoethoxyethoxy (also known in the art as 2'-O-dimethyl-amino-ethoxy-ethyl or 2'-DMAEOE), i.e., 2'-O-CH2-O-CH2-N(CH3)2, as described in the examples below.

[0093] Other suitable sugar substituents include methoxy (-O-CH), aminopropoxy (-OCHCHNH), allyl (-CH-CH=CH), -O-allylCH-CH=CH), and fluoro (F). The 2'-sugar substituent may be in the arabino (up) or ribo (down) position. A preferred 2'-arabino modification is 2'-F. Similar modifications may also be made at other positions in oligomeric compounds, particularly the 3' position of the sugar of the 3'-terminal nucleoside or 2'-5'-linked oligonucleotide and the 5' position of the 5'-terminal nucleotide. Oligomeric compounds may also have sugar mimetics, such as cyclobutyl moieties, in place of the pentofuranosyl sugar.

[0094] Base modifications and substitutions The subject nucleic acids may also contain nucleobase (often simply referred to in the art as "base") modifications or substitutions. As used herein, "unmodified" or "natural" nucleobases include the purine bases adenine (A) and guanine (G), and the pyrimidine bases thymine (T), cytosine (C), and uracil (U). Modified nucleobases include other synthetic and natural nucleobases, such as 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil and cytosine, 5-propynyl (-C=C-CH3) uracil and cytosine, and other alkynyl derivatives of pyrimidine bases, 6-azo Uracil, cytosine, and thymine, 5-uracil (pseudouracil), 4-thiouracil, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxyl, and other 8-substituted adenines and guanines, 5-halo, especially 5-bromo, 5-trifluoromethyl, and other 5-substituted uracils and cytosines, 7-methylguanine and 7-methyladenine, 2-F-adenine, 2-aminoadenine, 8-azaguanine and 8-azaadenine, 7-deazaguanine and 7-deazaadenine, and 3-deazaguanine and 3-deazaadenine. Further modified nucleobases include tricyclic pyrimidines such as phenoxazine cytidine (1H-pyrimido(5,4-b)(1,4)benzoxazin-2(3H)-one), phenothiazine cytidine (1H-pyrimido(5,4-b)(1,4)benzothiazin-2(3H)-one), G-clamps such as substituted phenoxazine cytidines (e.g., 9-(2-aminoethoxy)-H-pyrimido(5,4-(b)(1,4)benzoxazin-2(3H)-one), carbazole cytidine (2H-pyrimido(4,5-b)indol-2-one), pyridoindole cytidine (H-pyrido(3',2':4,5)pyrrolo(2,3-d)pyrimidin-2-one).

[0095] Heterocyclic base moieties can also include those in which purine or pyrimidine base is replaced with other heterocycles, such as 7-deazaadenine, 7-deazaguanosine, 2-aminopyridine, and 2-pyridone.Additional nucleobases include those disclosed in U.S. Patent No. 3,687,808; those disclosed in The Concise Encyclopedia of Polymer Science and Engineering, pages 858-859, edited by Kroschwitz, JI, John Wiley & Sons, 1990; those disclosed by Englisch et al., Angewandte Chemie, International Edition, 1991, vol. 30, page 613; and those disclosed by Sanghvi, YS, Chapter 15, Antisense Research and Applications, pages 289-302, edited by Crooke, ST and Lebleu, B., CRC Press, 1993.Some of these nucleobases are useful for increasing the binding affinity of oligomeric compounds. These include 5-substituted pyrimidines, 6-azapyrimidines, and N-2, N-6, and O-6 substituted purines, including 2-aminopropyladenine, 5-propynyluracil, and 5-propynylcytosine. 5-Methylcytosine substitutions have been shown to increase nucleic acid duplex stability by 0.6-1.2°C (Sanghvi et al., eds., Antisense Research and Applications, CRC Press, Boca Raton, 1993, pp. 276-278), and are preferred base substitutions, for example, when combined with 2'-O-methoxyethyl sugar modifications.

[0096] c. Effector: non-nuclease DNA-modifying enzyme The third component of the platform disclosed herein is a non-nuclease effector. The effector is not a nuclease and does not have any nuclease activity, but may have the activity of other types of DNA-modifying enzymes. Examples of enzymatic activity include, but are not limited to, deamination activity, methyltransferase activity, demethylase activity, DNA repair activity, DNA damage activity, dismutase activity, nickase activity, alkylation activity, depurination or depyrimidination activity, oxidation activity, pyrimidine dimer formation activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, or glycosylase activity. In some embodiments, the effector has the activity of cytidine deaminase (e.g., AID, APOBEC3G, and APOBEC1), adenosine deaminase (e.g., ADA), DNA methyltransferase, and DNA demethylase. In some embodiments, the effectors are derived from different vertebrate species and have unique activity profiles.

[0097] In a preferred embodiment, this third component is a conjugate or fusion protein having an RNA-binding domain and an effector domain, the two domains optionally being joined via a linker.

[0098] In some embodiments, effectors are not required in some cell types (e.g., cancer lines that overexpress deaminases). In these cases, endogenous effectors (such as APOBEC, AID, etc.) can be genetically edited to include a recruitment module, eliminating the need for an exogenous editor. This is true for cell types that express the editor of interest, such as lymphoid (B+ T cells) and certain cancer cells. In addition, the nickase activity does not have to be derived from the Cas module, but can be recruited from an effector; for example, dCas9 may have an aptamer for recruiting both the nickase and the editor through the same gRNA recruitment.

[0099] RNA-binding domain Although various RNA-binding domains can be used in the present invention, the RNA-binding domain of a Cas protein (such as Cas9) or its variants (such as dCas9) should not be used. As mentioned above, direct fusion with Cas9 tethers it to DNA in a defined conformation, preventing the formation of a functional oligomeric enzyme complex at the correct location. Instead, the present invention utilizes various other RNA motif-RNA-binding protein binding pairs. Examples include those listed in Table 2.

[0100] Thus, the ability of an RNA-binding domain to bind to a recruitment RNA motif allows the recruitment of effector proteins to target sites. The flexibility of RNA scaffold-mediated recruitment allows for the relative ease of forming functional monomers, as well as dimers, tetramers, or oligomers, near target DNA or RNA sequences.

[0101] Effector domain The effector component comprises an active portion, i.e., an effector domain. In some embodiments, the effector domain comprises a naturally occurring active portion of a non-nuclease protein (e.g., a deaminase). In other embodiments, the effector domain comprises a modified amino acid sequence (e.g., a substitution, deletion, or insertion) of a naturally occurring active portion of a non-nuclease protein. The effector domain has enzymatic activity. Examples of this activity include deamination activity, methyltransferase activity, demethylase activity, DNA repair activity, DNA damage activity, dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer formation activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, glycosylase activity, DNA methylation, histone acetylation activity, or histone methylation activity. Some modifications of non-nuclease proteins (e.g., deaminases) can help reduce off-target effects. For example, as described below, mutating Ser38 of AID to Ala can reduce recruitment of AID to off-target sites.

[0102] Linker The two domains mentioned above, as well as other domains disclosed herein, may be joined by a linker, such as, but not limited to, chemical modification, a peptide linker, a chemical linker, a covalent or non-covalent bond, or a protein fusion, or by any means known to those skilled in the art. The joining may be permanent or reversible. See, for example, U.S. Pat. Nos. 4,625,014, 5,057,301, and 5,514,363, U.S. Patent Application Publication Nos. 20150182596 and 20100063258, and International Publication No. 2012142515. The contents of these documents are incorporated herein by reference in their entirety. In some embodiments, several linkers may be included to exploit the desired properties of each linker and each protein domain of the conjugate. For example, flexible linkers and linkers that increase the solubility of the conjugate are contemplated for use alone or in combination with other linkers. The peptide linker can be linked to one or more protein domains of the conjugate by expressing the DNA encoding the linker.The linker can be acid-cleavable, photocleavable, or thermosensitive linker.Methods for conjugation are well known to those skilled in the art and are included for use in the present invention.

[0103] In some embodiments, the RNA binding domain and the effector domain may be joined by a peptide linker. The peptide linker may be linked by expressing a nucleic acid that encodes the two domains and the linker in frame. Optionally, the linker peptide may be joined at either or both of the amino and carboxy termini of the domains. In some examples, the linker is an immunoglobulin hinge region linker, such as those disclosed in U.S. Patent Nos. 6,165,476, 5,856,456, U.S. Patent Application Publication Nos. 20150182596 and 2010 / 0063258, and WO 2012 / 142515. Each of these documents is incorporated herein by reference in its entirety.

[0104] Other domains The effector fusion protein may contain other domains. In certain embodiments, the effector fusion protein may contain at least one nuclear localization signal (NLS). Generally, an NLS contains a stretch of basic amino acids. Nuclear localization signals are known in the art (see, for example, Lange et al., J. Biol. Chem., 2007, vol. 282:5101-5105). The NLS may be located at the N-terminus, C-terminus, or internal position of the fusion protein.

[0105] In some embodiments, the fusion protein may comprise at least one cell-penetrating domain to facilitate delivery of the protein to target cells. In one embodiment, the cell-penetrating domain may be a cell-penetrating peptide sequence. Various cell-penetrating peptide sequences are known in the art, including the sequence of the HIV-1 TAT protein, the TLM of human HBV, Pep-1, VP22, and polyarginine peptide sequences.

[0106] In still other embodiments, the fusion protein may comprise at least one marker domain. Non-limiting examples of marker domains include fluorescent proteins, purification tags, and epitope tags. In some embodiments, the marker domain may be a fluorescent protein. In other embodiments, the marker domain may be a purification tag and / or an epitope tag. See, for example, U.S. Patent Application Publication No. 20140273233.

[0107] In one embodiment, AID is used as an example to explain how the system works. AID is a cytidine deaminase that can catalyze the deamination of cytidine in the context of DNA or RNA. When AID is delivered to a targeted site, it changes a C base to a U base. In dividing cells, this can lead to a C to T point mutation. Alternatively, the C to U change can trigger cellular DNA repair pathways, primarily the excision repair pathway, which removes the mismatched UG base pair and replaces it with a TA, AT, CG, or GC pair. This results in the generation of a point mutation at the targeted CG site. Because the excision repair pathway is present in most, if not all, somatic cells, recruiting AID to the target site can correct the CG base pair to something else. In this case, if the CG base pair is the underlying disease-causing genetic mutation in somatic tissues / cells, the above-described approach can be used to correct the mutation and thereby treat the disease.

[0108] Similarly, if the underlying disease-causing genetic mutation is an AT base pair at a specific site, the same approach can be used to recruit adenosine deaminase to that specific site, where it can modify the AT base pair to something else. Other effector enzymes are expected to generate other types of changes in base pairing. A non-exhaustive list of examples of DNA / RNA-modifying enzymes is detailed in Table 3.

[0109] [Table 3]

[0110] The three specific components described above comprise the present technology platform, each of which can be individually selected from the lists in Tables 1-3 to achieve a specific therapeutic / utilitarian goal.

[0111] In one example, the CasRcure system was constructed using (i) dCas9 from S. pyogenes as the sequence-targeting protein, (ii) an RNA scaffold containing a guide RNA sequence, a CRISPR RNA motif, and an MS2 operator motif, and (iii) an effector fusion containing human AID fused to the MS2 operator-binding protein MCP. The sequences of the components are listed below.

[0112] [ka]

[0113] [ka]

[0114] [ka]

[0115] [ka]

[0116] [ka]

[0117] [ka]

[0118] [ka]

[0119] [ka]

[0120] [ka]

[0121] [ka]

[0122] [ka]

[0123] The RNA scaffold shown above contains one MS2 loop (1xMS2). Below is an exemplary sequence encoding an RNA scaffold containing two MS2 loops (2xMS2), with the MS2 scaffold underlined.

[0124] [ka]

[0125] [ka]

[0126] [ka]

[0127] [ka]

[0128] Similar to the Cas proteins described above, non-nuclease effectors can also be obtained as recombinant polypeptides. Techniques for producing recombinant polypeptides are known in the art. See, for example, Creighton, "Proteins: Structures and Molecular Principles," W.H. Freeman & Co., New York, 1983; Ausubel et al., "Current Protocols in Molecular Biology," John Wiley & Sons, 2003; and Sambrook et al., "Molecular Cloning, A Laboratory Manual," Cold Spring Harbor Press, Cold Spring Harbor, New York, 2001).

[0129] As described herein, the recruitment of AID to off-target sites can be reduced by mutating Ser38 of AID to Ala. The DNA and protein sequences of both wild-type AID and AID_S38A (phosphorylation-null, pnAID) are listed below.

[0130] [ka]

[0131] [ka]

[0132] Exemplary Sequences Shown below are a few exemplary RNA sequences of the gRNA constructs used in this study, each containing, from 5' to 3', a customizable target, a gRNA scaffold, and one or two copies of the MS2 aptamer.

[0133] [ka]

[0134] The above three components of the platform / system disclosed herein can be expressed using one, two, or three expression vectors. The system can be programmed to target virtually any DNA or RNA sequence. In addition to the second-generation CRC base editors described above, similar second-generation CRC base editors can be generated by modifying the modular components of the system, including any suitable Cas orthologs, deaminase orthologs, and other DNA-modifying enzymes.

[0135] Expression system To use the above-described platform, it may be desirable to express one or more of the protein and RNA components from the nucleic acids encoding them. This can be done in various ways. For example, the nucleic acid encoding the RNA scaffold or protein can be cloned into one or more intermediate vectors for introduction into prokaryotic or eukaryotic cells for replication and / or transcription. The intermediate vector is typically a prokaryotic vector, such as a plasmid, or a shuttle vector, or an insect vector, for storing or manipulating the nucleic acid encoding the RNA scaffold or protein to produce the RNA scaffold or protein. The nucleic acid can also be cloned into one or more expression vectors for administration to plant cells, animal cells, preferably mammalian or human cells, fungal cells, bacterial cells, or protozoan cells. Thus, the present invention provides nucleic acids encoding any of the above-mentioned RNA scaffolds or proteins. Preferably, the nucleic acid is isolated and / or purified.

[0136] The present invention also provides a recombinant construct or vector having a sequence encoding one or more of the above-described RNA scaffolds or proteins. Examples of constructs include vectors, such as plasmids or viral vectors, into which the nucleic acid sequence of the present invention is inserted in a forward or reverse orientation. In a preferred embodiment, the construct further comprises a regulatory sequence, including a promoter, operably linked to the sequence. Many suitable vectors and promoters are known to those skilled in the art and are commercially available. Suitable cloning and expression vectors for use in prokaryotic and eukaryotic hosts are also described, for example, in Sambrook et al. (2001, Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Press).

[0137] A vector refers to a nucleic acid molecule capable of transporting another nucleic acid to which it is linked. A vector may be capable of autonomous replication or integration into host DNA. Examples of vectors include plasmids, cosmids, or viral vectors. The vectors of the present invention contain nucleic acids in a form suitable for nucleic acid expression in a host cell. Preferably, the vector contains one or more regulatory sequences operably linked to the nucleic acid sequence to be expressed. "Regulatory sequences" include promoters, enhancers, and other expression control elements (e.g., polyadenylation signals). Regulatory sequences include those that direct constitutive expression of a nucleotide sequence as well as inducible regulatory sequences. The design of an expression vector can depend on factors such as the choice of host cell to be transformed, transfected, or transduced, and the desired level of RNA or protein expression.

[0138] Examples of expression vectors include chromosomal, non-chromosomal, and synthetic DNA sequences, bacterial plasmids, phage DNA, baculovirus, yeast plasmids, vectors derived from combinations of plasmids and phage DNA, and viral DNA such as vaccinia, adenovirus, fowlpox, and pseudorabies. However, any other vector can be used as long as it is replicable and viable in the host. The appropriate nucleic acid sequence can be inserted into the vector by a variety of procedures. In general, a nucleic acid sequence encoding one of the above-described RNAs or proteins can be inserted into an appropriate restriction endonuclease site by procedures known in the art. Such procedures and related subcloning procedures are within the skill of those in the art.

[0139] The vector may contain appropriate sequences for amplifying expression. In addition, the expression vector preferably contains one or more selectable marker genes that provide a phenotypic trait for selection of transformed host cells, such as dihydrofolate reductase or neomycin resistance in eukaryotic cell culture, or tetracycline or ampicillin resistance in E. coli.

[0140] The vector for expressing RNA may contain an RNA Pol III promoter, such as HI, U6, or 7SK promoter, that drives the expression of RNA. Such human promoters allow the expression of RNA in mammalian cells after plasmid transfection. Alternatively, for example, the T7 promoter can be used for in vitro transcription, and RNA can be transcribed and purified in vitro.

[0141] A vector containing an appropriate nucleic acid sequence, such as those described above, and an appropriate promoter or control sequence can be used to transform, transfect, or infect a suitable host, allowing the host to express the RNA or protein described above. Examples of suitable expression hosts include bacterial cells (e.g., Escherichia coli, Streptomyces, Salmonella typhimurium), fungal cells (yeast), insect cells (e.g., Drosophila and Spodoptera frugiperda (Sf9)), animal cells (e.g., CHO, COS, and HEK293), adenovirus, and plant cells. The selection of an appropriate host is within the skill of the art. In some embodiments, the present invention provides a method for producing the above-mentioned RNA or protein by transforming, transfecting, or infecting a host cell with an expression vector containing a nucleotide sequence encoding one of the RNAs, polypeptides, or proteins. The host cell is then cultured under suitable conditions that allow expression of the RNA or protein.

[0142] Any of the procedures known in the art for introducing foreign nucleotide sequences into host cells can be used, including calcium phosphate transfection, polybrene, protoplast fusion, electroporation, nucleofection, liposomes, microinjection, the use of naked DNA, plasmid vectors, viral vectors, both episomal and integrative, and any of the other well-known methods for introducing cloned genomic DNA, cDNA, synthetic DNA, or other foreign genetic material into host cells.

[0143] method Another aspect of the present invention encompasses a method for modifying a target DNA sequence (e.g., a chromosomal sequence) or a target RNA sequence in a cell, embryo, human, or non-human animal. The method comprises introducing into a cell or embryo (i) a sequence-targeting protein or a polynucleotide encoding the same, (ii) an RNA scaffold or a DNA polynucleotide encoding the same, and (iii) a non-nuclease effector fusion protein or a polynucleotide encoding the same, as described above. The RNA scaffold guides the sequence-targeting protein and the fusion protein to the target polynucleotide at the target site, and the effector domain of the fusion protein modifies the sequence. As disclosed herein, the sequence-targeting protein, such as the Cas9 protein, has been modified to eliminate endonuclease activity.

[0144] In certain embodiments, the effector protein functions as a monomer. In that case, the system of the present invention can target a single site either upstream (left) or downstream (right) of the target site, as shown, for example, in Figure 1C of WO2018129129. In other embodiments, the effector protein requires dimerization for proper catalytic function. To that end, the system can be simultaneously multiplexed against target sequences upstream and downstream of the target site, thus enabling dimerization of the effector protein (e.g., as shown on the left side of Figure 1D of WO2018129129). Alternatively, recruitment of the effector protein to a single site may be sufficient to increase its affinity for adjacent effector proteins and promote dimerization (e.g., as shown on the right side of Figure 1D of WO2018129129). In yet some other embodiments, tetrameric effector enzymes can be recruited and positioned at target sites, for example, as shown in Figure 1E of WO2018129129. This can be achieved by dual targeting or single targeting (e.g., as shown on the left and right sides of Figure 1E of WO2018129129). The system disclosed herein can also be used to edit RNA targets (e.g., retroviral inactivation). In this case, if the effector protein requires the assembly of functional oligomers, single targeting to an RNA molecule can promote oligomerization, for example, as shown in WO2018129129.

[0145] There is no sequence restriction on the target polynucleotide, except that the PAM sequence immediately follows (downstream or 3') the sequence. Examples of PAM include, but are not limited to, NGG, NGGNG, and NNAGAAW (wherein N is defined as any nucleotide, and W is defined as either A or T). Other examples of PAM sequences are provided above, and those skilled in the art will be able to identify additional PAM sequences for use with a given CRISPR protein. The target site may be located in the coding region of a gene, introns of a gene, intergenic control regions, etc. The gene may be a protein-coding gene or an RNA-coding gene.

[0146] The target polynucleotide may be any polynucleotide that is endogenous or exogenous to a cell. For example, the target polynucleotide may be a polynucleotide present in the nucleus of a eukaryotic cell. The target polynucleotide may be a sequence that encodes a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide).

[0147] The protein components of this system of the present invention may be introduced into cells or embryos as isolated proteins. Alternatively, the components may be introduced via nucleic acids, such as DNA or RNA (e.g., in vitro transcribed RNA), encoding such components. In one embodiment, each protein may contain at least one cell-permeable domain that facilitates cellular uptake of the protein. In other embodiments, mRNA or DNA molecules encoding one or more proteins may be introduced into cells or embryos. Generally, the DNA sequence encoding the protein is operably linked to a promoter sequence that will function in the cell or embryo of interest. The DNA sequence may be linear, or the DNA sequence may be part of a vector. In yet other embodiments, the protein may be introduced into cells or embryos as an RNA-protein complex comprising the protein and an RNA scaffold as described above.

[0148] In alternative embodiments, the DNA encoding the protein may further comprise one or more sequences encoding components of an RNA scaffold.Generally, the DNA sequences encoding the protein and the RNA scaffold are operably linked to an appropriate promoter control sequence that allows the expression of the protein and the RNA scaffold in cells or embryos, respectively.The DNA sequences encoding the protein and the RNA scaffold may further comprise additional expression control, regulation, and / or processing sequences.The DNA sequences encoding the protein and the guide RNA may be linear or part of a vector.

[0149] In embodiments in which the RNA is introduced into a cell via a DNA molecule encoding the RNA, the RNA coding sequence may be operably linked to a promoter regulatory sequence for expressing the guide RNA in a eukaryotic cell. For example, the RNA coding sequence may be operably linked to a promoter sequence recognized by RNA polymerase III (Pol III). Examples of suitable Pol III promoters include, but are not limited to, mammalian U6 or H1 promoters. In exemplary embodiments, the RNA coding sequence is linked to a mouse or human U6 promoter. In other exemplary embodiments, the RNA coding sequence is linked to a mouse or human H1 promoter.

[0150] The DNA molecule encoding the protein and / or RNA may be linear or circular. In some embodiments, the DNA sequence may be part of a vector, such as a polycistronic vector. Suitable vectors include plasmid vectors, phagemids, cosmids, artificial / minichromosomes, transposons, and viral vectors. In exemplary embodiments, the DNA encoding the protein and / or RNA is present in a plasmid vector. Non-limiting examples of suitable plasmid vectors include pUC, pBR322, pET, pBluescript, and variants thereof. The vector may also include additional expression control sequences (e.g., enhancer sequences, Kozak sequences, polyadenylation sequences, transcription termination sequences, etc.), selectable marker sequences (e.g., antibiotic resistance genes), and origins of replication.

[0151] The protein components (or nucleic acids encoding them) and RNA components (or DNA encoding them) of the systems of the present invention can be introduced into cells or embryos by a variety of means. Typically, the embryo is a fertilized one-cell stage embryo of the species of interest. In some embodiments, the cells or embryos are transfected. Suitable transfection methods include calcium phosphate-mediated transfection, nucleofection (or electroporation), cationic polymer transfection (e.g., DEAE-dextran or polyethyleneimine), viral transfection, virosome transfection, virion transfection, liposome transfection, cationic liposome transfection, immunoliposome transfection, nonliposomal lipid transfection, dendrimer transfection, heat shock transfection, magnetofection, lipofection, gene gun delivery, impalefection, sonoporation, phototransfection, gold nanoparticle-mediated transfection, and proprietary agent-enhanced nucleic acid uptake. Transfection methods are well known in the art (see, e.g., "Current Protocols in Molecular Biology," Ausubel et al., John Wiley & Sons, New York, 2003, or "Molecular Cloning: A Laboratory Manual," Sambrook & Russell, Cold Spring Harbor Press, Cold Spring Harbor, NY, 3rd ed., 2001). In other embodiments, molecules are introduced into cells or embryos by microinjection. For example, molecules can be injected into the pronuclei of one-cell embryos.

[0152] The protein components (or nucleic acids encoding them) and RNA components (or DNA encoding them) of this system of the present invention can be introduced into cells or embryos simultaneously or sequentially. The ratio of protein (or its encoding nucleic acid) to RNA (or DNA encoding the RNA) will generally be approximately stoichiometric so that an RNA-protein complex can form. Similarly, the ratio of two different proteins (or encoding nucleic acids) will be approximately stoichiometric. In one embodiment, the protein components and RNA components (or DNA sequences encoding them) are delivered together in the same nucleic acid or vector.

[0153] The method further includes maintaining the cell or embryo under appropriate conditions such that the guide RNA guides the effector protein to the target site of the target sequence and the effector domain modifies the target sequence.

[0154] Generally, cell can be maintained under the conditions suitable for cell growth and / or maintenance.Suitable cell culture conditions are well known in the art, and are described, for example, in "Current Protocols in Molecular Biology" Ausubel et al., John Wiley & Sons, New York, 2003 or "Molecular Cloning: A Laboratory Manual" Sambrook & Russell, Cold Spring Harbor Press, Cold Spring Harbor, New York, 3rd edition, 2001), Santiago et al. (2008) PNAS 105:5809-5814; Moehle et al. (2007) PNAS 104:3055-3060; Urnov et al. (2005) Nature 435:646-651; and Lombardo et al. (2007) Nat.Biotechnology 25:1298-1306. Those skilled in the art will appreciate that methods for culturing cells are known in the art and can and will vary depending on the cell type, and in any case, routine optimization can be used to determine the best technique for a particular cell type.

[0155] Embryos can be cultured in vitro (e.g., in cell culture). Typically, embryos are cultured at an appropriate temperature and in an appropriate medium with the required O2 / CO2 ratio to allow for the expression of proteins and RNA scaffolds as needed. Suitable, non-limiting examples of medium include M2, M16, KSOM, BMOC, and HTF medium. Those skilled in the art will understand that culture conditions can and will vary depending on the species of the embryo. In any case, routine optimization can be used to determine the best culture conditions for a particular species of embryo. In some cases, cell lines can be derived from in vitro cultured embryos (e.g., embryonic stem cell lines).

[0156] Alternatively, the embryo may be cultured in vivo by transferring the embryo into the uterus of a female host. Generally speaking, the female host is from the same or similar species as the embryo. Preferably, the female host is pseudopregnant. Methods for preparing pseudopregnant female hosts are known in the art. Additionally, methods for transferring embryos into female hosts are known. In vivo embryo culture allows the embryo to develop and can result in the birth of an animal derived from the embryo. Such an animal will contain the modified chromosomal sequence in every cell of its body.

[0157] Various eukaryotic cells are suitable for use in this method. For example, the cells may be human cells, non-human mammalian cells, non-mammalian vertebrate cells, invertebrate cells, insect cells, plant cells, yeast cells, or single-cell eukaryotes. Various embryos are suitable for use in this method. For example, the embryos may be one-cell, two-cell, or four-cell human or non-human mammalian embryos. Exemplary mammalian embryos, including one-cell embryos, include, but are not limited to, mouse, rat, hamster, rodent, rabbit, cat, dog, sheep, pig, cow, horse, and primate embryos. In yet other embodiments, the cells may be stem cells. Suitable stem cells include, but are not limited to, embryonic stem cells, embryonic stem cell-like stem cells, fetal stem cells, adult stem cells, pluripotent stem cells, induced pluripotent stem cells, multipotent stem cells, oligopotent stem cells, and unipotent stem cells. In exemplary embodiments, the cells are mammalian cells or the embryos are mammalian embryos.

[0158] As shown in International Publication No. 2018129129, this Cis Double Nicking Technology was applied to conduct a study to enhance the transformation efficiency of a bacterial gene transformation model. Experimentally, nCas9 (nCas9 D10A or nCas9 H840) were programmed to target two adjacent positions on the same DNA strand. Double nicking of the same strand by two gRNAs does not induce double-stranded DNA breaks or activate DSB repair pathways. This is therefore a safe technique. A schematic of the procedure is provided in Figure 8 of WO2018129129. To test this technique, the bacterial gene encoding the RNA polymerase β subunit (rpoB) was targeted using gRNAs TS-2 and TS-3. This is a negative selection system, allowing the antibiotic rifampicin to be used to select specific rpoB mutants, as the mutants are resistant to this drug (Rifampicin). R ) Results in prokaryotic cells suggest that targeting efficiency can be enhanced up to 100-fold.

[0159] Furthermore, by utilizing the modular design of CRC, the present invention provides a method for recruiting two effectors (which may be the same or different) to a target sequence and synergistically enhancing gene conversion. Such a design is illustrated in Figure 10 of International Publication No. 2018129129. For example, both gRNAs can be engineered to have the same recruitment RNA motif (e.g., an MS2 scaffold), allowing the CRC effector fused to an MCP protein to recruit to both nicking sites. This allows two identical effectors to be recruited to the target sequence, increasing the local concentration of the effector or promoting dimerization or multimerization required for effector function.

[0160] Similarly, the present invention also provides a method for recruiting or excluding CRC effectors from either nicking site by selecting gRNAs with or without recruitment RNA motifs, respectively, which recruits one effector but exposes single-stranded DNA, allowing effector function to be promoted.

[0161] In another example, the present invention provides a method for recruiting two different functional effectors to the same target sequence. The two effectors work synergistically together to promote gene conversion. For example, to further increase targeting efficiency, the CRC can be programmed to recruit a deaminase (e.g., AID) to a nicking site closer to the target nucleotide, and a local DNA repair inhibitor (e.g., UNG inhibitor, UGI) to a second nicking site. AID, for example, promotes C to T conversion in the target sequence, while UGI locally inhibits the endogenous repair pathway. Thus, these two effectors specifically cooperate at the target site to enhance conversion efficiency. To avoid crosstalk between the CRC recruitment site and the inhibitor recruitment site, an orthogonal recruitment RNA motif can be used for each of these modules (e.g., MS2-MCP recruits the CRC effector AID fused with MCP, and PP7-PCP recruits UGI fused with PCP).

[0162] In some embodiments, heterodimerization can be applied when heterodimerization is required for proper effector activity. Heterodimerization can also be applied to any gene convertase system that requires at least two components to function effectively. A non-exhaustive list of mobilized RNA scaffolds and their RNA-binding protein partners is summarized in Table 2. Finally, if there are PAM sequence restrictions for cis double nicking, it is also possible to program Cas9 orthologs from species other than S. pyogenes, depending on the PAM sequence available near the target site. A non-exhaustive list of Cas9 orthologs from various species is summarized in Table 1.

[0163] The fundamental difference between BE and CRC is the mechanism by which the effector DNA-modifying enzyme is recruited to the target site. BE is mediated by direct fusion of Cas9 with the effector, whereas CRC is mediated through an RNA aptamer 3′ to the gRNA, which then recruits a cognate aptamer ligand fused to the effector. An attractive feature of the CRC system is its modular design. The functionalities for DNA recognition and effector action reside on separate molecules, and the interaction between the two functional modules is encoded by a gRNA molecule that can be easily reprogrammed. In this way, the CRISPR protein module and the effector module can be individually engineered and optimized without interfering with each other, as demonstrated in this study. Additionally, the CRC design may facilitate simultaneous targeting of different sites using different types of effectors (multiplexing). For example, an A to G effector (adenine deaminase) and a C to T effector (cytidine deaminase) can be introduced into the same cell to target different sequences, or to target one site for transcriptional activation (transiently) and a second site for stop codon knockout (permanently).

[0164] In the study shown in the example below, the two best CRC constructs, Gen2 CRC_AID( A CRCnu.2) and Gen2_CRC_APOBEC1( A1 CRCnu.2), both of which were codon-optimized Cas9 fused to 2×UGI. D10A It consists of a nickase, a gRNA with two copies of the MS2 aptamer linked to its 3' end, and a codon-optimized MCP-cytidine deaminase fusion protein (Figures 9B and 9C). A CRCnu.2 and A1The cytidine deaminases in CRCnu.2 are human AID and rat APOBEC1, respectively. The effector modules of both Gen2 CRC systems contain a nuclear localization signal and a flexible hinge linker that separates the cytidine deaminase from the RNA-aptamer ligand.

[0165] For example, at the tested target sites, the base editing activity of the two CRC constructs, although different, exceeds 10%, and can even reach 50%, and off-target activity is generally absent or low depending on the guide sequence used. These CRC constructs have reached a common benchmark and can be further tested and optimized in therapeutic settings, such as patient-derived cells and animal disease models. These CRC constructs can be used in at least three different therapeutic modes: (1) base conversion (including correction of disease-causing mutations and introduction of second site-suppressing mutations), (2) premature stop codon knockout, and (3) exon skipping.

[0166] In this study, we tested the therapeutic mode of base correction of loss-of-function mutations in a reporter GFP gene and the mode of stop codon knockout with high efficiency using a wild-type GFP transgene and the endogenous PDCD1 gene. Because the 3' splice site of almost all genes contains an AG consensus sequence (46, 47), exon skipping is feasible for some disease genes if an optimal PAM motif is available near the target splice site (48). Therefore, the base editing platform can provide a powerful therapy for permanently correcting disease-causing mutations (e.g., beta-thalassemia), permanently knocking out gene expression (e.g., CAR-T cell engineering), and permanently skipping the expression of disease-causing exons (e.g., Duchenne muscular dystrophy) in both ex vivo and in vivo therapeutic settings.

[0167] The core of the CRC platform is based on the ability of nuclease-deficient CRISPR complexes to serve as DNA or RNA sequence-specific targeting modules. This platform is also the basis for numerous other engineered systems, based on either RNA- or protein-based recruitment. In addition to the BE base editing system, the Feng Zhang group (16) and the Stanley Qi group (15) used gRNA components and RNA aptamers to recruit transcriptional regulatory effectors and reprogram transcriptional networks. The Bassik group placed a recruiting RNA aptamer in the tetraloop and stem-loop 2 of the gRNA to recruit a mutant hyperactive cytidine deaminase (CRISPR-X system) (20). Interestingly, when the RNA aptamer is placed at this location rather than at the 3' end of the gRNA, as in the CRC system, CRISPR-X exhibits a unique activity profile in which cytidine deamination activity is low-efficient and extends widely around and beyond the target protospacer sequence (20, 21). This property of CRISPRx, in conjunction with a highly active variant of the deaminase (AID), has been used to generate permutations and protein evolution / genetic engineering in cells and in vitro. This system is particularly useful for generating antibody diversity (21). It is expected that systems using CRISPR DNA / RNA sequence recognition modules will be further expanded to rewrite genomes or reprogram cellular programs. Therefore, the same strategy can be used in the CRC system described herein.

[0168] Practicality and Applications The systems and methods disclosed herein have a wide variety of practical applications, including modifying and editing (e.g., inactivating and activating) target polynucleotides in numerous cell types.Therefore, the systems and methods have a wide range of applications, for example, in research and therapy.For example, the systems and methods can be used in high-throughput screening, targeting multiple different gene loci with multiple systems having different guide RNAs to obtain multiple different phenotypic outcomes and screening them (e.g., screening for better growth or lethality in cell lines).In another example, the systems and methods can be used for gene mutagenesis (similar to CRISPR tiling) to create novel proteins.

[0169] Many devastating human diseases share a common cause: genetic alterations or mutations. Disease-causing mutations in patients are either acquired through inheritance from parents or caused by environmental factors. Such diseases include, but are not limited to, the following categories: First, some genetic disorders are caused by germline mutations. One example is cystic fibrosis, which is caused by a mutation in the CFTR gene inherited from a parent. A second, inhibitory mutation in the mutant CFTR can partially restore the function of the CFTR protein in somatic tissues. Other exemplary genetic diseases caused by point gene mutations that can be corrected by the present invention include Gaucher disease, alpha trypsin deficiency, and sickle cell anemia, to name a few. Second, some diseases, such as chronic viral infectious diseases, are caused by exogenous environmental factors and resulting genetic alterations. One example is AIDS, which is caused by the insertion of the human HIV viral genome into the genome of infected T cells. Third, some neurodegenerative diseases involve genetic alterations. One example is Huntington's disease, which is caused by the CAG trinucleotide expansion in the Huntingtin gene of affected individuals.Other examples include lysosomal storage diseases, epidermolysis bullosa, and retinal degeneration.Finally, cancer is caused by various somatic mutations accumulated in cancer cells.Therefore, correcting the genetic mutations that cause disease or functionally correcting the sequence provides an attractive therapeutic opportunity for treating these diseases.

[0170] Somatic gene editing is an attractive therapeutic strategy for many human diseases. To achieve successful therapeutic gene editing, three key factors are considered essential: (i) how to achieve sequence-specific recognition (the "sequence recognition module"); (ii) how to correct the underlying mutation (the "correction module"); and (iii) how to link the "correction module" and the "sequence recognition module" together to achieve sequence-specific correction. There are numerous ways to accomplish each individual task. However, none of the currently existing platforms or technologies can achieve optimal and practical somatic gene editing. More specifically, current gene-specific editing technologies are mostly based on nuclease-induced DNA DSBs and the resulting DSB-induced homologous recombination, and the activity of DSB-induced homologous recombination is low or absent in most somatic cells. Therefore, these technologies are limited in their use for therapeutic correction of pathological gene mutations in somatic tissues of most diseases.

[0171] In contrast, the system and method disclosed in the present invention enable DNA sequence-directed editing of genes or RNA transcripts without relying on nuclease activity. The system and method do not generate DSBs or rely on DSB-mediated homologous recombination. Furthermore, the modular design of this system allows for a highly flexible and convenient way to target any desired DNA or RNA sequence. Essentially, this approach allows for the guidance of DNA or RNA editing enzymes to virtually any DNA or RNA sequence in somatic cells, including stem cells. By precisely editing the target DNA or RNA sequence, the enzyme can correct mutated genes in genetic disorders, inactivate viral genomes in infected cells, generate stop codons for inactivation, eliminate the expression of disease-causing proteins in diseases including neurodegenerative disorders, silence oncogenic proteins in cancer, mutate splice consensus sites to eliminate disease-causing exons, or mutate regulatory sequences to restore therapeutic expression / inactivation of genes. Thus, the systems and methods disclosed herein can be used to correct the underlying genetic alterations in diseases, including the above-mentioned genetic disorders, chronic infectious diseases, neurodegenerative diseases, and cancer. Importantly, the systems and methods disclosed herein can be used to genetically engineer cells for both the generation of research tools or cell-based therapies.

[0172] genetic disorders It is estimated that over 6,000 genetic diseases are caused by known gene mutations. Correcting the underlying disease-causing mutation in pathological tissues / organs can provide disease alleviation or cure. For example, cystic fibrosis affects 1 in 3,000 people in the United States. It is caused by inheritance of a mutant CFTR gene, and 70% of patients have the same mutation, namely, a trinucleotide deletion resulting in the deletion of phenylalanine at position 508 (called ΔPhe508). ΔPhe508 leads to mislocalization and degradation of CFTR. Using the system and method disclosed in the present invention, the Val509 residue (GTT) in affected tissues (lungs) can be converted to Phe509 (TTT), thereby functionally correcting the ΔPhe508 mutation. In addition, a second inhibitory mutation in mutant ΔPhe508 CFTR (such as R553Q or R553M or V510D) can partially restore function of the CFTR protein in somatic tissues.

[0173] chronic infectious diseases The systems and methods disclosed herein can also be used to specifically inactivate any gene in a viral genome integrated into human cells / tissues. For example, the systems and methods disclosed herein can create stop codons to prematurely terminate the translation of essential viral genes, thereby correcting or curing chronic, debilitating infectious diseases. For example, current AIDS therapies can reduce viral load but cannot completely eliminate dormant HIV from positive T cells. The systems and methods disclosed herein can be used to permanently inactivate expression of essential HIV genes in an HIV genome integrated into human T cells by introducing one or more stop codons. Another example is hepatitis B virus (HBV). The systems and methods disclosed herein can be used to specifically inactivate essential HBV genes integrated into the human genome and silence the HBV life cycle.

[0174] Neurodegenerative diseases Some neurodegenerative diseases are caused by gain-of-function mutations. For example, SOD1G93A is linked to the development of amyotrophic lateral sclerosis (ALS). The system and method disclosed in the present invention can be used to either correct the mutation or eliminate mutant protein expression by introducing a stop codon or changing the splice site. For example, alternative splicing forms of tau protein containing exon 10 play a causal role in Alzheimer's disease. Changing the CG base pair of the consensus exon 10 splice site will result in the disappearance of the alternative splicing version of tau.

[0175] cancer Numerous genes, including tumor suppressor genes, oncogenes, and DNA repair genes, contribute to the development of cancer. Mutations in these genes are often linked to various cancers. Using the systems and methods disclosed in the present invention, these mutations can be specifically targeted and corrected. As a result, the causative oncogenic proteins can be functionally suppressed or their expression can be eliminated by introducing point mutations into either the catalytic site or the splice site.

[0176] Somatic gene knockout In some embodiments, protein expression of genes in somatic cells of human and non-human organisms can be eliminated by generating premature stop codons. This approach can be used for therapeutic purposes or to generate research tools.

[0177] Modification of regulatory elements This method can be used to change the sequence of DNA and RNA regulatory elements.As a result, this method provides a method for changing, silencing or activating gene expression by changing various mechanisms involved in gene expression.This can be used for therapeutic purposes and to generate research tools.

[0178] Stem cell genetic modification In some embodiments, the systems and methods disclosed herein can be used to genetically modify cells that are to be reprogrammed into different cell types. Suitable cells include, for example, stem cells (e.g., adult stem cells, embryonic stem cells, induced pluripotent stem cells, mesenchymal stem cells, etc., as referenced in Stem cells: past, present, and future. Zakrzewski et al., Stem Cell Res Ther. 2019 Feb. 26; 10(1):68), and progenitor cells (e.g., cardiac progenitor cells, neural progenitor cells, etc.), or mature cells used for conversion into different cell types (e.g., using algorithms such as those referenced in Molecular Interaction Networks to Select Factors for Cell Conversion. Ouyang JF et al., Methods Mol Biol. 2019; 1975:333-361). Suitable cells may be derived from any multicellular organism, including, for example, mammals (including, for example, rodents, humans, horses, camels, pigs), insects, birds (including, for example, chickens, ducks), etc. Suitable host cells include in vitro or ex vivo host cells, e.g., isolated host cells.

[0179] In some embodiments, the present invention can be used for ex vivo targeting and precise genetic modification of cells or tissues to correct the underlying genetic defect. After ex vivo correction, the tissue can be returned to the patient. Furthermore, this technology can be widely used in cell-based therapy to correct genetic diseases.

[0180] The term "stem cell" as used herein refers to a cell that can differentiate into a diverse range of specialized cell types under suitable conditions and that can self-renew and remain in an essentially undifferentiated, pluripotent state under other suitable conditions. The term "stem cell" also encompasses pluripotent cells, multipotent cells, precursor cells, and progenitor cells. Exemplary human stem cells can be derived from hematopoietic or mesenchymal stem cells obtained from bone marrow tissue, embryonic stem cells obtained from embryonic tissue, or embryonic germ cells obtained from fetal reproductive tissue. Exemplary pluripotent stem cells can also be produced from somatic cells by reprogramming them into a pluripotent state through the expression of certain transcription factors associated with pluripotency. Such cells are referred to as "induced pluripotent stem cells" or "iPScs or iPS cells."

[0181] "Embryonic stem (ES) cells" are obtained from an early embryo, such as the inner cell mass of the blastocyst stage, or produced by artificial means (e.g., nuclear transfer), and are undifferentiated pluripotent cells that can give rise to any differentiated cell type in the embryo or adult, including germ cells (such as sperm and eggs).

[0182] "Induced pluripotent stem cells (iPScs or iPS cells)" are cells generated by reprogramming somatic cells by expressing or inducing the expression of a combination of factors (hereinafter referred to as reprogramming factors). iPS cells can be generated using fetal, live-born, neonatal, juvenile, or adult somatic cells. Factors that can be used to reprogram somatic cells into pluripotent stem cells include, for example, Oct4 (sometimes referred to as Oct3 / 4), Sox2, c-Myc, Klf4, Nanog, and Lin28. In some embodiments, somatic cells are reprogrammed by expressing at least two reprogramming factors, at least three reprogramming factors, at least four reprogramming factors, at least five reprogramming factors, at least six reprogramming factors, or at least seven reprogramming factors to reprogram the somatic cells into pluripotent stem cells.

[0183] "Hematopoietic progenitor cells" or "hematopoietic precursor cells" refer to cells that are committed to the hematopoietic lineage but are capable of further hematopoietic differentiation, including hematopoietic stem cells, multipotent hematopoietic stem cells, common myeloid progenitor cells, megakaryocytic progenitor cells, erythroid progenitor cells, and lymphoid progenitor cells. Hematopoietic stem cells (HSCs) are multipotent stem cells that give rise to all blood cell types, including bone marrow tissue (monocytes and macrophages, granulocytes (neutrophils, basophils, eosinophils, and mast cells), erythrocytes, megakaryocytes / platelets, dendritic cells), and lymphoid lineages (T cells, B cells, NK cells).

[0184] "Pluripotent stem cells" refer to stem cells that have the potential to differentiate into all cells that make up one or more tissues or organs, or preferably any of the three germ layers: endoderm (inner stomach lining, gastrointestinal tract, lungs), mesoderm (muscle, bone, blood, urogenital tract), or endoderm (epidermal tissue and nervous system).

[0185] As used herein, the term "somatic cell" refers to any cell other than a germ cell, such as an egg or sperm, that does not directly transfer its DNA to the next generation. Typically, somatic cells have limited or no pluripotency. As used herein, somatic cells may be naturally occurring or genetically modified.

[0186] Cellular and Ex Vivo Therapies Various embodiments of the present invention also provide cell lines for use in therapy, produced or used in accordance with any of the other embodiments of the present invention. In one embodiment, the present invention relates to a method for generating therapeutic cells, such as T cells genetically engineered to express a chimeric antigen receptor (CAR-T) or a T cell receptor (TCR-T). CAR-T / TCR-T cells may be derived from primary T cells or differentiated from stem cells. Suitable stem cells include, but are not limited to, mammalian stem cells, such as human stem cells, including, but not limited to, hematopoietic stem cells, neural stem cells, embryonic stem cells, induced pluripotent stem cells (iPSCs), mesenchymal stem cells, mesodermal stem cells, liver stem cells, pancreatic stem cells, muscle stem cells, and retinal stem cells. Other stem cells include, but are not limited to, mammalian stem cells, such as mouse stem cells, e.g., mouse embryonic stem cells.

[0187] In various embodiments, the present invention can be used to knock down, modify, or increase the expression of a single gene or multiple genes in various types of cells or cell lines, including, but not limited to, cells of mammalian origin. This technology can be used for numerous applications, including, but not limited to, knocking down genes to prevent graft-versus-host disease by making non-host cells non-immunogenic to the host, or to prevent host-versus-graft disease by making non-host cells resistant to attack by the host. Such approaches also relate to the generation of allogeneic (off-the-shelf) or autologous (patient-specific) cell-based therapies. Such genes include, but are not limited to, T cell receptor (TRAC), B2M, major histocompatibility complex (MHC class I and class II) genes including co-receptors (HLA-F, HLA-G), innate immune response (MICA, MICB, HCP5), inflammation (NKBBiL, LTA, TNF, LTB, LST1, NCR3, AIF1), immunoreceptors (LY6), heat shock proteins (HSPA1L, HSPA1A, HSPA1B), complement cascade, regulatory receptors (NOTCH4), antigen processing (TAP, HLA-DM, HLA-DO), peptide transport (RING1), increased efficacy or persistence (PD-1 , CTLA-4, FOXP3, B7, etc.), genes involved in T cell interactions with the tumor microenvironment (including but not limited to, receptors for cytokines such as TGFB, interleukin (IL)-4, IL-7, IL-2, IL-4, and repressors of IL-15, IL-12, IL-18, IL-2, IFN gamma), genes involved in contributing to cytokine release syndrome (including but not limited to, GMCSF), genes encoding the antigen targeted by the CAR / TCR (e.g., endogenous CS1 where the CAR is designed against CS1), or other genes found to be beneficial in CAR-T / TCR-T or other cell-based therapies including but not limited to CAR-NK, CAR-B, etc.See, e.g., DeRenzo et al., Genetic Modification Strategies to Enhance CAR T Cell Persistence for Patients With Solid Tumors. Front. Immunol., February 15, 2019.

[0188] This technology can also be used to knock down or modify genes involved in the fratricide of immune cells such as T cells and NK cells, or genes that alert the quarantine system of a patient or animal that a foreign cell, particle, or molecule has entered the patient or animal, or genes that encode proteins that are current therapeutic targets used to impair or boost the immune response, such as CD52 and PD1, respectively.

[0189] One application is genetically engineering the HLA alleles of bone marrow cells to increase haplotype matching. Genetically engineered cells can be used in bone marrow transplants to treat leukemia. Another application is genetically engineering the negative regulatory element of the fetal hemoglobin gene in hematopoietic stem cells to treat sickle cell anemia and beta-thalassemia. Mutation of the negative regulatory element reactivates expression of the fetal hemoglobin gene in hematopoietic stem cells, compensating for the loss of function due to mutations in the adult alpha or beta hemoglobin gene. A further application is genetically engineering iPS cells to generate allogeneic therapy cells for various degenerative diseases, including Parkinson's disease (neuronal cell loss) and type 1 diabetes (pancreatic beta cell loss). Another exemplary application is genetic engineering of HIV-resistant T cells by inactivating the CCR5 gene and other genes encoding receptors required for HIV cell entry.

[0190] This technique can also be used to generate transgenic animals that can be used as disease models or for gene function studies.

[0191] As used herein, the term "immune cells" generally includes white blood cells (leukocytes) derived from hematopoietic stem cells (HSCs) produced in the bone marrow. Examples of immune cells include, but are not limited to, lymphocytes (T cells, B cells, and natural killer (NK) cells) and bone marrow-derived cells (neutrophils, eosinophils, basophils, monocytes, macrophages, dendritic cells).

[0192] Immune cells can be isolated from subjects, particularly human subjects.Immune cells can be obtained from the subject of interest, such as a subject suspected of having a certain disease or condition, a subject suspected of having a predisposition to a certain disease or condition, or a subject undergoing therapy for a certain disease or condition.Immune cells can be collected from any location present in a subject, including but not limited to blood, umbilical cord blood, spleen, thymus, lymph node, and bone marrow.Immune cells isolated can be used directly or can be stored for a certain period of time, such as by freezing.

[0193] Immune cells can be enriched / purified from any tissue in which they reside, including, but not limited to, blood (including blood collected from a blood bank or umbilical cord blood bank), spleen, bone marrow, tissues removed and / or exposed during surgical procedures, and tissues obtained by biopsy procedures. The tissues / organs from which immune cells are enriched, isolated, and / or purified can be isolated from both living and non-living subjects, with non-living subjects being organ donors. In certain embodiments, immune cells are isolated from blood, such as peripheral blood or umbilical cord blood. In some aspects, immune cells isolated from umbilical cord blood have enhanced immunomodulatory capabilities, such as measured by CD4 or CD8 positive T cell suppression. In certain aspects, immune cells are isolated from pooled blood, particularly pooled umbilical cord blood, to enhance immunomodulatory capabilities. Pooled blood may be derived from two or more sources, such as 3, 4, 5, 6, 7, 8, 9, 10, or more sources (e.g., donor subjects).

[0194] The population of immune cells can be obtained from a subject who needs or suffers from a disease associated with reduced immune cell activity.Therefore, the cells can be autologous to the subject who needs therapy.Alternatively, the population of immune cells can be obtained from a donor, preferably a histocompatible donor.The population of immune cells can be collected from peripheral blood, umbilical cord blood, bone marrow, spleen, or any other organ / tissue in which immune cells exist in the subject or donor.The immune cells can be isolated from a pool of subjects and / or donors, such as from pooled umbilical cord blood.

[0195] When the population of immune cells is obtained from a donor different from the subject, the donor is preferably allogeneic, but the obtained cells are subject-compatible in that they can be introduced into the subject.Allogeneic donor cells may or may not be human leukocyte antigen (HLA) compatible.To be compatible with the subject, allogeneic cells can be treated to reduce immunogenicity.

[0196] In some embodiments, the immune cells are T cells (e.g., regulatory T cells, CD4 + T cells, CD 8 The cells may be T cells, or gamma-delta T cells), NK cells, invariant NK cells, NKT cells, stem cells (e.g., mesenchymal stem cells (MSCs) or induced pluripotent stem (iPSC) cells). In some embodiments, the cells are monocytes or granulocytes, e.g., myeloid cells, macrophages, neutrophils, dendritic cells, mast cells, eosinophils, and / or basophils. Also provided herein are methods for producing and genetically engineering immune cells, as well as methods for using and administering cells for adoptive cell therapy, where the cells may be autologous or allogeneic. Thus, immune cells can be used as immunotherapies, such as to target cancer cells.

[0197] Gene editing in animals and plants The systems and methods described above can be used to generate transgenic non-human animals or plants with one or more genetic modifications of interest. In some embodiments, the transgenic non-human animals are homozygous for the genetic modifications. In some embodiments, the transgenic non-human animals are heterozygous for the genetic modifications. In some embodiments, the transgenic non-human animals are vertebrates, such as fish (e.g., zebrafish, goldfish, pufferfish, cavefish, etc.), amphibians (frogs, salamanders, etc.), birds (e.g., chickens, turkeys, etc.), reptiles (e.g., snakes, lizards, etc.), mammals (e.g., ungulates, such as pigs, cows, goats, sheep, etc.); lagomorphs (e.g., rabbits); rodents (e.g., rats, mice); or non-human primates.

[0198] The present invention can be used to treat animal diseases in a manner similar to that used to treat human diseases as described above. Alternatively, the present invention can be used to generate knock-in animal disease models with specific gene mutations for research, drug discovery, and target validation. The systems and methods described above can also be used to introduce point mutations into ES cells or embryos of various organisms to breed and improve animal breeds and crop quality.

[0199] The method for introducing exogenous nucleic acid into plant cells is well known in the art.Suitable methods include virus infection (such as double-stranded DNA virus), transfection, conjugation, protoplast fusion, electroporation, gene gun technology, calcium phosphate precipitation, direct microinjection, silicon carbide whisker technology, and Agrobacterium-mediated transformation.The choice of method generally depends on the type of cell to be transformed and the situation in which transformation occurs (that is, in vitro, ex vivo, or in vivo).

[0200] kit The present invention further provides kits containing reagents for carrying out the methods described above, including CRISPR / Cas-guided target binding or modification reactions. To this end, the reaction components for the methods disclosed herein, such as one or more of RNA, Cas protein, fusion effector protein, and related nucleic acid, can be provided in the form of a kit for use. In one embodiment, the kit includes a nucleic acid encoding a CRISPR protein or Cas protein, an effector protein, one or more of the RNA scaffolds described above, and a set of RNA molecules described above. In other embodiments, the kit may also include one or more other reaction components. In such kits, appropriate amounts of one or more reaction components are provided in one or more containers or supported on a substrate.

[0201] Examples of additional components of the kit include, but are not limited to, one or more host cells, one or more reagents for introducing exogenous nucleotide sequences into the host cells, one or more reagents for detecting RNA or protein expression or verifying the status of the target nucleic acid (e.g., probes or PCR primers), and reaction buffers or culture media (in 1× or concentrated form). The kit may also include one or more of the following components: supports, termination, modification, or digestion reagents, osmolytes, and detection equipment.

[0202] The reaction components used can be provided in various forms. For example, the components (e.g., enzymes, RNA, probes, and / or primers) can be suspended in an aqueous solution or can be present as freeze-dried or lyophilized powders, pellets, or beads. In the latter case, the components, upon reconstitution, form a complete mixture of components for use in the assay. The kits of the present invention can be provided at any suitable temperature. For example, for storage of kits containing protein components or their complexes in liquid, they are preferably provided and maintained at temperatures below 0°C, preferably -20°C or below, or otherwise frozen.

[0203] The kit or system may include any combination of components described herein in amounts sufficient for at least one assay. In some applications, one or more reaction components may be provided in pre-measured, single-use amounts in individual, typically disposable, tubes or equivalent containers. In such an arrangement, the RNA-guided reaction can be performed by directly adding the target nucleic acid, or a sample or cells containing the target nucleic acid, to the individual tubes. The amount of components provided in the kit may be any suitable amount and may depend on the target market for the product. The containers in which the components are provided may be any conventional container capable of retaining the provided form, such as a microcentrifuge tube, a microtiter plate, an ampoule, a bottle, or an integrated test device such as a fluidic device, cartridge, lateral flow, or other similar device.

[0204] The kit may also include packaging for holding the container or combination of containers. Typical packaging for such kits and systems includes solid matrices (e.g., glass, plastic, paper, foil, microparticles, etc.) that hold the reaction components or detection probes in any of a variety of configurations (e.g., vials, microtiter plate wells, microarrays, etc.). The kit may further include tangible recorded instructions for using the components.

[0205] definition Nucleic acid or polynucleotide refers to a DNA molecule (such as, but not limited to, cDNA or genomic DNA) or an RNA molecule (such as, but not limited to, mRNA), including DNA or RNA analogs. DNA or RNA analogs can be synthesized from nucleotide analogs. DNA or RNA molecules may contain non-naturally occurring moieties, such as modified bases, modified backbones, and deoxyribonucleotides of RNA. Nucleic acid molecules may be single-stranded or double-stranded.

[0206] The term "isolated," when referring to a nucleic acid molecule or polypeptide, means that the nucleic acid molecule or polypeptide is substantially free from at least one other component with which it is associated or found in nature.

[0207] As used herein, the term "guide RNA" generally refers to an RNA molecule (or collectively a group of RNA molecules) that can bind to a CRISPR protein and target the CRISPR protein to a specific location within a target DNA. A guide RNA may contain two segments: a DNA-targeting guide segment and a protein-binding segment. The DNA-targeting segment contains a nucleotide sequence that is complementary to (or at least capable of hybridizing under stringent conditions with) the target sequence. The protein-binding segment interacts with a CRISPR protein, such as Cas9 or a Cas9-associated polypeptide. These two segments may be located in the same RNA molecule or in two or more separate RNA molecules. When the two segments are present in separate RNA molecules, the molecule containing the DNA-targeting guide segment may be referred to as a CRISPR RNA (crRNA), and the molecule containing the protein-binding segment is referred to as a trans-activating RNA (tracrRNA).

[0208] As used herein, the term "target nucleic acid" or "target" refers to a nucleic acid containing a target nucleic acid sequence. A target nucleic acid can be single-stranded or double-stranded, and is often double-stranded DNA. As used herein, a "target nucleic acid sequence," "target sequence," or "target region" refers to a specific sequence or its complement that is to be bound or modified using a CRISPR system. A target sequence can be in any form of single-stranded or double-stranded nucleic acid, and can be present in an in vitro or in vivo nucleic acid in the genome of a cell.

[0209] "Target nucleic acid strand" refers to the strand of the target nucleic acid that is the target of base pairing with the guide RNA as disclosed herein. That is, the strand of the target nucleic acid that hybridizes with the crRNA and guide sequence is referred to as the "target nucleic acid strand." The other strand of the target nucleic acid that is not complementary to the guide sequence is referred to as the "non-complementary strand." In the case of a double-stranded target nucleic acid (e.g., DNA), each strand can be a "target nucleic acid strand" for designing the crRNA and guide RNA and can be used to practice the methods of the present invention as long as a suitable PAM site is present.

[0210] As used herein, the term "derived from" refers to a process in which a first component (e.g., a first molecule) or information from that first component is used to isolate, derive, or create a different second component (e.g., a second component that is different from the first). For example, mammalian codon-optimized Cas9 polynucleotides are derived from the wild-type Cas9 protein amino acid sequence. Also, variant mammalian codon-optimized Cas9 polynucleotides, including Cas9 single mutant nickases (nCas9 such as nCas9D10A) and Cas9 double mutant null-nucleases (dCas9 such as dCas9D10A H840A), are derived from polynucleotides encoding wild-type mammalian codon-optimized Cas9 proteins.

[0211] As used herein, the term "wild-type" is a term of the art understood by those skilled in the art and means the typical form of an organism, strain, gene, or trait occurring in nature, as distinguished from mutant or variant forms.

[0212] As used herein, the term "variant" refers to a first composition (e.g., a first molecule) relative to a second composition (e.g., a second molecule, also referred to as a "parent" molecule). A variant molecule may be derived from, isolated from, based on, or homologous to a parent molecule. For example, mutant forms of mammalian codon-optimized Cas9 (hspCas9), including a Cas9 single mutant nickase and a Cas9 double mutant null-nuclease, are variants of mammalian codon-optimized wild-type Cas9 (hspCas9). The term variant can be used to describe either a polynucleotide or a polypeptide.

[0213] When applied to polynucleotides, a variant molecule may have complete nucleotide sequence identity with the original parent molecule, or alternatively, may have less than 100% nucleotide sequence identity with the parent molecule. For example, a variant of a gene nucleotide sequence may be a second nucleotide sequence that is at least 50%, 60%, 70%, 80%, 90%, 95%, 98%, 99%, or more highly identical in nucleotide sequence to the original nucleotide sequence. Polynucleotide variants also include polynucleotides that contain the entire parent polynucleotide and further contain additional fused nucleotide sequences. Polynucleotide variants also include polynucleotides that are portions or subsequences of the parent polynucleotide, for example, unique subsequences of the polynucleotides disclosed herein (e.g., as determined by standard sequence comparison and alignment techniques) are also encompassed by the present invention.

[0214] In another aspect, polynucleotide variants include nucleotide sequences containing minor, trivial, or insignificant changes relative to the parent nucleotide sequence. For example, minor, trivial, or insignificant changes include (i) changes to a nucleotide sequence that do not change the amino acid sequence of the corresponding polypeptide, (ii) changes to a nucleotide sequence that occur outside the protein-encoding open reading frame of the polynucleotide, (iii) changes to a nucleotide sequence that result in deletions or insertions that may affect the corresponding amino acid sequence but have little or no effect on the biological activity of the polypeptide, and (iv) nucleotide changes that result in the substitution of an amino acid with a chemically similar amino acid. If the polynucleotide does not encode a protein (e.g., tRNA, crRNA, or tracrRNA), a variant of the polynucleotide may contain nucleotide changes that do not result in loss of function of the polynucleotide. In another aspect, conservative variants of the nucleotide sequences of the present disclosure that yield functionally identical nucleotide sequences are encompassed by the present invention. One of skill in the art will appreciate that numerous variants of the nucleotide sequences of the present disclosure are encompassed by the present invention.

[0215] When applied to proteins, a variant polypeptide may have complete amino acid sequence identity with the original parent polypeptide, or alternatively, may have less than 100% amino acid identity with the parent protein. For example, an amino acid sequence variant may be a second amino acid sequence that is at least 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, or more identical in amino acid sequence compared to the original amino acid sequence.

[0216] Polypeptide variants include polypeptides that contain the entire parent polypeptide and further contain additional fused amino acid sequences. Polypeptide variants also include polypeptides that are portions or subsequences of the parent polypeptide, for example, unique subsequences of the polypeptides disclosed herein (e.g., as determined by standard sequence comparison and alignment techniques) are also encompassed by the present invention.

[0217] In another embodiment, polypeptide variants include polypeptides containing minor, insignificant, or insignificant changes to the parent amino acid sequence. For example, minor, insignificant, or insignificant changes include amino acid changes (including substitutions, deletions, and insertions) that have little or no effect on the biological activity of the polypeptide and produce a functionally identical polypeptide, including the addition of non-functional peptide sequences. In other embodiments, variant polypeptides of the invention alter the biological activity of the parent molecule; for example, mutant variants of Cas9 polypeptides have modified or lost nuclease activity. Those of skill in the art will appreciate that numerous variants of the polypeptides of the present disclosure are encompassed by the present invention.

[0218] In some aspects, polynucleotide or polypeptide variants of the invention may include variant molecules in which a small percentage of the nucleotide or amino acid positions have been changed, added, or deleted, e.g., typically less than about 10%, less than about 5%, less than 4%, less than 2%, or less than 1%.

[0219] As used herein, the term "conservative substitution" in a nucleotide or amino acid sequence refers to a change in a nucleotide sequence that either (i) does not result in any corresponding change in the amino acid sequence due to redundancy in the triplet codon code, or (ii) results in the substitution of the original parent amino acid with an amino acid having a chemically similar structure. Conservative substitution tables providing functionally similar amino acids are well known in the art, where one amino acid residue is substituted with another amino acid residue having similar chemical properties (e.g., an aromatic side chain or a positively charged side chain), thus leaving the functional properties of the resulting polypeptide molecule substantially unchanged.

[0220] The following are groupings of naturally occurring amino acids with similar chemical properties, with substitutions within the group being "conservative" amino acid substitutions. The groupings shown below are not strict, as these naturally occurring amino acids may be placed in different groups when different functional properties are taken into account. Amino acids with nonpolar and / or aliphatic side chains include glycine, alanine, valine, leucine, isoleucine, and proline. Amino acids with polar, uncharged side chains include serine, threonine, cysteine, methionine, asparagine, and glutamine. Amino acids with aromatic side chains include phenylalanine, tyrosine, and tryptophan. Amino acids with positively charged side chains include lysine, arginine, and histidine. Amino acids with negatively charged side chains include aspartic acid and glutamic acid.

[0221] A "Cas9 mutant" or "Cas9 variant" refers to a protein or polypeptide derivative of a wild-type Cas9 protein, such as the S. pyogenes Cas9 protein (i.e., SEQ ID NO: 1), e.g., a protein having one or more point mutations, insertions, deletions, truncations, fusion proteins, or combinations thereof. A "Cas9 mutant" or "Cas9 variant" substantially retains the RNA targeting activity of the Cas9 protein. The protein or polypeptide may comprise, consist of, or consist essentially of a fragment of SEQ ID NO: 1. Generally, the mutant / variant is at least 50% identical to SEQ ID NO: 1 (e.g., any number between 50% and 100%). The mutant / variant can bind to an RNA molecule and target specific DNA sequences via the RNA molecule, and may further possess nuclease activity. Examples of such domains include the RuvC-like motif (amino acids 7-22, 759-766, and 982-989 of SEQ ID NO: 1) and the HNH motif (amino acids 837-863). See Gasiunas et al., Proc Natl Acad Sci USA. 2012 Sep 25;109(39):E2579-2586 and WO 2013176772.

[0222] "Complementarity" refers to the ability of a nucleic acid to form hydrogen bonds with another nucleic acid sequence, either by conventional Watson-Crick base pairing or other non-conventional methods. The percent complementarity indicates the percentage of residues in a nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with a second nucleic acid sequence (e.g., 5, 6, 7, 8, 9, 10 out of 10 are 50%, 60%, 70%, 80%, 90%, and 100% complementary). "Fully complementary" means that all contiguous residues of a nucleic acid sequence will hydrogen bond with the same number of contiguous residues in a second nucleic acid sequence. "Substantially complementary," as used herein, refers to a degree of complementarity that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% over a region of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, or more nucleotides, or refers to two nucleic acids that hybridize under stringent conditions.

[0223] As used herein, "stringent conditions" for hybridization refer to conditions under which a nucleic acid having complementarity to a target sequence hybridizes primarily to the target sequence and does not substantially hybridize to non-target sequences. Stringent conditions are generally sequence-dependent and vary depending on numerous factors. Generally, the longer the sequence, the higher the temperature at which that sequence specifically hybridizes to its target sequence. Non-limiting examples of stringent conditions are described in detail in Tijssen (1993), Laboratory Techniques in Biochemistry and Molecular Biology—Hybridization with Nucleic Acid Probes Part I, Chapter 2, "Overview of principles of hybridization and the strategy of nucleic acid probe assay," Elsevier, NY.

[0224] "Hybridization" or "hybridizing" refers to the process by which fully or partially complementary nucleic acid strands combine under specified hybridization conditions to form a double-stranded structure or region in which the two constituent strands are joined by hydrogen bonds. Typically, hydrogen bonds are formed between adenine and thymine or uracil (A and T or U) or cytidine and guanine (C and G), although other base pairs may also be formed (see, e.g., Adams et al., The Biochemistry of the Nucleic Acids, 11th ed., 1992).

[0225] As used herein, "expression" refers to the process by which a polynucleotide is transcribed from a DNA template (into mRNA or other RNA transcripts) and / or the process by which the transcribed mRNA is subsequently translated into a peptide, polypeptide, or protein. The transcript and the encoded polypeptide are sometimes collectively referred to as a "gene product." If the polynucleotide is derived from genomic DNA, expression may also include splicing of the mRNA in a eukaryotic cell.

[0226] The terms "polypeptide," "peptide," and "protein" are used interchangeably herein to refer to polymers of amino acids of any length. The polymers may be linear or branched, may comprise modified amino acids, and may be interrupted by non-amino acids. The terms also encompass amino acid polymers that have been modified by, for example, disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, pegylation, or any other manipulation, such as conjugation with a labeling component. As used herein, the term "amino acid" includes natural and / or unnatural or synthetic amino acids, including glycine and both D- and L-optical isomers, as well as amino acid analogs and peptidomimetics.

[0227] The term "fusion polypeptide" or "fusion protein" refers to a protein created by joining two or more polypeptide sequences together. Fusion polypeptides encompassed by the present invention include the translation product of a chimeric gene construct in which a nucleic acid sequence encoding one polypeptide, e.g., an RNA-binding domain, is joined to a nucleic acid sequence encoding a second polypeptide, e.g., an effector domain, so as to form a single open reading frame. In other words, a "fusion polypeptide" or "fusion protein" is a recombinant protein of two or more proteins joined by a peptide bond or via several peptides. Fusion proteins may also contain a peptide linker between the two domains.

[0228] The term "linker" refers to any means, entity, or moiety used to join two or more entities. Linkers can be covalent or non-covalent. Examples of covalent linkers include covalent bonds or linker moieties that are covalently attached to one or more of the proteins or domains to be linked. Linkers can also be non-covalent, e.g., organometallic bonds via a metal center such as a platinum atom. Covalent bonds can use a variety of functionalities, such as amide groups including carbonic acid derivatives, ethers, esters including organic and inorganic esters, amino, urethane, and urea. To provide linkages, domains can be modified by oxidation, hydroxylation, substitution, reduction, etc. to provide sites for coupling. Methods for conjugation are well known to those of skill in the art and are encompassed for use in the present invention. Linker moieties include, but are not limited to, chemical linker moieties or, for example, peptide linker moieties (linker sequences). It will be understood that modifications that do not significantly reduce the function of the RNA-binding domain and effector domain are preferred.

[0229] As used herein, the term "conjugate" or "conjugation" or "linked" refers to the joining of two or more entities to form one entity. Conjugates include both peptide-small molecule conjugates as well as peptide-protein / peptide conjugates.

[0230] The terms "subject" and "patient" are used interchangeably herein and refer to a vertebrate, preferably a mammal, more preferably a human. Mammals include, but are not limited to, rodents, monkeys, humans, livestock, sport animals, and pets. Also encompassed are tissues, cells, and their progeny of biological entities obtained in vivo or cultured in vitro. In some embodiments, the subject may be an invertebrate, such as an insect or nematode; in other cases, the subject may be a plant or a fungus.

[0231] As used herein, "treatment" or "treating," or "alleviating" or "ameliorating" are used interchangeably. These terms refer to an approach for obtaining a beneficial or desired result, including, but not limited to, therapeutic benefit and / or prophylactic benefit. Therapeutic benefit refers to any therapeutically relevant improvement or effect on one or more diseases, conditions, or symptoms being treated. In the case of prophylactic benefit, the composition can be administered to a subject at risk of developing a particular disease, condition, or symptom, or to a subject who experiences one or more physiological symptoms of a disease, even if the disease, condition, or symptom has not yet manifested.

[0232] The phrase "pharmaceutically or pharmacologically acceptable" refers to molecular entities and compositions that do not produce adverse, allergic, or other unexpected reactions when administered to an animal, such as a human, as appropriate. The preparation of pharmaceutical compositions containing therapeutic agents, such as cells, or additional active ingredients will be known to those of skill in the art in light of the present disclosure. Furthermore, it will be understood that for animal (e.g., human) administration, formulations should meet sterility, pyrogenicity, general safety, and purity standards required by the FDA Office of Biological Standards. As used herein, "pharmaceutically acceptable carriers" include any and all aqueous solvents (e.g., water, alcoholic / aqueous solutions, saline, parenteral vehicles such as sodium chloride, Ringer's dextrose, etc.), non-aqueous solvents (e.g., propylene glycol, polyethylene glycol, vegetable oils, and injectable organic esters such as ethyl oleate), dispersion media, coatings, surfactants, antioxidants, preservatives (e.g., antibacterial or antifungal agents, antioxidants, chelating agents, and inert gases), isotonic agents, absorption delaying agents, salts, drugs, drug stabilizers, gels, binders, excipients, disintegrants, lubricants, sweeteners, flavoring agents, dyes, liquids, and nutritional supplements, the like and combinations thereof that would be known to those skilled in the art. The pH and exact concentration of the various components in the pharmaceutical composition are adjusted according to well-known parameters.

[0233] As used herein, the term "contacting," when referring to any set of components, includes any process of mixing the components to be contacted into the same mixture (e.g., adding them to the same compartment or solution), and does not necessarily require actual physical contact between the listed components. The listed components may be contacted in any order or in any combination (or subcombination), and may include situations in which one or some of the listed components are subsequently removed from the mixture, optionally before adding other listed components. For example, "contacting A with B and C" includes any or all of the following situations: (i) mixing A with C and then adding B to the mixture; (ii) mixing A and B into a mixture, removing B from the mixture, and then adding C to the mixture; or (iii) adding A to a mixture of B and C. "Contacting" a target nucleic acid or cell with one or more reaction components, such as a Cas protein or guide RNA, includes any or all of the following situations: (i) contacting the target or cell with a first component of a reaction mixture to create a mixture, and then adding other components of the reaction mixture to the mixture in any order or combination; and (ii) completely forming the reaction mixture before mixing with the target or cell.

[0234] The term "mixture," as used herein, refers to a combination of elements that are dispersed and in no particular order. A mixture is heterogeneous and cannot be spatially separated into its different components. Examples of mixtures of elements include those in which several different elements are dissolved in the same aqueous solution, or those in which several different elements are bound to a solid support randomly or in no particular order, and the different elements are not spatially distinguished. In other words, a mixture is not locatable.

[0235] As disclosed herein, numerous numerical ranges are provided. Unless the context clearly dictates otherwise, each intervening value between the upper and lower limits of that range, to the tenth of the unit of the lower limit, is understood to be specifically disclosed. Each smaller range between any stated value or intervening value in a stated range and any other stated or intervening value within that stated range is encompassed within the invention. The upper and lower limits of such smaller ranges may independently be included or excluded, and each range where either or both limits are included in the smaller range, or where neither limit is included in the smaller range, is also encompassed within the invention, subject to any specifically excluded limit in the stated range. Where a stated range includes one or both limits, ranges excluding either or both of those included limits are also included within the invention. The term "about" generally refers to plus or minus 10% of the stated numerical value. For example, "about 10%" can indicate a range of 9% to 11%, and "about 20" can mean 18 to 22. Other meanings of "about" may be apparent from the context, such as rounding, so for example, "about 1" may mean 0.5 to 1.4. [Example]

[0236] Example 1 Materials and Methods This example describes the materials and methods used in Examples 2-12 below.

[0237] Bacterial strains E. coli DH5α competent cells were purchased from THERMO FISHER (catalog number 18265017) and used for basic cloning. The E. coli MG1655 strain used for rpoB gene targeting was kindly provided by Dr. Stanley Qi (Stanford University). MG1655 cells were made competent using a standard CaCl2 protocol.

[0238] Bacterial expression plasmids pgRNA-bacteria (pUC19, ampicillin resistance; ADDGENE Plasmid #44251) was engineered to contain two offset BbsI restriction sites for guide sequence cloning and one or two MS2 stem-loop sequences at the 3' end. These modifications were introduced using a standard gene synthesis service (GENEWIZ; South Plainfield, NJ, USA). The synthesized cassette was cloned into the pUC19 backbone using the SpeI and HindIII restriction sites. The effector module (AID-linker-MCP) was cloned into the pCDFDuet empty vector (DF13, streptomycin resistance; ADDGENE Plasmid #49796) using the BglII and BamHI restriction sites. The dCas9-bacteria plasmid (p15A, chloramphenicol resistance; ADDGENE #44249) and pwtCas9-bacteria (p15A; ADDGENE #44250) were used to generate nCas9 by exchanging portions of the wild-type HNH and RuvC active sites derived from dCas9 from pwtCas9, respectively. D10A and nCas9 H840A Nickase was generated. The HNH domain was cloned using the Acc65I and BamHI restriction sites. The RuvC domain was cloned using the XbaI and NheI restriction sites. The Cas9 and effector constructs are under the control of a tetracycline-inducible promoter.

[0239] bacterial gRNA design RpoB-targeting gRNAs were manually designed in SNAPGENE VIEWER (GSL BIOTECH) at or near the rifampicin resistance-determining region (RRDR) of the E. coli rpoB gene. (23) The gRNA sequence and PAM are summarized in Table S1. The guide sequence was designed to have a 5' overhang compatible with the overhang left by BbsI digestion (i.e., Fwd5'-CTAGN 20 -3' (SEQ ID NO: 84), Rev5'-AAACN 20 -3' (SEQ ID NO: 85), where N 20is the programmable guide sequence, which must be complementary between the Fwd and Rev oligos).

[0240] [Table 4]

[0241] Bacterial Treatment: Chemically competent E. coli MG1655 cells were transformed with 9 ng of a 1:1:1 combination of the appropriate plasmids encoding specific gRNA (ampicillin), AID_MCP (streptomycin), and Cas9 (chloramphenicol) constructs. After transformation, cells were selected overnight in liquid LB medium containing working concentrations of ampicillin, streptomycin, and chloramphenicol. The next day, cells were diluted into selective medium supplemented with 3 μM tetracycline to induce expression of the protein-coding modules. After overnight growth, OD was measured and serial dilutions were performed to obtain 10 8 ~10 3 Cells are plated onto LB agar containing rifampicin. Plates are incubated at 37° C. and monitored for 48 hours. Percent survival is calculated by dividing the number of surviving colonies by the number of cells plated.

[0242] Mutation analysis in bacterial experiments Genomic DNA was extracted from 8 to 12 colonies from the appropriate experiment. The target region of the rpoB gene (i.e., the RRDR region) was PCR amplified. The purified PCR products were sequenced by GENEWIZ (South Plainfield, NJ, USA) using Sanger chemistry. Primer sequences are summarized in the table below.

[0243] [Table 5]

[0244] Mammalian expression plasmids.

[0245] A CRCn, A CRCnu, and A1 To generate the CRCnu polycistronic construct, an AID_MCP fusion or an APOBEC1_MCP fusion was synthesized at GENWIZ (South Plainfield, NJ, USA) and cloned upstream of nCas9_UGI (13). The two modules are separated by a self-cleaving T2A peptide. A To generate CRCnu.2, the construct was codon-optimized and an additional copy of UGI was included downstream of concas9 (29). To generate the gRNA_2×MS2 vector, a gRNA scaffold (15) fused to two MS2 loops was synthesized at GENEWIZ (South Plainfield, NJ, USA) and cloned into phU6_gRNA (ADDGENE Plasmid #53188) (49). The nfEGFP gene, which contains an A to G mutation at nucleotide 200 of the GFP gene, was synthesized at GENEWIZ (South Plainfield, NJ, USA) and cloned into the pCMV_Sports6 vector using the SalI and NotI restriction sites.

[0246] gRNA design Targeting gRNAs were manually designed using SNAPGENE VIEWER (GSL BIOTECH). All gRNAs used in this study are listed in Tables S3 and S4.

[0247] [Table 6]

[0248] [Table 7]

[0249] cell culture HEK293T cells were purchased from ATCC (CRL-3216). Transgenic EGFP reporters were generated by standard lentiviral transduction of HEK293T cells and selected with puromycin. Cells expressing GFP variants were obtained by limiting dilution. Cells were grown and maintained at 37°C and 5% CO2 in Dulbecco's modified Eagle's medium (DMEM, THERMOFISHER) supplemented with 10% fetal bovine serum, 1x glutamine (THERMOFISHER), and 1x antibiotic-antimycotic (THERMOFISHER).

[0250] process HEK293T and its derivatives nf2.16 or 293_GFP cells were seeded into 6-well plates (3.5 × 10 per well) the day before the experiment. 5 Transfection was performed on 75-85% confluent cells using a combination of DNA from the CRC and gRNA constructs, each at a 3:1 ratio, totaling 2 µg. LIPOFECTAMINE 2000 (THERMOFISHER) or LIPOFECTAMINE 3000 was used as the transfection reagent according to the manufacturer's instructions. Fluorescent images were taken 72 hours after transfection, as needed, and GFP signals were quantified by flow cytometry on a Gallios flow cytometer (BECKMAN COULTER) at the Rutgers University Flow Cytometry Core Facility. For knockout experiments, cells were passaged and cultured for an additional 96 hours to allow for GFP turnover in the treated cells. After treatment, DNA was purified for downstream analysis using the DNEASY BLOOD AND TISSUE kit (QIAGEN).

[0251] FACS analysis nf2.16 cells, ACRCnu / nfEGFP_NT1 was transfected. 72 hours after transfection, GFP-positive cells were sorted using a BECKMAN COULTER MOFLO XDP cell sorting instrument at the Rutgers University Flow Cytometry Core Facility according to the manufacturer's instructions. Sorted cells expressing wild-type GFP were cultured, and DNA was extracted using the DNEASY BLOOD AND TISSUE kit (QIAGEN). The target region was amplified by PCR and subsequently Sanger sequenced by GENEWIZ (NJ, USA). The primers used for PCR were the same as those used for high-throughput sequencing analysis (see below and Table S6).

[0252] Whole exome sequencing (WES) WES was performed by GENEWIZ (South Plainfield, NJ, USA). WES libraries were constructed using the AGILENT SURESELECT HUMAN ALL EXON (V6r2) Library Preparation Kit and sequenced in a paired-end 2 × 150 bp format using the ILLUMINA HISEQ. To estimate potential CRC off-target activity, the raw data were analyzed as follows.

[0253] Variant calling and alternative reference structures WES raw reads were aligned to the human reference genome (hg38) using BWA (version 0.7.15). Variants were identified using the GENOME ANALYSIS TOOL Kit (GATK) version 3.8, largely following GATK best practices. Briefly, duplicate reads were first flagged using Picard MARKDUPLICATES. Base quality was recalibrated using BASERECALIBRATOR, and then variants for each sample were called using HAPLOTYPECALLER, followed by joint genotyping with GENOTYPEGVCFS. Variants detected in the resulting VCF files were further recalibrated using VARIANTRECALIBRATOR.

[0254] For downstream analysis, we focused only on exon regions as defined in "SURESELECT HUMAN ALL EXON V6r2." In the analysis, overlapping regions were merged using the merge function in bedtools.

[0255] To construct an alternative reference based on the parent cell line T6, we extracted all variants genotyped in T6. Using GATK3.8 FASTAALTERNATEREFERENCEMAKER with default options, we constructed alternative reference sequences for the exon regions specified in the merged exon target file.

[0256] Motif definition and mutation analysis The AID "WRCH" binding motif represents the product of ["AT", "AG", "C", "ACT"] and the coordinates of any four such consecutive nucleotides were saved. We used Python to identify and extract the genomic locations of WRCH motifs within a reference FASTA sequence (either hg38 or an alternative reference). The reference FASTA sequence was also scanned for sequences complementary to "WRCH", i.e., "DGYW", given by the product of ["AGT", "G", "CT", "AT"]. A non-WRCH motif was defined as a four-nucleotide sequence with a C in the third position that is not a WRCH. Similarly, a non-DGYW motif is any four-nucleotide sequence with a G in the second position that is not a DGYW. In total, there are 12 possible WRCH motifs, 12 DGYW motifs, 52 non-WRCH motifs, and 52 non-DGYW motifs. In the mutation analysis, WRCH and DGYW classifications were investigated separately. When searching for potential AID-derived mutation sites, C>T changes are classified as WRCH motif mutations or non-WRCH motif mutations based on their surrounding bases. Similarly, G>A changes are classified as DGYW motif mutations or non-DGYW motif mutations based on their surrounding bases.

[0257] Putative CRISPR off-target regions We scanned the reference genome hg38 for putative loci targeted by CRISPR gRNAs using CCTop (https: / / crispr.cos.uni-heidelberg.de / ) and CRISPRDesign (http: / / crispr.mit.edu / ). In total, we obtained 54 putative off-target regions and extracted variants within these regions.

[0258] High-throughput sequencing analysis The sequences of the primers used in this study are summarized in Table S6. All PCR amplifications were performed using high-fidelity PHUSION hot-start DNA polymerase (NEW ENGLAND BIOLABS) according to the manufacturer's instructions. PCR products were purified with a QIAQUICK PCR Purification Kit (QIAGEN) and submitted to GENEWIZ (South Plainfield, NJ, USA) for high-throughput sequencing analysis. Data analysis, particularly single-nucleotide polymorphism (SNP) and insertion / deletion (indel) frequencies, was performed by GENEWIZ personnel using a proprietary pipeline. Sequence output was used to generate SNP and indel frequency figures.

[0259] Whole-exome sequencing analysis DNA Library Preparation and HiSeq Sequencing. Initial DNA sample quality assessment, DNA library preparation, sequencing, and bioinformatics analysis were performed at GENEWIZ, Inc. (South Plainfield, NJ, USA). Genomic DNA samples were quantified using a QUBIT 2.0 fluorometer (LIFE TECHNOLOGIES, Carlsbad, CA, USA), and DNA integrity was examined on a 0.6% agarose gel with 50 ng of sample loaded per lane. The SURESELECTXT EXOME ENRICHMENT SYSTEM for ILLUMINA paired-end multiplexed sequencing libraries and the SURESELECT HUMAN ALL EXON V5 bait library were used for target-enriched DNA library preparation (for 200 ng starting material) according to the manufacturer's recommendations (AGILENT, Santa Clara, CA, USA) and standard low-input protocols. Briefly, genomic DNA was fragmented by acoustic shearing using a COVARIS LE200 focused sonication instrument. The fragmented DNA was cleaned, end-repaired, and 3'-end adenylated. Adapters were ligated to the DNA fragments, and the adapter-ligated DNA fragments were enriched by limited-cycle PCR. The adapter-ligated DNA fragments were verified using an AGILENT TAPESTATION system (AGILENT TECHNOLOGIES, Palo Alto, CA, USA) and quantified using a QUBIT 2.0 fluorometer. 750 ng of adapter-ligated DNA fragments were hybridized with biotinylated RNA baits at 65°C for 24 hours. The hybrid DNA was captured using streptavidin-coated magnetic beads. After extensive washing, the captured DNA was amplified and indexed using ILLUMINA indexing primers. The captured DNA library was verified using an AGILENT TAPESTATION system and quantified using a QUBIT 2.0 fluorometer and real-time PCR (APPLIED BIOSYSTEMS, Carlsbad, CA, USA).

[0260] DNA Library Sequencing: ILLUMINA reagents and kits for cluster generation and sequencing were used for enriched DNA sequencing. The captured DNA libraries were multiplexed at equimolar mass, and the pooled DNA libraries were clustered onto two lanes of a flow cell using an ILLUMINA cBOT. After clustering, the flow cell was loaded onto an ILLUMINA HISEQ instrument according to the manufacturer's instructions. Samples were sequenced using a 2x150 paired-end (PE) configuration. Image analysis and base calling were performed using the HISEQ instrument's HiSeq control software (HCS 2.0).

[0261] High-throughput sequencing analysis Library preparation. DNA library preparation and ILLUMINA sequencing. DNA library preparation, sequencing reactions, and initial bioinformatics analysis were performed at GENEWIZ, Inc. (South Plainfield, NJ, USA). DNA amplicons were indexed and enriched by limited-cycle PCR. DNA libraries were validated using TapeStation (Agilent Technologies, Palo Alto, CA, USA) and quantified using a QUBIT 2.0 fluorometer and real-time PCR (APPLIED BIOSYSTEMS, Carlsbad, CA, USA). Pooled DNA libraries were loaded into the ILLUMINA instrument according to the manufacturer's instructions. Samples were sequenced using a 2 × 250 paired-end (PE) configuration. Image analysis and base calling were performed using the ILLUMINA instrument's ILLUMINA Control Software (HCS).

[0262] Data Analysis. Raw ILLUMINA reads were examined for adapters and quality using FASTQC. Raw ILLUMINA sequencing reads were trimmed of poor quality adapters and nucleotides using TRIMMOMATIC v.0.36. Forward and reverse reads were then overlapped, and paired sequence reads were merged to form a single sequence if the overlapping region was identical using the reformat function within bbmap. Merged reads were aligned to a reference sequence, and variant detection was performed using GENEWIZ's proprietary AMPLICON-EZ program.

[0263] Example 2 CRC System: Modular Base Editing Platform The CRC base editing system consists of three functional modules, shown in Figures 1A and 1B: (1) a nuclease-deficient Cas9 protein; (2) a programmable chimeric RNA scaffold containing a gRNA (for sequence recognition [2.1] and Cas9 binding [2.2]) and a recruiting RNA aptamer (for effector module recruitment [2.3]); and (3) an effector module containing a cytidine deaminase (effector [3.1]) fused to an RNA aptamer ligand, a small RNA-binding protein [3.2] that specifically interacts with the recruiting RNA aptamer. The initial prototype system consisted of a catalytically inactive Cas9 protein (dCas9, containing mutations D10A and H840A that abolish its nuclease activity), an RNA aptamer derived from the operator stem-loop of bacteriophage MS2 (MS2) synthetically fused to the 3' end of a gRNA scaffold, and a bacterial vector expressing human activation-induced cytidine deaminase (AID) fused to the MS2 coat protein (MCP), which interacts with MS2 (Figure 7). While the effector is shown as a monomer in Figure 1, AID or other effectors can form functional oligomers at the site of action within cells.

[0264] Example 3 Proof of concept for CRC in prokaryotic cells We tested the system in bacteria using a negative selection approach with the antibiotic rifampicin. Rifampicin inhibits transcription by binding near the catalytic pocket of the β subunit of bacterial RNA polymerase, encoded by the rpoB gene, and physically blocking RNA elongation (22). We determined that mutations along a specific segment of the rpoB gene are associated with rifampicin resistance. This region is known to be the rifampicin resistance-determining region (RRDR; Figure 1C) (23).

[0265] For these experiments, we designed four gRNAs to target the template strand (TS1–TS4; Figure 1C, Table S1) using catalytically inactive Cas9 (dCas9) as the DNA targeting module and one MS2 motif as the recruitment module. We developed a system expressing AID_MCP and dCas9 as the effector and targeting modules, respectively. A Denoted as CRCd. Guided by gRNA TS4 A Treatment with CRCd resulted in a 35-fold higher survival rate than scrambled-treated cells (Figures 1D and 1E). A Sequence analysis of isolated colonies treated with CRCd / rpoB_TS4 revealed that this system introduced a targeted C-to-T mutation at codon 531, changing a serine to a phenylalanine, a mutation known to induce rifampicin resistance (23, 24) (Figure 1F). The higher efficiency observed in TS4-treated cells may be due to the location of the targeted C within the protospacer (the unpaired DNA strand within the CRISPR R-loop), in this case at position 8 from the 5' end of the protospacer. On the other hand, TS2 and TS3 targeted Cs at positions 12 and 14, respectively, suggesting a preference for positions distal to the PAM motif within the protospacer region.

[0266] Collectively, these data demonstrate that targeted nucleotide modification using an RNA-aptamer-based effector recruitment mechanism represents a potentially viable approach for targeted base editing.

[0267] Example 4: Manipulation of individual modules for system optimization The positive results of the above exploratory experiments prompted us to further engineer the CRC system to increase its targeting efficiency using gRNA rpoB_TS4 for comparison. First, we engineered the Cas9 module into a dCas9 ( A Nickase Cas9 creates a single-strand DNA break (nick) in the complementary strand of the base editing target from CRCd D10A By replacing it with A This resulted in a 4.6-fold increase in the number of viable colonies compared to CRCd (Figure 2A). H840A ( A CRC H840A ) processing is A Compared to CRCd, this resulted in a modest improvement in editing efficiency, with a less than two-fold increase in survival (Figure 2A). Notably, doubling the number of RNA aptamer sequences resulted in an enhanced survival rate, with colony numbers increasing by more than 360-fold compared to scrambled cells, and A This was increased by more than 16-fold compared to CRCd-treated cells (Fig. 2A).

[0268] A CRC H840A teeth, A Although it moderately increased the survival rate compared to CRCd (Figure 2A), sequence analysis of individual clones revealed that random mutations were frequently generated outside the targeted region (within the protospacer) (Figure 8A). The latter system consistently targeted residue C1592 at codon 531, but A CRC H840A nCas induced frequent mutations not only in the target region but also in several upstream nucleotides (Figure 8A). Therefore, we selected nCas for further genetic engineering and optimization. D10A It was decided to use only the mobilization module.

[0269] To continue the optimization process, A CRCd and A CRC D10A We decided to engineer the system by testing different spatial configurations of the effector module in it, and to this end, we used various linkers of different lengths and flexibilities to separate the AID and MCP (Table S2).

[0270] [Table 8]

[0271] A flexible 25 amino acid linker (L25) derived from the hinge region of immunoglobin gamma 3 (IgG3) showed the highest efficiency, but variability between different linkers was observed, particularly A CRC D10A The difference between the most and least efficient configurations was relatively small, with a factor of two (Fig. 2B). These results suggest that the spatial separation between AID and MCP in the effector module can be quite flexible.

[0272] Various types of cytidine deaminases can be incorporated into the CRC system as effectors. We have identified two other proteins related to AID from the APOBEC family of cytidine deaminases: APOBEC1 and APOBEC3G (respectively, A1 CRC D10A and A3G CRC D10A ) was tested. A 1 CRC D10A shows higher conversion efficiency, followed by A CRC D10A And finally A3G CRC D10A showed the lowest activity (Fig. 2C). A1 CRC D10A induced double mutants at a high rate, whereas A3G CRCD10A revealed that nucleotides outside the protospacer were frequently targeted (Fig. 8B). A3G CRC D10A was excluded from further optimization.

[0273] Example 5 The CRC system corrects loss-of-function mutations in the GFP gene in mammalian cells To determine whether the CRC system works in mammalian cells, we performed a multi-step analysis in HEK293 cells. A CRC D10A The system was tested. Mammalian expression of the various components generated a polycistronic vector under the control of a CMV promoter, containing the AID_MCP fusion and nCas9, separated by a self-cleaving 2A peptide. D10A This was achieved by expressing nCas9 (Figure 9A). In cells, uracil DNA glycosylase (UNG) initiates the repair of U:G mismatches induced by cytidine deamination (25-27). To enhance the nucleotide conversion efficiency at the target site, the bacterial UNG inhibitor peptide (UGI) (28) was fused to nCas9, thus inducing local UNG inhibition, a strategy to enhance the efficiency of BE base editors. This mammalian CRC expression construct A The gRNA construct is designated as CRCnu. The gRNA construct is driven by a U6 promoter and has two MS2 loops at the 3' end of the CRISPR scaffold (2xMS2; Figure 9B).

[0274] We designed a GFP reporter with an A-to-G point mutation along the chromophore sequence, resulting in a tyrosine-to-cysteine ​​mutation at position 66 (Y66C) (Figure 3A). This mutation renders the protein non-fluorescent (nfEGFP), thus mimicking a loss-of-function (LOF) mutation. We also designed a gRNA that targets the non-template strand (NT) surrounding the mutation region (nfEGFP_NT1; Figure 3A and Table S3).

[0275] First, we attempted to correct the LOF mutation in extrachromosomal DNA. To this end, we transfected the target nfEGFP construct into HEK293T cells. A CRCnu and nfEGFP_NT1 gRNAs were transiently expressed together (Figure 3B). For comparison, we used the third- and fourth-generation BE base editors, BE3 (13) and BE4max (29). A It was tested alongside CRCnu. A Higher GFP conversion was observed in CRCnu cells than in cells treated with BE4max and BE3 (Figure 3B). Quantification by flow cytometry revealed that A It was revealed that GFP positivity (GFP+) after CRCnu / nfEGFP_NT1 treatment was 62%, whereas BE4max / nfEGFP_NT1 and BE3 / nfEGFP_NT1 treatment resulted in 35% and 30% GFP+ cells, respectively ( Figure 3C ).

[0276] To investigate whether the system has base editing activity on chromosomal DNA sequences, a low-copy-number mutant nfEGFP gene was stably integrated into the HEK293 genome (the resulting cell line was designated nf2.16). A nf2.16 cells treated with CRCnu, BE4max, or BE3 showed correction efficiencies of 9.8%, 2.3%, and 1.3%, respectively (Figure 3D). After treatment, GFP-positive cells were sorted by fluorescence-activated cell sorting (FACS) followed by Sanger sequencing. These results confirmed G-to-A conversion at the targeted base and restoration of the wild-type sequence (Figure 3E).

[0277] Collectively, these results demonstrate that the CRC system can edit extrachromosomal and chromosomal sequences, and these data demonstrate that CRC-mediated base editing is feasible and efficient in mammalian as well as prokaryotic cells.

[0278] Example 6 Whole-exome analysis of potential off-target effects To assess potential CRC-mediated off-target activity at the whole-exon level, A CRCnu / nfEGFP_NT1, A nf2.16 cells treated with CRCnu / Scramble or left untreated were subjected to whole-exon sequencing, analyzing all exons with an average of 300x coverage across the genome. Analysis of point mutations showed no increase in global single-base mutations in treated cells compared to untreated controls (Figure 3F). AID is a WR C H / D G Because cytosine residues within the YW motif are preferentially mutated (underlined C and G are mutable positions) (30), to further confirm that expression of the effector (AID) does not increase point mutations, we investigated the mutation rates of AID motifs and non-motifs and compared them between treated and untreated cells. No differences were found between CRC-treated and untreated samples in both motif and non-motif sequences (Figure 3G). Collectively, these data indicate that the CRC system does not have a significant effect on inducing global mutagenesis in the genome.

[0279] Example 7: Base editing by CRC at endogenous target sequences To determine the ability of CRISPR to modify endogenous loci in the human genome, we targeted extensively studied regions (i.e., HEK293 sites 2, 3, and 4) using conventional nuclease-dependent CRISPR (31, 32) and BE base editing (13) to study on-target efficacy, on-target indel formation rates, and potential off-target effects on homologous sequences. These sites and their targeting gRNAs are listed in Table S4.

[0280] High-throughput sequencing analysis revealed that CRC targeting at these sites resulted in significant C to T conversions with high purity (i.e., low conversion frequency) (Figures 4A-4C). ACRCnu treatment resulted in efficient nucleotide conversion (Figures 4B and 4C, respectively). These observations indicate that CRC is capable of targeting endogenous genomic sequences.

[0281] For these targets, A The CRCnu construct (expressing AID as an effector) is an APOBEC1-based CRC editor. A1 Note that it appears to have a wider window of activity than CRCnu. A Upon CRCnu treatment, detectable editing was observed in Cs more distal to the PAM (C11 at site 2, C9 at site 3, and C8 at site 4; Figures 4A-4C). A1 CRCnu (Figures 4D-4F) does not show significant activity at these positions. Because base editing is highly constrained by PAM availability and the relative position of the target nucleotide within the protospacer, it may be advantageous to use systems with different activity window widths.

[0282] Example 8 Comparison of on-target indel formation rates and off-target activity between the CRC and BE systems Cas9 nickases are generally considered safe because single-strand breaks in DNA are well tolerated and efficiently repaired by cells (33-35). However, researchers have found that BE base editors containing nickases can still generate indels at target sites, albeit at a much lower rate than traditional CRISPR approaches (13, 29, 36). To determine the extent of indel formation after CRC treatment, we analyzed the data to estimate the frequency of such events in treated and untreated cells. Indels were detected at frequencies comparable to those induced by BE base editors after CRC treatment (13, 36), but both were significantly lower than those using traditional CRISPR approaches (36, 37). Meanwhile, untreated cells showed only background levels of indels. The distribution and frequency of indels in treated cells correlated with the gRNA target site. In conclusion, CRC induces detectable indels at levels similar to those of BE base editors. Both of these are significantly lower levels compared to conventional CRISPR methods.

[0283] To estimate the extent of off-target activity in CRC and compare it to the BE system, we investigated selected known off-target sites at sites 2, 3, and 4, which had previously been identified by chromatin immunoprecipitation of dCas9 bound to off-target sites ( 31 ), using the GUIDE-seq method to determine wild-type Cas9 off-target activity ( 32 ) and evaluate BE base editors ( 13 ). The off-target sites are summarized in Table S5.

[0284] [Table 9]

[0285] High-throughput sequencing analysis revealed that the majority of the analyzed off-target sites did not exhibit editing activity (Figure 11). In S4O1 (site 4 off-target site 1), we observed detectable C to T editing, but the frequency was much lower than that reported for the same site with BE3 (i.e., less than 1% in C3, C5, and C8 of CRC-treated cells, compared with 10% in C5 of BE3-treated cells (13)).

[0286] Example 9 Construction of second-generation CRCs by codon optimization and enhanced local UNG inhibition We generated second-generation CRC constructs by codon-optimizing to enhance construct expression and by adding an additional UGI copy to Cas9 to enhance local UNG inhibition, and tested base editing efficiency as well as on-target indel formation and off-target effects. A CRCnu.2 and A1 We named them CRCnu.2 (which have AID and APOBEC1 as effectors, respectively; Figure 9C).

[0287] For comparison, the inventors A CRCnu.2, A1 CRCnu.2, and BE4max ( 29 ) were used to target HEK293T site 2. A CRCnu.2 and A1 CRCnu.2 efficiency is A For CRCnu.2, it reached 37% C->T at C4 and 41% at C6 ( Figure 5A ). A1 At the same C after CRCnu.2 treatment, the editing efficiencies reached 10% and 43% (Figure 5B), a dramatic increase compared to their first-generation counterparts at the same sites, where the maximum editing efficiencies were only around 30% and 20%, respectively. A CRCnu.2 induced a 7% C->T at C11, confirming that AID has a broader window of activity as a CRC effector at this site than APOBEC1 ( Fig. 5A ).

[0288] Also, optimized A We evaluated the off-target activity of the CRCnu.2 system, which revealed a pattern similar to that of the first-generation CRC editor, in which base editing was undetectable at most off-target sites (Figure 11). Interestingly, A1 CRCnu.2 induced a mutation rate comparable to BE4 in C6 (43% vs. 44%), but a much lower mutation rate in C4 (10% vs. 21%). A1 We show that CRCnu.2 may have preferred mutation sites within the protospacer region that are different from BE4max, potentially leading to a more dispersed base editing pattern than BE4.

[0289] Furthermore, the present inventors A CRCnu.2 targeted sites 3 and 4, thereby targeting the same site. A This resulted in improved editing efficiency compared to CRCnu (Figure 13, compare with Figures 4B-4C), while the frequency of indel formation remained low (Figure 14).

[0290] Collectively, these data demonstrate that optimized second-generation CRC base editors exhibit higher efficacy compared to their first-generation CRC counterparts while maintaining low on-target indel formation rates and similar off-target profiles. Furthermore, these data support the idea that second-generation CRC base editors may perform at a similar level to the BE base editor BE4max, but have different activity windows and editing position preferences.

[0291] Example 10 CRC efficiently mediates targeted gene disruption by inducing premature stop codons The primary application of genome editing technology is targeted gene disruption, typically through DSB and NHEJ activation, ultimately inducing a frameshift mutation that introduces a premature stop codon into the target gene transcript (38). Targeted gene inactivation can be an effective therapeutic strategy for eliminating disease-causing gene products. CRC and other base editing strategies offer a safer alternative to gene inactivation by directly editing CAG (glutamine, Q), CAA (glutamine, Q), CGA (arginine, R), and TGG (tryptophan, W) codons into TAG, TAA, and TGA stop codons via C-to-T conversion. Cytidine deaminase-mediated base editing using the BE system has been exploited to induce premature stop codons in a targeted manner without the need for DSB generation (39, 40).

[0292] We attempted to test the ability of CRC to induce a stop codon in the EGFP reporter gene. One gRNA was designed to target Q157 (EGFP_TS1) to generate a stop codon at that position (Figure 6A, Table S3). HEK293 cells stably expressing EGFP were targeted with TS1, resulting in efficient disruption of GFP expression (Figures 6B and 6C). Flow cytometry analysis revealed that TS1 induced 17.8% of GFP-negative cells (Figure 6C). HTS analysis demonstrated that a stop codon was induced at the target site, confirming the flow cytometry observations. TS1 resulted in a 24% C to T mutation at codon 157 (Figure 6D). Low levels of indel formation, following a pattern similar to that observed in previous experiments, were detected in treated cells (Figure 15).

[0293] Finally, to assess the ability of CRC to induce premature stop codons in endogenous targets, we transfected the PDCD1 locus. AWe attempted to treat CRCnu.2. The PDCD1 gene encodes the immune checkpoint receptor PD1 (programmed cell death protein 1), a major target of immunotherapy strategies aimed at treating various types of cancer (41). We targeted codon 133, which encodes glutamine (Q133) in the PD1 protein, and designed one gRNA to induce a stop codon at this position (PDCD1_TS1; Figure 6F, Table S4). The PDCD1_TS1 gRNA targeted A CRCnu.2 caused a 14% C-to-T transversion in C3, converting codon Q133 (CAG) to a stop codon (TAG) (Figure 6G). We observed a bystander C edit at C8 with similar efficiency (Figure 6G). This mutation occurs at the third position of codon 134 and does not alter the isoleucine residue encoded by this codon. Together, these results provide proof-of-concept for the efficient induction of targeted gene knockouts by the CRC base editing approach.

[0294] Example 11 APOBEC1 from different species shows unexpectedly wider activity windows or higher activity at certain positions In this example, different CRC systems were created using APOBEC1 from different species, including APOBEC1 from rat, lizard (Myodrianorrhynchos gracilis), and bat (Myotis lucifugus). The effector proteins and DNA sequences are shown below.

[0295] [ka]

[0296] These systems were investigated in the same manner as described above. The results are shown in Figures 16A-16D. As shown, these CRC systems, using a wide variety of cytidine deaminases from different species and different deaminase families, such as the lizard Apobec1, exhibit distinctly different activity windows and preferred positions than any previously described base editing system. These CRC systems can be used to specifically target nucleotides close to PAM motifs for nucleic acid modifications (e.g., disease mutation correction) not achieved by other known effectors.

[0297] Example 12 Different species of AID or APOBEC1 exhibit unexpectedly different activity windows or higher activity at some positions In this example, we generated CRC systems using AID or APOBEC1 from several species, including rat, lizard (Myodrianorhynchus), and bat (Myotis lucifugus). The effector proteins and DNA sequences are shown below.

[0298] Anole (lizard) AID orthologue Shown below is the amino acid sequence of mydrianole single-stranded DNA cytosine deaminase (activation-induced cytidine deaminase, AID) fused to MS2 coat protein (MCP).

[0299] [ka]

[0300] In the sequence above, the AID sequence (bold) is linked to the MCP sequence (underlined) via a hinge linker (italic), and the N-terminal nuclear localization signal is also underlined. Below is the codon-optimized nucleotide sequence for expressing the protein in human cells.

[0301] [ka]

[0302] Anole (lizard) APOBEC1 ortholog Shown below is the amino acid sequence of the mydrianol single-stranded DNA apolipoprotein B mRNA editing enzyme complex (APOBEC1) fused to MCP.

[0303] [ka]

[0304] In the sequence above, the APOBEC1 sequence (bold) is linked to the MCP sequence (underlined) via a hinge linker (italic), and the N-terminal nuclear localization signal is also underlined. Below is the codon-optimized nucleotide sequence for expression of the protein in human cells.

[0305] [ka]

[0306] Myotis brandetii (bat) AID ortholog Shown below is the amino acid sequence of Myotis brandetii single-stranded DNA cytosine deaminase (activation-induced cytidine deaminase, AID) fused to MS2 coat protein (MCP).

[0307] [ka]

[0308] In the sequence above, the AID sequence (bold) is linked to the MCP sequence (underlined) via a hinge linker (italic), and the N-terminal nuclear localization signal is also underlined. Below is the codon-optimized nucleotide sequence for expressing the protein in human cells.

[0309] [ka]

[0310] gRNA sequence The complete gRNA construct coding sequence is shown below (target inserted at the underlined / bold site via BbsI restriction digestion).

[0311] [ka]

[0312] [ka]

[0313] [Table 10]

[0314] These lizard (Mydrianolus) and bat (Myotis lucifugus) AIDs or APOBEC1s were investigated in the same manner as described above. These effectors were assembled into second-generation CRC constructs (i.e., LizardA CRCnu.2, LizardA1 CRCnu.2, BatA CRCnu.2, and BatA1 CRCnu.2 construct, A refers to AID, and A1 refers to APOBEC1). The results are shown in Figures 17-20.

[0315] First, the lizard LizardA1 The CRCnu.2 system is a rat A1 It exhibited a broader activity window compared to CRCnu.2, and cytidine nucleotides outside the activity window (positions 3–9 of the protospacer), especially the cytidine proximal to the PAM, were found to be accessible to lizard APOBEC1 effectors.

[0316] Figure 17 shows a lizard LizardA1CRCnu.2, rat A1 CRCnu.2, lizard A The C to T conversion rates at the human fetal hemoglobin promoter locus in K562 cells by the CRCnu.2 and BE4max. systems were compared. Briefly, K562 cells were transfected with 1 μg of CRC expression vectors (gRNAs containing the MS2 aptamer, MCP fused to lizard APOBEC1, lizard AID, or rat APOBEC1, and expressing nCas9D10A or BE4max) using a Neon electroporation device. After transfection, cells were grown for 72 hours, genomic DNA was isolated, and target fragments were amplified by PCR and subjected to Sanger sequencing and high-throughput sequencing. Data show representative results from two independent experiments. Results showed that all four effectors exhibited high activity against cytidines at positions C6 and C7 of this locus (consistent with the literature-documented activity window of positions 3–9 of the protospacer). In contrast, LizardA1 CRCnu.2 (lizard Apobec1) showed high activity with C3; LizardA CRCnu.2 (lizard AID) showed high activity at C14 (outside the canonical activity window) in addition to high activity at C6 and C7.

[0317] Figure 18 shows LizardA1 CRCnu.2 and rat A1 Comparison of C to T conversion rates at the site 2 locus in HEK293 cells using the CRCnu.2 system is shown. Briefly, HEK293 cells were transfected with 1 µg of CRC expression vectors (expressing gRNA containing the MS2 aptamer, MCP fused to lizard APOBEC1 or rat APOBEC1, and nCas9D10A) using a Neon electroporation device. After transfection, cells were grown for 72 hours, genomic DNA was isolated, and target fragments were amplified by PCR and subjected to Sanger sequencing and high-throughput sequencing. LizardA1 CRCnu.2 and ratA1 CRCnu.2 were compared. Data are representative of three independent experiments. These experiments were performed in rats. A1 We showed that the CRCnu.2 construct exhibited high activity against cytidines at positions C4 and C6 of this locus (consistent with the documented activity window of positions 3-9 of the protospacer). LizardA1 CRCnu.2 showed high activity at C4 and C6 as well as high activity at C11 (outside the canonical activity window).

[0318] Since the PAM motif is at the 3' end of the sequence shown in the chart, the above results indicate that the cytidine proximal to the PAM motif is LizardA1 CRCnu.2 can be targeted by A1 This indicates that it cannot be targeted by CRCnu.2 or BE4max.

[0319] Second, we expressed AID as an effector. LizardA The CRCnu.2 system is A It exhibited a broader activity window compared to CRCnu.2, and cytidine nucleotides outside the activity window (positions 3–9 of the protospacer), especially the cytidine proximal to the PAM, were found to be accessible to lizard AID.

[0320] Figure 19 shows LizardA CRCnu.2 (lizard AID) and humans AComparison of C to T conversion rates at the site 3 locus in HEK293 cells using the CRCnu.2 (human AID) system is shown. HEK293 cells were transfected with 1 μg of CRC expression vectors (expressing gRNA containing the MS2 aptamer, MCP fused to lizard AID or rat AID, and nCas9D10A) using a Neon electroporation device. After transfection, cells were grown for 72 hours, genomic DNA was isolated, and target fragments were amplified by PCR and subjected to Sanger sequencing and high-throughput sequencing. Lizard LizardA CRCnu.2 (gray) and human A CRCnu.2 (orange) were compared. Data are representative of two independent experiments. These experiments were performed in humans. A We showed that CRCnu.2 exhibited high activity against cytidines at positions C3, C5, and C9 of this locus (consistent with the documented activity window of positions 3-9 of the protospacer). LizardA CRCnu.2 showed high activity at C14 (outside the canonical window) in addition to high activity at C3, C5, and C9. Because the PAM motif is at the 3' end of the sequence shown in the chart, these results suggest that the cytidine proximal to the PAM motif is LizardA Can be targeted by CRCnu.2, but in humans A This suggests that CRCnu.2 cannot be targeted.

[0321] Third, BatA The CRCnu.2 (bat AID) system, at a specific locus, A It was found that the base editing activity was higher than that of CRCnu.2 (human AID). BatA CRCnu.2 and Human AComparison of C to T conversion rates at the site 3 locus in HEK293 cells using the CRCnu.2 system is shown. HEK293 cells were transfected with 1 µg of DNA in total, expressing a CRC expression vector (gRNA containing the MS2 aptamer, MCP fused to bat AID or rat AID, and nCas9D10A) using a Neon electroporation device. After transfection, cells were grown for 72 hours, genomic DNA was isolated, and target fragments were amplified by PCR and subjected to Sanger sequencing and high-throughput sequencing. BatA CRCnu.2 and Human A CRCnu.2 were compared. Data show representative results of two independent experiments. BatA CRCnu.2 binds to cytidine at positions C3, C5, and C9, especially at C5. A It was shown that it exhibited higher activity than CRCnu.2.

[0322] Example 13 Comparison of inactive Cas (Dead Cas) and nickase in mammalian cells In this example, a study was performed to compare inactive Cas and nickase in HEK cells (Figure 21).

[0323] method Generation of dCas9 Site-directed mutagenesis (SDM) using the Q5 Site-Directed Mutagenesis Kit (NEB: Catalog No. E0554S) A A catalytically inactive Cas9 (dCas9) version of the CRCnu.2 construct was generated. The forward primer was designed to incorporate a 2-bp mismatch with the target nCas9 sequence, which changes codon 840 of nCas9 from CAT (histidine) to GCT (alanine) after PCR amplification. The H840A mutation inactivates the HNH catalytic domain of Cas9, thereby AIn combination with the D10A mutation already present in CRCnu.2, this generates a catalytically inactive Cas9 that can no longer cleave dsDNA. The primers used for SDM are detailed in the table below. For the forward primer, the lowercase "gc" represents a mismatch with the target sequence, generating a CAT-GCT mutation at codon 840 of nCas9.

[0324] [Table 11]

[0325] The PCR amplification was set up as follows.

[0326] [Table 12]

[0327] The PCR reaction conditions are as follows:

[0328] [Table 13]

[0329] Expression plasmids The components of the base editing system are expressed as a single polycistronic unit, whereby the Cas component and the MCP / deaminase fusion form two separate proteins via the T2A self-cleaving peptide.

[0330] The sgRNA components of the base editing system were expressed in a separate vector in which sgRNA expression was driven by an RNA polymerase III U6 promoter. The sgRNA was expressed as a single unit encompassing the crRNA and tracrRNA components of the Cas9 duplex RNA system linked by an artificial tetraloop. Additionally, two copies of the RNA aptamer MS2 were tethered to the 3′ of the sgRNA via a fold-back dsRNA linker to enable deaminase recruitment. As a control, an sgRNA lacking the MS2 motif (MS2less) was used. Because it lacks the MCP, the recruited aptamer should be unable to edit the target locus. A polyT termination signal was included 3′ of the sgRNA to catalyze transcription termination. A list of the sgRNAs used and their sequences is provided in the table below.

[0331] [Table 14]

[0332] In the table above, sequences in lowercase letters represent the target-specific protospacer component of the sgRNA, and sequences in uppercase letters represent the tracrRNA component of the sgRNA. Superscript numbers represent C residues present within the target base editing window. A protospacer consisting of a scrambled sequence (Scrambled_2×MS2) was used as a negative control.

[0333] Cell culture and transfection All transfection experiments were performed with HEK293 cells, which were cultured at 37°C and 5% CO2. HEK293 cells were maintained in Dulbecco's Modified Eagle's Medium (DMEM) supplemented with 10% FBS. To ensure 70% confluency for transfection, HEK293 cells were seeded at a cell density of 50,000 cells / well in 24-well culture plates 24 hours prior to transfection. After 24 hours, cells were lipid-transfected with 200 ng of plasmid DNA (150 ng of base-editing / BE4max vector and 50 ng of sgRNA expression vector) using LIPOFECTAMINE 3000 Reagent (THERMOFISHER SCIENTIFIC: Catalog No. L3000015).

[0334] Cell lysis and flow cytometry Seventy-two hours after transfection, the medium was aspirated and the cells were washed once with PBS. Cells were then detached from the well surface using 100 μl of TrypLE express enzyme (THERMOFISHER SCIENTIFIC: Catalog No. 12605010). The dissociated cells were then pelleted by centrifugation at 300×rpm for 5 minutes at room temperature and then resuspended in 100 μl of PBS. 20 μl of the cell suspension was transferred to a well of a 96-well plate containing 36 μl of Direct PCR Lysis Reagent (VIAGEN biotech: Catalog No. 302-C), and cell lysis was performed under the following conditions: 55°C for 30 minutes, followed by 95°C for 30 minutes. The remaining 80 μl of the resuspended cells was transferred to a 96-well plate and collected by centrifugation at 300×rpm for 5 minutes at room temperature. The supernatant was discarded, and the pelleted cells were resuspended in 50 μl of MACS buffer (MILTENYI BIOTEC) supplemented with 0.5% BSA in preparation for flow cytometry analysis. All flow cytometry was performed using iQue3 (SARTORIUS).

[0335] PCR amplification of the target region 1 μl of cell lysate was used per PCR reaction. Q5 High Fidelity 2x Master Mix (NEB: Catalog No.-M0491S) was used for amplification of sgRNA target sites, and the reaction mix was set up as follows:

[0336] [Table 15]

[0337] The PCR cycling parameters for amplification of target site 2 were as follows:

[0338] [Table 16]

[0339] result The Cas9 nickase (nCas9-D10A) is the construct of choice for base editing because nicking of the non-edited strand stimulates the cellular mismatch machinery, which uses the edited strand as a template for repair, thereby shifting the post-replicative probability balance toward C-to-T editing. Introducing the H840A mutation into nCas9 abolishes its nickase functionality and prevents nicking of the non-edited DNA strand. We measured the ability of the base editing system to achieve edits at target loci using catalytically inactive Cas9 (dCas9). A CRCnu.2 was used as a template to generate a dCas9 version of the base editor, and editing efficiency was measured at site 2.

[0340] As shown in Figure 21, these data indicate that the base editing system can achieve on-target editing when using dCas9. These data also show that the highest level of editing at both C residues in the target sequence was achieved with nCas9 ( A CRCnu.2) has been achieved, and C 1 showed 42% editing, and C 2 indicates 60% editing. ACRCdu.2), the editing activity was reduced, but still (C 1 =10%;C 2 = 14%), when MS2less sgRNA was used ( A CRCdu.2_MS2less) or non-targeting scrambled guide ( A CRCdu.2_scrambled). In conclusion, the use of catalytically inactive Cas9 is compatible with on-target editing using base editing systems.

[0341] References 1.Fu YF, et al. (2013) High-frequency off-target mutagenesis induced by CRISPR-Cas nucleases in human cells. Nature Biotechnology 31(9):822-+. 2.Singh P, Schimenti JC, & Bolcun-Filas E (2014) A Mouse Geneticist's Practical Guide to CRISPR Applications. Genetics. 3.Ran FA, et al. (2013) Double Nicking by RNA-Guided CRISPR Cas9 for Enhanced Genome Editing Specificity. Cell 154(6):1380-1389. 4.Tsai SQ, et al. (2014) Dimeric CRISPR RNA-guided FokI nucleases for highly specific genome editing. Nat Biotech 32(6):569-576. 5.Guilinger JP, Thompson DB, & Liu DR (2014) Fusion of catalytically inactive Cas9 to FokI nuclease improves the specificity of genome modification. Nat Biotechnol 32(6):577-582. 6.Kleinstiver BP, et al. (2016) High-fidelity CRISPR-Cas9 nucleases with no detectable genome-wide off-target effects. Nature 529(7587):490-495. 7.Slaymaker IM, et al. (2016) Rationally engineered Cas9 nucleases with improved specificity. Science 351(6268):84-88. 8.Kosicki M, Tomberg K, & Bradley A (2018) Repair of double-strand breaks induced by CRISPR-Cas9 leads to large deletions and complex rearrangements. Nature Biotechnology 36:765. 9.Rivera-Torres N, Banas K, Bialk P, Bloh KM, & Kmiec EB (2017) Insertional Mutagenesis by CRISPR / Cas9 Ribonucleoprotein Gene Editing in Cells Targeted for Point Mutation Repair Directed by Short Single-Stranded DNA Oligonucleotides. PloS one 12(1):e0169350. 10.Corrigan-Curay J, et al. (2015) Genome editing technologies: defining a path to clinic. Mol Ther 23(5):796-806. 11.Cox DB, Platt RJ, & Zhang F (2015) Therapeutic genome editing: prospects and challenges. Nature medicine 21(2):121-131. 12.Iyama T & Wilson DM (2013) DNA repair mechanisms in dividing and non-dividing cells. DNA repair 12(8):620-636. 13.Komor AC, Kim YB, Packer MS, Zuris JA, & Liu DR (2016) Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage. Nature 533(7603):420-+. 14.Cox DB, et al. (2017) RNA editing with CRISPR-Cas13. Science 358(6366):1019-1027. 15.Zalatan JG, et al. (2015) Engineering complex synthetic transcriptional programs with CRISPR RNA scaffolds. Cell 160(1-2):339-350. 16.Konermann S, et al. (2015) Genome-scale transcriptional activation by an engineered CRISPR-Cas9 complex. Nature 517(7536):583-588. 17.Wang S, Su J-H, Zhang F, & Zhuang X (2016) An RNA-aptamer-based two-color CRISPR labeling system. Scientific reports 6:26857. 18.Qin P, et al. (2017) Live cell imaging of low-and non-repetitive chromosome loci using CRISPR-Cas9. Nature communications 8:14725. 19.Jin S, Collantes, JC (2017) Nuclease-Independent Targeted Gene Editing Platform and Uses Thereof. PCT / US2016 / 042413 (Priority date: 15.07.2015) 20.Hess GT, et al. (2016) Directed evolution using dCas9-targeted somatic hypermutation in mammalian cells. Nature methods 13(12):1036. 21.Liu LD, et al. (2018) Intrinsic nucleotide preference of Diversifying Base editors guides antibody ex vivo affinity maturation. Cell Reports 25(4):884-892. e883. 22.Campbell EA, et al. (2001) Structural Mechanism for Rifampicin Inhibition of Bacterial RNA Polymerase. Cell 104(6):901-912. 23.Goldstein BP (2014) Resistance to rifampicin: a review. J Antibiot (Tokyo) 67(9):625-630. 24.Xu M, Zhou YN, Goldstein BP, & Jin DJ (2005) Cross-Resistance of Escherichia coli RNA Polymerases Conferring Rifampin Resistance to Different Antibiotics. Journal of Bacteriology 187(8):2783-2792. 25.Petersen-Mahrt SK, Harris RS, & Neuberger MS (2002) AID mutates E. coli suggesting a DNA deamination mechanism for antibody diversification. Nature 418(6893):99-104. 26.Krokan HE & Bjoras M (2013) Base excision repair. Cold Spring Harbor perspectives in biology 5(4):a012583. 27.Jacobs AL & Schar P (2012) DNA glycosylases: in DNA repair and beyond. Chromosoma 121(1):1-20. 28.Mol CD, et al. (1995) Crystal structure of human uracil-DNA glycosylase in complex with a protein inhibitor: protein mimicry of DNA. Cell 82(5):701-708. 29.Koblan LW, et al. (2018) Improving cytidine and adenine base editors by expression optimization and ancestral reconstruction. Nature Biotechnology. 30.Odegard VH & Schatz DG (2006) Targeting of somatic hypermutation. Nat Rev Immunol 6(8):573-583. 31.Kuscu C, Arslan S, Singh R, Thorpe J, & Adli M (2014) Genome-wide analysis reveals characteristics of off-target sites bound by the Cas9 endonuclease. Nature Biotechnology 32:677. 32.Tsai SQ, et al. (2014) GUIDE-seq enables genome-wide profiling of off-target cleavage by CRISPR-Cas nucleases. Nature Biotechnology 33:187. 33.Caldecott KW (2001) Mammalian DNA single-strand break repair: an X-ra(y)ted affair. Bioessays 23(5):447-455. 34.Caldecott KW (2008) Single-strand break repair and genetic disease. Nature Reviews Genetics 9(8):619-631. 35.Caldecott KW (2014) DNA single-strand break repair. Experimental cell research 329(1):2-8. 36.Rees HA & Liu DR (2018) Base editing: precision chemistry on the genome and transcriptome of living cells. Nature Reviews Genetics:1. 37.Chakrabarti AM, et al. (2019) Target-Specific Precision of CRISPR-Mediated Genome Editing. Molecular cell 73(4):699-713 e696. 38.Sander JD & Joung JK (2014) CRISPR-Cas systems for editing, regulating and targeting genomes. Nat Biotechnol 32(4):347-355. 39.Kuscu C, et al. (2017) CRISPR-STOP: gene silencing through base-editing-induced nonsense mutations. Nature methods 14:710. 40.Billon P, et al. (2017) CRISPR-Mediated Base Editing Enables Efficient Disruption of Eukaryotic Genes through Induction of STOP Codons. Molecular cell 67(6):1068-1079.e1064. 41.Pardoll DM (2012) The blockade of immune checkpoints in cancer immunotherapy. Nature Reviews Cancer 12:252. 42.Grunewald J, et al. (2019) Transcriptome-wide off-target RNA editing induced by CRISPR-guided DNA base editors. Nature 569(7756):433. 43.Duan D, Yue Y, & Engelhardt JF (2001) Expanding AAV packaging capacity with trans-splicing or overlapping vectors: a quantitative comparison. Molecular therapy 4(4):383-391. 44.Carvalho LS, et al. (2017) Evaluating efficiencies of dual AAV approaches for retinal targeting. Frontiers in neuroscience 11:503. 45.Grieger JC & Samulski RJ (2005) Packaging Capacity of Adeno-Associated Virus Serotypes: Impact of Larger Genomes on Infectivity and Postentry Steps. Journal of Virology 79(15):9933-9944. 46.Shapiro MB & Senapathy P (1987) RNA splice junctions of different classes of eukaryotes: sequence statistics and functional implications in gene expression. Nucleic acids research 15(17):7155-7174. 47. Baralle D & Baralle M (2005) Splicing in action: assessing disease causing sequence changes. Journal of medical genetics 42(10):737-748. 48.Gapinske M, et al. (2018) CRISPR-SKIP: programmable gene splicing with single base editors. Genome biology 19(1):107. 49. Kabadi AM, Ousterout DG, Hilton IB, & Gersbach CA (2014) Multiplex CRISPR / Cas9-based genome engineering from a single lentiviral vector. Nucleic acids research 42(19):e147-e147.

[0342] The above examples and description of preferred embodiments should be construed as illustrative, not limiting, of the invention defined by the claims. As will be readily understood, numerous variations and combinations of the features set forth above can be used without departing from the invention as set forth in the claims. Such variations are not considered a departure from the scope of the invention, and all such variations are intended to be included within the scope of the following claims. All references cited herein are incorporated by reference in their entirety.

Claims

1. (i) a sequence targeting component or a polynucleotide encoding the same, said component comprising: (a) a CRISPR protein, wherein the CRISPR protein comprises a dCas9 or nCas9 sequence of a species selected from the group consisting of Streptococcus pyogenes, Streptococcus agalactiae, Staphylococcus aureus, Streptococcus thermophilus, Streptococcus thermophilus, Neisseria meningitidis, and Treponema denticola; and (b) First Uracil DNA Glycosylase (UNG) Inhibitor Peptide (UGI) a sequence targeting component or a polynucleotide encoding the same, comprising a target fusion protein having: (ii) an RNA scaffold or a DNA polynucleotide encoding the same, wherein the scaffold comprises: (a) a nucleic acid targeting motif comprising a guide RNA sequence complementary to a target nucleic acid sequence; (b) an RNA motif capable of binding to the CRISPR protein; and (c) First recruitment RNA motif an RNA scaffold or a DNA polynucleotide encoding the same, comprising: (iii) a first effector fusion protein or a polynucleotide encoding the same, said protein comprising: (a) a first RNA binding domain capable of binding to said first recruitment RNA motif; (b) a linker, and (c) effector domain A first effector fusion protein or a polynucleotide encoding the same, comprising: wherein the first effector fusion protein or the effector domain has cytidine deamination activity or adenosine deamination activity.

2. The system of claim 1 , wherein the target fusion protein comprises two or more UGIs.

3. The system of claim 1 or 2, wherein the RNA scaffold comprises two or more recruitment RNA motifs.

4. The system described in any one of claims 1 to 3, wherein the target fusion protein further comprises a second UGI, and one or more of the polynucleotides encoding the CRISPR protein, the first UGI, the second UGI, the RNA binding domain, and the effector domain are optimized for expression in eukaryotic cells or mammalian cells.

5. The system of any one of claims 1 to 4, wherein the sequence targeting component or the first effector fusion protein comprises one or more nuclear localization signals (NLS).

6. The system of claim 5 , wherein the sequence targeting component comprises two NLSs.

7. 10. The system of claim 1, wherein the CRISPR protein does not have nuclease activity or does not have nuclease activity to generate a DNA double-strand break.

8. the first recruitment RNA motif and the first RNA binding domain comprise: Telomerase Ku-binding motif and Ku protein telomerase Sm7 binding motif and Sm7 protein; MS2 phage operator stem loop and MS2 coat protein (MCP), PP7 phage operator stem loop and PP7 coat protein (PCP), SfMu phage Com stem-loop and Com RNA-binding protein, The system according to any one of claims 1 to 7, wherein the pair is selected from the group consisting of:

9. 9. The system of any one of claims 1 to 8, wherein the effector domain of cytidine deaminating activity is wild-type AID, CDA, APOBEC1, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D, APOBEC3F, or other APOBEC family enzymes of a species selected from the group consisting of human, rat, mouse, bat, naked mole rat, elephant, chicken, lizard, giant tortoise, coelacanth, and other vertebrate species.

10. 10. The system of any one of claims 1 to 9, wherein the effector domain of adenosine deamination activity is a wild-type ADA, ADAR family enzyme, or tRNA adenosine deaminase from a species selected from the group consisting of bacteria, yeast, human, rat, mouse, bat, naked mole rat, elephant, chicken, lizard, giant tortoise, coelacanth, and other vertebrate species.

11. An isolated nucleic acid encoding component (i) of the system according to any one of claims 1 to 10.

12. An expression vector or host cell comprising the nucleic acid of claim 11.

13. 11. An in vitro method for site-specifically modifying a target nucleic acid sequence, comprising contacting the target nucleic acid sequence with the system of any one of claims 1 to 10.

14. 14. The method of claim 13, wherein the target nucleic acid is present in a cell (excluding a human cell).

15. 15. The method of claim 14, wherein the target nucleic acid is extrachromosomal DNA.

16. 15. The method of claim 14, wherein the target nucleic acid is genomic DNA on a chromosome.

17. 17. The method of any one of claims 14 to 16, wherein the cell is selected from the group consisting of an archaeal cell, a bacterial cell, a eukaryotic cell, a eukaryotic unicellular organism, a somatic cell, a germ cell, a stem cell, a plant cell, an algae cell, an animal cell, an invertebrate cell, a vertebrate cell, a fish cell, a frog cell, an avian cell, a mammalian cell, a porcine cell, a bovine cell, a caprine cell, a ovine cell, a rodent cell, a rat cell, a mouse cell, a horse cell, and a non-human primate cell.

18. 18. The method of claim 17, wherein the cell is a plant cell.

19. 18. The method of claim 17, wherein the cell is present in or derived from a non-human subject.

20. 20. The method of claim 19, wherein the non-human subject has a genetic mutation of the gene.

21. 21. The method of claim 20, wherein the non-human subject has or is at risk of having a disorder caused by the genetic mutation.

22. 22. The method of any one of claims 19 to 21, wherein the site-specific modification corrects a gene mutation, or inactivates expression of a gene, or alters the expression level of a gene, or alters intron-exon splicing.

23. 20. The method of claim 19, wherein the non-human subject has or is at risk of being exposed to a pathogen.

24. 24. The method of claim 23, wherein the site-specific modification inactivates a gene of the pathogen.

25. A kit for site-specifically modifying a target nucleic acid sequence, comprising the system according to any one of claims 1 to 10.

26. 26. The kit of claim 25, further comprising one or more components selected from the group consisting of reagents for renaturation and / or dilution, and reagents for introducing the nucleic acid or polypeptide into a host cell.

Citation Information

Patent Citations

  • Targeted gene editing platform independent of DNA double strand break and uses thereof

    WO2018129129A1

Cited By

  • Highly efficient DNA base editors mediated by RNA-aptamer recruitment for targeted genome modification and uses thereof

    JP2025160188A