Synthetic self-replicating RNA vectors encoding CRISPR proteins and their use

By replicating and expressing CRISPR proteins within cells using a self-replicating RNA vector, the risk of gene integration in CRISPR gene editing systems has been addressed, achieving stable expression and efficient gene editing results.

JP2026074192APending Publication Date: 2026-05-01EMD MILLIPORE CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
EMD MILLIPORE CORP
Filing Date
2026-02-05
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing CRISPR gene editing systems, such as lentivirus and plasmid systems, pose a risk of CRISPR coding sequences integrating into the host cell genome. A novel delivery system is needed that can stably express CRISPR proteins and avoid gene integration.

Method used

A self-replicating RNA vector is used to encode the CRISPR protein and its corresponding guide RNA. The CRISPR protein is replicated and expressed in cells using the self-replicating RNA vector, avoiding the use of DNA intermediates and thus preventing gene integration.

Benefits of technology

This study achieved stable expression of CRISPR proteins in multiple generations of cells, reduced the risk of gene integration, and improved the safety and efficiency of gene editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026074192000001
    Figure 2026074192000001
  • Figure 2026074192000002
    Figure 2026074192000002
  • Figure 2026074192000003
    Figure 2026074192000003
Patent Text Reader

Abstract

To provide a CRISPR gene delivery system that ensures robust expression of CRISPR proteins and avoids the risk of genomic integration. [Solution] A synthetic, non-infectious, self-replicating RNA vector encoding a CRISPR protein is provided. Each self-replicating RNA vector contains sequences encoding multiple non-structural replication complex proteins derived from alphaviruses, and sequences encoding a CRISPR protein. A method for genome editing is also provided, which involves introducing a synthetic self-replicating RNA vector into a cell along with at least one corresponding guide RNA.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Related applications This application claims the benefit of priority of U.S. Provisional Patent Application No. 62 / 847,032, filed on May 13, 2019, the content of which is hereby incorporated by reference in its entirety.

[0002] Sequence List This application includes a sequence listing submitted in ASCII format via EFS-Web, which is hereby incorporated by reference in its entirety. The ASCII copy created on May 13, 2020 is named P19-084_WO-PCT_SL.txt and has a size of 77,524 bytes.

[0003] field The present disclosure relates to synthetic self-replicating RNA vectors encoding CRISPR proteins, which can be transfected into cells together with corresponding guide RNAs for genome editing methods.

Background Art

[0004] Due to the recent development of clustered regularly interspaced short palindromic repeats (CRISPR) and CRISPR-associated (Cas) CRISPR / Cas systems as genome editing tools, the design of site-specific endonucleases for eukaryotic genome modification has become easier and simpler than ever. To deliver CRISPR into eukaryotic cells, many delivery systems have been developed. However, lentivirus, retrovirus, and plasmid systems have the risk that the CRISPR coding sequence is integrated into the host cell genome. Therefore, there is a need for a CRISPR gene delivery system that provides robust expression of CRISPR proteins and avoids the risk of genome integration.

Summary of the Invention

[0005] In various embodiments of this disclosure, a self-replicating RNA vector encoding a CRISPR protein is provided.

[0006] Certain aspects of this disclosure provide a self-replicating RNA vector comprising sequences encoding multiple non-structural replication complex proteins derived from alphaviruses, and sequences encoding CRISPR proteins. The CRISPR proteins may be type II Cas9 proteins, type V Cas12 proteins, type VI Cas13 proteins, CasX proteins, or CasY proteins. In one embodiment, the CRISPR protein may be Streptococcus pyogenes Cas9, Francisella novicida Cas9, Staphylococcus aureus Cas9, Streptococcus thermophilus Cas9, Streptococcus pasteurianus Cas9, Campylobacter jejuni Cas9, Neisseria meningitis Cas9, Neisseria cinerea Cas9, Francisella novicida Cas12, Acidaminococcus sp. Cas12, Lachnospiraceae bacterium ND2006 Cas12, Leptotrichia wadei Cas13a, Leptotrichia shahii Cas13a, Prevotella sp. P5-125 Cas13, or Ruminococcus flavefaciens Cas13d. In a particular iteration, the CRISPR protein is Streptococcus pyogenes Cas9 or Staphylococcus aureus Cas9.

[0007] In some circumstances, the sequence encoding a CRISPR protein may contain at least one nucleotide insertion, deletion, and / or substitution such that the CRISPR protein has altered catalytic activity, improved target site specificity, and / or reduced off-target effects. A CRISPR protein may have double-strand cleavage activity, be able to cleave one strand of a double-stranded sequence, or may have no cleavage activity at all.

[0008] CRISPR proteins may also be ligated to at least one heterologous domain. For example, a CRISPR protein may be ligated to at least one nuclear localization signal. A CRISPR protein may also be ligated to at least one fluorescent protein, at least one chromatin regulatory motif, at least one epigenetic modification domain, at least one transcriptional regulatory domain, at least one RNA aptamer binding domain, or a combination thereof.

[0009] The sequences of self-replicating RNA vectors encoding multiple non-structural replication complex proteins derived from alphaviruses may be derived from Auravirus, Babanki virus, Berma Forest virus, Beval virus, Buggy Creek virus, Chikungunya virus, Eastern Equine Encephalitis virus, Everglades virus, Fort Morgan virus, Getavirus, Highlands J virus, Kyzylagach virus, Mayarovirus, Middelburg virus, Mukambo virus, Ndumu virus, Pixna virus, Onyonnyon virus, Ross River virus, Sagiyama virus, Semliki Forest virus, Sindbis virus, Unavirus, Venezuelan Encephalitis virus, Western Equine Encephalitis virus, or Wataroa virus. These various sequences and their derivatives are publicly known and can be found in the scientific literature, which is incorporated herein by reference in their entirety. In certain embodiments, the sequences encoding multiple non-structural replication complex proteins are derived from Venezuelan Encephalitis virus. See, for example, Yoshioka et al. (Cell Stem Cell 13, 246-254, August 1, 2013) and / or Petrakova et al. (J. Virology; vol. 72, no. 12, June 2005, p. 7597-7608), which are incorporated herein by reference in their entirety. The self-replicating RNA vector may further include sequences encoding selectable markers and / or sequences encoding interferon response inhibitors.

[0010] Another aspect of the present disclosure provides a complex comprising the self-replicating RNA vector described above, and at least one guide RNA designed to form a complex with a CRISPR protein encoded by the self-replicating RNA vector.

[0011] This disclosure also provides eukaryotic cells or cell lines comprising the self-replicating RNA vector disclosed herein.

[0012] Further aspects of this disclosure include plasmid vectors encoding self-replicating RNA vectors.

[0013] A further aspect of this disclosure provides a method for targeted genome editing. This method comprises introducing one of the self-replicating RNA vectors disclosed herein and at least one guide RNA designed to form a complex with a CRISPR protein encoded by the self-replicating RNA vector into a eukaryotic cell.

[0014] Other aspects and variations of this disclosure are described in more detail below. [Brief explanation of the drawing]

[0015] This patent or application document includes at least one figure drawn in color. A copy of the publication of this patent or patent application containing the color figure will be provided by the Office upon request and payment of the required fees.

[0016] [Figure 1A] Figure 1A shows a schematic diagram illustrating the structure of the Simplicon® cloning vector and each VEE-Cas9 RNA described herein.

[0017] [Figure 1B] Figure 1B shows the agarose gel electrophoresis of VEE-Cas9 RNA. Lane 1: RNA marker (200-6000 nucleotides); Lane 2: eSpCas9 RNA; Lane 3: eSp-Cas9-GFP RNA; Lane 4: Cas9-GFP RNA.

[0018] [Figure 1C] Figure 1C shows images of GFP (TagGFP2) expression in cells transfected with VEE-Cas9. VEE-Cas9 RNA and B18R-E3L RNA were co-transfected into HFF on day 1, and GFP expression was imaged by fluorescence microscopy on day 2.

[0019] [Figure 1D] Figure 1D shows the Western blot of Cas9 protein. Lane 1: control; Lane 2: eSpCas9; Lanes 3 and 4: eSpCas9-GFP2; Lanes 5 and 6: Cas9-GFP2.

[0020] [Figure 2A] Figure 2A shows the images of GFP expression on the second day after co-introducing VEE-Cas9-TagGFP2 and one or two synthetic gRNAs targeting K-Ras into HFFs simultaneously or sequentially.

[0021] [Figure 2B] Figure 2B shows the efficiency of genome editing after co-introduction or sequential gene introduction of VEE-Cas9 RNA and gRNA. Cells were harvested on the third day for co-introduction (simultaneous), the fourth day, and the sixth day (after completion of puromycin selection) for sequential gene introduction. Indels were detected using the Guide-iT Mutation Detection Kit. + / - indicates resolvase treatment. The cleavage efficiency(%) is shown below the gel image.

[0022] [Figure 2C] Figure 2C shows the images of GFP expression of human iPSC cell lines on the second day after sequential gene introduction of VEE-Cas9-TagGFP2 and one K-Ras gRNA (25 nM).

[0023] [Figure 2D] Figure 2D shows the efficiency of genome editing in the human iPSCs shown in Figure 2C. Cells were harvested on the fifth day (3 days after gRNA gene introduction) and the seventh day (after completion of puromycin selection). Indels were detected using the Guide-iT Mutation Detection Kit. + / - indicates resolvase treatment. The cleavage efficiency(%) is shown below the gel image.

[0024] [Figure 3A]Figure 3A shows the efficiency of genome editing in HFF cells selected for 7 days after being transfected with VEE-Cas9 (S-Cas9) and B18R-E3L RNA, or infected with lentivirus Cas9 (LV-Cas9), and then treated with puromycin (0.8 μg / mL) or blastosidine (4 μg / mL), respectively. Cells were passaged on day 7, and one type of K-Ras or EMX-1 sgRNA was transfected on day 9. Cells were harvested on day 11, and genome editing was analyzed. + / - indicates resolverase treatment. % indicates cleavage efficiency.

[0025] [Figure 3B] Figure 3B shows Guava flow cytometry analysis of GFP expression on day 4 in HEK293T cells co-transfected with the Cas9 plasmid and the K-Ras gRNA plasmid, or in HEK293T cells co-transfected with VEE-Cas9-TagGFP2 RNA, B18R-E3L RNA, and one K-Ras-gRNA.

[0026] [Figure 3C] Figure 3C shows the efficiency of genome editing in the cells described in Figure 3B. + / - indicates resolverase treatment. % indicates cleavage efficiency.

[0027] [Figure 4A] Figure 4A shows GFP expression analyzed by Guava flow cytometry in HFF cells produced by puromycin selection for one month after gene transfer of VEE-Cas9-TagGFP2, and by genome editing after co-transfer of one type of K-Ras sgRNA.

[0028] [Figure 4B] Figure 4B shows GFP expression analyzed by Guava flow cytometry in HEK293T cells generated by genome editing after 1 month of puromycin selection following VEE-Cas9-TagGFP2 gene transfection, and after co-transfection of one type of EMX-1 sgRNA.

[0029] [Figure 4C] Figure 4C compares genome editing in HFF cells that were sequentially transfused with VEE-eSpCas9, VEE-eSpCas9-TagGFP2, or VEE-Cas9-TagGFP2 and one EMX-1 gRNA. Cells were harvested on day 5 and genome editing was analyzed.

[0030] [Figure 4D] Figure 4D compares genome editing in HEK293T cells sequentially transfused with VEE-Cas9-TagGFP2 or VEE-eSpCas9-TagGFP2, B18R-E3L RNA, and one EMX-1 gRNA. Cells were harvested on day 4 and genome editing was analyzed.

[0031] [Figure 5A] Figure 5A shows images of GFP or RFP expression on day 1 and day 3, respectively, in HEK293T cells transfected with VEE-Cas9-D10A-TagGFP2 and VEE-Cas9-D10A-TagRFP.

[0032] [Figure 5B] Figure 5B shows images of GFP or RFP expression after puromycin selection on day 14. GFP expression was also analyzed by Guava flow cytometry.

[0033] [Figure 5C] Figure 5C compares genome editing in HEK 293T cells that were sequentially transfused with VEE-Cas9-D10A-TagGFP2, VEE-Cas9-D10A-TagRFP, or VEE-Cas9-TagGFP2 on day 1, and EMX-1 sgRNA-1, 9, or both on day 2. Cells were harvested on day 5 and genome editing was analyzed.

[0034] [Figure 5D]Figure 5D compares genome editing in HEK 293T cell lines expressing VEE-Cas9-D10A-TagGFP2, VEE-Cas9-D10A-TagRFP, or VEE-Cas9-TagGFP2. Each cell line was transfected with EMX-1 sgRNA-1 and / or -9 on day 1. Cells were harvested on day 4 and genome editing was analyzed.

[0035] [Figure 6A] Figure 6A shows the insertion of the Rab11-BamHI oligo by Cas9 cleavage in pooled U2OS or HEK293T cells shown in the figure. VEE-Cas9-TagGFP2 or VEE-Cas9-D10A-GFP2 was sequentially introduced into U2OS or HEK 293T cells on day 1, and sgRNA and the Rab11-BamHI oligo on day 2. Cells were harvested on day 5 and analyzed for BamHI insertion. Gel images show the digestion of BamHI resulting from the insertion of the Rab-11-BamHI oligo at the cleavage sites of VEE-Cas9-TagGFP2 (left) or VEE-Cas9-D10A-GFP2 (right).

[0036] [Figure 6B] Figure 6B shows the insertion of the Rab11-BamHI oligo at the Cas9 cleavage site in isolated clones. VEE-Cas9-GFP2 was sequentially transfused into U2OS or human iPSCs on day 1, and sgRNA and the Rab11-BamHI oligo on day 2. Cells were selected with puromycin, and clones were isolated for analysis. Gel images show the digestion of BamHI resulting from the insertion of the Rab-11-BamHI oligo at the cleavage site of VEE-Cas9-TagGFP2.

[0037] [Figure 6C] Figure 6C shows the insertion of a GFP fragment at the Cas9 cleavage site. HEK293 cells were sequentially transfused with VEE-Cas9-RFP on day 1, and with sgRNA and a GFP fragment at the GAPDH site on day 2. GFP expression was recorded on day 5. [Modes for carrying out the invention]

[0038] This disclosure provides a synthetic, non-infectious, self-replicating RNA vector based on an alphavirus in which the sequence encoding the viral structural protein has been removed and replaced with the sequence encoding the CRISPR protein. The self-replicating RNA vector is a single-stranded RNA that mimics cellular mRNA having a 5' cap and a poly(A) tail. In certain embodiments, the self-replicating RNA vector does not utilize a DNA intermediate, and as a result there is no risk of genomic integration of the CRISPR sequence. The self-replicating RNA vector enables robust expression of the CRISPR protein across multiple cell generations. Co-introduction of the corresponding guide RNA into cells enables targeted genome editing. The level of CRISPR expression decreases over time due to dilution and degradation. This specification provides a method for introducing the CRISPR protein into cells together with the corresponding guide RNA for genome editing using a self-replicating RNA vector.

[0039] (I) RNA vector encoding a CRISPR protein Certain aspects of this disclosure provide a synthetic self-replicating RNA vector encoding a CRISPR protein. The self-replicating RNA vector contains a sequence that ensures replication of the RNA vector and translation of a heterologous protein sequence (e.g., at least one CRISPR protein) across multiple cell generations. The self-replicating RNA vector is based on a modified alphavirus that encodes multiple replication complex proteins, but in which the viral structural gene has been removed and replaced with a sequence encoding at least one CRISPR protein. Thus, upon entry into a cell, the RNA serves as a template for translation of the viral replication complex proteins and the CRISPR protein. The viral replication complex proteins form a replication complex, enabling further replication of the RNA vector in the cell's cytoplasm. Since the replicated RNA cannot recombine with cellular DNA, there is no risk of the CRISPR sequence being integrated into the cell's genome.

[0040] (a) Synthetic self-replicating RNA The synthetic self-replicating RNA (or replicon) contains all the sequence elements necessary for the translation of the encoded protein and the replication of the RNA vector. In particular, the replicon is based on a modified alphavirus in which the non-structural replicase gene is maintained and the structural gene (necessary for the creation of the infectious particle) is removed. In various embodiments, the modified alphavirus may be derived from Auravirus, Babanki virus, Berma Forest virus, Beval virus, Buggy Creek virus, Chikungunya virus, Eastern Equine Encephalitis virus, Everglades virus, Fort Morgan virus, Getavirus, Highland J virus, Kyzylagach virus, Mayarovirus, Middelburg virus, Mukambo virus, Ndumu virus, Pixna virus, Onyonnyon virus, Ross River virus, Sagiyama virus, Semliki Forest virus, Sindbis virus, Unavirus, Venezuelan Equine Encephalitis virus, Western Equine Encephalitis virus, or Wataroa virus. These various sequences and their derivatives are publicly known and can be found in the scientific literature, which are incorporated herein by reference in their entirety. In certain embodiments, the synthetic self-replicating RNA is based on a modified Venezuelan encephalitis (VEE) virus from which the structural gene has been removed. See, for example, Yoshioka et al. (Cell Stem Cell 13, 246-254, August 1, 2013) and / or Petrakova et al. (J. Virology; vol. 72, no. 12, June 2005, p. 7597-7608), which are incorporated herein by reference in their entirety, respectively.

[0041] The self-replicating RNA contains sequences encoding multiple non-structural replication complex proteins. In certain embodiments, the synthetic self-replicating RNA may encode four non-structural replication complex proteins (i.e., nsP1, nsP2, nsP3, and nsP4). The non-structural replication complex proteins may be encoded by a single open reading frame (ORF).

[0042] The self-replicating RNA vector further contains a sequence encoding at least one CRISPR protein, which is detailed in the following section (I)(b).

[0043] Generally, self-replicating RNA vectors contain a 5' cap, a 5' untranslated region (UTR) at the 5' end, and a 3' UTR and poly-A tail at the 3' end. Self-replicating RNA vectors typically contain a promoter upstream of the sequence encoding the CRISPR protein. The upstream promoter may be a 26S subgenome promoter.

[0044] In some embodiments, the self-replicating RNA vector may further include a sequence encoding at least one selectable marker and / or a sequence encoding an inhibitor of the interferon response. Suitable selectable markers, though not limited to these examples, include puromycin, geneticin, neomycin, hydromycin B, and blastidinin S. Suitable interferon response inhibitors, though not limited to these examples, include vaccinia virus protein E3L, vaccinia virus protein B18R, influenza virus protein NS1, or lymphocytic choriomeningitis virus nucleoprotein.

[0045] Various protein-coding sequences can be separated by internal ribosome entry sequences (IRESs) or sequences encoding 2A peptides. Examples of suitable 2A peptides, though not limited to thosea asigna virus 2A peptide or T2A, foot-and-mouth disease virus 2A peptide or F2A, equine rhinitis A virus 2A peptide or E2A, and porcine rhinitis virus-1 2A peptide or P2A.

[0046] In certain embodiments, the self-replicating RNA vector may be obtained based on a modified Venezuelan encephalitis (VEE) virus and may include, from 5' to 3', a 5' cap, a 5' UTR, sequences encoding four non-structural replicases derived from VEE, a promoter, a sequence encoding a CRISPR protein, optionally an IRES, optionally a sequence encoding an E3L protein, optionally an IRES, optionally a sequence encoding a selectable marker, an alphavirus 3' UTR, and a poly-A tail (see Figure 1A).

[0047] (b) CRISPR protein The self-replicating RNA vector also contains a sequence encoding at least one CRISPR protein. CRISPR proteins, which provide adaptive immunity against invading nucleic acids, are present in various bacteria and archaea. In various embodiments, the CRISPR protein may be a type II Cas9 protein, a type V Cas12 (formerly called Cpf1) protein, a type VI Cas13 (formerly called C2cd) protein, a CasX protein, or a CasY protein.

[0048] CRISPR proteins are found in Acaryochloris spp., Acetohalobium spp., Acidaminococcus spp., Acidithiobacillus spp., Acidothermus spp., Akkermansia spp., Alicyclobacillus spp., Allochromatium spp., Ammonifex spp., Anabaena spp., Arthrospira spp., Bacillus spp., Bifidobacterium spp., Burkholderiales spp., Caldicelulosiruptor spp., Campylobacter spp., Candidatus spp., Clostridium spp., Corynebacterium spp., Crocosphaera spp., Cyanothece spp., Deltaproteobacterium spp., Exiguobacterium spp., Finegoldia spp., Francisella spp., Ktedonobacter spp., Lachnospiraceae spp., Lactobacillus spp., Leptotrichia spp., Lyngbya spp., Marinobacter spp., Methanohalobium spp., Microscilla spp., Microcoleus spp., Microcystis spp., Mycoplasma spp., Natranaerobius spp., Neisseria spp., Nitratifractor spp., Nitrosococcus spp., Nocardiopsis spp., Nodularia spp., Nostoc spp., Oenococcus spp., Oscillatoria spp., Parasutterella spp., Pelotomaculum spp., Petrotoga spp., Planctomyces spp., Polaromonas spp., Prevotella spp., Pseudoalteromonas spp., Ralstonia spp., Ruminococcus spp.These may be derived from Staphylococcus spp., Streptococcus spp., Streptomyces spp., Streptosporangium spp., Synechococcus spp., Thermosipho spp., Verrucomicrobia spp., or Wolinella spp. Various CRISPR protein sequences and their derivatives are known and can be found in the scientific literature, which is incorporated herein by reference.

[0049] In some embodiments, the CRISPR protein is Streptococcus pyogenes Cas9, Francisella novicida Cas9, Staphylococcus aureus Cas9, Streptococcus thermophilus Cas9, Streptococcus pasteurianus Cas9, Campylobacter jejuni Cas9, Neisseria meningitis Cas9, Neisseria cinerea Cas9, Francisella novicida Cas12, Acidaminococcus sp. Cas12, Lachnospiraceae bacterium ND2006 Cas12a, Leptotrichia wadeii Cas13a, Leptotrichia shahii Cas13a, Prevotella sp. P5-125 Cas13, Ruminococcus flavefaciens Cas13d, Deltaproteobacterium CasX, Planctomyces CasX, or Candidatus CasY. In certain embodiments, the CRISPR protein is either Streptococcus pyogenes Cas9 or Staphylococcus aureus Cas9.

[0050] CRISPR proteins can be wild-type or naturally occurring proteins. Wild-type CRISPR proteins generally contain two nuclease domains (e.g., the Cas9 protein contains RuvC and HNH domains), each of which cleaves one strand of a double-stranded sequence. CRISPR proteins also contain domains that interact with guide RNA (e.g., REC1, REC2) or RNA / DNA heteroduplexes (e.g., REC3), and domains that interact with protospacer fringe motifs (PAMs) (i.e., PAM-interacting domains).

[0051] Alternatively, CRISPR proteins may be modified or designed to alter their activity, specificity, and / or stability. For example, a CRISPR protein may be designed to include one or more modifications / mutations (i.e., at least one amino acid substitution, deletion, and / or insertion). Modified CRISPR proteins may exhibit changes in catalytic (nuclease) activity, improved target site specificity, reduced off-target effects, altered PAM specificity, and increased stability.

[0052] In various embodiments, the CRISPR protein may be a nuclease (i.e., cleaving both strands of a double-stranded nucleotide sequence or cleaving a single-stranded nucleotide sequence). In other embodiments, the CRISPR protein may be a nickase that cleaves one strand of a double-stranded sequence. A nickase can be designed by inactivating one of the nuclease domains of the CRISPR protein. For example, to produce Cas9 nickase (e.g., nCas9), the RuvC domain of the Cas9 protein may be inactivated by mutations such as D10A, D8A, E762A and / or D986A, or the HNH domain of the Cas9 protein may be inactivated by mutations such as H840A, H559A, N854A, N856A and / or N863A (see the numbering system for Streptococcus pyogenes Cas9, SpyCas9). Equivalent mutations in other CRISPR proteins can produce nickases (e.g., nCas12). Inactivation of both nuclease domains produces CRISPR proteins that lack cleavage activity (i.e., non-catalytic proteins or nuclease-inactive proteins, such as dCas9 and dCas12).

[0053] CRISPR proteins can also be designed by substitutions, deletions, and / or insertions of one or more amino acids to have improved target specificity, improved fidelity, altered PAM specificity, reduced off-target effects, and / or increased stability. Not limited to one example of mutations that improve target specificity, improve fidelity, and / or reduce off-target effects, include N497A, R661A, Q695A, K810A, K848A, K855A, Q926A, K1003A, R1060A, and / or D1135E (see the SpyCas9 numbering system).

[0054] RNA vector sequences encoding CRISPR proteins can be codon-optimized for efficient translation into the target eukaryotic cell. Codon optimization programs are available as freeware or from commercial sources. In certain embodiments, sequences encoding CRISPR proteins can be codon-optimized for efficient expression in human cells.

[0055] Any heterogeneous domain In some embodiments, a CRISPR protein may be designed to contain at least one heterogeneous domain (i.e., a CRISPR protein may be ligated to one or more heterogeneous domains). The heterogeneous domain may be a nuclear localization signal (NLS), a cell permeability domain, a marker domain (e.g., a fluorescent protein), a chromatin regulatory motif, an epigenetic modification domain (e.g., a deaminase domain and a histone acetyltransferase domain), a transcriptional regulatory domain, an RNA aptamer binding domain, or a non-CRISPR nuclease domain. When two or more heterogeneous domains are fused to a CRISPR protein, the two or more heterogeneous domains may be the same or different. One or more heterogeneous domains may be ligated to a CRISPR protein at its N-terminus, C-terminus, internal position, or a combination thereof. Ligation may be direct via chemical bonding or indirect via one or more linkers. Suitable linkers are known in the art. In some embodiments, ligation may be via a 2A peptide sequence.

[0056] In some embodiments, one or more heterogeneous domains may be nuclear localization signals (NLS). Examples of nuclear localization signals that are not limited to these include PKKKRKV (SEQ ID NO: 1), PKKKRRV (SEQ ID NO: 2), KRPAATKKAGQAKKKK (SEQ ID NO: 3), YGRKKRRQRRR (SEQ ID NO: 4), RKKRRQRRR (SEQ ID NO: 5), PAAKRVKLD (SEQ ID NO: 6), RQRRNELKRSP (SEQ ID NO: 7), VSRKRPRP (SEQ ID NO: 8), PPKKARED (SEQ ID NO: 9), PQPKKKPL (SEQ ID NO: 10), SALIKKKKKMAP (SEQ ID NO: 11), This includes PKQKKRK (SEQ ID NO: 12), RKLKKKIKKL (SEQ ID NO: 13), REKKKFLKRR (SEQ ID NO: 14), KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 15), RKCLQAGMNLEARKTKK (SEQ ID NO: 16), NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 17), and RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 18).

[0057] In other embodiments, one or more heterogeneous domains may be cell-permeable domains. Examples of suitable cell-permeable domains include, but are not limited to, GRKKRRQRRRPPQPKKKRKV (SEQ ID NO: 19), PLSSIFSRIGDPPKKKRKV (SEQ ID NO: 20), GALFLGWLGAAGSTMGAPKKKRKV (SEQ ID NO: 21), GALFLGFLGAAGSTMGAWSQPKKKRKV (SEQ ID NO: 22), KETWWETWWTEWSQPKKKRKV (SEQ ID NO: 23), YARAAARQARA (SEQ ID NO: 24), THRLPRRRRRR (SEQ ID NO: 25), GGRRARRRRRR (SEQ ID NO: 26), RRQRRTSKLMKR (SEQ ID NO: 27), GWTLNSAGYLLGKINLKALAALAKKIL (SEQ ID NO: 28), KALAWEAKLAKALAKALAKHLAKALAKALKCEA (SEQ ID NO: 29), and RQIKIWFQNRRMKWKK (SEQ ID NO: 30).

[0058] In alternative embodiments, one or more heterogeneous domains may be marker domains. The marker domains include a fluorescent protein and a purified tag or epitope tag. While not limited to suitable fluorescent proteins, examples include: green fluorescent proteins (e.g., GFP, eGFP, GFP-2, tagGFP, turboGFP, Emerald, Azami Green, monomeric Azami Green, CopGFP, AceGFP, ZsGreen1), yellow fluorescent proteins (e.g., YFP, EYFP, Citrine, Venus, YPet, PhiYFP, ZsYellow1), blue fluorescent proteins (e.g., BFP, EBFP, EBFP2, Azurite, mKalama1, GFPuv, Sapphire, T-sapphire), cyan fluorescent proteins (e.g., ECFP, Cerulean, CyPet, AmCyan1, Midoriishi-Cyan), and red fluorescent proteins (e.g., mKate, mKate2, mPlum, DsRed). This includes monomers (mCherry, mRFP1, DsRed-Express, DsRed2, DsRed-Monomer, HcRed-Tandem, HcRed1, AsRed2, eqFP611, mRaspberry, mStrawberry, Jred) and orange fluorescent proteins (e.g., mOrange, mKO, Kusabira-Orange, monomeric Kusabira-Orange, mTangerine, tdTomato) or combinations thereof. Fluorescent proteins may consist of tandem repeats of one or more fluorescent proteins (e.g., Suntag). Not limited examples of suitable purification tags or epitope tags include 6xHis (SEQ ID NO: 40), FLAG®, HA, GST, Myc, and SAM.Examples of heterogeneous fusions that facilitate the detection or enrichment of CRISPR complexes include, but are not limited to, streptavidin (Kipriyanov et al., Human Antibodies, 1995, 6(3):93-101), avidin (Airenne et al., Biomolecular Engineering, 1999, 16(1-4):87-92), monomeric avidin (Laitinen et al., Journal of Biological Chemistry, 2003, 278(6):4010-4014), and peptide tags that promote biotinylation during recombinant production (Cull et al., Methods in Enzymology, 2000, 326:430-440).

[0059] In further other embodiments, one or more heterogeneous domains may be chromatin regulatory motifs (CMMs). Examples of CMMs that are not limited to these include nucleosome interacting peptides derived from high-mobility group (HMG) proteins (e.g., HMGB1, HMGB2, HMGB3, HMGN1, HMGN2, HMGN3a, HMGN3b, HMGN4, and HMGN5 proteins), the central globular domain of histone H1 variants (e.g., histone H1.0, H1.1, H1.2, H1.3, H1.4, H1.5, H1.6, H1.7, H1.8, H1.9, and H1.10), or chromatin remodeling complexes (e.g., SWI / SNF as described in U.S. Patent Application No. 16 / 031,819 (this disclosure is incorporated herein by reference)). SWI tch / S ucrose N on- F (ermentable), ISWI ( I mitation SWI tch), CHD( C hromodomain- H elicase- D NA binding), Mi-2 / NuRD( Nu cleosome R emodeling andD This includes the DNA-binding domain of the eacetylase, INO80, SWR1, and RSC complex. Suitable CMMs may also be derived from topoisomerases, helicases, or viral proteins. The sources of CMMs can be diverse. CMMs may be derived from humans, animals (i.e., vertebrates and invertebrates), plants, algae, or yeast.

[0060] In yet another embodiment, one or more heterogeneous domains may be epigenetic modification domains. Suitable epigenetic modification domains, not limited to these, include DNA deamination (e.g., cytidine deaminase, adenosine deaminase, guanine deaminase), DNA methyltransferase activity (e.g., cytosine methyltransferase), DNA demethylase activity, DNA amination, DNA oxidation activity, DNA helicase activity, histone acetyltransferase (HAT) activity (e.g., HAT domain from E1A-binding protein p300), histone deacetylase activity, histone methyltransferase activity, histone demethylase activity, and histone acetyltransferase activity. The epigenetic modification domains include those having enzyme activity, histone phosphatase activity, histone ubiquitin ligase activity, histone deubiquitination activity, histone adenylation activity, histone deadenylation activity, histone SUMOylation activity, histone deSUMOylation activity, histone ribosylation activity, histone deribosylation (deribosylation) activity, histone myristoylation activity, histone demyristoylation activity, histone citrullination activity, histone alkylation activity, histone dealkylation activity, or histone oxidation activity. In certain embodiments, the epigenetic modification domains may include cytidine deaminase activity, adenosine deaminase activity, histone acetyltransferase activity, or DNA methyltransferase activity.

[0061] In other embodiments, one or more heterogeneous domains may be transcriptional regulatory domains (i.e., transcriptional activating domains or transcriptional repressing domains). Suitable transcriptional activating domains include, but are not limited to, the herpes simplex virus VP16 domain, VP64 (i.e., four tandem copies of VP16), VP160 (i.e., ten tandem copies of VP16), the NFκB p65 activating domain (p65), the Epstein-Barr virus R trans-activator (Rta) domain, VPR (i.e., VP64 + p65 + Rta), the p300-dependent transcriptional activating domain, the p53 activating domains 1 and 2, the heat shock factor 1 (HSF1) activating domain, the Smad4 activating domain (SAD), the cAMP response element binding protein (CREB) activating domain, the E2A activating domain, the activated T cell nuclear factor (NFAT) activating domain, or a combination thereof. Appropriate transcriptional repression domains, though not limited to these, include Kruppel-associated box (KRAB) repression domains, Mxi repression domains, inducible cAMP early repressor (ICER) domains, YY1 glycine-rich repression domains, Sp1-like repressors, E(spl) repressors, IκB repressors, Sin3 repressors, methylated CpG-binding protein 2 (MeCP2) repressors, or combinations thereof. Transcriptional activating or repressing domains can be genetically fused to the Cas9 protein or bound via non-covalent protein-protein, protein-RNA, or protein-DNA interactions.

[0062] In a further embodiment, one or more heterogeneous domains may be RNA aptamer-binding domains (Konermann et al., Nature, 2015, 517(7536):583-588; Zalatan et al., Cell, 2015, 160(1-2):339-50). Examples of suitable RNA aptamer protein domains include MS2 coat protein (MCP), PP7 bacteriophage coat protein (PCP), Mu bacteriophage Com protein, lambda bacteriophage N22 protein, stem-loop binding protein (SLBP), fragile X intellectual disability-related protein 1 (FXR1), and proteins, fragments thereof, or derivatives thereof derived from bacteriophages (such as AP205, BZ13, f1, f2, fd, fr, ID2, JP34 / GA, JP501, JP34, JP500, KU1, M11, M12, MX1, NL95, PP7, φCb5, φCb8r, φCb12r, φCb23r, Qβ, R17, SP-β, TW18, TW19, and VK).

[0063] In further other embodiments, one or more heterogeneous domains may be non-CRISPR nuclease domains. Suitable nuclease domains may be obtained from any endonuclease or exonuclease. Not limited examples of endonucleases from which the nuclease domain may be derived include, but are not limited to, restriction endonucleases and homing endonucleases. In some embodiments, the nuclease domain may be derived from a II-S type restriction endonuclease. II-S type endonucleases typically cleave DNA at a site several base pairs away from the recognition / binding site and, as such, have separable binding and cleavage domains. These enzymes are generally monomers that transiently associate to form dimers and cleave each strand of DNA at twisted positions. Not limited examples of suitable II-S type endonucleases include BfiI, BpmI, BsaI, BsgI, BsmBI, BsmI, BspMI, FokI, MboII, and SapI. In some embodiments, the nuclease domain may be a FokI nuclease domain or a derivative thereof. The II-S type nuclease domain may be modified to promote dimerization of two different nuclease domains. For example, the cleavage domain of FokI may be modified by mutating specific amino acid residues. In some, but not limited to, amino acid residues at positions 446, 447, 479, 483, 484, 486, 487, 490, 491, 496, 498, 499, 500, 531, 534, 537, and 538 of the FokI nuclease domain are targets for modification. In certain embodiments, the FokI nuclease domain may comprise a first FokI half-domain containing Q486E, I499L, and / or N496D mutations, and a second FokI half-domain containing E490K, I538K, and / or H537R mutations.

[0064] (II) A complex containing a self-replicating RNA vector and guide RNA. Another aspect of the present disclosure encompasses a complex comprising one of the self-replicating RNA vectors described above in Section (I) and at least one guide RNA, wherein the guide RNA is designed to form a complex with a CRISPR protein encoded by the self-replicating RNA vector.

[0065] Guide RNA interacts with the CRISPR protein and the target sequence in the target nucleic acid, guiding the CRISPR protein to the target sequence. The target sequence has no sequence restrictions except that it is adjacent to a protospacer-adjacent motif (PAM) sequence. CRISPR proteins can recognize various PAM sequences. For example, the PAM sequences of the Cas9 protein include 5'-NGG, 5'-NGGNG, 5'-NNAGAAW, 5'-NNNNGATT, and 5-NNNNRYAC, while the PAM sequences of the Cas12 protein include 5'-TTN and 5'-TTTV, where N is defined as any nucleotide, R is defined as either G or A, W is defined as either A or T, Y is defined as either C or T, and V is defined as A, C, or G. Generally, Cas9 PAMs are located at 3' of the target sequence, and Cas12 PAMs are located at 5' of the target sequence.

[0066] Guide RNAs are designed to form a complex with a specific CRISPR protein. Generally, guide RNAs contain (i) a CRISPR RNA (crRNA) with a guide sequence at its 5' end that hybridizes with the target sequence, and (ii) a trans-acting crRNA (tracrRNA) sequence that interacts with the CRISPR protein. The crRNA guide sequence of each guide RNA is different (i.e., sequence-specific). The tracrRNA sequence is generally identical in guide RNAs designed to form a complex with a specific CRISPR protein.

[0067] The crRNA guide sequence is designed to hybridize with the target sequence (i.e., protospacer) in the sequence of interest. Generally, the complementarity between the crRNA and the target sequence is at least 80%, at least 85%, at least 90%, at least 95%, or at least 99%. In certain embodiments, the complementarity is perfect (i.e., 100%). In various embodiments, the length of the crRNA guide sequence can range from about 15 nucleotides to about 25 nucleotides. For example, the crRNA guide sequence may be about 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides long. In certain embodiments, the crRNA is about 19, 20, or 21 nucleotides long. In one embodiment, the crRNA guide sequence has a length of 20 nucleotides.

[0068] Guide RNA contains a repetitive sequence that forms at least one stem-loop structure that interacts with the CRISPR protein, and a 3' sequence that generally remains single-stranded. The length of each loop and stem can vary. For example, loops can range from about 3 to about 10 nucleotides in length, and stems can range from about 6 to about 20 base pairs in length. Stems can contain one or more bulges of 1 to about 10 nucleotides. The length of the single-stranded 3' region can vary. The tracrRNA sequence in guide RNA is generally based on the sequence of wild-type tracrRNA that interacts with the wild-type CRISPR protein. The wild-type sequence can be modified to promote secondary structure formation, increase the stability of secondary structures, and promote expression in eukaryotic cells, among other things. For example, one or more nucleotide changes can be introduced into the guide RNA coding sequence. The tracrRNA sequence can range in length from about 50 to about 300 nucleotides. In various embodiments, tracrRNA may be in the range of length of approximately 50 to approximately 90 nucleotides, approximately 90 to approximately 110 nucleotides, approximately 110 to approximately 130 nucleotides, approximately 130 to approximately 150 nucleotides, approximately 150 to approximately 170 nucleotides, approximately 170 to approximately 200 nucleotides, approximately 200 to approximately 250 nucleotides, or approximately 250 to approximately 300 nucleotides.

[0069] In some embodiments, the guide RNA may be a single molecule (e.g., a single guide RNA (sgRNA) or one type of sgRNA), with the crRNA sequence ligated to a tracrRNA sequence. In some embodiments, the guide RNA may be two distinct molecules (e.g., two types of gRNA). The first molecule comprises a crRNA containing a 3' sequence (containing about 6 to about 20 nucleotides) that can base-pair with the 5' end of the second molecule, and the second molecule comprises a tracrRNA containing a 5' sequence (containing about 6 to about 20 nucleotides) that can base-pair with the 3' end of the first molecule.

[0070] In some embodiments, the tracrRNA sequence of the guide RNA may be modified to include one or more aptamer sequences (Konermann et al., Nature, 2015, 517(7536):583-588; Zalatan et al., Cell, 2015, 160(1-2):339-50). Suitable aptamer sequences include aptamer sequences that bind to aptamer proteins selected from MCP, PCP, Com, SLBP, FXR1, AP205, BZ13, f1, f2, fd, fr, ID2, JP34 / GA, JP501, JP34, JP500, KU1, M11, M12, MX1, NL95, PP7, φCb5, φCb8r, φCb12r, φCb23r, Qβ, R17, SP-β, TW18, TW19, VK, fragments thereof, or derivatives thereof. Those skilled in the art are aware that the length of aptamer sequences can vary.

[0071] In other embodiments, the guide RNA may further include at least one detectable label. The detectable label may be a fluorophore (e.g., FAM, TMR, Cy3, Cy5, Texas Red, Oregon Green, Alexa Fluor, Halo tag, or appropriate fluorescent dye), a detection tag (e.g., biotin and digoxigenin), a quantum dot, or a gold particle.

[0072] Guide RNA may comprise standard ribonucleotides and / or modified ribonucleotides. In some embodiments, guide RNA may comprise standard deoxyribonucleotides or modified deoxyribonucleotides. In some embodiments where guide RNA is synthesized enzymatically (i.e., in vivo or in vitro), guide RNA generally comprises standard ribonucleotides. In some embodiments where guide RNA is synthesized chemically, guide RNA may comprise standard or modified ribonucleotides and / or deoxyribonucleotides. Modified ribonucleotides and / or deoxyribonucleotides may comprise base modifications (e.g., pseudouridine, 2-thiouridine, and N6-methyladenosine, etc.) and / or sugar modifications (e.g., 2'-O-methyl, 2'-fluoro, 2'-amino, and locked nucleic acids (LNAs, etc.). The guide RNA backbone may also be modified to include phosphorothioate bonds, boranophosphate bonds, or peptide nucleic acids.

[0073] (III) Eukaryotic cells Another aspect of this disclosure includes a eukaryotic cell or cell line comprising one of the self-replicating RNA vectors described above in Section (I). That is, the eukaryotic cell or cell line is transfected with one of the self-replicating RNA vectors. In some embodiments, the eukaryotic cell or cell line may further comprise at least one guide RNA designed to form a complex with the CRISPR protein encoded by the RNA vector.

[0074] Eukaryotic cells or cell lines may be human cells, non-human mammalian cells, non-mammalian vertebrate cells, invertebrate cells, plant cells, or single-celled eukaryotes. Examples of suitable eukaryotic cells are detailed in section (V)(c) below. Eukaryotic cells may be in vitro, ex vivo, or in vivo.

[0075] (IV) Plasmid vector encoding self-replicating RNA Further aspects of the present disclosure provide plasmid vectors encoding self-replicating RNA as described above in Section (I). In particular, the plasmid vectors include sequences encoding alphavirus-derived non-structural replication complex proteins, sequences encoding CRISPR proteins, and additional viral sequences (such as the 5'UTR, subgenome promoter, and 3'UTR), optionally selectable marker sequences, optionally interferon inhibitor sequences, optionally IRES, etc.

[0076] Generally, plasmid vectors encoding self-replicating RNA are DNA vectors. The sequence encoding the self-replicating RNA may be operably ligated to a promoter sequence recognized by phage RNA polymerase for RNA synthesis in vitro. For example, the promoter sequence may be a T7, T3, or SP6 promoter sequence, or a variant of a T7, T3, or SP6 promoter sequence. The promoter sequence may be wild-type or modified for more efficient or effective expression. The plasmid vector may further include at least one transcription termination sequence, as well as at least one origin of replication and / or selectable marker sequences (e.g., antibiotic resistance genes) for proliferation in bacterial cells. The plasmid vector may originate from pUC, pBR322, pET, pBluescript, or variants thereof. Further information regarding the vector and its use can be found in "Current Protocols in Molecular Biology" Ausubel et al., John Wiley & Sons, New York, 2003 or "Molecular Cloning: A Laboratory Manual" Sambrook & Russell, Cold Spring Harbor Press, Cold Spring Harbor, NY, 3 rd It can be found in the 2001 edition.

[0077] During the in vitro synthesis of self-replicating RNA, the RNA can be purified, 5' capped, and polyadenylated using standard procedures or commercially available kits.

[0078] (V) Methods for genome editing Further aspects of this disclosure encompass methods for genome editing in eukaryotic cells. Generally, the methods include introducing one of the self-replicating RNA vectors described above in Section (I) and at least one guide RNA designed to form a complex with the CRISPR protein encoded by the RNA vector into a eukaryotic cell of interest. The methods may also include introducing the complexes identified in Section (II) above into a eukaryote. The CRISPR system comprises the CRISPR protein and the guide RNA.

[0079] In some embodiments in which the CRISPR protein comprises nuclease activity or nickase activity, genome editing may comprise substitution of at least one nucleotide, deletion of at least one nucleotide, and / or insertion of at least one nucleotide. In some iterations, the method comprises introducing one CRISPR system comprising nuclease activity or two CRISPR systems comprising nickase activity into a eukaryotic cell, without introducing a donor polynucleotide, such that the CRISPR system introduces a double-strand break at a target site of a sequence of interest, and the repair of the double-strand break by the cell's DNA repair process inactivates the sequence (i.e., gene knockout) by introducing at least one nucleotide change (i.e., an indel). In other iterations, the method involves introducing a CRISPR system containing nuclease activity or two CRISPR systems containing nickase activity, along with a donor polynucleotide, into eukaryotic cells, such that the CRISPR system introduces a double-strand break at a target site of the sequence of interest, and the double-strand break is repaired by the cell's DNA repair process, resulting in the insertion or exchange of the sequence in the donor polynucleotide at the target site of the sequence of interest (i.e., gene modification or gene knock-in).

[0080] In one embodiment in which the CRISPR protein has epigenetic modification activity or transcriptional regulatory activity, genome editing may include the conversion (i.e., base editing) of at least one nucleotide in or near the target site, modification of at least one nucleotide in or near the target site, modification of at least one histone protein in or near the target site, and / or transcriptional changes in or near the target site in the chromosomal sequence.

[0081] (a) Introduction into cells As described above, the method involves introducing at least one self-replicating RNA vector, as described in Section (I), at least one guide RNA designed to form a complex with the CRISPR protein encoded by the RNA vector, and optionally a donor polynucleotide into a eukaryotic cell. These molecules can be introduced into the target cell by a variety of means.

[0082] For example, self-replicating RNA vectors, guide RNAs, and any donor polynucleotides can be used to deliver genes into target cells. Suitable gene delivery methods include nucleofection (or electroporation), calcium phosphate-mediated gene delivery, cationic polymer gene delivery (e.g., DEAE-dextran or polyethyleneimine), viral transduction, virosomal gene delivery, virion gene delivery, liposome gene delivery, cationic liposome gene delivery, immunoliposome gene delivery, non-liposomal lipid gene delivery, dendrimer gene delivery, heat shock gene delivery, magnetofection, lipofection, gene gun delivery, impalefection, sonoporation, optical gene delivery, and nucleic acid uptake enhanced by proprietary drugs. Gene transfer methods are well known in this field (see, for example, "Current Protocols in Molecular Biology" Ausubel et al., John Wiley & Sons, New York, 2003 or "Molecular Cloning: A Laboratory Manual" Sambrook & Russell, Cold Spring Harbor Press, Cold Spring Harbor, NY, 3rd edition, 2001). In other embodiments, molecules may be introduced into cells by microinjection. For example, molecules may be injected into the cytoplasm or nucleus of the target cell. The amount of each molecule introduced into the cell may vary, but those skilled in the art are familiar with means of determining appropriate amounts.

[0083] Various molecules can be introduced into cells simultaneously or sequentially. For example, a self-replicating RNA vector and a guide RNA can be introduced at the same time. Alternatively, one may be introduced into the cell first, and the other afterward.

[0084] Generally, cells are maintained under conditions suitable for cell proliferation and / or maintenance. Appropriate cell culture conditions are well known in the art and are described, for example, in Santiago et al., Proc. Natl. Acad. Sci. USA, 2008, 105:5809-5814; Moehle et al. Proc. Natl. Acad. Sci. USA, 2007, 104:3055-3060; Urnov et al., Nature, 2005, 435:646-651; and Lombardo et al., Nat. Biotechnol., 2007, 25:1298-1306. Those skilled in the art will recognize that methods for culturing cells are well known in the art and can vary depending on the cell type. In all cases, standard optimization may be used to determine the best technique for a particular cell type.

[0085] (b) Any donor polynucleotide In some embodiments, where the CRISPR protein comprises nuclease activity or nickase activity, the method may further include introducing at least one donor polynucleotide into the cell. The donor polynucleotide may be single-stranded or double-stranded, linear or circular, and / or RNA or DNA. In some embodiments, the donor polynucleotide may be a vector (e.g., a plasmid vector).

[0086] A donor polynucleotide contains at least one donor sequence. In some embodiments, the donor sequence of a donor polynucleotide may be an endogenous or modified version of a native chromosome sequence. For example, the donor sequence may be essentially identical to a portion of the chromosome sequence targeted by the CRISPR system or near it, but containing at least one nucleotide exchange. Thus, by integration into or exchange with the native sequence, the sequence at the target chromosomal location contains at least one nucleotide exchange. For example, this exchange may be one or more nucleotide insertions, one or more nucleotide deletions, one or more nucleotide substitutions, or a combination thereof. As a result of the integration of the “gene modification” of the modified sequence, the cell may produce a modified gene product from the target chromosome sequence.

[0087] In other embodiments, the donor sequence of a donor polynucleotide may be an exogenous sequence. In this specification, “exogenous” sequence means a sequence that is not native to the cell, or a sequence whose native location in the cell’s genome is different. For example, an exogenous sequence may include a protein-coding sequence, which may be operably linked to an exogenous promoter-regulatory sequence so that integration into the genome allows the cell to express the protein encoded by the integrated sequence. Alternatively, an exogenous sequence may be integrated into a chromosomal sequence so that its expression is regulated by an endogenous promoter-regulatory sequence. In other iterations, the exogenous sequence may include transcriptional regulatory sequences, other expression regulatory sequences, and RNA-coding sequences. As described above, the integration of an exogenous sequence into a chromosomal sequence is called “knock-in.”

[0088] As can be seen by those skilled in the art, the length of the donor sequence can vary. For example, the length of the donor sequence can range from a few nucleotides to several hundred nucleotides, or from several hundred nucleotides to several hundred thousand nucleotides.

[0089] Typically, the donor sequence in a donor polynucleotide is adjacent to both the upstream and downstream sequences, and these sequences have substantial sequence identity with the sequences located upstream and downstream, respectively, of the sequence targeted by the CRISPR system. Due to this sequence similarity, the upstream and downstream sequences of the donor polynucleotide enable homologous recombination between the donor polynucleotide and the target chromosomal sequence, allowing the donor sequence to be incorporated into (or replaced by) the chromosomal sequence.

[0090] In this specification, the upstream sequence refers to a nucleic acid sequence having substantial sequence identity with a chromosomal sequence upstream of the sequence targeted by the CRISPR system. Similarly, the downstream sequence refers to a nucleic acid sequence having substantial sequence identity with a chromosomal sequence downstream of the sequence targeted by the CRISPR system. In this specification, the term "substantial sequence identity" refers to sequences having at least about 75% sequence identity. Thus, the upstream and downstream sequences in a donor polynucleotide may have about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with the upstream or downstream sequence of the target sequence. In exemplary embodiments, the upstream and downstream sequences in a donor polynucleotide may have approximately 95% or 100% sequence identity with the upstream or downstream chromosomal sequence of the sequence targeted by the CRISPR system.

[0091] In some embodiments, the upstream sequence has substantial sequence identity with a chromosomal sequence located immediately upstream of the sequence targeted by the CRISPR system. In other embodiments, the upstream sequence has substantial sequence identity with a chromosomal sequence located within approximately 100 nucleotides upstream of the target sequence. For example, the upstream sequence may have substantial sequence identity with a chromosomal sequence located approximately 1 to 20, 21 to 40, 41 to 60, 61 to 80, or 81 to 100 nucleotides upstream of the target sequence. In some embodiments, the downstream sequence has substantial sequence identity with a chromosomal sequence located immediately downstream of the sequence targeted by the CRISPR system. In other embodiments, the downstream sequence has substantial sequence identity with a chromosomal sequence located within approximately 100 nucleotides downstream of the target sequence. Therefore, for example, downstream sequences may have substantial sequence identity with chromosomal sequences located approximately 1 to 20, 21 to 40, 41 to 60, 61 to 80, or 81 to 100 nucleotides downstream from the target sequence.

[0092] Each upstream or downstream sequence may be approximately 20 to 5000 nucleotides long. In some embodiments, the upstream and downstream sequences may contain approximately 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2800, 3000, 3200, 3400, 3600, 3800, 4000, 4200, 4400, 4600, 4800, or 5000 nucleotides. In specific embodiments, the upstream and downstream sequences may be approximately 50 to 1500 nucleotides long.

[0093] (c) Cell type A variety of eukaryotic cells are suitable for use in the methods disclosed herein. For example, cells may be human cells, non-human mammalian cells, non-mammalian vertebrate cells, invertebrate cells, insect cells, plant cells, yeast cells, or single-celled eukaryotes. In some embodiments, cells may be primary cells isolated directly from a particular tissue. In other embodiments, cells may be cell lines. In yet another embodiment, cells may be a single cell embryo. For example, non-human mammalian embryos include rat, hamster, rodent, rabbit, cat, dog, sheep, pig, cattle, horse, and primate embryos. In yet another embodiment, cells may be stem cells (such as embryonic stem cells, ES-like stem cells, fetal stem cells, adult stem cells, and induced pluripotent stem cells). In some embodiments, the stem cells are not human embryonic stem cells. Furthermore, stem cells may include stem cells produced by techniques disclosed in WO2003 / 046141 (which is incorporated in its entirety herein) or Chung et al. (Cell Stem Cell, 2008, 2:113-117). In various embodiments, cells may be in vitro (e.g., in a culture), ex vivo (i.e., in tissue isolated from a living organism) or in vivo (i.e., in a living organism). In exemplary embodiments, cells are mammalian cells or mammalian cell lines. In specific embodiments, cells are human cells or human cell lines.

[0094] Appropriate mammalian cell or cell line examples include, but are not limited to, human embryonic kidney cells (HEK293, HEK293T); human cervical cancer cells (HELA); human lung cells (W138); human hepatocytes (Hep G2); human U2-OS osteosarcoma cells, human A549 cells, human A-431 cells and human K562 cells; Chinese hamster ovary (CHO) cells, baby hamster kidney (BHK) cells; mouse myeloma NS0 cells, human primary fibroblasts, human precipitous fibroblasts, mouse embryonic fibroblast 3T3 cells (NIH3T3), mouse B lymphoma A20 cells; mouse melanoma B16 cells; mouse myoblast C2C12 cells; mouse myeloma SP2 / 0 cells; mouse embryonic mesenchymal C3H-10T1 / 2 cells; mouse This includes mouse cancer CT26 cells, mouse prostate DuCuP cells, mouse mammary gland EMT6 cells, mouse liver cancer Hepa1c1c7 cells, mouse myeloma J5582 cells, mouse epithelial MTD-1A cells, mouse cardiomyocyte MyEnd cells, mouse kidney RenCa cells, mouse pancreatic RIN-5F cells, mouse melanoma X64 cells, mouse lymphoma YAC-1 cells, rat glioblastoma 9L cells, rat B lymphoma RBL cells, rat neuroblastoma B35 cells, rat liver cancer cells (HTC), buffalo rat liver BRL 3A cells, canine kidney cells (MDCK), canine mammary gland (CMT) cells, rat osteosarcoma D17 cells, rat monocyte / macrophage DH82 cells, monkey kidney SV-40 transformed fibroblast (COS7) cells, monkey kidney CVI-76 cells, and African green monkey kidney (VERO-76) cells. A comprehensive list of mammalian cell lines can be found in the catalog of the American Type Culture Collection (ATCC, Manassas, VA).

[0095] (VI) Applications The compositions and methods disclosed herein may be used in a variety of therapeutic, diagnostic, industrial, and research applications. In some embodiments, the disclosure may be used to modify any chromosomal sequence of interest in cells, animals, or plants in order to model and / or study gene function, study a desired genetic or epigenetic state, or study biochemical pathways involved in various diseases or disorders. For example, a transgenic organism may be created to model a disease or disorder in which the expression of one or more nucleic acid sequences associated with the disease or disorder is altered. Disease models may be used to study the effects of mutations on organisms, study the onset and / or progression of diseases, study the effects of pharmaceutically active compounds on diseases, and / or evaluate the efficacy of potential gene therapy strategies.

[0096] In other embodiments, compositions and methods may be used to perform efficient and cost-effective functional genome screening (which can be used to investigate the function of genes involved in specific biological processes and what changes in gene expression may affect those processes), or to perform saturated mutagenesis or deep scanning mutagenesis of genomic loci in combination with cellular phenotype. For example, saturated mutagenesis or deep scanning mutagenesis may be used to determine the critical minimum features and distinct vulnerabilities of functional elements necessary for gene expression, drug resistance, and disease recovery.

[0097] In further embodiments, the compositions and methods disclosed herein may be used for diagnostic tests used to establish the presence of a disease or disorder and / or to determine treatment options. Examples of appropriate diagnostic tests include the detection of specific mutations in cancer cells (e.g., specific mutations in EGFR and HER2, etc.), the detection of specific mutations associated with specific diseases (e.g., trinucleotide repeats, mutations in β-globin associated with sickle cell anemia, specific SNPs, etc.), the detection of hepatitis, and the detection of viruses (e.g., Zika virus).

[0098] In further embodiments, the compositions and methods disclosed herein may be used to correct gene mutations associated with specific diseases or disorders (for example, to correct mutations in the globin gene associated with sickle cell disease or thalassemia, to correct mutations in the adenosine deaminase gene associated with severe combined immunodeficiency (SCID), to reduce the expression of HTT, the disease-causing gene for Huntington's disease, or to correct mutations in the rhodopsin gene for the treatment of retinitis pigmentosa). Such modifications may be carried out in ex vivo cells.

[0099] In further embodiments, the compositions and methods disclosed herein may be used to produce crops with improved traits or increased resistance to environmental stress. The disclosure may also be used to produce livestock or production animals with improved traits. For example, pigs have many attractive features as biomedical models (particularly in regenerative medicine or xenotransplantation).

[0100] definition Unless otherwise defined, all technical and scientific terms used herein have the meanings generally understood by those skilled in the art of the present invention. The following references provide general definitions of many terms used herein: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd Ed. 1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). In this specification, unless otherwise specified, the following terms have the meanings derived therefrom.

[0101] Where introducing elements of this disclosure or its preferred embodiments, the articles “a,” “an,” “the,” and “the” are intended to mean that there are one or more elements. The terms “include,” “contain,” and “have” are inclusive and are intended to mean that there may be further elements other than those described.

[0102] The term "approximately," when used in relation to a numerical value x, means, for example, x ± 5%.

[0103] In this specification, the terms “complementary” or “complementarity” refer to the relationship between double-stranded nucleic acids through base pairing via specific hydrogen bonds. Base pairing can be standard Watson-Crick base pairing (e.g., 5'-AGTC-3' pairs with the complementary sequence 3'-TCAG-5'). Base pairing can also be Hougsteen hydrogen bonds or reverse Hougsteen hydrogen bonds. Complementarity is usually measured with respect to the double-stranded region, so overhangs are excluded, for example. Complementarity between the two strands of a double-stranded region can be partial, and if only some (e.g., 70%) of the bases are complementary, it may be expressed as a percentage (e.g., 70%). Non-complementary bases are “mismatched”. If all the bases of the double-stranded region are complementary, the complementarity can be complete (i.e., 100%).

[0104] In this specification, the terms “CRISPR / Cas system” or “Cas9 system” refer to a complex comprising a Cas9 protein (i.e., a nuclease, nicasse, or catalytically inactive protein) and guide RNA.

[0105] In this specification, the term "endogenous sequence" refers to a chromosomal sequence that is native to a cell.

[0106] In this specification, the term "exogenous" refers to a sequence that is not natural to a cell, or a chromosomal sequence whose natural position in the cell's genome is different from that of the cell.

[0107] The term "expression" in relation to genes or polynucleotides refers to the transcription of a gene or polynucleotide, and, if necessary, the translation of the mRNA transcript into a protein or polypeptide. Therefore, as is clear from the context, protein or polypeptide expression results from the transcription and / or translation of the open reading frame.

[0108] In this specification, “gene” means any DNA region (including exons and introns) that codes for a gene product and any DNA region that controls the generation of a gene product, regardless of whether such a regulatory sequence is adjacent to the coding sequence and / or the sequence being transcribed. Thus, genes include, but are not limited to, promoter sequences, terminators, translational regulatory sequences (such as ribosome binding sites and internal ribosome entry sites), enhancers, silencers, insulators, boundary elements, origins of replication, matrix binding sites, and locus regulatory regions.

[0109] The term "heterogeneous" refers to an entity that is not endogenous or native to the target cell. For example, a heterogeneous protein is a protein that originates from an exogenous source or is originally derived from it (such as an exogenously introduced nucleic acid sequence). In some cases, heterogeneous proteins are not normally produced by the target cell.

[0110] The term "nickase" refers to an enzyme that cleaves one strand of a double-stranded nucleic acid sequence (i.e., makes a break in the double-stranded sequence). For example, a nuclease with double-strand cleavage activity can be modified by mutation and / or deletion to function as a nickase and cleave only one strand of a double-stranded sequence.

[0111] In this specification, the term "nuclease" refers to an enzyme that cleaves both strands of a double-stranded nucleic acid sequence or a single-stranded nucleic acid sequence.

[0112] The terms “nucleic acid” and “polynucleotide” refer to deoxyribonucleotide polymers or ribonucleotide polymers in a linear or cyclic structure, either single-stranded or double-stranded. For the purposes of this disclosure, these terms should not be construed as limiting with respect to the length of the polymer. These terms may include known analogues of natural nucleotides, as well as nucleotides modified in the base, sugar, and / or phosphate moieties (e.g., phosphorothioate skeletons). Generally, analogues of a particular nucleotide have the same base-pairing specificity (i.e., analogues of A base-pair with T).

[0113] The term "nucleotide" refers to a deoxyribonucleotide or ribonucleotide. A nucleotide can be a standard nucleotide (i.e., adenosine, guanosine, cytidine, thymidine, and uridine), a nucleotide isomer, or a nucleotide analog. A nucleotide analog refers to a nucleotide having a modified purine or pyrimidine base or a modified ribose moiety. A nucleotide analog can be a naturally occurring nucleotide (e.g., inosine, pseudouridine, etc.) or a nucleotide that does not exist in nature. Not limited examples of modifications to the sugar or base moiety of a nucleotide include the addition (or removal) of acetyl, amino, carboxy, carboxymethyl, hydroxy, methyl, phosphoryl, and thiol groups, as well as the substitution of carbon and nitrogen atoms of a base with other atoms (e.g., 7-deazapurine). Nucleotide analogs also include dideoxynucleotides, 2'-O-methylnucleotides, locked nucleic acids (LNAs), peptide nucleic acids (PNAs), and morpholinos.

[0114] The terms "polypeptide" and "protein" are used interchangeably and refer to polymers of amino acid residues.

[0115] The terms "target sequence" and "target site" are used interchangeably and refer to a specific sequence in the target nucleic acid (e.g., chromosomal DNA or cellular RNA) that the CRISPR system targets, and the site in the nucleic acid or nucleic acid-related protein that the CRISPR system modifies.

[0116] Techniques for determining the sequence identity of nucleic acids and amino acids are well known in the art. Typically, such techniques involve determining the nucleotide sequence of mRNA for a gene and / or the amino acid sequence encoded thereby, and comparing these sequences with a second nucleotide or amino acid sequence. Genomic sequences can also be determined and compared in this manner. Generally, identity refers to the exact nucleotide-to-nucleotide or amino acid-to-amino acid match of two polynucleotide or polypeptide sequences. Two or more sequences (polynucleotides or amino acids) can be compared by determining their proportion identity. The proportion identity of two sequences, whether nucleic acid or amino acid sequences, is calculated by dividing the number of exact matches between the two aligned sequences by the length of the shorter sequence and multiplying by 100. Approximate alignment for nucleic acid sequences is provided by the local homology algorithm in Smith and Waterman, Advances in Applied Mathematics 2:482-489 (1981). This algorithm, developed by Dayhoff, Atlas of Protein Sequences and Structure, MO Dayhoff ed., 5 suppl. 3:353-358, National Biomedical Research Foundation, Washington, DC, USA, can be applied to amino acid sequences by using a scoring matrix standardized by Gribskov, Nucl. Acids Res. 14(6):6745-6763 (1986). An exemplary implementation of this algorithm for determining the proportion identity of sequences is provided by the Genetics Computer Group (Madison, Wis.) in practical applications of "BestFit". Other programs suitable for calculating the proportion identity or similarity between sequences are commonly known in the field; for example, another alignment program used for initial parameters is BLAST.For example, BLASTN and BLASTP can be used with the following initial parameters: Gene code = standard; filter = none; strand = both; cutoff = 60; expected value = 10; matrix = BLOSUM62; description = 50 sequences; sorting method = HIGH SCORE; database = non-redundant, GenBank+EMBL+DDBJ+PDB+GenBank CDS translation+Swiss protein+Spupdate+PIR. Details of these programs can be found on the GenBank website.

[0117] Since various modifications can be made to the cells and methods described above without departing from the scope of the present invention, everything contained in the above description and the examples given below is intended to be interpretable rather than restrictive. [Examples]

[0118] The following embodiments illustrate specific aspects of the present disclosure.

[0119] Example 1. Preparation of a self-replicating Cas9 RNA vector The polycistronic synthetic self-replicating RNA encoding Cas9 is modified from the Venezuelan encephalitis (VEE) virus (i.e., Simplicon) with the structural gene removed. (商標) The cDNAs were prepared based on the cloning vector E3L (MilliporeSigma). The cDNAs of Cas9-T2A-TagGFP2, eSpCas9-T2A-GFP2, and eSpCas9 were amplified by PCR using the pVAV-Cas9-2A-GFP plasmid (encoding wild-type SpCas9), CMV-eSpCas9-2A-GFP plasmid, and SpCas9-blastosidine Lenti plasmid as templates, respectively. eSpCas9 is a modified version of wild-type SpCas9, modified to enhance on-target fidelity without reducing cleavage efficiency (Slaymaker et al., Science. 2016, 351(6268):84-8). Each cDNA was then processed using Simplicon(商標) The VEE-Cas9 RNAs were cloned into the NdeI / NotI region of the vector and named T7-VEE-Cas9-TagGFP2 (SEQ ID NO: 31), T7-VEE-eSpCas9-TagGFP2 (SEQ ID NO: 32), and T7-VEE-eSpCas9 (SEQ ID NO: 33), respectively. Schematic diagrams of each VEE-Cas9 RNA are shown in Figure 1A.

[0120] For RNA synthesis, each VEE-Cas9 plasmid and B18R-E3L plasmid was linearized by MluI digestion and BamHI digestion, respectively, to create templates for RNA synthesis. RNA synthesis and 5' capping were performed using CleanCap. (登録商標) The process was carried out at 37°C for 2 hours using the RiboMAX Large Scale RNA Production System-T7 (Promega) kit in the presence of Reagent AG (Trilink). For B18R-E3L RNA synthesis, an additional poly(A) tail of approximately 150 bases was added by poly(A) polymerase (CELLSCRIPT) at 37°C for 30 minutes. After purification and precipitation with 2.5 M ammonium acetate, the RNA was resuspended in RNA storage solution (Ambion) at a concentration of 1 μg / μl and stored at -80°C. VEE-Cas9 RNA was analyzed by agarose gel electrophoresis (Figure 1B). The VEE-Cas9 RNA band showed the strongest intensity at the predicted size (approximately 15 kb) with minimal degradation of the band.

[0121] Example 2. Gene transfer of self-replicating Cas9 RNA These VEE-Cas9 RNAs were co-introduced into cells along with B18R-E3L RNA, which inhibits the interferon (IFN) response induced by RNA gene transfer and replication.

[0122] Human preputial fibroblasts (HFF) and HEK293T cells were cultured in DMEM containing 10% FBS, MEM non-essential amino acids (NEAA), pyruvate, penicillin, and streptomycin. Cells were passaged one day prior to transduction to a culture density of 30-60% on the day of transduction. Each VEE-Cas9 RNA was co-transfected into the cells with B18R-E3L RNA in a 1:1 ratio (1 microgram each per well in a 6-well plate) using lipofectamine MessengerMax transduction reagent (Thermofisher).

[0123] As shown in Figure 1C, TagGFP2 (GFP) expression was observed in cells transfected with Cas9-TagGFP2 and eSpCas9-TagGFP2, and the expression of each Cas9 protein was confirmed by Western blotting (Figure 1D). These data demonstrate that self-replicating RNA technology enables Cas9 expression in human cells.

[0124] Example 3. Non-embedded Cas9 genome editing Next, we investigated targeted genome editing using VEE-Cas9 RNA in combination with chemically synthesized gRNA. Two commercially available synthetic RNAs (crRNA:tracrRNA complexes) and one single synthetic gRNA (sgRNA) targeting K-Ras and EMX-1 were purchased from Thermofisher Scientific. The sequences of the crRNAs containing the PAM sequences (underlined) for K-Ras and EMX-1 targets were 5'-TAGTTGGAGCTGGTGGCGT, respectively. AGG (Sequence ID 34) and 5'-GAGTCCGAGCAGAAGAAGAA GGG This is (Sequence ID 35).

[0125] First, VEE-Cas9-TagGFP2 and B18R-E3L RNA (mentioned above) were co-introduced into HFF cells along with either one or two gRNAs. Cells were harvested 2–6 days after gene transfer, and indels were detected using the Guide-iT mutation detection kit (Clontech). (ChemiDoc)(商標) DNA cleavage efficiency was calculated using an imaging system. The target regions of the KRas gene (340 bp) and the EMX-1 gene (410 bp) were amplified by PCR using the primer sets 5'-GATACACGTCTGCAGTCAACTG (SEQ ID NO: 36) / 5'-GCATATTACTGGTGCAGGACC (SEQ ID NO: 37) and 5'-GCCTGAGTGTTGAGGCCCCA (SEQ ID NO: 38) / 5'-GTCCCTCTGTCAATGGCGGC (SEQ ID NO: 39), respectively.

[0126] As shown in Figure 2A, GFP expression was observed using one or two gRNAs, but GFP expression was reduced in cells transfected with two gRNAs. On day 3 (two days after transfection), more than 15% cleavage was observed in cells transfected with one sgRNA, but no cleavage was detected in cells transfected with two gRNAs (Figure 2B). Puromycin selection was performed to remove Cas9-negative cells, and the efficiency of cleavage increased in cells transfected with one sgRNA, but no cleavage was yet detected in cells transfected with two gRNAs (Figure 2B, day 6).

[0127] To increase editing efficiency, VEE-Cas9 RNA and gRNA were sequentially introduced into cells. In this experiment, VEE-Cas9 and B18R-E3L RNA were co-introduced on day 1, and gRNA was introduced on day 2 using the lipofectamine RNAiMax gene transduction reagent (Thermofisher). As shown in Figure 2B, sequential gene transduction using one or two gRNAs resulted in higher genome editing efficiency on day 4 (two days after gRNA gene transduction), with further increases in efficiency after puromycin selection (day 6). Two different amounts of gRNA (25 nM and 50 nM) were tested, but no significant difference was observed (Figure 2B).

[0128] Next, VEE-Cas9 genome editing was tested in human iPSC cell lines using a sequential gene transfer method with a single sgRNA. Epitherial-1 iPSCs or PBMC-iPSCs (CD34+ umbilical cord blood iPSCs) were cultured in laminin-coated wells in the presence of mTeSRTM-1 culture medium (Stemcell Technologies). The iPSC cells were then transfused as described above. GFP expression was obtained in both human iPSC cell lines (Figure 2C), and the genome editing efficiency ranged from 16% to 32% in both cell lines (Figure 2D).

[0129] Example 4. Comparison of self-replicating Cas9 RNA with other expression systems The efficiency of genome editing was compared between VEE-Cas9 RNA and lentiviral Cas9. Lentiviral vectors possess a blastosidine selection marker. Therefore, lentiviral Cas9 (LV-Cas9) (MOI=3) or VEE-Cas9-TagGFP2 (S-Cas9) was introduced into HFF, and the cells were then selected for 1 week with blastosidine (4 μg / mL) or puromycin (0.8 μg / mL), respectively. After selection, both cells were passaged the following day for gene transfer of a single sgRNA. Similar genome editing efficiencies were observed for the K-Ras target (35% and 36%, respectively) and the EMX-1 target (47% and 40%, respectively) using both Cas9 expression methods (Figure 3A).

[0130] Next, VEE-Cas9 RNA was compared to the plasmid Cas9 encoding Cas9-TagGFP2. HEK293T cells were co-transfected with plasmid Cas9 and a plasmid encoding K-Ras-gRNA (all DNA components), or with VEE-Cas9-TagGFP2, B18R-E3L RNA, and one K-Ras sgRNA (all RNA components). As shown in Figure 3B, gene transfection with either all DNA components or all RNA components resulted in similar proportions of GFP-positive cells (65–74%). The efficiency of genome editing was also similar in both methods (20–29%) (Figure 3C). These data indicate that genome editing with VEE-Cas9 RNA is comparable to genome editing given by lentiviruses and plasmid Cas9 expression tools.

[0131] Example 5. Long-term expression of self-replicating Cass9 RNA VEE-Cas9-TagGFP2 RNA and B18R-E3L RNA were co-introduced into either HFF cells or HEK293T cells, and the cells were maintained under puromycin selectivity in the presence of B18R protein. After 1 month and 4 passaging cycles, one K-Ras sgRNA was co-introduced into HFF cells. As shown in Figure 4A, approximately 40% of HFF cells were GFP-positive, and a genome editing efficiency of approximately 10% was achieved. After 1 month and 8 passaging cycles, one EMX-1 sgRNA was co-introduced into HEK293T cells. Approximately 81% of HEK293T cells were GFP-positive, and a genome editing efficiency of approximately 26% was found (Figure 4B). These data suggest that VEE-Cas9 RNA enables the creation of cell lines expressing Cas9 without manipulating the host cell genome, and that these cell lines can be used for targeted genome editing.

[0132] The genome editing efficiencies of the three VEE-Cas9 RNA vectors prepared in Example 1 were compared. As shown in Figures 4C and 4D, similar genome editing efficiencies were obtained in HFF and HEK293T cells using eSpCas9, eSpCas9-TagGFP2, and Cas9-TagGFP2 (e.g., 5-12% in HFF and 20-26% in HEK293T cells).

[0133] Example 6. Genome editing using self-replicating D10A-Cas9 RNA The D10A mutation in the Cas9 protein results in single-strand breaks instead of double-strand breaks. Therefore, it is thought to enable genome editing with fewer off-target breaks and is considered useful for precise genome editing. To investigate the availability of D10A-Cas9 with self-replicating RNA, a novel construct of Cas9 with the D10A point mutation was created. Figure 5A shows the expression of D10A-Cas9-TagGFP2 and D10A-Cas9-TagRFP in 293T cells on day 1 and day 3, and a cell line expressing Cas9 (293T) was created (Figure 5B). To test the availability of D10A-Cas9, two sgRNAs (sgRNA-1 and 9) were introduced into the EMX1 locus to create double-strand breaks. As shown in Figure 5C, genome editing was observed when two types of sgRNAs were introduced into cells expressing the Cas9-D10A mutant, but not when only one type of sgRNA was introduced. The same result was observed in the D10A-Cas9 cell line (Figure 5D). These data indicate that self-replicating D10A-Cas9 is available for precise genome editing.

[0134] Example 7. Insertion of DNA oligos and GFP DNA fragments by self-replicating Cas9 RNA. Next, DNA insertion at Cas9 cleavage sites was investigated. First, a DNA oligo containing a BamHI restriction enzyme site (BamHI oligo) was inserted into the Rab11 locus. On day 1, self-replicating Cas9 or D10A-Cas9 was transfected, followed by co-transfection of BamHI oligo with sgRNA. Cells were harvested for analysis 3 days after sgRNA transfection, or clones were isolated after puromycin selection. As shown in Figure 6A, BamHI digestion of PCR products at the Rab11A locus revealed insertion of BamHI oligo in both Cas9 and D10A-Cas9 cleaved samples. Cell clones were isolated, and BamHI oligo insertion was confirmed. As shown in Figure 6B, BamHI oligo was detected in 4 out of 6 iPSC clones and 8 out of 11 U2OS cell clones. Next, a TagGFP2 fragment was inserted into the GAPDH locus. On day 1, HEK293 cells were transfected with the self-replicating Cas9-TagRFP gene, followed by co-transfection with PCR-amplified GFP fragments and sgRNA on day 2. GFP-positive cells were observed 3 days after sgRNA transduction (Figure 6C). These data suggest that self-replicating Cas9 functions for DNA insertion in genome editing.

Claims

1. A self-replicating RNA vector containing sequences encoding multiple non-structural replication complex proteins derived from alphaviruses and sequences encoding the CRISPR protein.

2. The self-replicating RNA vector according to claim 1, wherein the CRISPR protein is a type II Cas9 protein, a type V Cas12 protein, a type VI Cas13 protein, a CasX protein, or a CasY protein.

3. CRISPR protein is present in Streptococcus pyogenes Cas9, Francisella novicida Cas9, Staphylococcus aureus Cas9, Streptococcus thermophilus Cas9, Streptococcus pasteurianus Cas9, Campylobacter jejuni Cas9, Neisseria meningitis Cas9, Neisseria cinerea Cas9, Francisella novicida Cas12, Acidaminococcus A self-replicating RNA vector according to claim 1 or 2, wherein the RNA vector is sp. Cas12, Lachnospiraceae bacterium ND2006 Cas12, Leptotricia wadei Cas13a, Leptotricia shahii Cas13a, Prevotella sp. P5-125 Cas13, or Ruminococcus flavefaciens Cas13d.

4. The self-replicating RNA vector according to claim 3, wherein the CRISPR protein is Streptococcus pyogenes Cas9 or Staphylococcus aureus Cas9.

5. A self-replicating RNA vector according to any one of claims 1 to 4, wherein the sequence encoding the CRISPR protein comprises at least one nucleotide insertion, deletion and / or substitution such that the CRISPR protein has altered catalytic activity, improved target site specificity and / or reduced off-target effects.

6. A self-replicating RNA vector according to any one of claims 1 to 5, wherein the CRISPR protein is a nuclease or nicasse, or does not have cleavage activity.

7. A self-replicating RNA vector according to any one of claims 1 to 6, wherein the CRISPR protein is linked to at least one nuclear localization signal.

8. A self-replicating RNA vector according to any one of claims 1 to 7, wherein the CRISPR protein is ligated to at least one fluorescent protein, at least one chromatin regulatory motif, at least one functional domain, or a combination thereof.

9. The self-replicating RNA vector according to claim 8, wherein at least one functional domain is an epigenetic modification domain, a transcriptional activation domain, or a transcriptional repression domain.

10. A self-replicating RNA vector according to any one of claims 1 to 9, wherein the sequence encoding the CRISPR protein is codon-optimized for expression in human cells.

11. A self-replicating RNA vector according to any one of claims 1 to 10, wherein the alphavirus is Auravirus, Babanki virus, Berma Forest virus, Beval virus, Buggy Creek virus, Chikungunya virus, Eastern Equine Encephalitis virus, Everglades virus, Fort Morgan virus, Getavirus, Highlands J virus, Kyzylagach virus, Mayarovirus, Middleburg virus, Mukambo virus, Ndumu virus, Pixnavirus, Onyonnyon virus, Ross River virus, Sagiyama virus, Semiliki Forest virus, Sindbis virus, Unavirus, Venezuelan Equine Encephalitis virus, Western Equine Encephalitis virus, or Wataroa virus.

12. The self-replicating RNA vector according to claim 11, wherein the alphavirus is Venezuelan encephalitis virus.

13. A self-replicating RNA vector according to any one of claims 1 to 12, wherein the vector further comprises a sequence encoding at least one selectable marker.

14. A self-replicating RNA vector according to any one of claims 1 to 13, wherein the vector further comprises a sequence encoding an E3L protein.

15. A self-replicating RNA vector according to any one of claims 1 to 14, wherein the vector is based on a modified Venezuelan encephalitis (VEE) virus and comprises, from 5' to 3', a 5' cap, a 5' UTR, a sequence encoding multiple non-structural replication complex proteins encoding four non-structural replication complex proteins derived from the VEE virus, a promoter, a sequence encoding the CRISPR protein, optionally IRES, optionally a sequence encoding the E3L protein, optionally IRES, optionally a sequence encoding a selectable marker, an alphavirus 3' UTR, and a poly A tail.

16. A self-replicating RNA vector according to any one of claims 1 to 15, and a complex comprising at least one guide RNA designed to form a complex with a CRISPR protein encoded by the self-replicating RNA vector.

17. A eukaryotic cell or cell line comprising a self-replicating RNA vector according to any one of claims 1 to 15.

18. The eukaryotic cell or cell line according to claim 17, further comprising at least one guide RNA designed to form a complex with a CRISPR protein encoded by a self-replicating RNA vector.

19. A plasmid vector encoding a self-replicating RNA vector as specified in any one of claims 1 to 14.

20. The plasmid vector according to claim 16, further comprising a T7 or SP6 promoter for in vitro transcription.

21. A method for targeted genome editing, comprising introducing a self-replicating RNA vector according to any one of claims 1 to 15, and at least one guide RNA designed to form a complex with a CRISPR protein encoded by the self-replicating RNA vector, into a eukaryotic cell.

22. The method according to claim 21, further comprising introducing at least one donor polynucleotide into a cell.

23. The method according to claim 21 or 22, wherein the eukaryotic cell is a human cell.