RNA scaffold

By designing an optimized RNA scaffold-mediated recruitment system, the problems of targeting inaccuracy and off-target effects in genome targeting and editing of CRISPR systems have been solved, enabling efficient and precise gene modification and editing, which is suitable for therapy and genetic engineering.

CN116507629BActive Publication Date: 2026-04-21RUIFUDI EXPLORATION CO LTD +1
View PDF 24 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
RUIFUDI EXPLORATION CO LTD
Filing Date
2021-07-09
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing CRISPR systems have the potential for inaccurate targeting and off-target effects during genome targeting and editing, making it difficult to achieve efficient and precise gene modification.

Method used

An optimized RNA scaffold was designed, comprising tracrRNA, crRNA, and RNA motifs with extended sequences, which are linked or fused via adapters to combine aptamer molecules and effector modules to form an RNA scaffold-mediated recruitment system for the precise delivery of effectors to genomic targets.

Benefits of technology

It improves the flexibility, stability, and targeting accuracy of gene editing, reduces off-target effects, and enhances the efficiency and specificity of genome modification, making it suitable for therapeutic and genetic engineering applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116507629B_ABST
    Figure CN116507629B_ABST
Patent Text Reader

Abstract

Disclosed are RNA scaffolds comprising a tracrRNA; and a recruiting RNA motif with an extension sequence for targeted gene editing and related uses. The method enables precise modification of the genome while minimizing the potential for off-target effects, making the method particularly suitable for therapeutic applications.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This invention relates to RNA scaffolds for CRISPR systems. Background Technology

[0002] CRISPR-Cas technology is rapidly developing, and the scope of CRISPR applications continues to expand (Lau, The CRISPR Journal, Vol 1, No 6). A key component of the CRISPR system is the guide RNA (gRNA), which forms part of the RNA scaffold. This guide RNA first targets the CRISPR system to the desired target in the genome, and then delivers the bioactive effector to the target to perform the desired function. The RNA scaffold must precisely deliver the effector with the correct orientation and spatial conformation to effectively perform the function in a specific manner to produce the desired results without causing off-target effects. Therefore, precise genome-targeting effector systems require optimized RNA scaffolds.

[0003] The inventors have designed optimized RNA scaffolds for enhanced targeting performance. The RNA scaffolds, systems, and methods provided herein enable precise modification of the genome while minimizing the possibility of off-target effects, making the methods and systems particularly suitable for therapeutic applications. Summary of the Invention

[0004] In a first aspect, the present invention provides an RNA scaffold comprising:

[0005] (a) tracrRNA; and

[0006] (b) RNA motifs with extended sequences.

[0007] In one embodiment, the RNA scaffold according to the first aspect further comprises crRNA containing a guide RNA sequence. The RNA scaffold according to the first aspect includes one or more modifications. The RNA motif is attached to the 3' end of the tracrRNA via a linker. In a preferred embodiment, the linker is single-stranded RNA or chemically linked. The single-stranded RNA linker comprises 0-10 nucleotides, preferably 2-6 nucleotides.

[0008] In one embodiment, the RNA scaffold according to the first aspect comprises tracrRNA, which is fused with crRNA containing a guide RNA sequence to form a single RNA molecule. In other embodiments, the RNA scaffold according to the first aspect comprises tracrRNA synthesized as separate RNA molecules and crRNA containing a guide RNA sequence. In any embodiment, tracrRNA hybridizes with crRNA via a repeat:anti-repeat region. When synthesized as such... Figure 10When a single RNA molecule, as shown in B, is synthesized, tracrRNA includes an anti-repeat region, a tetracyclic ring, and the 3' constant region of the gRNA. When synthesized as separate RNA molecules, tracrRNA contains an anti-repeat region and the 3' constant region of the sgRNA, and as shown in B. Figure 10 The quadruple ring is absent as shown in D. The anti-repeat region of tracrRNA hybridizes with the repeat region of crRNA. In a preferred embodiment, the repeat:anti-repeat region is extended.

[0009] The RNA scaffold of the present invention comprises one or more RNA motifs, wherein the one or more RNA motifs include one or more modifications. The one or more modifications may be at the 5' end and / or the 3' end of the one or more RNA motifs. The RNA scaffold of the present invention may include one or more modifications, including the substitution of the A base at position 10 with 2-aminopurine (2AP). The RNA scaffold may use 2' deoxy-2-aminopurine or 2' ribose-2-aminopurine. The RNA scaffold of the present invention may have one or more modifications to the backbone and / or sugar portion of the RNA scaffold. The extended sequence of the RNA motif is a double-stranded extension, wherein the extended sequence of the RNA motif comprises 2-24 nucleotides. In one embodiment, a 4-nucleotide extension produces a stem of 23 nucleotides in total length. In another embodiment, a 10-nucleotide extension produces a stem of 29 nucleotides in total length. In another embodiment, a 16-nucleotide extension produces a stem of 35 nucleotides in total length. In another embodiment, a 26-nucleotide extension produces a stem of 45 nucleotides in total length.

[0010] The RNA scaffold of the present invention comprises one or more RNA motifs that bind to aptamer-binding molecules. The one or more RNA motifs are selected from the following aptamers: MS2, Ku, PP7, SfMu, and Sm7. For example, the MS2 aptamer binds to MCP proteins. In a preferred embodiment, the RNA scaffold comprises one recruiting MS2 RNA motif. In other embodiments, the RNA scaffold comprises two recruiting MS2 RNA motifs. In a preferred embodiment, the MS2 aptamer is wild-type MS2, mutant MS2, or a variant thereof. The mutant MS2 used herein is a C-5, F-5 heterozygote, and / or an F-5 mutant. The RNA motif recruitment effector module of the RNA scaffold according to the present invention. As disclosed herein, the effector module comprises an RNA-binding domain and an effector domain capable of binding to the RNA motif. Suitable effector domains are selected from: reporter molecules, tags, molecules, proteins, microparticles, and nanoparticles. In a preferred embodiment, the effector domain is a DNA-modifying enzyme. Suitable DNA modifying enzymes are selected from: AID, CDA, APOBEC1, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D, APOBEC3F or other APOBEC family enzymes, ADA, ADAR family enzymes or tRNA adenosine deaminase.

[0011] In a second aspect, the present invention provides a system for genetic modification, comprising:

[0012] (a) CRISPR protein;

[0013] (b) the crRNA of the present invention as defined above;

[0014] (c) The RNA scaffold of the present invention as described above;

[0015] (d) Aptamer-bound molecules;

[0016] (e) Effects submodule;

[0017] The system according to the second aspect includes components (a)-(e) delivered in the form of nucleic acids, protein complexes and / or expressed by any suitable expression vector.

[0018] The system provided herein may comprise a CRISPR protein fused to one or more uracil DNA glycosyltransferase (UNG) inhibitor peptides (UGIs). In a preferred embodiment, the CRISPR used in the system according to the second aspect is a type II CRISPR protein, such as Cas9. The CRISPR protein and / or effector submodule used in the system according to the second aspect may comprise one or more nuclear localization signals (NLS). The CRISPR protein may be a type II Cas protein, which is a nuclease nuclease or has nicking enzyme activity.

[0019] The effector module used in the system according to the second aspect can be an effector fusion protein comprising an RNA-binding domain capable of binding to an RNA motif and an effector domain. The system according to the second aspect can use an RNA motif and an effector module containing pairs of RNA-binding domains selected from the group consisting of:

[0020] Telomerase Ku-binding motifs and Ku protein or its RNA-binding portion,

[0021] Telomerase Sm7 binding motif and Sm7 protein or its RNA binding portion.

[0022] The stem-loop of the MS2 phage operator gene and the MS2 capsid protein (MCP) or its RNA-binding portion; the stem-loop of the PP7 phage operator gene and the PP7 capsid protein (PCP) or its RNA-binding portion.

[0023] SfMu phage Com stem loop and Com RNA-binding protein or its RNA-binding portion.

[0024] In a third aspect, the present invention provides a method for genetically modifying cells, wherein the method comprises introducing and / or expressing the system according to the second aspect into cells. The method according to the third aspect can be used for genetically modifying cells, including but not limited to correcting gene mutations or inactivating gene expression or altering gene expression levels or intron-exon splicing. The genetic modification according to the method provided in the third aspect is a point mutation, optionally wherein the point mutation introduces a prematurely maturing stop codon, disrupts a start codon, disrupts a splicing site, or corrects a gene mutation. Attached Figure Description

[0025] Figure 1 : Figure 1 A shows a system comprising three structural and functional components: (1) a sequence-targeting component (e.g., a Cas protein); (2) an RNA scaffold for sequence recognition and effector recruitment, comprising crRNA, tracrRNA, and an RNA motif; and (3) an effector module (e.g., a non-nuclease DNA-modifying enzyme, such as AID, fused to a small protein that binds to the RNA motif). More specifically, as Figure 1 As shown in Figure A, the components of the RNA scaffold-mediated recruitment platform include: sequence-targeting component 1 (e.g., dCas9 or nCas9). D10A); RNA scaffold 2, which includes cRNA 2.1 containing guide RNA (and repeats: anti-repetitive stem repeats) for sequence targeting, tracrRNA 2.2 for Cas protein binding, and RNA motif 2.3 for recruiting effector modules, and effector module 3, which includes effector domain 3.1 (e.g., cytidine deaminase) fused to RNA aptamer ligand 3.2. Figure 1 B shows a schematic diagram of the RNA scaffold-mediated recruitment complex at the target sequence: Cas9 (or dCas9 or nCas9) binds to tracrRNA, and RNA motifs (e.g., aptamers) recruit effector submodules to form an active RNA scaffold-mediated recruitment system capable of editing target residues on unpaired DNA within the CRISPR R loop.

[0026] Figure 2 (A) MS2 hairpin sequence with C-5 substitution and (B) MS2 hairpin sequence containing F-5 mutant sequence, wherein A at position A-10 is additionally substituted with d2AP.

[0027] Figure 3 : The RNA motif of MS2 stem extension contains (A) 4nt (B) 10nt (C) 16nt and (D) (26nt) compared to wild-type MS2.

[0028] Figure 4 : A module of an RNA scaffold that contains tracrRNA, an RNA motif with an extended sequence, and crRNA containing a guide RNA sequence.

[0029] Figure 5 : phenotypic disruption of the TRAC Ex3 SA splicing site due to synthetic aptamers with altered cytosine-to-thymine bases. Synthetic crRNA: tracrRNA (with and without aptamers) with electroporated nCas9-UGI-UGI and rApobec1 and hAID deaminases.

[0030] Figure 6 : Base variations at the TRAC Ex3 SA splicing site due to synthetic aptamers with altered cytosine to thymine bases. Synthetic crRNA: tracrRNA (with and without aptamers) with electroporated nCas9-UGI-UGI and rApobec1 and hAID deaminases. Data are shown as the percentage of T sequenced at the indicated target C residues, as measured by Sanger sequencing.

[0031] Figure 7HEK Site2 was edited with tracrRNA containing a 4nt or 16nt extension of the MS2 hairpin sequence with nCas9-UGI-UGI and rApobec1 deaminase. Data were presented as the percentage of T sequenced at the indicated target C residues, as measured by Sanger sequencing.

[0032] Figure 8 HEK Site2 and HEK Site3 were edited with tracrRNA containing one or two MS2 hairpins at the 3' end of an RNA motif with nCas9-UGI-UGI and hAID deaminase. Data were presented as the percentage of T sequenced at the indicated target C residues, as measured by Sanger sequencing.

[0033] Figure 9 The efficiency of base editing at different target loci was determined by using various RNA scaffolds. Figure 9 AC: The effect of MS2 aptamer location and number, and the extension of the anti-repeating upper stem, on APOBEC-1-mediated base editing. Base editing was measured at 3x target loci, and the sequences and C residues within the base editing target window are shown in Table 5 within Example 1. RNA scaffolds introduced a single copy of the MS2 aptamer (1xMS2) or two copies of the MS2 aptamer (2xMS2) at the 3' of a tetraloop (TL), stem-loop 2 (SL2), or RNA scaffold (3'). Additionally, some designs introduced a 14-base extension (7 bp-extended US) of the anti-repeating upper stem. Data are shown as the percentage of T sequenced at the indicated target C residues, as measured by Sanger sequencing. Error bars represent the standard deviation of the mean from three replicate experiments. Figure 9 DH: APOBEC-1-induced editing was measured at five additional loci, where the previously optimal 1xMS2_3'7bp extended US was tested together with 2xMS2_3'7bp extended US. Sequences and C residues within the base editing target window are shown in Table 5 of Example 1. Data are presented as the percentage of T sequenced at the indicated target C residues, as measured by Sanger sequencing. Error bars represent the standard deviation of the mean from three replicates. Figure 9I: Repetitions: Comparison of the effects of different lengths of antirepetitive upper stem extensions on aptamer-dependent APOBEC-1-mediated base editing. The analysis included upper stem extensions of 2 bp, 5 bp, 7 bp, and 10 bp, as well as sgRNAs with extended upper stems and sgRNAs with non-extended upper stems (1xMS2_3'). Data are presented as the percentage of T sequenced at the indicated target C residues, as measured by Sanger sequencing. Error bars represent the standard deviation of the means from three replicates.

[0034] Figure 10 The annotated diagram illustrates the different parts of the RNA scaffold when synthesized as a single molecule or as separate molecules. Figure 10 A: An RNA scaffold synthesized as a single molecule having two MS2 molecules, as disclosed in prior art WO2017011721. Figure 10 B: Synthesize an RNA scaffold having a single molecule of MS2 as described herein. Figure 10 C: Synthesized as a single molecule of RNA scaffold with one MS2, which has a 7bp extension on either side of the anti-repetition:repetition region. Figure 10 D: The RNA scaffold is synthesized into separate molecules, in which the tetracycle is absent. Figure 10 E: The RNA scaffold synthesized into separate molecules, which have a 2AP modification at position 10 of the MS2 stem-loop. Figure 10 F: The RNA scaffold synthesized into separate molecules, which has a 2AP modification at position 10 in the MS2 stem-loop F-5 mutant.

[0035] Figure 11 In nCas9-UGI-UGI U2OS-stable cells, base editing was performed using chemically synthesized C-5 or F-51xMS2_3′tracrRNAs containing both crRNA and rApobec1 deaminase mRNAs. The gene sites targeted by each cRNA were (A)CR0118_PDCD1, (B)CR0107_PDCD1, (C)CR0057-TRAC_EX3, (D)CR0151_CD2, (E)HEK Site2, (F)CR0121_PDCD1, and (G)CR0165_CIITA. Data are presented as the percentage of T sequences sequenced at the indicated target C residues, as measured by Sanger sequencing.

[0036] Figure 12Base editing was performed in nCas9-UGI-UGI U2OS-stable cells using chemically synthesized C-5 or F-51xMS2_3′tracrRNA, which contains both crRNA and hAID deaminase mRNA. The gene sites targeted by each cRNA were (A)CR0151_CD2, (B)CR0121_PDCD1, and (C)CR0165_CIITA. Data are presented as the percentage of T sequences sequenced at the indicated target C residues, as measured by Sanger sequencing.

[0037] Figure 13 Base editing was performed in nCas9-UGI-UGI U2OS-stabilized cells using chemically synthesized 1xMS2_3′sgRNAs (C-5), 1xMS2_3′_7bp-extended_US sgRNAs (C-5) containing a repeat: anti-repeat upper stem with 7 base pairs, or 1xMS2_3′tracrRNA (C-5) containing crRNA and hAID deaminase mRNA. The gene sites targeted by each crRNA were (A) TRAC_22550571, (B) PDCD1_241852953, and (C) CTNNB1. Data are shown as the percentage of T sequenced at the indicated target C residues, as measured by Sanger sequencing.

[0038] Figure 14 Base editing was performed on chemically synthesized C-5 or F-51xMS2_3′tracrRNAs with crRNA and variable levels of rApobec1 deaminase mRNAs in nCas9-UGI-UGI U2OS-stable cells. The target sites for each cRNA were (A) HEK Site 2 and (B) CR0107_PDCD1. Detailed Implementation

[0039] This invention relates to novel RNA scaffolds for targeting the genome and delivering functional effectors. Such functional effectors include enzymes, reporter molecules, tags, molecules, proteins, microparticles, and nanoparticles.

[0040] One application of this invention relates to CRISPR gene editing and screening. This invention can be used in any CRISPR gene editing system. Another application of this invention involves using RNA scaffolds to recruit effector modules to target DNA sequences in the genome. This invention has specific applications in CRISPR base editing systems, such as RNA scaffold-mediated recruitment systems.

[0041] Examples of RNA scaffold-mediated recruitment systems include the following functional components: (1) a CRISPR / Cas-based module designed for sequence targeting; (2) an RNA scaffold-based module for guiding the platform to the target sequence and for recruiting effector submodules; and (3) effector submodules, such as cytidine ammonia-lyases (e.g., activation-induced cytidine deaminase, AID).

[0042] In a first aspect, this paper provides an RNA scaffold comprising: (a) a tracrRNA; and (b) an RNA motif with an extended sequence. As disclosed herein, the RNA scaffold is optimized for enhanced gene editing. An RNA scaffold-mediated recruitment system is a complex of many components, including an RNA scaffold that needs to be assembled in a specific manner to perform a precise function. This complex must locate specific parts of the genome and precisely achieve the correct orientation and spatial conformation to enable efficient genome editing in a specific manner to obtain the desired output. Furthermore, the complex must efficiently recruit and deliver biologically active effector submodules, such as enzymes in the correct orientation / conformation, to preserve enzyme activity and edit the genome without causing significant off-target effects. Previous base editing systems have been associated with poor or limited editing in multiple regions.

[0043] To overcome these problems, the inventors have introduced one or more modifications into RNA scaffold-mediated recruitment systems, particularly RNA scaffolds identified through trial and error.

[0044] While not wishing to be bound by any theory, it is believed that some of these modifications cause conformational changes to components of the RNA scaffold-mediated recruitment system. Improvements in labeling were observed using the RNA scaffolds disclosed herein. Advantageously, the optimized system incorporating the RNA scaffold (which itself contains an RNA motif with an extended sequence) exhibits greater flexibility, stability, localization, and affinity, thereby enabling efficient editing of previously resistant regions, including treatment-related loci, while maintaining performance. The novel RNA scaffolds expand the library of editable targets and improve the efficiency of gene editing.

[0045] RNA scaffold-mediated recruitment system

[0046] Conventional nuclease-dependent precise genome editing for correcting mutations typically requires the introduction of DNA double-strand breaks (DSBs) and activation of homology-dependent repair (HDR) pathways.

[0047] Recently, RNA-mediated base editing systems have also been developed. These systems recruit base-editing enzymes to target DNA sequences via the RNA component of the CRISPR complex. The system comprises a modified gRNA with a reprogrammable RNA aptamer at its 3' end, which recruits a homologous aptamer ligand fused to an effector (e.g., a deaminase effector). Using this system, targeted nucleotide modifications are achieved with high precision in prokaryotic cells and eukaryotic cells, including mammalian cells; see WO2018129129 and WO2017011721. A novel, second-generation RNAi-mediated base editing system with increased specificity and efficiency was tested in prokaryotic cells and further improved in mammalian cells. The second-generation system / platform exhibits high specificity, high efficiency, and low off-target potential. Utilizing a modular design that completely separates the nucleic acid modification module from the nucleic acid recognition module, RNA-mediated base editing systems offer an alternative to recruiting effectors through fusion or direct interaction with sequence-targeting proteins, which cannot effectively separate sequence-targeting function from nucleic acid modification function. The invention disclosed herein is an RNA scaffold-mediated recruitment system, an improved version of the modular design of RNA-mediated base editing systems. Various modifications have been introduced into the system's components, thereby improving the system's flexibility, specificity, and efficiency. This novel RNA scaffold-mediated recruitment system is not limited to base editing but has many potential applications, such as genome editing, genome screening, and genome markers, providing a powerful tool for genetic engineering and therapeutic development.

[0048] Figure 1 Figures A and 1B illustrate a schematic diagram of an exemplary RNA scaffold-mediated recruitment system for the methods provided herein. The system comprises three structural and functional components: (1) a sequence-targeting component (e.g., a Cas protein); (2) an RNA scaffold for sequence recognition and effector recruitment, comprising crRNA, tracrRNA, and an RNA motif; and (3) an effector module (e.g., a non-nuclease DNA-modifying enzyme, such as AID fused to a small protein that binds to the RNA motif). More specifically, as shown in the appendix… Figure 1 As shown in Figure A, the components of the RNA scaffold-mediated recruitment platform include: sequence targeting component 1 (e.g., dCas9 or nCas9). D10A ); RNA scaffold 2, which includes cRNA 2.1 containing guide RNA (and repeat: anti-repeat stem) for sequence targeting, tracrRNA 2.2 for Cas protein binding and RNA motif 2.3 for recruiting effector modules, and effector module 3, which includes effector domain 3.1 (e.g., cytidine deaminase) fused to RNA aptamer ligand 3.2. Figure 1B shows a schematic diagram of the RNA scaffold-mediated recruitment complex at the target sequence: Cas9 (or dCas9 or nCas9) binds to tracrRNA, and RNA motifs (e.g., aptamers) recruit effector submodules, forming an active RNA scaffold-mediated recruitment system capable of editing target residues on unpaired DNA within the CRISPR R loop. These three components can be constructed in a single expression vector or multiple separate expression vectors, or introduced in a DNA-free form (mRNA or protein and chemically synthesized RNA molecules). The combination of all three specific components constitutes the enablement of the technology platform. Although Figure 1 B displays the three components of the RNA scaffold in a specific 5' to 3' order, but the components can also be arranged in a different order when needed, such as for optimization for different Cas protein variants.

[0049] As disclosed herein, there are several key differences between recruitment mechanisms: RNA scaffold-mediated recruitment systems compared to the direct fusion of Cas9 with effector protein systems (BE systems). The modular design of RNA scaffold-mediated recruitment systems allows for flexible systems engineering. Modules are interchangeable, and many combinations of different modules can be achieved by simply swapping the nucleotide sequences of the recruited RNA aptamers and homologous ligands. On the other hand, recruiting effectors through direct fusion or direct interaction with protein components of sequence-targeting units always requires redesigning new fusion proteins, which is technically more difficult and has less predictable results. Furthermore, RNA scaffold-mediated recruitment systems may promote the oligomerization of effector proteins, while direct fusion may hinder the formation of oligomers due to steric hindrance.

[0050] Due to their relative ease of use and scalability, CRISPR / Cas-like gene systems are poised to dominate the therapeutic landscape, making them an attractive gene-editing technology for developing novel therapeutic applications. As disclosed herein, RNA scaffold-mediated recruitment systems leverage certain advantages of CRISPR / Cas systems. To overcome the limitations associated with the DSB and HDR requirements of conventional CRISPR / Cas gene-editing systems, an elegant gene-editing approach called base editing (BE) has been developed, which combines the DNA-targeting capabilities of Cas9 lacking double-strand cleavage activity (e.g., dCas9 or nCas9) with the DNA-editing capabilities of APOBCE-1 (an enzyme member of the APOBEC family of DNA / RNA cytidine deaminases). By directly fusing deaminase effectors to a nuclease-deficient Cas9 protein called dCas9, these tools (called base editors) can introduce targeted point mutations in genomic DNA or RNA without generating DSB or requiring HDR activity. Essentially, the BE system utilizes the lack of a CRISPR / Cas9 complex as a DNA targeting mechanism, where mutant Cas9 acts as an anchor to recruit cytidine or adenine deaminase through direct protein-protein fusion.

[0051] On the other hand, RNA scaffold-mediated recruitment systems employ a different approach. More specifically, in RNA scaffold-mediated recruitment systems, the RNA component of the CRISPR / Cas9 complex acts as an anchor for effector recruitment by incorporating RNA motifs (e.g., aptamers) into the RNA molecule. The RNA aptamer then recruits effector modules, such as effectors fused with RNA aptamer ligands. Compared to recruitment via direct protein fusion of protein components or other recruitment methods, RNA scaffold-mediated recruitment systems possess several unique characteristics that are advantageous for systems engineering and achieving better functionality. For example, they feature a modular design, where nucleic acid sequence targeting functions and effector functions reside in different molecules, allowing for independent reprogramming of functional modules and system reuse. Reprogramming of RNA scaffold-mediated recruitment systems requires only changes to the RNA aptamer sequence in the gRNA and the exchange of homologous RNA aptamer ligands for effectors. It does not require redesigning individual functional Cas9 fusion proteins. Furthermore, the smaller size of the effector modules may potentially allow for more efficient oligomerization of functional effectors. Furthermore, since RNA scaffold-mediated recruitment does not require the production of Cas9 fusion proteins, which further increases the gene / transcriptional size of Cas9, the system can potentially be constructed in a more efficient manner for packaging and delivery via viral vectors, non-viral vectors, mRNA molecules, mechanical devices, or protein components.

[0052] As disclosed herein, this invention provides further engineering of an RNA scaffold-mediated recruitment system for precise gene editing. As demonstrated herein, the optimized RNA scaffold recruitment system exhibits several important different features compared to previous RNA-mediated base editing systems described in WO2018129129 and WO2017011721 (the full text of which is incorporated herein by reference). First, compared to first- and second-generation RNA-mediated base editing systems, the optimized RNA scaffold recruitment system exhibits significantly increased target-hit efficiency while maintaining low or no detectable off-target effects. Second, the optimized RNA scaffold recruitment system has greater flexibility due to modifications introduced into various components of the system (e.g., 3' end extension sequences of the RNA motif). Third, the optimized RNA scaffold has improved steric hindrance due to the positioning of the RNA motif relative to the tracrRNA.

[0053] a. Sequence Targeting Module

[0054] The sequence targeting components of the methods and systems presented in this article typically utilize Cas proteins from the CRISPR / Cas system derived from bacterial species as sequence targeting proteins.

[0055] In some embodiments, the Cas protein is a mutant Cas protein, for example, a dCas protein with a mutation in its nuclease catalytic domain and therefore lacking nuclease activity, or an nCas protein with a partial mutation in one of its catalytic domains and therefore lacking nuclease activity for generating DSB. The Cas protein is specifically recognized by the tracrRNA component of the RNA scaffold, which guides the Cas protein to its target DNA or RNA sequence. The latter is side-mounted with a 3' PAM.

[0056] Cas protein

[0057] Various Cas proteins can be used in this invention. Interchangeable Cas proteins, CRISPR-related proteins, or CRISPR proteins refer to proteins from or derived from CRISPR-Cas class 1 or class 2 systems that have RNA-directed DNA binding. Non-limiting examples of suitable CRISPR / Cas proteins include Cas3, Cas4, Cas5, Cas5e (or CasD), Cas6, Cas6e, Cas6f, Cas7, Cas8a1, Cas8a2, Cas8b, Cas8c, Cas9, Cas10, Cas10d, CasF, CasG, CasH, Csy1, Csy2, Csy3, Cse1 (or CasA), Cse2 (or CasB), Cse3 (or CasE), Cse4 (or Ca... sC), Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csz1, Csx15, Csf1, Csf2, Csf3, Csf4 and Cu1966, e.g., Koonin and Makarova, 2019, origins and evolution of crispr cas systems, Review Philos Trans R Soc Lond B Biol Sci. 2019 May 13; 374 (1772).

[0058] In one embodiment, the Cas protein is derived from a type 2 CRISPR-Cas system. In a preferred embodiment, the Cas protein is a type 2 Cas system. In an exemplary embodiment, the Cas protein is a Cas9 protein or a Cas9-derived protein. Cas9 proteins can originate from *Streptococcus pyogenes*, *Streptococcus thermophilus*, *Streptococcus* sp., *Nocardiopsis dassonvillei*, *Streptomyces pristinaespiralis*, *Streptomyces viridochromogenes*, *Streptomyces viridochromogenes*, *Streptomyces porangium roseum*, *Streptosporangium roseum*, *Alicyclobacillus sacidocaldarius*, *Bacillus pseudomycoides*, *Bacillus selenitireducens*, *Exiguobacterium sibiricum*, and *Lactobacillus delbrueckii*. * *delbrueckii*, *Lactobacillus salivarius*, *Microscilla marina*, *Burkholderiales bacterium*, *Polaromonas naphthalenivorans*, *Polaromonas sp.*, *Crocosphaera watsonii*, *Cyanotheces sp.*, *Microcystis aeruginosa*, and *Synechococcus sp.*), Acetohalobium arabaticum, Ammonifex degensii, Caldicelulosiruptorbecscii, Candidatus desulforudis, Clostridium botulinum, Clostridium difficile, Finegoldia magna, Natranaerobius thermophilus, Pelotomaculum thermopropionicum, Acidithiobacillus caldus, Acidithiobacillus ferrooxidans, Allochromatium vinosum, Marinobacter sp., Nitrosococcus halophilus halophilus), Nitrosococcus watsoni, Pseudoalteromonashaloplanktis, Ktedonobacter racemifer, Methanohalobium evestigatum, Anabaena variabilis, Nodularia spumigena, Nostoc sp., Arthrospira maxima, Arthrospira platensis, Arthrospira sp., Lyngbyasp., Microcoleus chthonoplastes, Oscillatoria sp.The following bacteria were found in the strain: *Petrotoga mobilis*, *Thermosipho africanus*, *Acaryochloris marina*, *Legionella pneumophila*, *Francisella novicida*, *Gamma proteobacterium* HTCC5015, *Parasutterella excrementihominis*, *Sutter ellawadsworthensis*, *Sulfurospirillum sp.* SC ADC, *Ruminobacter sp.* RM87, *Burkholderiales bacterium* 1147, *Bacteroidetes oral taxon* strain 274 F0058, and *Wolinella*. *Succinogenes*, *Burkholderiales bacterium* YL45, *Ruminobacter amylophilus*, *Campylobacter sp.* P0111, *Campylobacter sp.* RM9261, *Campylobacter lanienae* strain RM8001, *Campylobacter lanienae* strain P0121, *Turicium typhimurium*, *Legionella londiniensis*, *Salinivibrio sharmensis*, *Leptospira* sp. isolate FW.030, *Moritella* sp. isolate NORP46, and *Fndozoicomonas* sp.S-B4-1U, *Tamilnaduibacter albinus*, *Vibrio natriegens*, *Arcobacter skirrowii*, *Francisella philomiragia*, *Francisella hispaniensis*, or *Parendozoicomonas haliclonae*.

[0059] Typically, Cas proteins contain at least one RNA-binding domain. This RNA-binding domain interacts with guide RNA. Cas proteins can be wild-type or modified forms, lacking nuclease activity or possessing only single-strand cleavage activity. Cas proteins can be modified to increase nucleic acid binding affinity and / or specificity, alter enzyme activity, and / or change other properties of the protein. For example, nuclease (i.e., DNase, RNase) domains of the protein can be modified, deleted, or inactivated. Alternatively, the protein can be truncated to remove domains not essential for its function. Proteins can also be truncated or modified to optimize activity.

[0060] In some embodiments, the Cas protein can be a mutant of the wild-type Cas protein (e.g., Cas9) or a fragment thereof. In other embodiments, the Cas protein can be derived from a mutant Cas protein. For example, the amino acid sequence of the Cas9 protein can be modified to alter one or more properties of the protein (e.g., nuclease activity, affinity, stability, etc.). Alternatively, domains of the Cas9 protein that are not involved in RNA targeting can be removed from the protein, resulting in a modified Cas9 protein smaller than the wild-type Cas9 protein. In some embodiments, the system utilizes a Cas9 protein from *Streptococcus pyogenes*, which is encoded in bacteria or codon-optimized for expression in mammalian cells.

[0061] A mutant Cas protein is a polypeptide derivative of a wild-type protein, such as a protein having one or more point mutations, insertions, deletions, truncations, fusion proteins, or combinations thereof. The mutant possesses at least one of RNA-directed DNA-binding activity or RNA-directed nuclease activity, or both. Typically, the modified version has at least 50% (e.g., any number between 50% and 100%, such as 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, and 99%) identity with the wild-type protein (e.g., SEQ ID NO: 1).

[0062] Cas proteins (and other protein components described herein) can be obtained as recombinant polypeptides. To prepare the recombinant polypeptide, the nucleic acid encoding it can be linked to another nucleic acid encoding a fusion chaperone (e.g., glutathione-methyltransferase (GST), a 6x-His epitope tag, or the M13 gene 3 protein). The resulting fusion nucleic acid expresses the fusion protein in a suitable host cell, which can be isolated using methods known in the art. The isolated fusion protein can be further processed, for example, by enzymatic digestion to remove the fusion chaperone and obtain the recombinant polypeptide of the present invention. Alternatively, the protein can be chemically synthesized using conventional methods known in the art or by the recombinant DNA techniques described herein and methods known in the art.

[0063] The Cas protein described in this invention may be provided in purified or isolated form, or may be part of a composition. Preferably, in the composition, the protein is first purified to a certain extent, more preferably a high level of purity (e.g., about 80%, 90%, 95%, or more than 99%). The compositions according to the invention may be any type of composition desired, but are generally aqueous compositions suitable for use as or included in compositions for RNA-guided targeting. Various substances that may be included in such nuclease reaction compositions are well known to those skilled in the art.

[0064] To implement the methods disclosed herein for modifying target nucleic acids, proteins can be generated in target cells via mRNA, protein-RNA complexes (RNPs), or any suitable expression vector. Examples of expression vectors include chromosomal, non-chromosomal, and synthetic DNA sequences, bacterial plasmids, microcircles, bacteriophage DNA, baculoviruses, yeast plasmids, vectors derived from combinations of plasmid and bacteriophage DNA, and viral DNA such as vaccinia virus, adenovirus, vaccinia virus, and pseudorabies virus. Further details are described in the Expression Systems and Methods section below.

[0065] As disclosed herein, either dead nuclease Cas9 (dCas9, such as the D10A and H840A mutant proteins from *Streptococcus pyogenes*) or nuclease-deficient nickase Cas9 (nCas9, such as the D10A mutant protein from *Streptococcus pyogenes*) can be used. dCas9 or nCas9 can also be derived from various bacterial species. Table 1 provides a non-exhaustive list of examples of Cas9 and their corresponding PAM requirements. Synthetic Cas substitutes can also be used, such as those described in Rauch et al., *Programmable RNA-Guided RNA Effector Proteins Built from Human Parts*, Cell Volume 178, Vol. 1, June 27, 2019, pp. 122-134, e12.

[0066] Table 1.

[0067]

[0068] N is any nucleotide (A, G, T, or C), R is A or G, and W is A or T.

[0069] UGI

[0070] In some embodiments of this disclosure, the sequence-targeting component comprises a fusion between (a) a CRISPR protein and (b) a first uracil DNA glycosyltransferase (UNG) inhibitory peptide (UGI). For example, the fusion protein may comprise a Cas protein (e.g., Cas9 protein) fused to the UGI. Such a fusion protein may exhibit increased nucleic acid editing efficiency compared to a fusion protein that does not contain a UGI domain. In some embodiments, the UGI comprises a wild-type UGI sequence or a sequence having the following amino acid sequence: sp|P14739|UNGI_BPPB2: Uracil-DNA glycosyltransferase inhibitor (UGI)MTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLL TSDAPEYKPWALVIQDSNGENKIKML SEQ ID NO:2.

[0071] In some embodiments, the UGI protein provided herein comprises a fragment of UGI or a UGI and protein homologous to UGI or a UGI fragment. For example, in some embodiments, UGI comprises a fragment of the amino acid sequence described above. In some embodiments, UGI comprises an amino acid sequence homologous to the amino acid sequence described above or an amino acid sequence homologous to a fragment of the amino acid sequence described above in the UGI sequence. In some embodiments, a homolog comprising UGI or a UGI fragment or a UGI fragment is referred to as a "UGI variant". A UGI variant is homologous to UGI or a fragment thereof. For example, a UGI variant is at least about 70% (e.g., at least about 80%, 90%, 95%, 96%, 97%, 98%, 99%) of the wild-type UGI or UGI sequence as described above.

[0072] This document provides suitable UGI protein and nucleotide sequences, and other suitable UGI sequences are known to those skilled in the art, including, for example, those disclosed below: Wang et al., Uracil-DNAglycosylase inhibitor gene of bacteriophage PBS2 encodes a binding protein specific for uracil-DNAglycosylase. J Biol. Chem. 264:1163-1171 (1989); Lundquist et al., Site-directed mutagenesis and characterization of uracil-DNAglycosylase inhibitor protein. Role of specific carboxylic amino acids in complex formation with Escherichia coli uracil-DNA glycosylase. J Biol. Chem. 272:21408-21419 (1997); Ravishankar et al., X-ray analysis of a complex of Escherichia coli uracil-DNAglycosylase (EcUDG) with a proteinaceous inhibitor. The structure elucidation of a prokaryotic UDG. Nucleic Acids Res. 26:4880-4887 (1998); and Putnam et al., Proteinmimicry of DNA from crystal structures of the uracil-DNAglycosylase inhibitor protein and its complex with Escherichia coli uracil-DNAglycosylase. JMol. Biol. 287:331-346 (1999), the entire contents of each of these articles are incorporated herein by reference.

[0073] b. RNA scaffolds for sequence recognition and effector recruitment

[0074] The second component of the platform disclosed in this paper is an RNA scaffold, which has three subcomponents: crRNA containing a guide RNA sequence, trans-activating CRISPR RNA (tracrRNA), and an RNA motif with an extended sequence. This scaffold can be a single RNA molecule or a complex of multiple RNA molecules. The crRNA containing the guide RNA of the RNA scaffold is linked to the tracrRNA via a repeat:anti-repeat region, which consists of a 7-bp lower stem and a 4-bp upper stem, interposed with 4-nucleotide bulge structures. When the RNA scaffold is expressed as a single molecule, the repeat:anti-repeat regions are linked by a four-nucleotide quadruple loop, such as... Figure 10 As shown in B. When the RNA scaffold is expressed as multiple RNA molecules, there is no quadruple loop, and it repeats: the anti-repetitive region connects the crRNA and tracrRNA molecules, as shown in Figure B. Figure 10 As shown in D.

[0075] As disclosed herein, a CRISPR / Cas-like module, comprising a programmable guide RNA, crRNA, tracrRNA, and a Cas protein, is formed together for sequence targeting and recognition. Simultaneously, the RNA motif recruits effector submodules, such as base-editing enzymes, via RNA-protein binding, which perform genetic modifications. Therefore, the RNA scaffold connects the effector submodule (e.g., base-editing enzyme) and the sequence recognition module (e.g., type II Cas protein). The RNA scaffold disclosed herein includes one or more modifications.

[0076] Programmable guide RNA (crRNA)

[0077] A key component is the programmable guide RNA. Due to its simplicity and efficiency, the CRISPR-Cas system has been used to perform genome editing in the cells of various organisms. The specificity of this system is determined by the base pairing between the target DNA and the custom-designed guide RNA. By engineering and modulating the base pairing properties of the guide RNA, any sequence of interest can be targeted, provided that a PAM sequence adjacent to the target sequence is present.

[0078] In the subcomponents of the RNA scaffold disclosed herein, the guide sequence provides target specificity. It comprises complementary regions capable of hybridizing with a pre-selected target site of interest. In various embodiments, the target-specific component of the guide sequence can comprise from about 10 nucleotides to more than about 25 nucleotides. For example, the length of the base-pairing region between the guide sequence and the corresponding target site sequence can be about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 22, 23, 24, 25, or more than 25 nucleotides. In an exemplary embodiment, the guide sequence is about 17-20 nucleotides long, for example, 20 nucleotides.

[0079] Additionally, crRNA has a 3' constant region of the target-specified sequence. This sequence forms a repeat: the antirepetitive stem, which connects the crRNA to the tracrRNA component of the RNA scaffold. The constant 3' sequence of crRNA is complementary to the 5' sequence of tracrRNA, thus forming a double-stranded stem. The repeat: antirepetitive region of the RNA scaffold can be divided into three parts: a lower stem, a protrusion, and an upper stem. The lower stem is a 7 bp length form with Watson-Crick and non-Watson-Crick base pairings; this is followed by a 4-nucleotide protrusion. The upper stem consists of a 4 bp structure. When synthesized as a single RNA molecule, tracrRNA includes the antirepetitive region, a tetraloop, and the 3' constant region of sgRNA. When synthesized as a single RNA molecule, tracrRNA includes the antirepetitive region and the 3' constant region of sgRNA, but the tetraloop is absent.

[0080] One requirement for selecting a suitable target nucleic acid is that it has a 3' PAM site / sequence. Each target sequence and its corresponding PAM site / sequence are referred to as the Cas target site in this paper. Class 2 CRISPR systems, such as type II enzymes, are among the most well-known systems, requiring only the Cas9 protein and a guide RNA complementary to the target sequence to achieve target cleavage. The class 2 type II CRISPR system for *Streptococcus pyogenes*, such as Cas9, uses a target site with N12-20NGG, where NGG represents the PAM site from *Streptococcus pyogenes*, and N12-20 represents 12-20 nucleotides directly from the 5' to the PAM site. Other PAM site sequences from bacteria of other species include NGNNG, NNNGATT, NNAGAAW, and NAAAAAC. See, for example, US 20140273233, WO2013176772, Cong et al., (2012), Science 339(6121):819–823, Jinek et al., (2012), Science 337(6096):816–821, Mali et al., (2013), Science 339(6121):823–826, Gasinas et al., (2012), Proc Natl Acad Sci US A.109(39):E2579–E2586, Cho et al., (2013) Nature Biotechnology 31,230–232, Hou et al., Proc Natl Acad Sci US A. 2013 Sep24; 110(39):15644-9, Mojica et al., Microbiology. 2009 Mar; 155(Pt 3):733-40, and www.addgene.org / CRISPR / . The contents of these references are incorporated herein by full citation.

[0081] The target nucleic acid strand can be either of the two strands of the genomic DNA in the host cell. Examples of such genomic dsDNA include, but are not limited to, host cell chromosomes, mitochondrial DNA, and stably maintained plasmids. However, it should be understood that this method can be applied to other dsDNA present in the host cell, such as unstable plasmid DNA, viral DNA, and bacteriophage DNA, as long as a Cas target site is present, regardless of the nature of the host cell dsDNA. This method can also be applied to RNA.

[0082] tracrRNA

[0083] In addition to the guide sequence described above, the RNA scaffold of the present invention includes additional active or inactive sub-components. In one example, the scaffold has tracrRNA. For example, the scaffold may be a hybrid RNA molecule comprising the above-described crRNA of a programmable guide RNA fused with tracrRNA to mimic the natural crRNA:tracrRNA duplex. An exemplary hybrid crRNA:tracrRNA gRNA sequence SEQ ID NO:3 is shown below:

[0084] 5'-(20nt guide)-

[0085] GUUUAAGAGCUAUGCUGGAAACAGCAUAGCAAGUUUAAAUAAGGCUAG UCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUUU-3'. Various tracrRNA sequences are known in the art, and examples include the following tracrRNAs and their active moieties. As used herein, the active moieties of tracrRNAs retain the ability to form complexes with Cas proteins (e.g., Cas9, dCas9, or nCas9). Methods for generating crRNA-tracrRNA mixed RNAs (also known as single guide RNAs or sgRNAs) are known in the art. In one embodiment in which crRNA and tracrRNA are provided as a single gRNA (sgRNA), the two components are linked together by a four-stem loop. In some embodiments, the repeat-anti-repeat region is extended. There is an extension of 2, 3, 4, 5, 6, 7, or more than 7 bases on either side of the repeat:anti-repeat region. In a preferred embodiment, the repeat:anti-repeat region has an extension of 7 nucleotides on either side of the upper stem, such as... Figure 10 C and Figure 10 As shown in D. The extension of 7 bases on either side of the upper stem creates a region 14 base pairs long. When the RNA scaffold is synthesized into a single RNA molecule, the 7-base extension on either side of the upper stem results in a total upper stem length of 11 bases on either side, and a total length of 22 nucleotides, as shown in Figure D. Figure 10 As shown in C. When the RNA scaffold is synthesized into two separate RNA molecules, the extension of 7 bases on either side of the upper stem results in the upper stem having a total of 11 bases on either side, and a total length of 25 nucleotides, as shown in Figure C. Figure 10 As shown in Figure D. In one embodiment, when the RNA scaffold is synthesized as a single RNA molecule, the total length of the upper stem of the repeat:anti-repeat region is 22 nucleotides. In other embodiments, when the RNA scaffold is synthesized as two separate RNA molecules, the total length of the upper stem of the repeat:anti-repeat region is 25 nucleotides. In other embodiments, the extension may be more than 7 basal regions.

[0086] See, for example, WO 2014099750, US 20140179006, and US 20140273226. The contents of these documents are incorporated herein by reference in their entirety.

[0087] TracrRNAs of Streptococcus pyogenes Cas9 with various truncations and extensions are shown below:

[0088] GGAACCAUUCAAAACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 4);

[0089] UAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUC GGUGC (SEQ IDNO: 5);

[0090] AGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGC (SEQ ID NO: 6);

[0091] CAAAACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 7);

[0092] UAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAGUG (SEQ ID NO: 8);

[0093] UAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCA(SEQ ID NO:9); and

[0094] UAGCAAGUUAAAAUAAGGCUAGUCCG (SEQ ID NO: 10).

[0095] In some implementations, the tracrRNA is derived from Streptococcus pyogenes.

[0096] In some implementations, tracrRNA and crRNA containing the guide sequence are two separate RNA molecules that together form part of a functional guide RNA and RNA scaffold. In this case, tracrRNA should be able to interact with crRNA having the guide sequence (typically through base pairing) to form a two-part guide crRNA: tracrRNA.

[0097] RNA motif

[0098] The third subcomponent of the RNA scaffold is the RNA motif, which effectively recruits the effector module (base-editing enzyme) to the target DNA. The RNA motif is also known as the recruiting RNA motif. This connection is crucial for the gene editing systems and methods disclosed herein. As disclosed herein, the RNA scaffold may possess one or more RNA motifs.

[0099] Existing techniques for recruiting effector / DNA editing enzymes to target sequences involve directly fusing effector proteins to dCas9. While direct fusion of effector enzymes to proteins required for sequence recognition (e.g., dCas9) has been successful in sequence-specific transcriptional activation or repression, protein-protein fusion designs can present steric hindrances, which is undesirable for enzymes that require the formation of multimeric complexes for their activity. In fact, most nucleotide editing enzymes (e.g., AID or APOBEC3G) require the formation of dimers, tetramers, or higher-order oligomers for their DNA editing catalytic activity. Direct fusion to dCas9 with a defined conformation anchored to DNA would prevent the formation of a functional oligomeric enzyme complex at the correct location.

[0100] In contrast, the RNA scaffold-mediated recruitment system and method provided herein are based on RNA scaffold-mediated recruitment of effector proteins. More specifically, the platform utilizes various RNA motif / RNA-binding protein binding pairs. For this purpose, the RNA scaffold is designed such that an RNA motif (e.g., the MS2 operator gene motif) that specifically binds to aptamer-binding molecules (e.g., RNA-binding proteins (e.g., MS2 capsid protein, MCP)) is linked to the RNA scaffold via a linker sequence at the 3' end of the tracrRNA. The linker can be single-stranded RNA or chemically linked. In one embodiment, the single-stranded linker contains 0-10 nucleotides, preferably 2-6 nucleotides. The single-stranded sequence can contain GC nucleotides. Advantageously, the linker, such as the single-stranded linker, separates the loop of the RNA motif from the bulky stem-loop of the tracrRNA. One or more RNA motifs as disclosed herein have an extension sequence. In a preferred embodiment, the extension sequence is a double-stranded extension. The length of the extension sequence ranges from 2-24 nucleotides. In some embodiments, the one or more RNA motifs include one or more modifications. The one or more modifications may be at the 5' end and / or the 3' end of one or more RNA motifs.

[0101] Therefore, the RNA scaffold component of the platform disclosed in this paper is a designed RNA molecule that includes not only crRNA for specific DNA / RNA sequence recognition and tracrRNA for Cas protein binding, but also RNA motifs for effector recruitment. Figure 1 B). In this way, recruited effector modules can be recruited to target sites through their ability to bind RNA motifs. Due to the flexibility of RNA scaffold-mediated recruitment, functional monomers, as well as dimers, tetramers, or oligomers, can be formed relatively readily near the target DNA or RNA sequence. These RNA motif / binding protein pairs can be derived from natural sources (e.g., RNA phages or yeast telomerases) or can be artificially designed (e.g., RNA aptamers and their corresponding binding protein ligands). A non-exhaustive list of examples of recruited RNA motif / RNA binding protein pairs that can be used in the methods and systems provided herein is summarized in Table 2.

[0102] Table 2. Examples of recruiting RNA motifs that can be used in this invention, and their paired RNA-binding protein / protein domains.

[0103]

[0104] * Recruited proteins are fused to effector proteins, as shown in Table 3.

[0105] The following lists the sequences of the above binding pairs.

[0106] 1. Telomerase-Ku dimeric motif / Ku heterodimer

[0107] a. Ku-binding hairpin

[0108] 5’-UUCUUGUCGUACUUAUAGAUCGCUACGUUAUUUCAAUUUUGAAAAUCUGAGUCCUGGGAGUGCGGA-3’ SEQ ID NO:11

[0109] Ku heterodimer SEQ ID NO:12

[0110] MSGWESYYKTEGDEEAEEEQEENLEASGDYKYSGRDSLIFLVDASKAMFESQSEDELTPFDMSIQCIQSVYISKIISSDRDLLAVVFYGTEKDKNSVNFKNIYVLQELDNPGAKRILELDQFKGQQGQKRFQDMMGHGSDYSLSEVLWVCANLFSDVQFKMSHKRIMLFTNEDNPHGNDSAKASRARTKAGDLRDTGIFLDLMHLKKPGGFDISLFYRDIISIAEDEDLRVHFEESSKLEDLLRKVRAKETRKRALSRLKLKLNKDIVISVGIYNLVQKALKPPPIKLYRETNEPVKTKTRTFNTSTGGLLLPSDTKRSQIYGSRQIILEKEETEELKRFDDPGLMLMGFKPLVLLKKHHYLRPSLFVYPEESLVIGSSTLFSALLIKCLEKEVAALCRYTPRRNIPPYFVALVPQEEELDDQKIQVTPPGFQLVFLPFADDKRKMPFTEKIMATPEQVGKMKAIVEKLRFTYRSDSFENPVLQQHFRNLEALALDLMEPEQAVDLTLPKVEAMNKRLGSLVDEFKELVYPPDYNPEGKVTKRKHDNEGSGSKRPKVEYSEEELKTHISKGTLGKFTVPMLKEACRAYGLKSGLKKQELLEALTKHFQD>

[0111] MVRSGNKAAVVLCMDVGFTMSNSIPGIESPFEQAKKVITMFVQRQVFAENKDEIALVLFGTDGTDNPLSGGDQYQNITVHRHLMLPDFDLLEDIESKIQPGSQQADFLDALIVSMDVIQHE TIGKKFEKRHIEIFTDLSSRFSKSQLDIIIHSLKKCDISERHSIHWPCRLTIGSNLSIRIAAYKSILQERVKKTWTVVDAKTLKKEDIQKETVYCLNDDDETEVLKEDIIQGFRYGSDIVP FSKVDEEQMKYKSEGKCFSVLGFCKSSQVQRRFFMGNQVLKVFAARDDEAAAVALSSLIHALDDLDMVAIVRYAYDKRANPQVGVAFPHIKHNYECLVYVQLPFMEDLRQYMFSSLKNSKK YAPTEAQLNAVDALIDSMSLAKKDEKTDTLEDLFPTTKIPNPRFQRLFQCLLHRALHPREPLPPIQQHIWNMLNPPAEVTTKSQIPLSKIKTLFPLIEAKKKDQVTAQEIFQDNHEDGPTAK

[0112] '>' Separates the two dimers.

[0113] 2. Telomerase Sm7 binding motif / Sm7 homoheptamer

[0114] a. Sm consensus site (single strand)

[0115] 5'-AAUUUUUGGA-3'SEQ ID NO:13

[0116] Monomer Sm-like protein (archaea) SEQ ID NO:14

[0117] GSVIDVSSQRVNVQRPLDALGNSLNSPVIIKLKGDREFRGVLKSFDLHMNLVLNDAEELEDGEVTRRLGTVLIRGDNIVYISP

[0118] 3. MS2 phage operator stem loop / MS2 capsid protein

[0119] a.MS2 phage operator gene stem loop

[0120] 5'-ACAUGAGGAUCACCCAUGU-3'SEQ ID NO:15

[0121] MS2 capsid protein SEQ ID NO:16

[0122] MASNFTQFVLVDNGGTGDVTVAPSNFANGIAEWISSNSRSQAYKVTCSVRQSSAQNRKYTIKVEVPKGAWRSYLNMELTIPIFATNSDCELIVKAMQGLLKDGNPIPSAIAANSGIY

[0123] 4. PP7 phage operator stem loop / PP7 capsid protein

[0124] a.PP7 phage operator gene stem loop

[0125] 5'-aUAAGGAGUUUAUAUGGAAAACCCUUA-3'SEQ ID NO:17

[0126] PP7 capsid protein (PCP) SEQ ID NO:18

[0127] MSKTIVLSVGEATRTLTEIQSTADRQIFEEKVGPLVGRLRLTASLRQNGAKTAYRVNLKLDQADVVDCSTSVCGELPKVRYTQVWSHDVTIVANSTEASRKSLYDLTKSLVATSQVEDLVVNLVPLGR

[0128] 5. SfMu Com stem-loop / SfMu Com binding protein

[0129] a.SfMu Com stem ring

[0130] 5'-CUGAAUGCCUGCGAGCAUC-3'SEQ ID NO:19

[0131] a.SfMu Com-binding protein SEQ ID NO:20

[0132] MKSIRCKNCNKLLFKADSFDHIEIRCPRCKRHIIMLNACEHPTEKHCGKREKIT HSDETVRY

[0133] RNA scaffolds can be single RNA molecules or complexes of multiple RNA molecules. For example, guide RNA, tracrRNA, and RNA motifs can be three segments of a long, single RNA molecule. Alternatively, one, two, or all three can be on separate molecules. In the latter case, the three components can be linked together covalently or nonvalently (including, for example, Watson-Crick base pairing) to form a scaffold.

[0134] In one example, the RNA scaffold may comprise two separate RNA molecules. The first RNA molecule may comprise a crRNA containing a programmable guide RNA and a region capable of forming a stem duplex structure with complementary regions. The second RNA molecule may also comprise complementary regions in addition to the tracrRNA and RNA motif. The first and second RNA molecules form the RNA scaffold of the present invention via this stem duplex structure. In one embodiment, the first and second RNA molecules each comprise a sequence (about 6 to about 20 nucleotides) that pairs with another sequence of bases. Similarly, the tracrRNA and RNA motif may also be on separate RNA molecules and linked together with another stem duplex.

[0135] The RNA and related scaffolds of the present invention can be prepared by various methods known in the art, including cell-based expression, in vitro transcription, and chemical synthesis or combinations thereof. The ability to chemically synthesize relatively long RNAs (up to 200 or more) allows for the production of RNAs with special characteristics superior to those achieved by four basic ribonucleotides (A, C, G, and U).

[0136] The Cas protein-guide RNA scaffold complex can be prepared using recombinant techniques utilizing host cell systems or in vitro translation-transcription systems known in the art. Details of such systems and techniques can be found, for example, in WO2014144761, WO2014144592, WO2013176772, US20140273226, and US20140273233, the contents of which are incorporated herein by reference in their entirety. The complex can be isolated or purified, at least to a certain extent, from cellular material or the in vitro translation-transcription system in which it is produced.

[0137] Modification

[0138] The RNA scaffolds disclosed herein may include one or more modifications.

[0139] Such modifications may include the inclusion and / or removal of at least one non-naturally occurring nucleotide or modified nucleotide or its analogue. Examples of such modifications include, but are not limited to, adding nucleotides to extend the sequence, substituting nucleotides, adding adapter sequences, removing nucleotides, and modifying the localization of various components of the RNA scaffold. One or more modifications are targeted at the backbone and / or sugar portion of the RNA scaffold.

[0140] Nucleotides can be modified at the ribose, phosphate linker, and / or base moiety. Modified nucleotides may include 2′-O-methyl analogs, 2′-fluoro analogs, 2′-deoxy analogs, or 2′-ribose analogs. The nucleic acid backbone can be modified; for example, a phosphate thioester backbone can be used. Locked nucleic acids (LNAs) or bridged nucleic acids (BNAs) are also feasible. Other examples of modified bases include, but are not limited to, 2-aminopurine, 5-bromouridine, 5-methylcytidine, 5-methoxyuridine, pseudouridine, inosine, and 7-methylguanosine. These modifications can be applied to any component of the RNA scaffold. These modifications can be applied to any component of the CRISPR system. In a preferred embodiment, these modifications are made to the RNA component (e.g., the guide RNA sequence).

[0141] In some embodiments, the RNA scaffold or its sub-parts may include one or more modifications, such as base modifications, backbone modifications, etc., to provide nucleic acids with new or enhanced characteristics (e.g., improved stability).

[0142] Modified backbone and modified nucleoside interchain

[0143] Examples of suitable nucleic acids containing modifications include those with modified backbones, bases, sugars, or non-natural nucleoside bonds. Nucleic acids (with modified backbones) include those that retain phosphorus atoms in the backbone and those that do not have phosphorus atoms in the backbone.

[0144] Suitable modified oligonucleotide backbones containing phosphorus atoms include, for example, thiophosphates, chiral thiophosphates, dithiophosphates, phosphonate triesters, aminoalkylphosphonate triesters, methyl and other alkylphosphonates, including 3'-alkylene phosphonates, 5'-alkylene phosphonates and chiral phosphonates, aminophosphates, including 3'-aminophosphates and aminoalkylphosphamides, phosphorodiamidates, thiophosphamides, thioalkylphosphonates, thioalkylphosphate triesters, selenophosphates and borophosphates with normal 3'-5' bonds, their 2'-5' linked analogs, and those with reverse polarity, wherein one or more nucleotide internucleotide bonds are 3'-3', 5'-5', or 2'-2' bonds. Suitable oligonucleotides with reverse polarity include a single 3' to 3' bond at the 3' internucleotide bond, i.e., a single inverted nucleoside residue that may be basic (with a nucleobase deletion or a hydroxyl group substituted for it). It also includes various salts (such as potassium or sodium), mixed salts, and free acid forms.

[0145] In some embodiments, the subject nucleic acid comprises one or more thiophosphate and / or heteroatom nucleoside bonds, particularly —CH2—NH—O—CH2—, —CH2—N(CH3)—O—CH2- (referred to as methylene (methylimino) or MMI backbone), —CH2—O—N(CH3)—CH2—, —CH2—N(CH3)—N(CH3)—CH2— and —O—N(CH3)—CH2—CH2— (wherein the native phosphodiester nucleoside bond is represented as —O—P(═O)(OH)—O—CH2—). MMI-type nucleoside bonds are disclosed in U.S. Patent No. 5,489,677, cited above. Suitable amide nucleoside bonds are disclosed in U.S. Patent No. 5,602,240.

[0146] Also suitable are nucleic acids having a morpholine backbone structure, such as those described, for example, in U.S. Patent No. 5,034,506. For example, in some embodiments, the subject nucleic acid comprises a 6-membered morpholine ring instead of a ribose ring. In some of these embodiments, phosphodiesterase or other non-phosphodiester nucleoside internucleotide bonds replace phosphodiester bonds.

[0147] Suitable modified polynucleotide backbones, excluding phosphorus atoms, have backbones formed by short-chain alkyl or cycloalkyl nucleoside bonds, mixed heteroatom and alkyl or cycloalkyl nucleoside bonds, or one or more short-chain heteroatom or heterocyclic nucleoside bonds. These include backbones having: morpholine bonds (partially formed by the sugar moiety of the nucleoside); siloxane backbones; sulfide, sulfoxide, and sulfone backbones; formacetyl and thioformyl backbones; methyleneformyl and thioformyl backbones; riboacetyl backbones; olefin-containing backbones; aminosulfonate backbones; methyleneimino and methylenehydrazine backbones; sulfonate and sulfonamide backbones; amide backbones; and other backbones having mixed N, O, S, and CH2 component moieties.

[0148] Simulation

[0149] The subject nucleic acid can be a nucleic acid mimic. The term "mimic" applied to polynucleotides is intended to include polynucleotides in which only the furanose ring or both the furanose ring and the intermolecular bond are replaced by a non-furanose group; replacement of only the furanose ring is also referred to in the art as a sugar substitute. Heterocyclic bases or modified heterocyclic bases are maintained for hybridization with a suitable target nucleic acid. One such nucleic acid is a polynucleotide mimic that has shown excellent hybridization properties and is called a peptide nucleic acid (PNA). In PNAs, the sugar backbone of the polynucleotide is replaced by an amide-containing backbone, particularly an aminoethylglycine backbone. The nucleotide is retained and directly or indirectly bound to the nitrogen atom of the amide moiety of the backbone.

[0150] One polynucleotide mimic with excellent hybridization properties has been reported as peptide nucleic acid (PNA). The backbone of a PNA compound consists of two or more linked aminoethylglycine units, providing the amide-containing backbone. Heterocyclic bases are directly or indirectly bound to the nitrogen atom of the amide moiety in the backbone. Representative U.S. patents describing the preparation of PNA compounds include, but are not limited to: U.S. Patent Nos. 5,539,082; 5,714,331 and 5,719,262.

[0151] Another class of polynucleotide mimics that have been studied is based on morpholino units (morpholinonucleotides) with links to heterocyclic bases attached to a morpholino ring. Several linking groups linking morpholino monomer units in morpholinonucleotides have been reported. One class of linking groups has been selected to yield nonionic oligomers. Nonionic morpholino oligomers are unlikely to have undesirable interactions with cellular proteins. Morpholine-based polynucleotides are nonionic mimics of oligonucleotides that are unlikely to form undesirable interactions with cellular proteins (Dwaine A. Braasch and David R. Corey, Biochemistry, 2002, 41(14), 4503-4510). Morpholine-based polynucleotides are disclosed in U.S. Patent No. 5,034,506. Various compounds within morpholino polynucleotides have been prepared, having a variety of different linking groups linking monomer subunits.

[0152] Another class of polynucleotide mimics is called cyclohexene nucleic acid (CeNA). The furanose ring, normally present in DNA / RNA molecules, is replaced by a cyclohexene ring. CeNADMT-protected phosphoramidite monomers have been prepared and used in the synthesis of oligomers following classical phosphoramidite chemistry. Fully modified CeNA oligomers and oligonucleotides with specific positions modified with CeNA have been prepared and studied (see Wang et al., J. Am. Chem. Soc., 2000, 122, 8595-8602). Generally, binding CeNA monomers to the DNA strand improves the stability of its DNA / RNA hybrid. CeNA oligoadenylates form complexes complementary to both RNA and DNA, exhibiting similar stability to the native complex. Studies of CeNA structures bound to native nucleic acid structures have been conducted using NMR and circular dichroism to facilitate conformational adaptation.

[0153] Further modifications include locked nucleic acids (LNAs), in which a 2′-hydroxyl group is attached to the 4′ carbon atom of the sugar ring, forming a 2′-C, 4′-C-oxymethylene bond, thus forming a bicyclic sugar moiety. This bond can be a methylene group (—CH2—), a group bridging the 2′ oxygen atom and the 4′ carbon atom, where n is 1 or 2 (Singh et al., Chem. Commun., 1998, 4, 455-456). LNAs and LNA analogs exhibit very high dual thermal stability (Tm = +3 to +10 °C) for complementary DNA and RNA, stability against 3′-external nuclear degradation, and good solubility properties. Toxic and non-toxic antisense oligonucleotides containing LNAs have been described (Wahlestedt et al., Proc. Natl. Acad. Sci. USA, 2000, 97, 5633-5638).

[0154] The synthesis and preparation of LNA monomers adenine, cytosine, guanine, 5-methylcytosine, thymine, and uracil, along with their oligomerization and nucleic acid recognition properties, have been described (Koshkin et al., Tetrahedron, 1998, 54, 3607-3630). LNA and its preparation are also described in WO 98 / 39352 and WO 99 / 14226.

[0155] Modified sugar portion

[0156] The subject nucleic acid may also include one or more substituted sugar moieties. Suitable polynucleotides include sugar substituents selected from the following: OH; H; F; O-, S-, or N-alkyl; O-, S-, or N-alkenyl; O-, S-, or N-alkynyl; or O-alkyl-Co-alkyl, wherein the alkyl, alkenyl, and alkynyl groups may be substituted or unsubstituted C1-C. 10 Alkyl or C2-C 10 Alkenyl and ynyl groups. O((CH2)) is particularly suitable. n O) m CH3, O(CH2) n OCH3, O(CH2) n NH2, O(CH2) n CH3, O(CH2) n ONH2 and O(CH2) n ON((CH2) n CH3)2, where n and m are 1 to approximately 10. Other suitable polynucleotides include sugar substituents selected from the following: C1 to C2. 10Lower alkyl groups, substituted lower alkyl groups, alkenyl groups, alkynyl groups, aryl groups, O-alkaneyl groups or O-aryl groups, SH, SCH3, OCN, Cl, Br, CN, CF3, OCF3, SOCH3, SO2CH3, ONO2, NO2, N3, NH2, heterocyclic alkyl groups, heterocyclic alkaneyl groups, aminoalkylamino groups, polyalkylamino groups, substituted silyl groups, RNA cleaving groups, reporter groups, intercalators, groups used to improve the pharmacokinetic properties of oligonucleotides, or groups used to improve the pharmacokinetic properties of oligonucleotides, and other substituents with similar properties. Suitable modifications include 2′-methoxyethoxy (2′—O—CH2 CH2OCH3, also known as 2′-O-(2-methoxyethyl) or 2′-MOE) (Martin et al., Helv. Chim. Acta, 1995, 78, 486-504), i.e. alkoxyalkoxy. Further suitable modifications include 2′-dimethylaminoethoxy, i.e., the O(CH2)2ON(CH3)2 group, also known as 2′-DMAOE, as described in the examples below, and 2′-dimethylaminoethoxyethoxy (also known in the art as 2′-O-dimethyl-amino-ethoxy-ethyl or 2′-DMAEOE), i.e., 2′—O—CH2—O—CH2—N(CH3)2.

[0157] Other suitable sugar substituents include methoxy (—O—CH3), aminopropoxy (—O CH2 CH2CH2NH2), allyl (—CH2—CH═CH2), -O-allyl (CH2—CH═CH2), and fluorine (F). The 2′-sugar substituent can be in the arabinose (top) or ribose (bottom) position. A suitable 2′-arabinose modification is 2′-F. Similar modifications can also be made at other positions on the oligomer, particularly at the 3′ position of the sugar on the 3′ terminal nucleotide or the 5′ position of the 5′ terminal nucleotide on the 2′-5′ linked oligonucleotide. The oligomer can also have sugar mimics, such as a cyclobutyl moiety replacing the pentofuranose sugar.

[0158] Base modification and substitution

[0159] Nucleic acids may also include nucleobase modifications or substitutions (often simply referred to as "bases" in the art). As used herein, "unmodified" or "natural" nucleobases include purine bases adenine (A) and guanine (G), and pyrimidine bases thymine (T), cytosine (C), and uracil (U). Modified nucleobases include other synthetic and natural nucleobases, such as 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil and cytosine, 5-propynyl (—C═C—CH3)uracil and other alkynyl derivatives of pyrimidine bases, 6-azouracil, cytosine, and other alkynyl derivatives of pyrimidine bases. Pyrimidines and thymines, 5-uracil (pseudouracil), 4-thiouracil, 8-halogenated, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxy and other 8-substituted adenines and guanines, 5-halogenated, especially 5-bromo, 5-trifluoromethyl and other 5-substituted uracils and cytosines, 7-methylguanine and 7-methyladenine, 2-F-adenine, 2-amino-adenine, 8-azaguanine and 8-azaadenine, 7-deazoguanine and 7-deazoadenine and 3-deazoguanine and 3-deazoadenine. Further modifications to the nucleobases include tricyclic pyrimidines such as benzoxazincytidine (1H-pyrimido(5,4-b)(1,4)benzoxazin-2(3H)-one), phenothiazincytidine (1H-pyrimido(5,4-b)(1,4)benzothiazoline-2(3H)-one), G-clamps such as substituted phenothiazincytidines (e.g., 9-(2-aminoethoxy)-H-pyrimido(5,4-(b)(1,4)benzoxazin-2(3H)-one), carbazocytidine (2H-pyrimido(4,5-b)indole-2-one), and pyridoindolecytidine (H-pyrido(3′,2′:4,5)pyrrolo(2,3-d)pyrimido-2-one).

[0160] The heterocyclic base moiety may also include those in which the purine or pyrimidine base is substituted by other heterocycles, such as 7-deazono-adenine, 7-deazono-guanosine, 2-aminopyridine, and 2-pyridone. Additional nucleobases include those disclosed in U.S. Patent No. 3,687,808, *The Concise Encyclopedia of Polymer Science and Engineering*, pages 858-859, Kroschwitz, JI, ed. John Wiley & Sons, 1990, Englisch et al., *Angewandte Chemie*, International Edition, 1991, 30, 613, and Sanghvi, YS, Chapter 15, *Antisense Research and Applications*, pages 289-302, Crooke, Stan Lebleu, B., ed., CRC Press, 1993. Some of these nucleobases can be used to increase the binding affinity of oligomeric compounds. These include 5-substituted pyrimidines, 6-azapyrimidines, and N-2, N-6, and O-6 substituted purines, including 2-aminopropyladenine, 5-propynyluracil, and 5-propynylcytosine. 5-methylcytosine substitution has been shown to increase the stability of nucleic acid duplexes by 0.6–1.2 °C (Sanghvi et al., eds., Antisense Research and Applications, CRC Press, Boca Raton, 1993, pp. 276–278) and is suitable for methyl substitution, for example, when combined with 2′-O-methoxyethyl sugar modifications.

[0161] Modifications disclosed herein can be incorporated into various locations on the RNA scaffold, such as at the quadruple loop of sgRNA, in repeat-antirepetitive regions of crRNA:tracrRNA components, at any location on tracrRNA, such as at the 5' end, 3' end, stem-loop 1, 2, or 3, and on the RNA motif. Modifications disclosed herein include, but are not limited to, extensions of repeat-antirepetitive sgRNA or two-part crRNA:tracrRNA components that localize the RNA motif at the 3' end of the tracrRNA motif, attaching RNA motif linkers to CRISPR motifs, modifying nucleotides of the RNA motif, and extending the RNA motif.

[0162] Localization of RNA motifs

[0163] RNA motifs can be located at various positions on the RNA scaffold, as described in Example 1. The RNA scaffold of the present invention can have one MS2 RNA motif or can have two MS2 RNA motifs. The RNA motif (e.g., the MS2 aptamer) can be located at the 3' end of the tracrRNA, at the quadruple loop of the sgRNA, at the stem loop 2 of the tracrRNA, and at the stem loop 3 of the tracrRNA. The localization of the aptamer (e.g., the MS2 aptamer) is crucial due to the steric hindrance that can be created by large loops. In a preferred embodiment, the MS2 aptamer is located at the 3' end of the CRISPR motif. Advantageously, the MS2 aptamer is located at the 3' end of the CRISPR motif, thus spatially reducing the steric hindrance of other large loops on the RNA scaffold.

[0164] connector

[0165] RNA motifs can be linked to tracrRNA motifs via adapters. Adapters can be single-stranded RNA or chemically linked. Single-stranded RNA adapters can be 2, 3, 4, 5, 6, 7, or more than 7 nucleotides. Advantageously, the adapter sequence provides flexibility for the RNA scaffold. Adapter sequences may include GC nucleotides.

[0166] Nucleotides that modify RNA motifs

[0167] RNA motifs (e.g., aptamer sequences) can be modified. In a preferred embodiment, the RNA motif includes one or more modifications. For example, suitable modifications refer to C-5 and F-5 aptamer mutants. In a preferred embodiment, the modification of the aptamer is the substitution of adenine at position 10 with 2-aminopurine (2-AP). Advantageously, the substitution induces a conformational change, resulting in greater affinity compared to wild-type MS2. While not wishing to be bound by any theory, it is believed that the conformational change induced by 2-AP results in the formation of a hydrogen bond between the exocyclic amino group of the 2-AP nucleotide at position 10 and the carbonyl group B59 in the backbone. It is believed that replacing the MS2 hairpin sequence with the higher-affinity MS2 sequence will lead to increased gene editing efficiency because the substituted amino acid helps to align the RNA stem-loop into a conformation that is better recognized by capsid proteins.

[0168] The above lists suitable modifications to RNA motifs, such as 2′deoxy-2-aminopurine, 2′ribose-2-aminopurine, phosphate thioester modifications (mods), 2′-O-methyl modifications (mods), 2′-fluorine modifications (mods), and LNA modifications (mods). Advantageously, these modifications contribute to increased stability and promote stronger bond / fold structures for desired hairpins.

[0169] Other suitable modifications can be made at the 5' end and / or 3' end of one or more RNA motifs.

[0170] RNA motif extension

[0171] The length of RNA motif extensions can be variable. RNA motif extensions can range from 2 to 24 nucleotides. RNA motif extensions can be greater than 24 nucleotides. Figure 3 AD shows multiple extensions of the recruited RNA motif relative to wild-type MS2, and the extended sequences are shown below. Figure 3 A is a 4-nucleotide (2bp) extension that produces a stem of 23 nucleotides in total length (SEQ ID NO:21). Figure 3 B is a 10-nucleotide (5bp) extension that produces a stem of 29 nucleotides in total length (SEQ ID NO:22). Figure 3 C is a 16-nucleotide (8 bp) extension that produces a stem of 35 nucleotides in total length (SEQ ID NO:23). Figure 3 D is a 26-nucleotide (13 bp) extension that produces a stem of 45 nucleotides in total length (SEQ ID NO:24). Advantageously, the extension of the RNA motif increases the flexibility of the motif. The extension of the RNA motif can be double-stranded or single-stranded. Double-stranded extension provides greater stability to the RNA scaffold. In a preferred embodiment, the extension of the RNA motif is double-stranded.

[0172] RNA motif extension sequence

[0173] Extension (SEQ ID NO:21)

[0174] Extension (SEQ ID NO:22)

[0175] Extension (SEQ IDNO:23)

[0176] Extension (SEQ ID NO:24)

[0177] illustrate

[0178] GC linkers are underlined, nucleotide extensions are shown in bold, and aptamers are in italics.

[0179] Repetition: Anti-repetition region

[0180] crRNA and tracrRNA can be provided as sgRNA or as two separate components. crRNA hybridizes with tracrRNA via repeat:anti-repetition regions. The repeat region of crRNA hybridizes with the anti-repetition region of tracrRNA. Repeat:anti-repetition regions can be extended to increase the flexibility, proper folding, and stability of the components. Repeat:anti-repetition regions can be extended by 2, 3, 4, 5, 6, 7, or more bases on either side of the region. Repeat:anti-repetition regions can be extended by a total of 14 nucleotides. Repeat:anti-repetition can also include other modifications as described above.

[0181] Combination of Modifications

[0182] The RNA scaffold may have one or more of the modifications described above. One or more modifications to the RNA scaffold are one or more of the modifications described above, such as repetition: extension of the anti-repetition region, extension of the recruiting RNA motif, or substitution of a nucleotide by 2AP. One or more modifications may be on different components of the RNA scaffold, such as sgRNA repetition: extension of the anti-repetition region, or 2-part crRNA:tracrRNA, and extension of the RNA motif. One or more modifications may also be on the same component of the RNA scaffold, such as extension of the RNA motif and substitution of RNA motif nucleotides. Modifications may be two or more, three or more, four or more, or five or more. In one embodiment, the modification may be an extension of the RNA motif and / or a substitution of RNA motif nucleotides. For example, the modification may be an extension of the RNA motif or a substitution of RNA motif nucleotides. In other cases, the RNA motif may have an extended length and nucleotide substitutions.

[0183] fit

[0184] In some embodiments, the aptamer-binding protein can be a wild-type protein, a mutant of a wild-type protein, or a variant thereof. An example of an RNA motif used herein is the MS2 aptamer. The RNA motif binds to the aptamer-binding molecule. The MS2 motif specifically binds to the MS2 phage capsid protein (MCP). Repeated in vitro selection processes generate a family of aptamers. Two members of this aptamer family include the MS2 C-5 mutant and the MS2 F-5 mutant. One significant difference between wild-type MS2 and the C-5 and F-5 mutants is the substitution of uracil nucleotides for cytosine at position 5 of the aptamer loop. The F-5 mutant has been reported to have a higher affinity for the capsid protein compared to wild-type and other members of the aptamer family. Appropriately, both the C-5 and F-5 mutants are used as aptamers in this invention. In one embodiment, the MS2 aptamer is wild-type MS2, a mutant MS2, or a variant thereof. In another embodiment, the MS2 aptamer comprises the C-5 and / or F-5 mutations. The MS2 protein linked to the CRISPR motif can be a single copy (i.e., one MS2 loop) or a double copy (i.e., two MS2 loops). In a preferred embodiment, the RNA scaffold has one RNA motif. In other embodiments, the RNA scaffold has more than one, more than two, or more than three RNA motifs. In other embodiments, the RNA scaffold has two RNA motifs.

[0185] c. Effects submodule

[0186] The third component of the platform disclosed in this invention is a non-nuclease effector. As disclosed herein, the effector module includes an RNA-binding domain capable of binding RNA motifs and effector domains. Effector domains as used herein include, but are not limited to, enzymes, reporter molecules, tags, molecules, proteins, microparticles, and nanoparticles. In one embodiment, the effector domain is a DNA-modifying enzyme.

[0187] Effectors are not nucleases and do not possess any nuclease activity, but may possess other types of DNA-modifying enzyme activity, such as base editing. Examples of enzyme activity include, but are not limited to, deamination activity, methyltransferase activity, deacetylase activity, DNA repair activity, DNA damage activity, dismutase activity, nicking enzyme activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer formation activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, or glycosylation activity. In some embodiments, effectors possess cytidine deaminase (e.g., AID, APOBEC3G), adenosine deaminase (e.g., ADA), DNA methyltransferase, and DNA deacetylase activities. In some embodiments, effectors are derived from different vertebrate species and possess different activity characteristics.

[0188] In a preferred embodiment, the third component is a conjugate or fusion protein having an RNA-binding domain and an effector domain. These two domains can be linked via a linker.

[0189] In some implementations, effectors are not required in certain cell types (e.g., cancer lines overexpressing deaminases). In this case, endogenous effectors (e.g., APOBEC, AID, etc.) can be genetically edited to include recruitment modules, thus eliminating the need for exogenous editors. This is suitable for cell types expressing editors of interest, such as lymphoid (B+T cells) and certain cancer cells. Furthermore, the nicking enzyme activity does not necessarily derive from the Cas module but can be recruited from the effector – for example, dCas9 can have aptamers to recruit both the nicking enzyme and the editor via the same gRNA. Effector proteins, as used herein, can be wild-type, genetically engineered, or chimeric enzymes.

[0190] RNA-binding domain

[0191] Although various RNA-binding domains can be used in this invention, RNA-binding domains of Cas proteins (e.g., Cas9) or variants thereof (e.g., dCas9) should not be used. As mentioned above, direct fusion with dCas9 (which anchors to DNA in a defined conformation) will prevent the formation of a functional oligomeric complex at the correct location. Instead, this invention has the advantage of various other RNA motif-RNA-binding protein binding pairs. Examples include those listed in Table 2.

[0192] In this way, effector proteins can be recruited to target sites by binding to RNA motifs via RNA-binding domains. Due to the flexibility of RNA scaffold-mediated recruitment, functional monomers, as well as dimers, tetramers, or oligomers, can be formed relatively easily near the target DNA or RNA sequence.

[0193] Effect subdomain

[0194] Effector components include an active portion, i.e., an effector domain. In one embodiment, an effector domain as used herein includes, but is not limited to, enzymes, reporter molecules, tags, molecules, proteins, microparticles, and nanoparticles. In some embodiments, the effector domain includes the naturally occurring active portion of a non-nuclease protein (e.g., deaminase). In other embodiments, the effector domain includes a modified amino acid sequence (e.g., substitution, deletion, insertion) of the naturally occurring active portion of a non-nuclease protein. The effector domain has enzymatic activity. Examples of this activity include deamination activity, methyltransferase activity, demethylase activity, DNA repair activity, DNA damage activity, dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer formation activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, glycosylation activity, DNA methylation, histone acetylation activity, or histone methylation activity. Some modifications in non-nuclease proteins (e.g., deaminases) can help reduce off-target effects. For example, as described below, the recruitment of AID to off-target sites can be reduced by mutating Ser38 in AID to Ala.

[0195] connector

[0196] The two domains described above, as well as other domains disclosed herein, can be joined by adapters, such as, but not limited to, chemical modifications, peptide adapters, chemical adapters, covalent or non-covalent bonds, protein fusion, or any method known to those skilled in the art. The joining can be permanent or reversible. See, for example, U.S. Patent Nos. 4625014, 5057301, and 5514363, US Application Nos. 20150182596 and 20100063258, and WO2012142515, the contents of which are incorporated herein by reference in their entirety. In some embodiments, several adapters may be included to utilize the desired properties of each adapter and each protein domain in the conjugate. For example, flexible adapters and adapters intended to increase the solubility of the conjugate may be used alone or in combination with other adapters. Peptide adapters can be joined by expressing DNA encoding the adapter into one or more protein domains in the conjugate. Adapters can be acid-cleavable, photocleavable, and heat-sensitive adapters. Methods for conjugation are well known to those skilled in the art and are included in the uses of this invention.

[0197] In some embodiments, the RNA-binding domain and the effector domain can be linked by a peptide linker. The peptide linker can be linked by expressing a nucleic acid encoding both domains and the linker within the frame. Optionally, the linker peptide can bind at either the N-terminus or C-terminus of the domain. In some instances, the linker is an immunoglobulin hinge region linker as disclosed in U.S. Patent Nos. 6,165,476, 5,856,456, U.S. Applications Nos. 20150182596 and 2010 / 0063258, and International Application WO2012 / 142515, each of which is incorporated herein by reference in its entirety.

[0198] Other structural domains

[0199] The effector fusion protein may contain other domains. In some embodiments, the effector fusion protein may contain at least one nuclear localization signal (NLS). Typically, the NLS comprises a stretch of a basic amino acid. Nuclear localization signals are known in the art (see, for example, Lange et al., J. Biol. Chem., 2007, 282:5101-5105). The NLS may be located at the N-terminus, C-terminus, or internal location of the fusion protein.

[0200] In some embodiments, the fusion protein may include at least one cell-penetrating domain to facilitate protein delivery to target cells. In one embodiment, the cell-penetrating domain may be a cell-penetrating peptide sequence. Various cell-penetrating peptide sequences are known in the art, and examples include cell-penetrating peptide sequences of HIV-1 TAT protein, and TLM, Pep-1, VP22, and polyarginine peptide sequences of human HBV.

[0201] In other embodiments, the fusion protein may include at least one labeling domain. Non-limiting examples of labeling domains include fluorescent proteins, purification tags, and epitope tags. In some embodiments, the labeling domain may be a fluorescent protein. In other embodiments, the labeling domain may be a purification tag and / or an epitope tag. See, for example, US 20140273233.

[0202] In one implementation, AID is used as an example to illustrate how the system works. AID is a cytidine deaminase, which catalyzes the deamination of cytidine in the context of DNA or RNA. When brought to a target site, AID changes a C base to a U base. In dividing cells, this can result in a C-to-T point mutation. Alternatively, the C-to-U change can trigger cellular DNA repair pathways, primarily the excision repair pathway, which removes the mismatched UG base pair and replaces it with a TA, AT, CG, or GC pair. As a result, a point mutation is generated at the target CG site. Since the excision repair pathway is present in most (if not all) somatic cells, recruiting AID to the target site can correct the CG base pair to another. In this case, if the CG base pair is a genetic mutation in somatic tissue / cells that causes a potential disease, the above method can be used to correct the mutation and thereby treat the disease.

[0203] For the same reason, if the underlying disease causing the genetic mutation is at a specific site of the AT base pair, the same method can be used to recruit adenosine deaminase to that specific site, where the adenosine deaminase can correct the AT base pair to something else. Other effector enzymes are expected to produce other types of changes in base pairing. A non-exhaustive list of examples of DNA / RNA modifying enzymes is detailed in Table 3.

[0204] Table 3. Examples of effector proteins that can be used in this invention

[0205]

[0206] full name of effector protein

[0207] AID: Activation-induced cytidine deaminase, also known as AICDA

[0208] APOBEC1: Apolipoprotein B mRNA editing enzyme that catalyzes peptide-like 1.

[0209] APOBEC3A: Apolipoprotein B mRNA editing enzyme, catalyzing peptide-like 3A.

[0210] APOBEC3B: Apolipoprotein B mRNA editing enzyme that catalyzes peptide-like 3B editing.

[0211] APOBEC3C: Apolipoprotein B mRNA editing enzyme, catalyzing peptide-like 3C...

[0212] APOBEC3D: Apolipoprotein B mRNA editing enzyme that catalyzes peptide-like 3D editing.

[0213] APOBEC3F: Apolipoprotein B mRNA editing enzyme that catalyzes peptide-like 3F.

[0214] APOBEC3G: Apolipoprotein B mRNA editing enzyme that catalyzes peptide-like 3G.

[0215] APOBEC3H: Apolipoprotein B mRNA editing enzyme that catalyzes peptide-like 3H.

[0216] CDA: Cytidine deaminase

[0217] ADA: Adenosine deaminase

[0218] ADAR1: Adenosine deaminase acting on RNA1

[0219] ADAR2: Adenosine deaminase acting on RNA2

[0220] ADAR3: Adenosine deaminase acting on RNA3

[0221] tadA: tRNA-specific adenosine deaminase

[0222] Dnmt1: DNA (cytosine-5-)-methyltransferase 1

[0223] Dnmt3a: DNA (cytosine-5-)-methyltransferase 3α

[0224] Dnmt3b: DNA (cytosine-5-)-methyltransferase 3β

[0225] Tet1: 10-11 Transposition 1

[0226] Tet2: 10-11 Transposition 2

[0227] Tdg: Thymine DNA glycosylation enzyme

[0228] The three specific components described above constitute the technology platform. Each component can be selected from the list in Tables 1-3 to achieve a specific therapeutic / utility goal.

[0229] The following can be used to construct an RNA scaffold-mediated recruitment system: (i) dCas9 / nCas9 from *Streptococcus pyogenes* as a sequence-targeting protein, (ii) an RNA scaffold containing crRNA comprising a guide RNA sequence, tracrRNA, and an RNA motif such as an MS2 operator motif, and (iii) an effector module containing human AID fused to the MS2 operator gene that binds to the MCP protein. The sequences of the components are listed below:

[0230] The dCas9 protein sequence of Streptococcus pyogenes (SEQ ID NO:25)

[0231] MDKKYSIGL AIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAILR RQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREM IEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIGELGSQILKEHPVENTQLQNEKLYYLQNGRDMYVDQELDINRLSDYDVD AIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD

[0232] (Underlined residues: D10A (D→A), H840A (H→A) active site mutations)

[0233] Cas9 D10A protein (underlined residues: D10A,)(SEQ ID NO:26)

[0234]

[0235]

[0236] DNA encoding Cas9 D10A protein (29A>C)(SEQ ID NO:27)

[0237]

[0238]

[0239]

[0240]

[0241] RNA scaffold expression cassette (Streptococcus pyogenes), containing a 20-nucleotide programmable sequence, a CRISPR RNA motif (tracrRNA), and an MS2 operator gene motif:

[0242] SEQ ID NO:28

[0243]

[0244] (N 20 :: Programmable sequence; Underlined: CRISPR RNA motif (tracrRNA); Bold: MS2 motif; Italic: Terminator; Bold and italic: GC adapter; Bold and underlined: MS2 extension)

[0245] The RNA scaffold described above contains one MS2 loop (1xMS2). Below is an example of an RNA scaffold containing two MS2 loops (2xMS2), where the MS2 scaffold is underlined:

[0246] SEQ ID NO:29

[0247]

[0248] Effector AID–MCP fusion:

[0249] SEQ ID NO:30

[0250]

[0251]

[0252] Symbol explanation:

[0253]

[0254]

[0255] Similar to the Cas proteins described above, non-nuclease effectors can also be obtained as recombinant peptides. Techniques for preparing recombinant peptides are known in the art.

[0256] As described in this article, mutating Ser38 in AID to Ala can reduce the recruitment of AID to off-target sites. The DNA and protein sequences of wild-type AID and AID_S38A (phosphorylation null, pnAID) are listed below:

[0257] wtAID cDNA (Ser38 codon is in bold and underlined):

[0258] SEQ ID NO:31

[0259]

[0260] wtAID protein (Ser38 is in bold and underlined):

[0261] SEQ ID NO:32

[0262] EGLHENSVRLSRQLRRILLPLYEVDDLRDAFRTLGL

[0263] AID_S38A cDNA (S38A mutation is in bold and underlined)

[0264] SEQ ID NO:33

[0265]

[0266] AID_S38A protein (S38A mutation is shown in bold and underlined)

[0267] SEQ ID NO:34

[0268]

[0269] Exemplary sequence

[0270] Several exemplary sequences developed in this study are shown below.

[0271] A Protein sequence of the nu construct of the RNA scaffold-mediated recruitment system (SEQ ID NO:35):

[0272]

[0273]

[0274]

[0275] Symbol explanation:

[0276]

[0277] A The protein sequence (SEQ ID NO:36) of the RNA scaffold-mediated recruitment system nu.2 construct:

[0278]

[0279]

[0280] Symbol explanation:

[0281]

[0282] Protein sequence of the RNA scaffold-mediated recruitment system (SEQ ID NO:37):

[0283]

[0284]

[0285] Symbol explanation:

[0286]

[0287] The 2xUGI base editor sequence is represented by SEQ ID NO:186.

[0288] The following are several exemplary RNA sequences of the gRNA constructs used in this study. Each contains a customizable target, a gRNA scaffold, and one or two copies of the MS2 aptamer from the 5' end to the 3' end.

[0289] 1. Sequence of the gRNA_MS2 construct (SEQ ID NO:38):

[0290]

[0291] Symbol explanation:

[0292] 2. Sequence of the gRNA_2xMS2 construct (SEQ ID NO:39):

[0293]

[0294] Symbol explanation:

[0295] The three components of the platform / system disclosed herein can be expressed using one, two, or three expression vectors. The system can be programmed to substantially target any DNA or RNA sequence. Similar RNA scaffold recruitment systems can be generated by altering the modular components of the system, including any suitable Cas orthologs, deaminase orthologs, and other DNA-modifying enzymes.

[0296] Cell type / therapeutic use

[0297] The RNA scaffold recruitment system of the present invention can be used for genetically modified cells, including but not limited to animal cells, fungal cells, and plant cells. In a preferred embodiment, the RNA scaffold of the present invention can be used for genetically modified human cells. The present invention can be applied to master cell lines, immortalized cell lines, and primary cells isolated from humans. Examples of human cells include, but are not limited to, differentiated cells or cells in differentiation or stem cells. Suitable human cells include those derived from any one of the three germ cell layers of embryonic development (i.e., endoderm, mesoderm, and ectoderm). For example, human cells are those found in the following organs: skeletal muscle, bone, dermis, connective tissue, urogenital system, heart, blood (lymph nodes), and spleen (mesoderm); stomach, colon, liver, pancreas, bladder; urethral lining, trachea, lungs, pharynx, thyroid gland, parathyroid gland, epithelial portion of intestine (endoderm); or central nervous system, retina and lens, cranium and sensory organs, ganglia and nerves, pigment cells, connective tissue of the head, epidermis, hair, and mammary glands (ectoderm). In a preferred embodiment, the RNA scaffold is used to genetically modify primary immune cells or immune cell lines. Immune cells include T cells, NK cells, B cells, CD34+ hematopoietic stem cells (HSCs), and other cells involved in the production of lymphocytes and cells of the blood, bone marrow, spleen, lymph nodes, and thymus. Immune cells, particularly primary immune cells naturally present in the host animal or patient or derived from induced pluripotent stem cells (iPSCs), can be genetically modified. Immune cells include T cells, NK cells, B cells, pluripotent cells such as hematopoietic stem cells (HSCs), which are pluripotent cells capable of differentiating into immune cells, and other cells involved in the production of lymphocytes and cells of the blood, bone marrow, spleen, lymph nodes, and thymus.

[0298] This article provides methods for genome engineering in vitro, in vivo, or in vitro cells (e.g., methods for altering or manipulating the expression of one or more genes or gene products). Specifically, the methods provided herein can be used for targeted base editing disruptions in mammalian cells.

[0299] On the other hand, this article provides methods for targeting diseases to perform base editing correction. The target sequence can be any disease-related polynucleotide or gene already established in the art. Examples of useful applications of mutations or "corrections" of endogenous gene sequences include alterations to disease-related gene mutations, alterations to sequences encoding splice sites, alterations to regulatory sequences, alterations to sequences causing gain-of-function mutations, and / or alterations to sequences causing loss-of-function mutations, as well as targeted alterations to sequences encoding structural properties of proteins.

[0300] In some cases, it may be advantageous to genetically modify cells using the methods described herein to enable the cells to express chimeric antigen receptors (CARs) and / or T-cell receptors (TCRs). "Chimeric antigen receptor (CAR)" is sometimes referred to as "chimeric receptor," "T-body," or "chimeric immune receptor (CIR)." As used herein, the term "chimeric antigen receptor (CAR)" refers to an artificially constructed hybrid protein or polypeptide comprising an antibody extracellular antigen-binding domain (e.g., a single-chain variable fragment (scFv)) operably linked to a transmembrane domain and at least one intracellular domain. Typically, the antigen-binding domain of a CAR is specific to a particular antigen expressed on the surface of a target cell of interest. For example, T cells may be engineered to express a CAR specific to CD19 on B-cell lymphoma. For allogeneic antitumor cell therapies not limited by donor matching, cells may be engineered to isolate the nucleic acid encoding the CAR but also knock out the genes responsible for donor matching (TCRs and HLA markers).

[0301] As used herein, the terms “genetically modified” and “genetically engineered” are used interchangeably and refer to prokaryotic or eukaryotic cells comprising exogenous polynucleotides, regardless of the method used for insertion. In some cases, the effector cells have been modified to contain non-naturally occurring nucleic acid molecules that have been artificially produced or modified (e.g., using recombinant DNA technology) or derived from such molecules (e.g., through transcription, translation, etc.). Effector cells containing exogenous, recombinant, synthetic, and / or otherwise modified polynucleotides are considered engineered cells.

[0302] Cell therapy and ex vivo therapy

[0303] Various embodiments of the present invention also provide therapeutic cells produced or used according to any other embodiment of the invention. In one embodiment, the invention relates to a method for producing therapeutic cells, such as T cells engineered to express chimeric antigen receptor (CAR-T) or T cell receptor (TCR-T). CAR-T / TCR-T cells may be derived from primary T cells or differentiated from stem cells. Suitable stem cells include, but are not limited to, mammalian stem cells, such as human stem cells, including, but not limited to, hematopoietic, neural, embryonic, induced pluripotent stem cells (iPSCs), mesenchymal, mesoderm, liver, pancreas, muscle, and retinal stem cells. Other stem cells include, but are not limited to, mammalian stem cells, such as mouse stem cells, such as mouse embryonic stem cells.

[0304] In various embodiments, the present invention can be used to knock out, alter, or modify the expression of single or multiple genes in various types of cells or cell lines, including but not limited to cells derived from eukaryotes such as human cells. The present invention can be used for multiple modifications, i.e., one or more base edits, which can be introduced simultaneously or sequentially. This technology can be used in many applications, including but not limited to preventing graft-versus-host disease by knocking out genes from a gene to render non-host cells immunogenic to the host, or preventing host-versus-transplant disease by making non-host cells tolerant to host attack. These methods are also related to cell-based therapies that generate allogeneic (off-the-shelf) or autologous (patient-specific) genes. These genes include, but are not limited to, T-cell receptor (TRAC), major histocompatibility complex (MHCI and II) genes, including B2M, co-receptors (HLA-F, HLA-G), genes involved in innate immune responses (MICA, MICB, HCP5, STING, DDX41, and Toll-like receptors (TLRs)), inflammation (NKBBiL, LTA, TNF, LTB, LST1, NCR3, AIF1), heat shock proteins (HSPA1L, HSPA1A, HSPA1B), complement cascades, regulatory receptors (NOTCH family members), antigen processing (TAP, HLA-DM, HLA-DO), increased potency or persistence (such as PD-1, CTLA-4, and other members of the B7 family of checkpoint proteins), genes involved in immunosuppressive immune cells (such as FOXP3 and interleukin (IL)-10), and T-cell genes. Genes involved in interactions with the tumor microenvironment (including, but not limited to, receptors for cytokines such as TGFB, IL-4, IL-7, IL-2, IL-15, IL-12, IL-18, IFNγ), genes contributing to cytokine release syndrome (including, but not limited to, IL-6, IFNγ, IL-8 (CXCL8), IL-10, GM-CSF, MIP-1α / β, MCP-1 (CCL2), CXCL9, and CXCL10 (IP-10), genes encoding antigens targeted by CAR / TCR (e.g., CARs engineered to control CS1 endogenous CS1), or other genes found to be beneficial to CAR-T / TCR-T (such as TET2) or other cell-based therapeutics, including, but not limited to, CAR-NK, CAR-B, etc. See, for example, DeRenzo et al., Genetic Modification Strategies to Enhance CAR T Cell Persistence for Patients With Solid Tumors. Front. Immunol., February 15, 2019.

[0305] This technology can also be used to knock down or modify genes involved in the killing of immune cells (such as T cells and NK cells) or genes that warn patients or animals that foreign cells, particles, or molecules have entered the patient or animal's immune system, or genes encoding proteins that are current therapeutic targets (e.g., CD52 and PD1) for impairing or enhancing immune responses, respectively.

[0306] One application is to engineer HLA alleles in bone marrow cells to increase haplotype matching. Engineered cells can be used in bone marrow transplantation to treat leukemia. Another application is to engineer negative regulatory elements of the fetal hemoglobin gene in hematopoietic stem cells for the treatment of sickle cell anemia and β-thalassemia. The negative regulatory element will be mutated and will reactivate the expression of the fetal hemoglobin gene in hematopoietic stem cells, thereby compensating for functional losses due to mutations in the adult α or β hemoglobin gene. Another application is to engineer iPS cells to produce allogeneic therapeutic cells for various degenerative diseases, including Parkinson's disease (neuronal cell loss) and type 1 diabetes (pancreatic β-cell loss). Other exemplary applications include engineering HIV-resistant T cells by inactivating the CCR5 gene and other genes encoding the receptor required for HIV entry into cells.

[0307] Types of genetic modifications

[0308] Therefore, this paper provides methods for targeted disruption of transcription or translation of target genes. Specifically, the methods include targeted disruption of transcription or translation of target genes by interfering with start codons, introducing premature stop codons, and / or targeting and disrupting intron / exon sites.

[0309] Using the methods described herein, one or more genes of interest in primary cells can be enhanced and / or knocked out with improved efficiency and reduced rates of off-target insertions or deletions. In a preferred embodiment, the method is used for multiple base editing including gene knock-in, gene knockout, and missense mutations.

[0310] As described in the following paragraphs and examples, the inventors' flow cytometry method for genome engineering employs a base editor (e.g., third- and fourth-generation base editors, adenine base editor) for targeted gene interference via knockout and missense mutations, as well as targeted gene knock-in in the presence of a DNA donor template. The methods described herein are well-suited for studying hematopoietic cell biology and gene function, modeling diseases (e.g., major immunodeficiency), correcting disease-causing point mutations, and generating novel cell products (e.g., T-cell products) for therapeutic applications.

[0311] Delivery of components into cells

[0312] The following examples provide suitable methods for delivering base-editing components into cells.

[0313] In the embodiments provided herein, the RNA scaffold is chemically synthesized RNA and introduced into the cell via any suitable technique (e.g., electroporation). Base-editing enzyme components and class 2 Cas enzyme components can be introduced into the cell as mRNA or protein.

[0314] In some embodiments, the components including the base editor and guide molecule can be delivered to cells in vitro, ex vivo, or in vivo. In some cases, viral or plasmid vector systems are used to deliver the base editing components described herein. Preferably, the vector is a viral vector, such as a lentivirus or baculovirus, or preferably an adenovirus / adeno-associated virus (AAV) vector, but other delivery methods are known (e.g., yeast systems, microvesicles, gene guns / means of attaching vectors to gold nanoparticles) and are contemplated. In some embodiments, nucleic acids encoding gRNA and a base editor fusion protein are packaged for delivery to cells in one or more viral delivery vectors. Suitable viral delivery vectors include, but are not limited to, adenovirus / adeno-associated virus (AAV) vectors and lentiviral vectors. In some cases, non-viral transfer methods known in the art can be used to introduce nucleic acids or proteins into mammalian cells. Nucleic acids and proteins can be delivered together with pharmaceutically acceptable vectors or, for example, encapsulated in liposomes. Other delivery methods are known (e.g., yeast systems, microvesicles, gene guns / means of attaching vectors to gold nanoparticles) and are contemplated. In some cases, cells are electroporated for uptake of gRNA and base editors (e.g., BE3, BE4, ABE). In other cases, after introducing gRNA, base editors, and vectors via electroporation, DNA donor templates are delivered as adeno-associated virus type 6 (AAV6) vectors by adding viral supernatant to the culture medium.

[0315] The rate of insertion or deletion (indel) formation can be determined by appropriate methods. For example, Sanger sequencing or next-generation sequencing (NGS) can be used to detect the rate of indel formation. Preferably, the contact results in less than 20% off-target indel formation during base editing. This contact results in a desired product to undesired product ratio of at least 2:1 during base editing.

[0316] Expression System

[0317] To use the platforms described above, it may be necessary to express one or more protein and RNA components from the nucleic acids encoding them. This can be done in various ways. For example, nucleic acids encoding RNA scaffolds or proteins can be cloned into one or more intermediate vectors for introduction into prokaryotic or eukaryotic cells for replication and / or transcription. Intermediate vectors are typically prokaryotic vectors, such as plasmids or shuttle vectors, or insect vectors, for storing or manipulating the nucleic acids encoding RNA scaffolds or proteins for the production of RNA scaffolds or proteins. Nucleic acids can also be cloned into one or more expression vectors for application to plant cells, animal cells, preferably mammalian or human cells, fungal cells, bacterial cells, or protozoan cells. Therefore, the present invention provides nucleic acids encoding any of the aforementioned RNA scaffolds or proteins. Preferably, the nucleic acids are isolated and / or purified.

[0318] The present invention also provides recombinant constructs or vectors having sequences encoding one or more of the aforementioned RNA scaffolds or proteins. Examples of such constructs include vectors, such as plasmids or viral vectors, into which the nucleic acid sequences of the present invention are inserted in either a forward or reverse orientation. In a preferred embodiment, the construct further includes a regulatory sequence, comprising a promoter operatively linked to the sequence. A wide range of suitable vectors and promoters are known to those skilled in the art and are commercially available. Suitable cloning and expression vectors are used with prokaryotic and eukaryotic hosts known in the art.

[0319] A vector is a nucleic acid molecule capable of delivering another nucleic acid linked to it. Vectors can autonomously replicate or integrate into host DNA. Examples of vectors include plasmids, granules, or viral vectors. The vectors of the present invention comprise nucleic acids in a form suitable for expression in host cells. Preferably, the vector comprises one or more regulatory sequences operatively linked to a nucleic acid sequence to be expressed. "Regulatory sequences" include promoters, enhancers, and other expression control elements (e.g., polyadenylation signals). Regulatory sequences include constitutive expression of nucleotide sequences as well as inducible regulatory sequences. The design of the expression vector can depend on factors such as the choice of host cells to be transformed, transfected, or transduced, and the desired expression level of RNA or protein.

[0320] Examples of expression vectors include chromosomal, non-chromosomal, and synthetic DNA sequences, bacterial plasmids, bacteriophage DNA, baculoviruses, yeast plasmids, vectors derived from combinations of plasmid and bacteriophage DNA, and viral DNA such as vaccinia virus, adenovirus, fowlpox virus, and pseudorabies virus. However, any other vector may be used, provided it is reproducible and viable in the host. Suitable nucleic acid sequences can be inserted into the vector using various procedures. Typically, nucleic acid sequences encoding one of the aforementioned RNAs or proteins can be inserted with appropriate restriction endonuclease sites using procedures known in the art. Such procedures and related subcloning procedures are within the scope of those skilled in the art.

[0321] The vector may include a suitable sequence for amplifying expression. Furthermore, the expression vector preferably contains one or more optional marker genes to provide phenotypic traits for selecting host cells for transformation, such as dihydrofolate reductase or neomycin resistance for eukaryotic cell culture, or tetracycline or ampicillin resistance, for example, in *E. coli*.

[0322] Vectors for expressing RNA may include RNApol III promoters, such as HI, U6, or 7SK promoters, to drive RNA expression. These human promoters allow RNA expression in mammalian cells after plasmid transfection. Alternatively, the T7 promoter may be used, for example, for in vitro transcription, and the RNA can be transcribed and purified in vitro.

[0323] Vectors containing the appropriate nucleic acid sequences as described above, along with appropriate promoters or control sequences, can be used to transform, transfect, or infect appropriate hosts to allow the host to express the RNA or protein described above. Examples of suitable expression hosts include bacterial cells (e.g., *Escherichia coli*, *Streptomyces*, *Salmonella typhimurium*), fungal cells (yeast), insect cells (e.g., fruit flies and fall armyworms (Spodoptera frugiperda, Sf9)), animal cells (e.g., CHO, COS, and HEK293), adenoviruses, and plant cells. The selection of appropriate hosts is within the scope of those skilled in the art. In some specific embodiments, the present invention provides a method for producing the aforementioned RNA or protein by transforming, transfecting, or infecting host cells with an expression vector having a nucleotide sequence encoding one of RNA, a polypeptide, or a protein. The host cells are then cultured under appropriate conditions, which allows for the expression of the RNA or protein.

[0324] Any procedure known in the art for introducing foreign nucleotide sequences into host cells may be used. Examples include transfection with calcium phosphate, polybrene, protoplasmic fusion, electroporation, nuclear transfection, liposomes, microinjection, naked DNA, plasmid vectors, viral vectors, free and integrated methods, and any other well-known methods for introducing cloned genomic DNA, cDNA, synthetic DNA, or other foreign genetic material into host cells.

[0325] Cultured cells

[0326] The method further includes maintaining the cell under appropriate conditions such that the guide RNA directs the effector protein to a target site in the target sequence, and the effector domain modifies the target sequence.

[0327] Typically, cells can be maintained under conditions suitable for cell growth and / or maintenance. Suitable cell culture conditions are well known to those skilled in the art, and those skilled in the art understand that methods used for culturing cells are known in the art and can vary depending on the cell type. In all cases, routine optimization can be used to determine the optimal technique for a particular cell type.

[0328] Cells that can be used in the methods provided herein may be freshly isolated primary cells or obtained from frozen aliquots of primary cell cultures. In some cases, cells are electroporated for uptake of gRNA and base-editing fusion proteins. As described in the examples below, electroporation conditions for some assays (e.g., for T cells) may include 1400 volts, a pulse width of 10 ms, and 3 pulses. After electroporation, the electroporated T cells are allowed to recover in cell culture medium and then cultured in T cell expansion medium. In some cases, the electroporated cells are allowed to recover in cell culture medium for approximately 5 minutes to approximately 30 minutes (e.g., approximately 5 minutes, 10 minutes, 15 minutes, 20 minutes, 25 minutes, 30 minutes). Preferably, the recovered cell culture medium is free of antibiotics or other selective agents. In some cases, the T cell expansion medium is a complete CTS OpTmizer T-cell expansion medium.

[0329] application

[0330] The RNA scaffold of the present invention can be used for the following applications: genome editing, genome screening, generation of therapeutic cells, genome labeling, epigenome editing, karyotype engineering, chromatin imaging, transcriptome and metabolic pathway engineering, gene circuit engineering, cell signal transduction sensing, cell event recording, lineage information reconstruction, gene drive, DNA genotyping, miRNA quantification, in vivo cloning, site-guided mutagenesis, genome diversification, and in situ proteomics analysis.

[0331] Applications also include research on human diseases, such as cancer immunotherapy, antiviral therapy, phage therapy, cancer diagnosis, pathogen screening, microbiome remodeling, stem cell reprogramming, immune genome engineering, vaccine development, and antibody production.

[0332] definition

[0333] Nucleic acid or polynucleotide refers to a DNA molecule (e.g., but not limited to cDNA or genomic DNA) or an RNA molecule (e.g., but not limited to mRNA), and includes DNA or RNA analogs. DNA or RNA analogs can be synthesized from nucleotide analogs. DNA or RNA molecules may include parts that are not naturally occurring, such as modified bases, modified backbones, deoxyribonucleotides in RNA, etc. Nucleic acid molecules can be single-stranded or double-stranded. Those skilled in the art will understand that uracil is a nucleotide that replaces thymine in the RNA format. The DNA sequences disclosed herein will contain thymine nucleotides, and the corresponding RNA sequences will contain uracil nucleotides at the same positions.

[0334] When referring to nucleic acid molecules or polypeptides, the term "isolated" means that the nucleic acid molecule or polypeptide is substantially free of at least one other component associated with it or found together with it in terms of properties.

[0335] As used herein, the term "guide RNA" generally refers to an RNA molecule (or a group of RNA molecules) that can bind to CRISPR proteins and target them to a specific location within target DNA. Guide RNA can contain two segments: a DNA-targeting guide segment and a protein-binding segment. The DNA-targeting segment contains a nucleotide sequence complementary to (or at least capable of hybridizing under stringent conditions) the target sequence. The protein-binding segment interacts with a CRISPR protein (such as Cas9 or a Cas9-related polypeptide). These two segments can reside in the same RNA molecule or in two or more separate RNA molecules. When the two segments are in separate RNA molecules, the molecule containing the DNA-targeting guide segment is sometimes called CRISPR RNA (crRNA), while the molecule containing the protein-binding segment is called trans-activating RNA (tracrRNA).

[0336] As used herein, the term “target nucleic acid” or “target” refers to a nucleic acid containing a target nucleic acid sequence. Target nucleic acids can be single-stranded or double-stranded, and are typically double-stranded DNA. As used herein, “target nucleic acid sequence,” “target sequence,” or “target region” refers to a specific sequence or its complementary sequence that is intended to be bound or modified using the CRISPR system. Target sequences can be located within the genome of a cell, in vitro, or in vivo, and can be in any form of single-stranded or double-stranded nucleic acid.

[0337] "Target nucleic acid strand" refers to the strand of the target nucleic acid that pairs with the guide RNA as disclosed herein. In other words, the strand of the target nucleic acid that hybridizes with the crRNA and the guide sequence is called the "target nucleic acid strand." The other strand of the target nucleic acid, which is not complementary to the guide sequence, is called the "non-complementary strand." In the case of double-stranded target nucleic acids (e.g., DNA), each strand can be a "target nucleic acid strand," provided there is a suitable PAM site, thus allowing the design of crRNA and guide RNA for implementing the methods of this invention.

[0338] As used herein, the term "derived from" refers to a process in which a first component (e.g., a first molecule) or information derived from the first component is used to isolate, derive, or prepare a different second component (e.g., a second molecule different from the first molecule). For example, mammalian codon-optimized Cas9 polynucleotides are derived from the amino acid sequence of wild-type Cas9 proteins. Additionally, variant mammalian codon-optimized Cas9 polynucleotides, including Cas9 single-mutant nickases (nCas9, e.g., nCas9D10A) and Cas9 double-mutant null-nucleases (dCas9, e.g., dCas9D10AH840A), are derived from polynucleotides encoding wild-type mammalian codon-optimized Cas9 proteins.

[0339] As used herein, the term "wild type" is a term in the art as understood by those skilled in the art, and refers to the typical form of an organism, strain, gene, or characteristic as it naturally exists, distinguished from mutant or variant forms.

[0340] As used herein, the term "variant" refers to a first component (e.g., a first molecule) associated with a second component (e.g., a second molecule, also known as a "parent" molecule). Variant molecules can be derived from, isolated from, based on, or homologous to a parent molecule. For example, mutant forms of mammalian codon-optimized Cas9 (hspCas9), including nickases of single-mutant Cas9 and null-nucleases of double-mutant Cas9, are variants of mammalian codon-optimized wild-type Cas9 (hspCas9). The term "variant" can also be used to describe polynucleotides or polypeptides.

[0341] When applied to polynucleotides, variant molecules can have an entire nucleotide sequence identity with the original parent molecule, or they can have less than 100% nucleotide sequence identity with the parent molecule. For example, a variant of a gene nucleotide sequence can be a second nucleotide sequence that is at least 50%, 60%, 70%, 80%, 90%, 95%, 98%, or 99% identical to the original nucleotide sequence. Polynucleotide variants also include polynucleotides comprising the entire parent polynucleotide and additional fusion nucleotide sequences. Polynucleotide variants also include polynucleotides that are portions or subsequences of the parent polynucleotide, such as unique subsequences of the polynucleotides disclosed herein (e.g., as determined by standard sequence comparison and alignment techniques).

[0342] In another aspect, polynucleotide variants include nucleotide sequences containing minor, insignificant, or negligible alterations to the parental nucleotide sequence. For example, minor, insignificant, or negligible alterations include changes to the nucleotide sequence that (i) do not change the amino acid sequence of the corresponding polypeptide, (ii) occur outside the protein-coding open reading frame of the polynucleotide, (iii) result in deletions or insertions that may affect the corresponding amino acid sequence but have little or no effect on the biological activity of the polypeptide, and (iv) result in the substitution of an amino acid with a chemically similar amino acid. In cases where the polynucleotide does not encode a protein (e.g., tRNA, crRNA, or tracrRNA), the variant of the polynucleotide may include nucleotide changes that do not result in a loss of function of the polynucleotide. In another aspect, the present invention covers conserved variants of the disclosed nucleotide sequences that produce functionally identical nucleotide sequences. Those skilled in the art will appreciate that the present invention covers many variants of the disclosed nucleotide sequences.

[0343] When applied to proteins, variant polypeptides can have an entire amino acid sequence identity with the original parent polypeptide, or they can have less than 100% amino acid identity with the parent protein. For example, a variant amino acid sequence can be a second amino acid sequence that is at least 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to the original amino acid sequence.

[0344] Peptide variants include peptides comprising the entire parent peptide and also comprising additional fusion amino acid sequences. Peptide variants also include peptides that are portions or subsequences of the parent peptide, such as unique subsequences of the peptides disclosed herein (e.g., as determined by standard sequence comparison and alignment techniques).

[0345] In another aspect, peptide variants include peptides containing minor, insignificant, or negligible alterations to the parental amino acid sequence. For example, minor, insignificant, or negligible changes include amino acid variations that have little or no effect on the bioactivity of the peptide (including substitution, deletion, and insertion) and produce functionally identical peptides, including the addition of a nonfunctional peptide sequence. In other aspects, the variant peptides of the present invention alter the bioactivity of the parent molecule, for example, mutant variations of Cas9 peptides that modify or lose nuclease activity. Those skilled in the art will appreciate that the present invention covers many variants of the disclosed peptides.

[0346] In some aspects, the polynucleotide or polypeptide variants of the present invention may include variant molecules that alter, add or delete small percentages of nucleotide or amino acid positions (e.g., typically less than about 10%, less than about 5%, less than 4%, less than 2% or less than 1%).

[0347] As used herein, the term “conservative substitution” in a nucleotide or amino acid sequence refers to a change in the nucleotide sequence that: (i) results in redundancy due to the triplet codon encoding the amino acid sequence without any corresponding change in the amino acid sequence, or (ii) results in the substitution of the original parent amino acid with an amino acid having a chemically similar structure. Conservative substitution of functionally similar amino acids is well known in the art, whereby one amino acid residue is substituted with another amino acid residue having similar chemical properties (e.g., an aromatic side chain or a positively charged side chain), and therefore substantially does not alter the functional properties of the resulting polypeptide molecule.

[0348] The following is a grouping of natural amino acids containing similar chemical properties, where the substitutions within a group are "conserved" amino acid substitutions. This grouping is not rigid, as these natural amino acids can be placed in different groups when considering different functional properties. Amino acids with nonpolar and / or aliphatic side chains include: glycine, alanine, valine, leucine, isoleucine, and proline. Amino acids with polar, uncharged side chains include: serine, threonine, cysteine, methionine, asparagine, and glutamine. Amino acids with aromatic side chains include: phenylalanine, tyrosine, and tryptophan. Amino acids with positively charged side chains include: lysine, arginine, and histidine. Amino acids with negatively charged side chains include aspartate and glutamate.

[0349] A “Cas9 mutant” or “Cas9 variant” refers to a protein or polypeptide derivative of the wild-type Cas9 protein, such as the *Streptococcus pyogenes* Cas9 protein, for example, a protein having one or more point mutations, insertions, deletions, truncations, fusion proteins, or combinations thereof. It essentially retains the RNA-targeting activity of the Cas9 protein. The protein or polypeptide may contain, be composed of, or be essentially composed of fragments of the *Streptococcus pyogenes* Cas9 protein. Typically, a mutant / variant is at least 50% (e.g., any number between 50% and 100%, including 0 and 100%) of identity with the *Streptococcus pyogenes* Cas9 protein. The mutant / variant can bind to and target specific DNA sequences via RNA molecules and may additionally possess nuclease activity. Examples of these domains include the RuvC-like motif (amino acids 7-22, 759-766, and 982-989 of the *Streptococcus pyogenes* Cas9 protein) and the HNH motif (amino acids 837-863). See Gasinas et al., Proc Natl Acad Sci US A. 25 Sep 2012; 109(39):E2579–E2586 and WO2013176772.

[0350] "Complementarity" refers to the ability of a nucleic acid to form hydrogen bonds with another nucleic acid sequence through conventional Watson-Crick base pairing or other non-conventional types. Percentage complementarity indicates the percentage of residues in a nucleic acid molecule that are complementary to a second nucleic acid sequence (e.g., 50%, 60%, 70%, 80%, 90%, and 100% complementary to Watson-Crick base pairing, respectively). "Perfect complementarity" means that all consecutive residues in the nucleic acid sequence will form hydrogen bonds with the same number of consecutive residues in the second nucleic acid sequence. As used herein, “substantially complementary” means a degree of complementarity of at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% in regions of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, or 50 nucleotides, or refers to two nucleic acids that hybridize under stringent conditions.

[0351] As used herein, “strict conditions” for hybridization refer to conditions under which nucleic acids complementary to the target sequence hybridize primarily with the target sequence and substantially with no hybridization with non-target sequences. Strict conditions are generally sequence-dependent and vary depending on many factors. Typically, the longer the sequence, the higher the temperature at which it hybridizes specifically with its target sequence. Non-restrictive examples of strict conditions are detailed in Tijssen (1993), Laboratory Techniques in Biochemistry and Molecular Biology—Hybridization with Nucleic Acid Probes, Part I, Chapter 2, “Overview of principles of hybridization and the strategy of nucleic acid probe assay”, Elsevier, NY.

[0352] "Hybridization" refers to the process by which completely or partially complementary nucleic acid chains together form a double-stranded structure or region under specified hybridization conditions, in which the two constituent chains are linked by hydrogen bonds. Although hydrogen bonds are usually formed between adenine and thymine or uracil (A and T or U) or cytidine and guanine (C and G), other base pairs can be formed (e.g., Adams et al., The Biochemistry of the Nucleic Acids, 11th ed., 1992).

[0353] As used herein, “expression” refers to the process of transcribing polynucleotides from a DNA template (e.g., into mRNA or other RNA transcripts) and / or the subsequent conversion of transcribed mRNA into peptides, polypeptides, or proteins. Transcription and the encoded polypeptides can be collectively referred to as “gene products.” If the polynucleotides originate from genomic DNA, expression may include splicing mRNA into eukaryotic cells.

[0354] The terms “polypeptide,” “peptide,” and “protein” are used interchangeably herein to refer to polymers of amino acids of any length. Polymers may be linear or branched, may contain modified amino acids, and may be interrupted by non-amino acid components. The terms also cover modified amino acid polymers; for example, those involving disulfide bond formation, glycosylation, esterification, acetylation, phosphorylation, PEGylation, or any other manipulation, such as conjugation with a labeled component. As used herein, the term “amino acid” includes natural and / or non-natural or synthetic amino acids, including glycine and D or L optical isomers, as well as amino acid analogs and peptide mimics.

[0355] The terms "fusion polypeptide" or "fusion protein" refer to a protein created by linking two or more polypeptide sequences together. The fusion polypeptides covered by this invention comprise the translational product of a chimeric gene construct that links a nucleic acid sequence encoding a first polypeptide (e.g., an RNA-binding domain) with a nucleic acid sequence encoding a second polypeptide (e.g., an effector domain) to form a single open reading frame. In other words, a "fusion polypeptide" or "fusion protein" is a recombinant protein of two or more proteins linked by peptide bonds or via several peptides. Fusion proteins may also include a peptide linker between the two domains.

[0356] The term "connector" refers to any means, entity, or portion used to connect two or more entities. A connector can be a covalent or non-covalent connector. Examples of covalent bonds include covalent bonds or connector portions covalently attached to one or more proteins or domains to be connected. Connectors can also be non-covalent, such as organometallic bonds through a metal center, such as a platinum atom. For covalent bonds, various functional groups can be used, such as amide groups, including carbonate derivatives, ethers, esters (including organic and inorganic esters), amino groups, carbamates, ureas, etc. To provide a connection, the domain can be modified by oxidation, hydroxylation, substitution, reduction, etc., to provide a coupling site. Methods for conjugation are well known to those skilled in the art and are covered in this invention. Connector portions include, but are not limited to, chemical connector portions, or, for example, peptide connector portions (connector sequences). It should be understood that modifications that do not significantly reduce the function of RNA-binding domains and effector domains are preferred.

[0357] As used herein, the terms “conjugate”, “conjugation”, or “linkage” refer to the connection of two or more entities to form a single entity. Conjugates encompass peptide-small molecule conjugates as well as peptide-protein / peptide conjugates.

[0358] The terms “subject” and “patient” are used interchangeably herein to refer to vertebrates, preferably mammals, and more preferably humans. Mammals include, but are not limited to, rodents, apes, humans, farm animals, sporting animals, and pets. It also encompasses tissues, cells, and their progeny of biological entities obtained in vivo or cultured in vitro. In some embodiments, the subject may be an invertebrate, such as an insect or nematode; while in other embodiments, the subject may be a plant or fungus.

[0359] As used herein, the terms “treatment,” “relief,” or “improvement” are used interchangeably. These terms refer to methods used to obtain a beneficial or desired outcome, including, but not limited to, therapeutic benefits and / or preventive benefits. A therapeutic benefit is any treatment-related improvement or effect on one or more diseases, conditions, or symptoms during treatment. For preventive benefits, the composition may be administered to a subject at risk of developing a particular disease, condition, or symptom, or to a subject who reports one or more physiological symptoms of a disease, even if the disease, condition, or symptom has not yet manifested.

[0360] As used herein, the term “contact” in any grouping context includes any process in which the components to be contacted are mixed into the same mixture (e.g., added to the same compartment or solution), and does not necessarily require actual physical contact between the components. The components may be contacted in any order or in any combination (or sub-combination), and may include cases where one or more of the components are subsequently optionally removed from the mixture before the addition of other said components. For example, “contacting A with B and C” includes any and all of the following: (i) A is mixed with C, and then B is added to the mixture; (ii) A and B are mixed to form a mixture; B is removed from the mixture, and then C is added to the mixture; and (iii) A is added to a mixture of B and C. “Contacting” a target nucleic acid or cell with one or more reaction components (e.g., Cas protein or guide RNA) includes any and all of the following: (i) contacting the target or cell with a first component of the reaction mixture to produce a mixture; then adding other components of the reaction mixture to the mixture in any order or combination; and (ii) fully forming the reaction mixture before mixing with the target or cell.

[0361] As used herein, the term "mixture" refers to a combination of elements that are dispersed and not in any particular order. Mixtures are multiphase and not spatially separable into their distinct components. Examples of mixtures of elements include many different elements dissolved in the same aqueous solution, or many different elements attached to a solid carrier in a random or unpredictable order, wherein the different elements are not spatially distinct. In other words, mixtures are not addressable.

[0362] As disclosed herein, multiple value ranges are provided. It should be understood that, unless the context explicitly states otherwise, each intermediate value between the upper and lower limits of the range, up to one-tenth of the lower limit unit, is also specifically disclosed. Each smaller range between any stated value or intermediate value within the range and any other stated value or intermediate value within the range is included in the invention. The upper and lower limits of these smaller ranges may be independently included in or excluded from the range, and each range that includes any one or both of the smaller ranges or excludes neither of the smaller ranges is also included in the invention, subject to any specific exclusions from the range. Where the range includes one or both limits, ranges that exclude any one or both of those included limits are also included in the invention. The term “about” generally refers to ±10% of the indicated quantity. For example, “about 10%” may represent a range of 9% to 11%, and “about 20” may represent 18 to 22. Other meanings of “about” may be apparent from the context, such as rounding, and thus, for example, “about 1” may also mean from 0.5 to 1.4.

[0363] Various exemplary embodiments of the compositions and methods according to the present invention are now described in the following examples.

[0364] Example

[0365] Example 1 - Modification of RNA scaffold

[0366] sgRNA sequence design

[0367] A complete list of the sgRNA designs used and their sequences is shown in Table 4. All sgRNA designs were based on *Streptococcus pyogenes* sgRNA, consisting of a target-specific 20 nt spacer sequence, a 76 nt b constant region sgRNA sequence, and a 7 nt poly-T U6 termination signal. All modifications were made to the constant components of the sgRNA and consisted of repetitions containing RNA aptamer hairpins and / or stems: anti-repetition extensions. A single copy (1xMS2) or two copies (2xMS2) of the MS2 hairpin sequence (C5 variant) was introduced into the 2' or 3' of the sgRNA, either in the quadruple loop or the stem loop. For 2xMS2tracrRNA, two designs were pursued, one of which integrated two copies of the C5MS2 variant into the 3' of the sgRNA, and the second design consisted of the C5 variant located at the 2' of the stem loop and an engineered MCP protein-binding f6 aptamer assimilated into the 3' of the sgRNA. The f6 aptamer is a different variant used for the 2x MS2 plasmid design. Repetition: Various extensions of the upper stem of the anti-repetition are introduced into either side of the stem, and in each case, the extension introduces a natural Streptococcus pyogenes sequence.

[0368] Table 4. Design and sequences of different sgRNAs. N indicates a 20 nt target-specific spacer sequence. Constant sgRNA sequences as previously described are highlighted in bold. Extended repeats: antirepetitive sequences are underlined. MS2 (C5 variant) or f6 aptamer sequences are shown in italics, while extensions of aptamer and adapter sequences are shown in italics and underlined. US = repeat: upper stem of antirepetition; TL = quadruple loop; SL2 = stem-loop 2.

[0369]

[0370] plasmid design

[0371] In addition to the sgRNA, all components of the base editing system are encoded on a single vector and represented as a single polycistronic unit from the CMV promoter. The vector encodes the expression of the APOBEC-1-MCP fusion protein fused to UGI at its C-terminus and nCas9 (D10A) – the nCas9-UGI fusion protein has two copies of SV40 NLS side-joined at the C-terminus of nCas9 and the N-terminus of UGI. Additionally, the vector encodes the expression of turboRFP to allow for monitoring of transfection efficiency.

[0372] The sgRNA components of the base editing system were expressed on separate vectors, with expression driven by the RNA polymerase III U6 promoter. The sgRNA was expressed as a single unit encompassing the crRNA and tracrRNA components of *Streptococcus pyogenes* Cas9 linked via the aforementioned artificial tetracycle. A list of target sgRNA sequences is shown in Table 5; if the target does not have a 5' G, a G must be added for expression from the U6 promoter.

[0373] The design of the expression in the BE4max base editor is as described above.

[0374] Table 5. gRNA target site sequences used for base editing. Cs located within the editing window are shown in bold.

[0375] Target Name target sequence Site 2 <![CDATA[GAAC 1 AC 2 AAAGCATAGACTGC(SEQ ID NO:53)]]> Site 3 <![CDATA[GGC 1 C 2 C 3 AGACTGAGCACGTGA(SEQ ID NO:54)]]> CTNNB1 <![CDATA[CTGGAC 2 TC 3 TGGAATCCATTC(SEQ ID NO:55)]]> EGFR <![CDATA[ATC 1 AC 2 GCAGCTCATGCCCTT(SEQ ID NO:56)]]> PCSK9 <![CDATA[CAGGTTC 2 C 3 ACGGGATGCTCT(SEQ ID NO:57)]]> FANCF <![CDATA[GGAATC 1 C 2 C 3 TTCTGCAGCACC(SEQ ID NO:58)]]> TRAC <![CDATA[TTC 1 GTATC 2 TGTAAAACCAAG(SEQ ID NO:59)]]> B2M <![CDATA[CTTAC 2 C 3 C 4 C 5 ACTTAACTATCT(SEQ ID NO:60)]]> CR0118_PDCD1 CAGTTCCAAACCCTGGTGGT(SEQ ID NO:61) CR0107_PDCD1 GGGGGTTCCAGGGCCTGTCT(SEQ ID NO:62) CR0057-TRAC_EX3 TTCGTATCTGTAAAACCAAG(SEQ ID NO:63) CR0151_CD2 GTTCAGCCAAAACCTCCCCA(SEQ ID NO:64) CR0121_PDCD1 GGAGTCTGAGAGATGGAGAG(SEQ ID NO:65) CR0165_CIITA CAGCTCACAGTGTGCCACCA(SEQ ID NO:66) TRAC_22550571 TTCAAAACCTGTCAGTGATT(SEQ ID NO:67) PDCD1_241852953 GGGGGTTCCAGGGCCTGTCT(SEQ ID NO:68) CTNNB1 CTGGACTCTGGAATCCATTC(SEQ ID NO:69)

[0376] Cell culture and transfection

[0377] HEK293 cells were cultured in DMEM (Dürbeco Modified Eagle Medium) supplemented with 10% FBS. 24 hours prior to transfection, 50,000 cells were seeded into individual wells of a 24-well plate to achieve 70% confluence for transfection. 24 hours later, cells were transfected using Lipofectamine 3000 reagent (Thermo Fisher Scientific) with 200 ng plasmid DNA (150 ng pin-point / BE4max vector and 50 ng sgRNA expression vector).

[0378] Cell lysis and flow cytometry

[0379] 72 hours after transfection, the culture medium was removed, cells were washed 1x with PBS and detached from the wells with 100 μl of TrypLE expression enzyme (ThermoFisher Scientific). The detached cells were then centrifuged at 300x rpm for 5 minutes at room temperature, and the supernatant was decanted. The cell pellet was washed 1x with PBS and centrifuged again at 300x rpm for 5 minutes, and the supernatant was discarded. The cell pellet was then resuspended in 100 μl of PBS. 20 μl of the resuspended cells were transferred to a 96-well plate and incubated with 36 μl of DirectPCR lysis reagent (Viagen Biotech) at 55°C for 0 minutes, followed by incubation at 95°C for 30 minutes. The cell lysate was stored at -20°C. The remaining 80 μl of the resuspended cells were transferred to a 96-well plate and centrifuged at 300x rpm for 5 minutes at room temperature. Decant the supernatant and resuspend the cells in 50 μl of MACS buffer (Miltenyi Biotec) supplemented with 0.5% BSA for flow cytometry analysis. All flow cytometry analyses were performed using an iQue3 (Sartorius) analyzer.

[0380] PCR amplification of the target region

[0381] Each PCR reaction used 1 μl of cell lysate obtained using DirectPCR lysis reagent. Q5 High Fidelity 2x Master Mix (NEB) was used for sgRNA target site amplification, and the reaction mixture was set as follows:

[0382]

[0383] The PCR reaction was carried out under the following thermal cycling conditions:

[0384]

[0385] The primers used and their annealing temperatures are detailed in Table 6 below:

[0386]

[0387] Unpurified PCR amplicon was sequenced using Sanger sequencing via Genewiz.

[0388] Example 2 - Base editing efficiency of modified RNA scaffolds

[0389] RNA synthesis

[0390] All crRNAs and tracrRNAs were synthesized by Horizon Discovery using chemistries protected with 2'-acetoxyethyl orthocyanin (2'-ACE) or 2'-tert-butyldimethylsilyl (2'-TBDMS). Chemical modifications were included at the noted locations, including two 2'-O-methyl nucleotides and two phosphate-thioester bonds at the 5' end of crRNA and the 3' end of tracrRNA (2xMS modification). The RNA oligonucleotides were 2'-deprotected / desalted and purified by high-performance liquid chromatography (HPLC) or polyacrylamide gel electrophoresis (PAGE). Prior to electroporation, the oligonucleotides were resuspended in 10 mM Tris pH 7.5 buffer.

[0391] Electroporation

[0392] Using Invitrogen TM Neon TM A 10 μL transfection kit was used to electroporate HEK 293T cells (ATCC, #CRL-11268). 50,000 cells, 1 μg mRNA, and a mixture of 6 μM synthetic crRNA:tracrRNA were electroporated at 1150V for 20 ms and 2 pulses. The mRNA (obtained from TriLink or in-house via standard methods) was mixed with nCas9-UGI at a 3:1 molar ratio with MCP-AID or MCP-APOBEC. Cell plates were seeded in 96-well plates containing whole serum growth medium and harvested after 72 hours for further processing.

[0393] Cell treatment

[0394] Cells were lysed at 56°C for 30 minutes in 100 μL buffer containing proteinase K (Thermo Scientific, #FEREO0492), RNase A (Thermo Scientific, #FEREN0531), and Phusion HF buffer (Thermo Scientific, #F-518L), followed by heat inactivation at 95°C for 5 minutes. The cell lysate was used to generate 200–400 nucleotide PCR amplicons spanning the region containing the base editing site. Unpurified PCR amplicons were sequenced using Sanger sequencing via Genewiz.

[0395] Editor's Analysis

[0396] Base editing efficiency was calculated from AB1 files using the Chimera analysis tool (an adaptation of the open-source tool BEAT). Chimera determines editing efficiency by first subtracting background noise to define the expected variability in the sample. This allows for estimation of editing efficiency without the need for normalization relative to control samples. Following this, Chimera uses the Median Absolute Deeviation (MAD) method to filter out any outliers from the noise, and then evaluates the base editor's editing efficiency across a 20 bp span of the input wizard sequence.

[0397] Example 3: Application of lentivirally integrated sgRNA in a base editing system for human primary immune cells

[0398] In this embodiment, primary human Pan T lymphocytes were used to demonstrate the utility of base-editing mRNA components in primary immune cells under the control of the PolIII promoter in the presence of constitutively expressed sgRNA with RNA aptamers. PanT cells were activated with anti-CD3 and anti-CD28 and then transduced using enriched and concentrated lentiviral particles. Puromycin selection was used to select successfully transduced cells to ensure that >95% of the population had at least one copy of the lentiviral insert. During the reactivation of selected T cells with anti-CD3 and anti-CD28, cells were electroporated with mRNA components for both deaminase-MCP and nCas9-UGI-UGI components. Cells were then incubated for an additional 72–96 hours, and cell surface knockout was examined by flow cytometry, while base editing was examined by targeted PCR amplification and Sanger sequencing.

[0399] Example 4: Application of synthetic crRNA and tracrRNA-aptamer guides to the alkali of human primary immune cells Basic editing system

[0400] In this embodiment, primary human Pan T lymphocytes were used to demonstrate the efficacy of a base editing system with crRNA and aptamer-modified tracrRNA components in primary immune cells. Pan T cells were activated using anti-CD3 and anti-CD28, and then electroporated with the mRNA components for deaminase-MCP, nCas9-UGI-UGI components, tracrRNA-aptamer, and crRNA. Cells were then incubated for an additional 72–96 hours, and cell surface knockout was examined by flow cytometry, while base editing was examined by targeted PCR amplification and Sanger sequencing.

[0401] Data show that base editing systems can edit primary immune cells without integrating DNA into the genome (via lentiviral boxes), utilizing different crRNA and tracrRNA aptamers with mRNA components. Results revealed different RNA aptamers and deaminase specificities for Apobec1 with a preference for single RNA motifs, while AID deaminase preferred dual RNA motifs in this context. The results demonstrate the high efficiency of base editing systems for altering specific bases for functional protein knockout via surface staining and flow cytometry, as well as through DNA-level modifications.

[0402] Materials and methods

[0403] Guidelines

[0404] Internally generated data was used to specify the base editing window calculated at a set distance from the PAM motif (NGG). This data was used to develop algorithms to predict phenotypes or gene KO applicable guide sequences for the following genes: TRAC, TRBC1, TRBC2, PDCD-1, B2M, and CD52 (Table 7). crRNA and tracrRNA were synthesized by Horizon Discovery (formerly Dharmacon) and Agilent.

[0405] Synthesized crRNA sequence (SEQ ID NO:86):

[0406] mN*mN*NNNNNNNNNNNNNNNNNNGUUUUAGAGCUAUGCUGUUUUG

[0407] 2'OMe(m) and thiophosphate (*) modified residues

[0408] Synthesized 1xMS2 tracrRNA aptamer sequence (SEQ ID NO:87):

[0409] AACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAA AGUGGCACCGAGUCGGUGCGCGCACAUGAGGAUCACCCAUGUGCUUUUmU*mU*U

[0410] Synthetic 2xMS2 tracrRNA aptamer sequence modified with 2'OMe(m) and thiophosphate (*) (SEQ ID NO:88):

[0411] AACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCGGGAGCACAUGAGGAUCACCCAUGUGCCACGAGCGACAUGAGGAUCACCCAUGUCGCUCGUGUUCCCUUUUmU*mU*U

[0412] 2'OMe(m) and thiophosphate (*) modified residues

[0413] Lentiviral sgRNA sequence (SEQ ID NO:89):

[0414] NNNNNNNNNNNNNNNNNNGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCgggagcACAUGAGGAUCACCCAUGUgccacgagcgACAUGAGGAUCACCCAUGUcgcUcgUgUUcccUUUU

[0415] mRNA component generation

[0416] The messenger RNA molecule was custom-designed using Trilinks with modified nucleotides: pseudouridine and 5-methylcytosine. The mRNA components were converted into the following proteins: deaminase AID=NLS-hAID-linker-MCP, deaminase Apobec1=NLS-rApobec1-linker-MCP, and the plasmid Cas9-UGI-UGI=NLS-nCas9-UGI-UGI-NLS was constructed.

[0417] Lentiviral constructs include additional selectable markers (e.g., antibiotics, fluorescent proteins) to ensure a single integrated copy within the genome of the target cell population. A sequence for a specific guide sequence (using T4 DNA ligase technology) is cloned into overhangs generated by IIS-type restriction enzyme sites. The target construct ensures the guide sequence is fully within the frame for efficient transcription from the human U6 PolIII promoter (including G nucleotides if not at the 5' of the sequence) and extends to the Cas9 scaffold and aptamer sequence prior to the termination sequence. Plasmid clones are checked by Sanger sequencing and restriction digestion QC before amplification of large-scale plasmid preparation (e.g., mass production).

[0418] Lentiviral particle generation

[0419] Using a third-generation plasmid system (Horizon Discovery), sgRNA-aptamer lentiviral constructs were prepared in functional lentiviral particles. The viral particles were then concentrated by percolation and aliquoted for transduction.

[0420] Lentiviral transduction

[0421] T cells were activated for >48 hours, transduced in plates treated with Retrotronectin (T100B, Takara-bio) at an MOI of 0.1, and incubated overnight at 37°C and 5% CO2.

[0422] Frozen T cell culture

[0423] The frozen CD3+ T cells (Hemacare) source was thawed and then cultured at 37°C and 5% CO2 in Immunocult XT medium with 1x penicillin / streptomycin (Thermofisher) (STEMCELL Technologies).

[0424] T-cell electroporation

[0425] T cells were electroporated 48–72 hours after activation using a Neon electroporator (Thermofisher) or a 4D nucleating agent (Lonza). Neon electroporator conditions were 1600V / 10ms / 3 pulses, with a 10µl tip and 250k cells, a total mRNA dose of 1–5µg for both deaminase-MCP and nCas9-UGC-UGI, and 0.2–1.8µl of compound crRNA (tracrR or sgRNA) for application. 4D nucleating agent conditions were EO-115 with a 20µl cuvette, 500kJ, a total mRNA dose of 1–5µg for both deaminase-MCP and nCas9-UGC-UGI (synthesized by Trilink), and 0.2–1.8µl of compound crRNA (tracrR or sgRNA) (Horizon Discovery). After electroporation, cells were transferred to Immunocult XT medium containing 100 U IL-2, 100 U IL-7 and 100 U IL-15 (STEMCELL Technologies) and cultured at 37°C and 5% CO2 for 48–72 hours.

[0426] CD3+ T cell activation

[0427] T cells were activated using Dynabeads Human T Activator CD3 / CD28 beads (Thermofisher) cultured in Immunocult XT medium (STEMCELL Technologies) at 37°C and 5% CO2 in the presence of 100 U / ml IL-2 (STEMCELL Technologies) and 1x penicillin / streptomycin (Thermofisher) at a 1:1 bead-to-cell ratio. After activation, the beads were removed by placing the cells on a magnet and transferring them back into the culture.

[0428] Flow cytometry

[0429] T cell identity and QC were confirmed by CD3 antibody staining (Biolegend). T cell activation was confirmed by CD25 staining. Phenotypic genes KO: TRAC was confirmed by CD3 and TCRab antibody staining (Biolegend), and B2M was confirmed by B2M-antibody (Biolegend); any phenotypic data are percentage changes relative to reference material on surviving cells determined solely by DAPI staining (BD Bioscience).

[0430] Genomic DNA Analysis

[0431] Genomic DNA was released from lysed cells 48–72 hours after electroporation. Loci of interest were amplified by PCR, and the products were then sent for Sanger sequencing (Genewiz). Data were analyzed using proprietary in-house software.

[0432] Table 7: Single-guide RNAs (sg RNAs) used for functional knockout base editing of TRAC, TRBC1, TRBC2, PDCD1, CD52, and B2M

[0433]

[0434]

[0435]

[0436]

[0437]

[0438] An exemplary list of guidelines for designing sgRNA and crRNA formats that can be used to create functional knockouts using the illustrated base editing techniques. The list includes specific guidelines for introducing premature stop codons and splicing interruption sites generated by proprietary in-house software.

[0439] Example 5: Base editing efficiency of modified RNA scaffolds in crRNA:tracrRNA and sgRNA

[0440] RNA synthesis

[0441] All chemically synthesized crRNA, tracrRNA, and sgRNA were protected using 2′-acetoxyethyl orthocyanin (2′-ACE) or 2′-tert-butyldimethylsilyl (2′-TBDMS). Chemical modifications, as noted, included two 2′-O-methyl nucleotides and two phosphate thioesters (2xMS modifications) at the 5′ end of crRNA and the 3′ end of tracrRNA. 2′OMe is indicated as m, and phosphate thioesters as *. RNA oligonucleotides were 2′-deprotected / desalted and purified by high-performance liquid chromatography (HPLC) or polyacrylamide gel electrophoresis (PAGE). The oligonucleotides were resuspended in 10 mM Tris pH 7.5 buffer prior to transfection. The gene sites targeted by each cRNA are (A)CR0118_PDCD1, (B)CR0107_PDCD1, (C)CR0057-TRAC_EX3, (D)CR0151_CD2, (E)Site 2, (F)CR0121_PDCD1, and (G)CR0165_CIITA, as follows. Figure 11As shown in AG, A) CR0151_CD2, (B) CR0121_PDCD1 and (C) CR0165_CIITA, as Figure 12 As shown in AC, and (A)TRAC_22550571, (B)PDCD1_241852953 and (C)CTNNB1, as Figure 13 As shown in AC. The sgRNA target site sequences used for base editing are listed in Table 5.

[0442] transfection

[0443] U2OS nCas9 stably transfected cells were transfected with a mixture of 25 nM of synthesized crRNA:tracrRNA and 200 ng of (a)rAPOBEC or (b)hAID mRNA. Cells were harvested after 72 hours.

[0444] Cell treatment

[0445] Cells were lysed at 56°C for 30 minutes in 100 μL of buffer containing proteinase K (Thermo Scientific, #FEREO0492), RNase A (Thermo Scientific, #FEREN0531), and Phusion HF buffer (Thermo Scientific, #F-518L), followed by heat inactivation at 95°C for 5 minutes. This cell lysate was used to generate 200–400 nucleotide PCR amplicons spanning the region containing the base editing site. Unpurified PCR amplicons were sequenced using Sanger sequencing via Genewiz.

[0446] Editor's Analysis

[0447] Base editing efficiency was calculated from AB1 files (an adaptation of the open-source tool BEAT) using the Chimera analysis tool. Chimera determines editing efficiency by first subtracting background noise to define the expected variability in the sample. This allows for estimation of editing efficiency without requiring normalization relative to control samples. Following this, Chimera uses the Median Absolute Deviation (MAD) method to filter out any outliers from the noise, and then evaluates the base editor's editing efficiency across a 20 bp span of the input wizard sequence. Figure 11Comparative editing efficiencies of base editing systems showing the introduction of a single copy of the C-5 or F-5MS2 variant at the 3' end of a tracrRNA. Data are presented for the following crRNAs: (A) CR0118_PDCD1, (B) CR0107_PDCD1, (C) CR0057-TRAC_EX3, (D) CR0151_CD2, (E) Site 2, (F) CR0121_PDCD1, and (G) CR0165_CIITA. The percentage of C-to-T edits detected indicates that the C-5 and F-5 variants provide comparable levels of base editing at all loci studied, and the editing window is equivocal between the two MS2 variants. Figure 12 Comparative editing efficiencies of base editing systems that introduce a single copy of the C-5 or F-5MS2 variant into the 3' end of the tracrRNA used for the following crRNAs: (A) CR0151_CD2, (B) CR0121_PDCD1, and (C) CR0165_CIITA. Figure 13 This indicates the level of base editing with chemically synthesized 1xMS2_3′sgRNAs (C-5) or 1xMS2_3′_7bp-extension_US sgRNAs (C-5) containing 7-base pair extensions of the repeat: anti-repeat upper stem. Figure 14 This indicates that when the amount of MCP-deaminase was reduced to 20 ng, the higher affinity F-5MS2 tracrRNA resulted in a higher percentage of C-to-T editing compared to C-5MS2.

[0448] The synthesized 1xMS2 tracrRNA-aptamer sequence used in Example 5:

[0449] 1x MS2_3′tracrRNA(C-5)SEQ ID NO:87

[0450] AACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCGCGCACAUG A GGAUCACCCAUGUGCUUUUmU*mU*U

[0451] 1x MS2_3′tracrRNA(F-5)SEQ ID NO:185

[0452] AACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCGCGGCCCGG-2AdP-GGAUCACCACGGGCCUUUU mU*mU*U

[0453] 2'OMe is represented as m, and thiophosphate is represented as *.

[0454] Protein sequences (2xUGI) of RNA scaffold-mediated recruitment systems:

[0455] SEQ ID NO:186

[0456]

[0457]

[0458] sequence list <110> Horizon Exploration Ltd. <120> RNA scaffold <130> P33290WO1 <160> 186 <170> PatentIn version 3.5 <210> 1 <211> 1368 <212> PRT <213> Artificial sequence <220> <223> Wild-type Cas9 protein <400> 1 Met Asp Lys Lys Tyr Ser Ile Gly Leu Asp Ile Gly Thr Asn Ser Val 1 5 10 15 Gly Trp Ala Val Ile Thr Asp Glu Tyr Lys Val Pro Ser Lys Lys Phe 20 25 30 Lys Val Leu Gly Asn Thr Asp Arg His Ser Ile Lys Lys Asn Leu Ile 35 40 45 Gly Ala Leu Leu Phe Asp Ser Gly Glu Thr Ala Glu Ala Thr Arg Leu 50 55 60 Lys Arg Thr Ala Arg Arg Tyr Thr Arg Arg Lys Asn Arg With Cys 65 70 75 80 Tyr Leu Gln Glu Ile Phe Ser Asn Glu Met Ala Asp Ser 85 90 95 Phe Phe His Arg Leu Glu Glu Ser Phe Leu Val Glu Glu Asp Lys Lys 100 105 110 His Glu Arg His Pro Ile Phe Gly Asn Ile Val Asp Glu Val Ala Tyr 115 120 125 His Glu Lys Tyr Pro Thr Ile Tyr His Leu Arg Lys Leu Val Asp 130 135 140 Ser Thr Asp Lys Ala Asp Leu Arg Leu Ile Tyr Leu Ala Leu Ala His 145 150 155 160 Met Ile Lys Phe Arg Gly His Phe Leu Ile Glu Gly Asp Leu Asn Pro 165 170 175 Asp Asn Ser Asp Val Asp Lys Leu Phe Ile Gln Leu Val Gln Thr Tyr 180 185 190 Asn Gln Leu Phe Glu Glu Asn Pro Ile Asn Ala Ser Gly Val Asp Ala 195 200 205 Lys Ser Arg Ser Arg Glu Asn 210 215 220 Leu Ile Ala Gln Leu Pro Gly Glu Lys Lys Asn Gly Leu Phe Gly Asn 225 230 235 240 Leu Ile Ala Leu Ser Leu Gly Leu Thr Pro Asn Phe Lys Ser Asn Phe 245 250 255 Asp Leu Ala Glu Asp Ala Lys Leu Gln Leu Ser Lys Asp Thr Tyr Asp 260 265 270 Asp Asp Leu Asp Asn Leu Leu Ala Gln Ile Gly Asp Gln Tyr Ala Asp 275 280 285 Leu Phe Leu Ala Ala Lys Asn Leu Ser Asp Ala Ile Leu Leu Ser Asp 290 295 300 Ile Leu Arg Val Asn Thr Glu Ile Thr Lys Ala Pro Leu Ser Ala Ser 305 310 315 320 Met Ile Lys Arg Tyr Asp Glu His His Gln Asp Leu Thr Leu Leu Lys 325 330 335 Ala Leu Val Arg Gln Gln Leu Pro Glu Lys Tyr Lys Glu Ile Phe Phe 340 345 350 Asp Gln Ser Lys Asn Gly Tyr Ala Gly Tyr Ile Asp Gly Gly Ala Ser 355 360 365 Gln Glu Glu Phe Tyr Lys Phe Ile Lys Pro Ile Leu Glu Lys Met Asp 370 375 380 Gly Thr Glu Glu Leu Leu Val Lys Leu Asn Arg Glu Asp Leu Leu Arg 385 390 395 400 Lys Gln Arg Thr Phe Asp Asn Gly Ser Ile Pro His Gln Ile His Leu 405 410 415 Gly Glu Leu His Ala Ile Leu Arg Arg Gln Glu Asp Phe Tyr Pro Phe 420 425 430 Leu Lys Asp Asn Arg Glu Lys Ile Glu Lys Ile Leu Thr Phe Arg Ile 435 440 445 Pro Tyr Tyr Val Gly Pro Leu Ala Arg Gly Asn Ser Arg Phe Ala Trp 450 455 460 Met Thr Arg Lys Ser Glu Glu Thr Ile Thr Pro Trp Asn Phe Glu Glu 465 470 475 480 Val Val Asp Lys Gly Ala Ser Ala Gln Ser Phe Ile Glu Arg Met Thr 485 490 495 Asn Phe Asp Lys Asn Leu Pro Asn Glu Lys Val Leu Pro Lys His Ser 500 505 510 Leu Leu Tyr Glu Tyr Phe Thr Val Tyr Asn Glu Leu Thr Lys Val Lys 515 520 525 Tyr Val Thr Glu Gly Met Arg Lys Pro Ala Phe Leu Ser Gly Glu Gln 530 535 540 Lys Lys Ala Ile Val Asp Leu Leu Phe Lys Thr Asn Arg Lys Val Thr 545 550 555 560 Val Lys Gln Leu Lys Glu Asp Tyr Phe Lys Lys Ile Glu Cys Phe Asp 565 570 575 Ser Val Glu Ile Ser Gly Val Glu Asp Arg Phe Asn Ala Ser Leu Gly 580 585 590 Thr Tyr His Asp Leu Leu Lys Ile Ile Lys Asp Lys Asp Phe Leu Asp 595 600 605 Asn Glu Glu Asn Glu Asp Ile Leu Glu Asp Ile Val Leu Thr Leu Thr 610 615 620 Leu Phe Glu Asp Arg Glu Met Ile Glu Glu Arg Leu Lys Thr Tyr Ala 625 630 635 640 His Leu Phe Asp Asp Lys Val Met Lys Gln Leu Lys Arg Arg Arg Tyr 645 650 655 Thr Gly Trp Gly Arg Leu Ser Arg Lys Leu Ile Asn Gly Ile Arg Asp 660 665 670 Lys Gln Ser Gly Lys Thr Ile Leu Asp Phe Leu Lys Ser Asp Gly Phe 675 680 685 Ala Asn Arg Asn Phe Met Gln Leu Ile His Asp Asp Ser Leu Thr Phe 690 695 700 Lys Glu Asp Ile Gln Lys Ala Gln Val Ser Gly Gln Gly Asp Ser Leu 705 710 715 720 His Glu His Ile Ala Asn Leu Ala Gly Ser Pro Ala Ile Lys Lys Gly 725 730 735 Ile Leu Gln Thr Val Lys Val Val Asp Glu Leu Val Lys Val Met Gly 740 745 750 Arg His Lys Pro Glu Asn Ile Val Ile Glu Met Ala Arg Glu Asn Gln 755 760 765 Thr Thr Gln Lys Gly Gln Lys Asn Ser Arg Glu Arg Met Lys Arg Ile 770 775 780 Glu Glu Gly Ile Lys Glu Leu Gly Ser Gln Ile Leu Lys Glu His Pro 785 790 795 800 Val Glu Asn Thr Gln Leu Gln Asn Glu Lys Leu Tyr Leu Tyr Tyr Leu 805 810 815 Gln Asn Gly Arg Asp Met Tyr Val Asp Gln Glu Leu Asp Ile Asn Arg 820 825 830 Leu Ser Asp Tyr Asp Val Asp His Ile Val Pro Gln Ser Phe Leu Lys 835 840 845 Asp Asp Ser Ile Asp Asn Lys Val Leu Thr Arg Ser Asp Lys Asn Arg 850 855 860 Gly Lys Ser Asp Asn Val Pro Ser Glu Glu Val Val Lys Lys Met Lys 865 870 875 880 Asn Tyr Trp Arg Gln Leu Leu Asn Ala Lys Leu Ile Thr Gln Arg Lys 885 890 895 Phe Asp Asn Leu Thr Lys Ala Glu Arg Gly Gly Leu Ser Glu Leu Asp 900 905 910 Lys Ala Gly Phe Ile Lys Arg Gln Leu Val Glu Thr Arg Gln Ile Thr 915 920 925 Lys His Val Ala Gln Ile Leu Asp Ser Arg Met Asn Thr Lys Tyr Asp 930 935 940 Glu Asn Asp Lys Leu Ile Arg Glu Val Lys Val Ile Thr Leu Lys Ser 945 950 955 960 Lys Leu Val Ser Asp Phe Arg Lys Asp Phe Gln Phe Tyr Lys Val Arg 965 970 975 Glu Ile Asn Asn Tyr His His Ala His Asp Ala Tyr Leu Asn Ala Val 980 985 990 Val Gly Thr Ala Leu Ile Lys Lys Tyr Pro Lys Leu Glu Ser Glu Phe 995 1000 1005 Val Tyr Gly Asp Tyr Lys Val Tyr Asp Val Arg Lys Met Ile Ala 1010 1015 1020 Lys Ser Glu Gln Glu Ile Gly Lys Ala Thr Ala Lys Tyr Phe Phe 1025 1030 1035 Tyr Ser Asn Ile Met Asn Phe Phe Lys Thr Glu Ile Thr Leu Ala 1040 1045 1050 Asn Gly Glu Ile Arg Lys Arg Pro Leu Ile Glu Thr Asn Gly Glu 1055 1060 1065 Thr Gly Glu Ile Val Trp Asp Lys Gly Arg Asp Phe Ala Thr Val 1070 1075 1080 Arg Lys Val Leu Ser Met Pro Gln Val Asn Ile Val Lys Lys Thr 1085 1090 1095 Glu Val Gln Thr Gly Gly Phe Ser Lys Glu Ser Ile Leu Pro Lys 1100 1105 1110 Arg Asn Ser Asp Lys Leu Ile Ala Arg Lys Lys Asp Trp Asp Pro 1115 1120 1125 Light Light Tyr Gly Gly Phe Asp Ser Pro Thr Val Ala Tyr Ser Val 1130 1135 1140 Leu Val Val Ala Lys Val Glu Lys Gly Lys Ser Lys Lys Leu Lys 1145 1150 1155 Ser Val Lys Glu Leu Leu Gly Ile Thr Ile Met Glu Arg Ser Ser 1160 1165 1170 Phe Glu Lys Asn Pro Ile Asp Phe Leu Glu Ala Lys Gly Tyr Lys 1175 1180 1185 Glu Val Lys Lys Asp Leu Ile Ile Lys Leu Pro Lys Tyr Ser Leu 1190 1195 1200 Phe Glu Leu Glu Asn Gly Arg Lys Arg Met Leu Ala Ser Ala Gly 1205 1210 1215 Glu Leu Gln Lys Gly Asn Glu Leu Ala Leu Pro Ser Lys Tyr Val 1220 1225 1230 Asn Phe Leu Tyr Leu Ala Ser His Tyr Glu Lys Leu Lys Gly Ser 1235 1240 1245 Pro Glu Asp Asn Glu Gln Lys Gln Leu Phe Val Glu Gln His Lys 1250 1255 1260 His Tyr Leu Asp Glu Ile Ile Glu Gln Ile Ser Glu Phe Ser Lys 1265 1270 1275 Arg Val Ile Leu Ala Asp Ala Asn Leu Asp Lys Val Leu Ser Ala 1280 1285 1290 Tyr Asn Lys His Arg Asp Lys Pro Ile Arg Glu Gln Ala Glu Asn 1295 1300 1305 Ile Ile His Leu Phe Thr Leu Thr Asn Leu Gly Ala Pro Ala Ala 1310 1315 1320 Phe Lys Tyr Phe Asp Thr Thr Ile Asp Arg Lys Arg Tyr Thr Ser 1325 1330 1335 Thr Lys Glu Val Leu Asp Ala Thr Leu Ile His Gln Ser Ile Thr 1340 1345 1350 Gly Leu Tyr Glu Thr Arg Ile Asp Leu Ser Gln Leu Gly Gly Asp 1355 1360 1365 <210> 2 <211> 84 <212> PRT <213> Artificial sequence <220> <223> Uracil-DNA glycosylation enzyme inhibitor (UGI) <400> 2 Met Thr Asn Leu Ser Asp Ile Ile Glu Lys Glu Thr Gly Lys Gln Leu 1 5 10 15 Val Ile Gln Glu Ser Ile Leu Met Leu Pro Glu Glu Val Glu Glu Val 20 25 30 Ile Gly Asn Lys Pro Glu Ser Asp Ile Leu Val His Thr Ala Tyr Asp 35 40 45 Glu Ser Thr Asp Glu Asn Val Met Leu Leu Thr Ser Asp Ala Pro Glu 50 55 60 Tyr Lys Pro Trp Ala Leu Val Ile Gln Asp Ser Asn Gly Glu Asn Lys 65 70 75 80 Ile Lys Met Leu <210> 3 <211> 93 <212> DNA <213> Artificial sequence <220> <223> Exemplary mixed crRNA: tracrRNA, gRNA sequence <400> 3 guuuaagagc uaugcuggaa acagcauagc aaguuuaaau aaggcuaguc cguuaucaac 60 uugaaaaagu ggcaccgagu cggugcuuuu uuu 93 <210> 4 <211> 79 <212> RNA <213> Artificial sequence <220> <223> tracrRNA 1 <400> 4 ggaaccauuc aaaacagcau agcaaguuaa aauaaggcua guccguuauc aacuugaaaa 60 aguggcaccg agucggugc 79 <210> 5 <211> 60 <212> RNA <213> Artificial sequence <220> <223> tracrRNA 2 <400> 5 uagcaaguua aaauaaggcu aguccguuau caacuugaaa aaguggcacc gagucggugc 60 <210> 6 <211> 64 <212> RNA <213> Artificial sequence <220> <223> tracrRNA 3 <400> 6 agcauagcaa guuaaaauaa ggcuaguccg uuaucaacuu gaaaaagugg caccgagucg 60 gugc 64 <210> 7 <211> 70 <212> RNA <213> Artificial sequence <220> <223> tracrRNA 4 <400> 7 caaaacagca uagcaaguua aaauaaggcu aguccguuau caacuugaaa aaguggcacc 60 gagucggugc 70 <210> 8 <211> 45 <212> RNA <213> Artificial sequence <220> <223> tracrRNA 5 <400> 8 uagcaaguua aaauaaggcu aguccguuau caacuugaaa aagug 45 <210> 9 <211> 32 <212> RNA <213> Artificial sequence <220> <223> tracrRNA 6 <400> 9 uagcaaguua aaauaaggcu aguccguuau ca 32 <210> 10 <211> 26 <212> RNA <213> Artificial sequence <220> <223> tracrRNA 7 <400> 10 uagcaaguua aaauaaggcu aguccg 26 <210> 11 <211> 66 <212> RNA <213> Artificial sequence <220> <223> Telomerase Ku binding motif <400> 11 uucuugucgu acuuauagau cgcuacguua uuucaauuuu gaaaaucuga gucugggag 60 ugcgga 66 <210> 12 <211> 1094 <212> PRT <213> Artificial sequence <220> <223> Telomerase Ku heterodimer <400> 12 Met Ser Gly Trp Glu Ser Tyr Tyr Lys Thr Glu Gly Asp Glu Glu Ala 1 5 10 15 Glu Glu Glu Gln Glu Glu Asn Leu Glu Ala Ser Gly Asp Tyr Lys Tyr 20 25 30 Ser Gly Arg Asp Ser Leu Ile Phe Leu Val Asp Ala Ser Lys Ala Met 35 40 45 Phe Glu Ser Gln Ser Glu Asp Glu Leu Thr Pro Phe Asp Met Ser Ile 50 55 60 Gln Cys Ile Gln Ser Val Tyr Ile Ser Lys Ile Ile Ser Ser Asp Arg 65 70 75 80 Asp Leu Leu Ala Val Val Phe Tyr Gly Thr Glu Lys Asp Lys Asn Ser 85 90 95 Val Asn Phe Lys Asn Ile Tyr Val Leu Gln Glu Leu Asp Asn Pro Gly 100 105 110 Ala Lys Arg Ile Leu Glu Leu Asp Gln Phe Lys Gly Gln Gln Gly Gln 115 120 125 Lys Arg Phe Gln Asp Met Met Gly His Gly Ser Asp Tyr Ser Leu Ser 130 135 140 Glu Val Leu Trp Val Cys Ala Asn Leu Phe Ser Asp Val Gln Phe Lys 145 150 155 160 Met Ser His Lys Arg Ile Met Leu Phe Thr Asn Glu Asp Asn Pro His 165 170 175 Gly Asn Asp Ser Ala Lys Ala Ser Arg Ala Arg Thr Lys Ala Gly Asp 180 185 190 Leu Arg Asp Thr Gly Ile Phe Leu Asp Leu Met His Leu Lys Lys Pro 195 200 205 Gly Gly Phe Asp Ile Ser Leu Phe Tyr Arg Asp Ile Ile Ser Ile Ala 210 215 220 Glu Asp Glu Asp Leu Arg Val His Phe Glu Glu Ser Ser Lys Leu Glu 225 230 235 240 Asp Leu Leu Arg Lys Val Arg Ala Lys Glu Thr Arg Lys Arg Ala Leu 245 250 255 Ser Arg Leu Lys Leu Lys Leu Asn Lys Asp Ile Val Ile Ser Val Gly 260 265 270 Ile Tyr Asn Leu Val Gln Lys Ala Leu Lys Pro Pro Pro Ile Lys Leu 275 280 285 Tyr Arg Glu Thr Asn Glu Pro Val Lys Thr Lys Thr Arg Thr Phe Asn 290 295 300 Thr Ser Thr Gly Gly Leu Leu Leu Pro Ser Asp Thr Lys Arg Ser Gln 305 310 315 320 Ile Tyr Gly Ser Arg Gln Ile Ile Leu Glu Lys Glu Glu Thr Glu Glu 325 330 335 Leu Lys Arg Phe Asp Asp Pro Gly Leu Met Leu Met Gly Phe Lys Pro 340 345 350 Leu Val Leu Leu Lys Lys His His Tyr Leu Arg Pro Ser Leu Phe Val 355 360 365 Tyr Pro Glu Glu Ser Leu Val Ile Gly Ser Ser Thr Leu Phe Ser Ala 370 375 380 Leu Leu Ile Lys Cys Leu Glu Lys Glu Val Ala Ala Leu Cys Arg Tyr 385 390 395 400 Thr Pro Arg Arg Asn Ile Pro Pro Tyr Phe Val Ala Leu Val Pro Gln 405 410 415 Glu Glu Glu Leu Asp Asp Gln Lys Ile Gln Val Thr Pro Pro Gly Phe 420 425 430 Gln Leu Val Phe Leu Pro Phe Ala Asp Asp Lys Arg Lys Met Pro Phe 435 440 445 Thr Glu Lys Ile Met Ala Thr Pro Glu Gln Val Gly Lys Met Lys Ala 450 455 460 Ile Val Glu Lys Leu Arg Phe Thr Tyr Arg Ser Asp Ser Phe Glu Asn 465 470 475 480 Pro Val Leu Gln Gln His Phe Arg Asn Leu Glu Ala Leu Ala Leu Asp 485 490 495 Leu Met Glu Pro Glu Gln Ala Val Asp Leu Thr Leu Pro Lys Val Glu 500 505 510 Ala Met Asn Lys Arg Leu Gly Ser Leu Val Asp Glu Phe Lys Glu Leu 515 520 525 Val Tyr Pro Pro Asp Tyr Asn Pro Glu Gly Lys Val Thr Lys Arg Lys 530 535 540 His Asp Asn Glu Gly Ser Gly Ser Lys Arg Pro Lys Val Glu Tyr Ser 545 550 555 560 Glu Glu Glu Leu Lys Thr His Ile Ser Lys Gly Thr Leu Gly Lys Phe 565 570 575 Thr Val Pro Met Leu Lys Glu Ala Cys Arg Ala Tyr Gly Leu Lys Ser 580 585 590 Gly Leu Lys Lys Gln Glu Leu Leu Glu Ala Leu Thr Lys His Phe Gln 595 600 605 Asp Met Val Arg Ser Gly Asn Lys Ala Ala Val Val Leu Cys Met Asp 610 615 620 Val Gly Phe Thr Met Ser Asn Ser Ile Pro Gly Ile Glu Ser Pro Phe 625 630 635 640 Glu Gln Ala Lys Lys Val Ile Thr Met Phe Val Gln Arg Gln Val Phe 645 650 655 Ala Glu Asn Lys Asp Glu Ile Ala Leu Val Leu Phe Gly Thr Asp Gly 660 665 670 Thr Asp Asn Pro Leu Ser Gly Gly Asp Gln Tyr Gln Asn Ile Thr Val 675 680 685 His Arg His Leu Met Leu Pro Asp Phe Asp Leu Leu Glu Asp Ile Glu 690 695 700 Ser Lys Ile Gln Pro Gly Ser Gln Gln Ala Asp Phe Leu Asp Ala Leu 705 710 715 720 Ile Val Ser Met Asp Val Ile Gln His Glu Thr Ile Gly Lys Lys Phe 725 730 735 Glu Lys Arg His Ile Glu Ile Phe Thr Asp Leu Ser Ser Arg Phe Ser 740 745 750 Lys Ser Gln Leu Asp Ile Ile Ile His Ser Leu Lys Lys Cys Asp Ile 755 760 765 Ser Glu Arg His Ser Ile His Trp Pro Cys Arg Leu Thr Ile Gly Ser 770 775 780 Asn Leu Ser Ile Arg Ile Ala Ala Tyr Lys Ser Ile Leu Gln Glu Arg 785 790 795 800 Val Lys Lys Thr Trp Thr Val Val Asp Ala Lys Thr Leu Lys Lys Glu 805 810 815 Asp Ile Gln Lys Glu Thr Val Tyr Cys Leu Asn Asp Asp Asp Glu Thr 820 825 830 Glu Val Leu Lys Glu Asp Ile Ile Gln Gly Phe Arg Tyr Gly Ser Asp 835 840 845 Ile Val Pro Phe Ser Lys Val Asp Glu Glu Gln Met Lys Tyr Lys Ser 850 855 860 Glu Gly Lys Cys Phe Ser Val Leu Gly Phe Cys Lys Ser Ser Gln Val 865 870 875 880 Gln Arg Arg Phe Phe Met Gly Asn Gln Val Leu Lys Val Phe Ala Ala 885,890,895 Arg Asp Asp Glu Ala Ala Val Ala Leu Ser Ser Leu Ile His Ala 900 905 910 Leu Asp Asp Leu Asp Met Val Ala Ile Val Arg Tyr Ala Tyr Asp Lys 915,920,925 Arg Ala Asn Pro Gln Val Gly Val Ala Phe Pro His Ile Lys His Asn 930,935,940 Tyr Glu Cys Leu Val Tyr Val Gln Leu Pro Phy Met Glu Asp Leu Arg 945 950 955 960 Gln Tyr Met Phe Ser Ser Leu Lys Asn Ser Lys Lys Tyr Ala Pro Thr 965,970,975 Glu Ala Gln Leu Asn Ala Val Asp Ala Leu Ile Asp Ser Met Ser Leu 980,985,990 Ala Lys Lys Asp Glu Lys Thr Asp Thr Leu Glu Asp Leu Phe Pro Thr 995 1000 1005 Thr Lys Ile Pro Asn Pro Arg Phe Gln Arg Leu Phe Gln Cys Leu 1010 1015 1020 Leu His Arg Ala Leu His Pro Arg Glu Pro Leu Pro Pro Ile Gln 1025 1030 1035 Gln His Ile Trp Asn Met Leu Asn Pro Pro Ala Glu Val Thr Thr 1040 1045 1050 Lys Ser Gln Ile Pro Leu Ser Lys Ile Lys Thr Leu Phe Pro Leu 1055 1060 1065 Ile Glu Ala Lys Lys Lys Asp Gln Val Thr Ala Gln Glu Ile Phe 1070 1075 1080 Gln Asp Asn His Glu Asp Gly Pro Thr Ala Lys 1085 1090 <210> 13 <211> 10 <212> DNA <213> Artificial sequence <220> <223> Telomerase Sm7 consensus site (single strand) <400> 13 aauuuuugga 10 <210> 14 <211> 83 <212> PRT <213> Artificial sequence <220> <223> Monomeric Sm-like protein (archaea) <400> 14 Gly Ser Val Ile Asp Val Ser Ser Gln Arg Val Asn Val Gln Arg Pro 1 5 10 15 Leu Asp Ala Leu Gly Asn Ser Leu Asn Ser Pro Val Ile Ile Lys Leu 20 25 30 Lys Gly Asp Arg Glu Phe Arg Gly Val Leu Lys Ser Phe Asp Leu His 35 40 45 Met Asn Leu Val Leu Asn Asp Ala Glu Glu Leu Glu Asp Gly Glu Val 50 55 60 Thr Arg Arg Leu Gly Thr Val Leu Ile Arg Gly Asp Asn Ile Val Tyr 65 70 75 80 Ile Ser Pro <210> 15 <211> 19 <212> DNA <213> Artificial sequence <220> <223> MS2 phage operator gene stem loop <400> 15 acaugaggau cacccaugu 19 <210> 16 <211> 117 <212> PRT <213> Artificial sequence <220> <223> MS2 capsid protein <400> 16 Met Ala Ser Asn Phe Thr Gln Phe Val Leu Val Asp Asn Gly Gly Thr 1 5 10 15 Gly Asp Val Thr Val Ala Pro Ser Asn Phe Ala Asn Gly Ile Ala Glu 20 25 30 Trp Ile Ser Ser Asn Ser Arg Ser Gln Ala Tyr Lys Val Thr Cys Ser 35 40 45 Val Arg Gln Ser Ser Ala Gln Asn Arg Lys Tyr Thr Ile Lys Val Glu 50 55 60 Val Pro Lys Gly Ala Trp Arg Ser Tyr Leu Asn Met Glu Leu Thr Ile 65 70 75 80 Pro Ile Phe Ala Thr Asn Ser Asp Cys Glu Leu Ile Val Lys Ala Met 85 90 95 Gln Gly Leu Leu Lys Asp Gly Asn Pro Ile Pro Ser Ala Ile Ala Ala 100 105 110 Asn Ser Gly Ile Tyr 115 <210> 17 <211> 26 <212> DNA <213> Artificial sequence <220> <223> PP7 phage operator gene stem loop <400> 17 auaaggaguuuauauggaaacccuua 26 <210> 18 <211> 128 <212> PRT <213> Artificial sequence <220> <223> PP7 capsid protein (PCP) <400> 18 Met Ser Lys Thr Ile Val Leu Ser Val Gly Glu Ala Thr Arg Thr Leu 1 5 10 15 Thr Glu Ile Gln Ser Thr Ala Asp Arg Gln Ile Phe Glu Glu Lys Val 20 25 30 Gly Pro Leu Val Gly Arg Leu Arg Leu Thr Ala Ser Leu Arg Gln Asn 35 40 45 Gly Ala Lys Thr Ala Tyr Arg Val Asn Leu Lys Leu Asp Gln Ala Asp 50 55 60 Val Val Asp Cys Ser Thr Ser Val Cys Gly Glu Leu Pro Lys Val Arg 65 70 75 80 Tyr Thr Gln Val Trp Ser His Asp Val Thr Ile Val Ala Asn Ser Thr 85 90 95 Glu Ala Ser Arg Lys Ser Leu Tyr Asp Leu Thr Lys Ser Leu Val Ala 100 105 110 Thr Ser Gln Val Glu Asp Leu Val Val Asn Leu Val Pro Leu Gly Arg 115 120 125 <210> 19 <211> 19 <212> DNA <213> artificial sequence <220> <223> SfMu Community <400> 19 cugaaugccu gcgagcauc 19 <210> 20 <211> 62 <212> PRT <213> artificial sequence <220> <223> SfMu Com binding protein <400> 20 Met Lys Ser Ile Arg Cys Lys Asn Cys Asn Lys Leu Leu Phe Lys Ala 1 5 10 15 Asp Ser Phe Asp His Ile Glu Ile Arg Cys Pro Arg Cys Lys Arg His 20 25 30 Ile Ile Met Leu Asn Ala Cys Glu His Pro Thr Glu Lys His Cys Gly 35 40 45 Lys Arg Glu Lys Ile Thr His Ser Asp Glu Thr Val Arg Tyr 50 55 60 <210> twenty one <211> 25 <212> RNA <213> Artificial sequence <220> <223> 4nt MS2 extension <220> <221> misc_feature <222> (1)..(2) <223> GC connector <220> <221> misc_feature <222> (3)..(4) <223> extend <220> <221> misc_feature <222> (4)..(23) <223> MS2 aptamer <220> <221> misc_feature <222> (24) (25) <223> extend <400> twenty one gcgcacauga ggaucacccaugugc 25 <210> twenty two <211> 31 <212> RNA <213> Artificial sequence <220> <223> 10nt MS2 extension <220> <221> misc_feature <222> (1)..(2) <223> connector <220> <221> misc_feature <222> (3)..(7) <223> extend <220> <221> misc_feature <222> (8)..(26) <223> MS2 aptamer <220> <221> misc_feature <222> (27) (31) <223> extend <400> twenty two gcgagcgaca ugaggaucac ccaugucgcu c 31 <210> twenty three <211> 37 <212> RNA <213> Artificial sequence <220> <223> 16 nt MS2 extension <220> <221> misc_feature <222> (1)..(2) <223> connector <220> <221> misc_feature <222> (3)..(10) <223> extend <220> <221> misc_feature <222> (11)..(29) <223> MS2 aptamer <220> <221> misc_feature <222> (30) (37) <223> extend <400> twenty three gccacgagcg acaugaggau cacccauguc gcucgug 37 <210> twenty four <211> 47 <212> RNA <213> Artificial sequence <220> <223> 26nt MS2 extension <220> <221> misc_feature <222> (1)..(2) <223> connector <220> <221> misc_feature <222> (3)..(15) <223> connector <220> <221> misc_feature <222> (16) (34) <223> MS2 aptamer <220> <221> misc_feature <222> (35) (47) <223> extend <400> twenty four gccgucagac gagcgacaug aggaucaccc augucgcucg ucugacg 47 <210> 25 <211> 1368 <212> PRT <213> Artificial sequence <220> <223> Streptococcus pyogenes dCas9 protein sequence <220> <221> MISC_FEATURE <222> (10)..(10) <223> Active site mutant D10A <220> <221> MISC_FEATURE <222> (840)..(840) <223> Active site mutant H840A <400> 25 Met Asp Lys Lys Tyr Ser Ile Gly Leu Ala Ile Gly Thr Asn Ser Val 1 5 10 15 Gly Trp Ala Val Ile Thr Asp Glu Tyr Lys Val Pro Ser Lys Lys Phe 20 25 30 Lys Val Leu Gly Asn Thr Asp Arg His Ser Ile Lys Lys Asn Leu Ile 35 40 45 Gly Ala Leu Leu Phe Asp Ser Gly Glu Thr Ala Glu Ala Thr Arg Leu 50 55 60 Lys Arg Thr Ala Arg Arg Arg Tyr Thr Arg Arg Lys Asn Arg Ile Cys 65 70 75 80 Tyr Leu Gln Glu Ile Phe Ser Asn Glu Met Ala Lys Val Asp Asp Ser 85 90 95 Phe Phe His Arg Leu Glu Glu Ser Phe Leu Val Glu Glu Asp Lys Lys 100 105 110 His Glu Arg His Pro Ile Phe Gly Asn Ile Val Asp Glu Val Ala Tyr 115 120 125 His Glu Lys Tyr Pro Thr Ile Tyr His Leu Arg Lys Lys Leu Val Asp 130 135 140 Ser Thr Asp Lys Ala Asp Leu Arg Leu Ile Tyr Leu Ala Leu Ala His 145 150 155 160 Met Ile Lys Phe Arg Gly His Phe Leu Ile Glu Gly Asp Leu Asn Pro 165 170 175 Asp Asn Ser Asp Val Asp Lys Leu Phe Ile Gln Leu Val Gln Thr Tyr 180 185 190 Asn Gln Leu Phe Glu Glu Asn Pro Ile Asn Ala Ser Gly Val Asp Ala 195 200 205 Lys Ala Ile Leu Ser Ala Arg Leu Ser Lys Ser Arg Arg Leu Glu Asn 210 215 220 Leu Ile Ala Gln Leu Pro Gly Glu Lys Lys Asn Gly Leu Phe Gly Asn 225 230 235 240 Leu Ile Ala Leu Ser Leu Gly Leu Thr Pro Asn Phe Lys Ser Asn Phe 245 250 255 Asp Leu Ala Glu Asp Ala Lys Leu Gln Leu Ser Lys Asp Thr Tyr Asp 260 265 270 Asp Asp Leu Asp Asn Leu Leu Ala Gln Ile Gly Asp Gln Tyr Ala Asp 275 280 285 Leu Phe Leu Ala Ala Lys Asn Leu Ser Asp Ala Ile Leu Leu Ser Asp 290 295 300 Ile Leu Arg Val Asn Thr Glu Ile Thr Lys Ala Pro Leu Ser Ala Ser 305 310 315 320 Met Ile Lys Arg Tyr Asp Glu His His Gln Asp Leu Thr Leu Leu Lys 325 330 335 Ala Leu Val Arg Gln Gln Leu Pro Glu Lys Tyr Lys Glu Ile Phe Phe 340 345 350 Asp Gln Ser Lys Asn Gly Tyr Ala Gly Tyr Ile Asp Gly Gly Ala Ser 355 360 365 Gln Glu Glu Phe Tyr Lys Phe Ile Lys Pro Ile Leu Glu Lys Met Asp 370 375 380 Gly Thr Glu Glu Leu Leu Val Lys Leu Asn Arg Glu Asp Leu Leu Arg 385 390 395 400 Lys Gln Arg Thr Phe Asp Asn Gly Ser Ile Pro His Gln Ile His Leu 405 410 415 Gly Glu Leu His Ala Ile Leu Arg Arg Gln Glu Asp Phe Tyr Pro Phe 420 425 430 Leu Lys Asp Asn Arg Glu Lys Ile Glu Lys Ile Leu Thr Phe Arg Ile 435 440 445 Pro Tyr Tyr Val Gly Pro Leu Ala Arg Gly Asn Ser Arg Phe Ala Trp 450 455 460 Met Thr Arg Lys Ser Glu Glu Thr Ile Thr Pro Trp Asn Phe Glu Glu 465 470 475 480 Val Val Asp Lys Gly Ala Ser Ala Gln Ser Phe Ile Glu Arg Met Thr 485 490 495 Asn Phe Asp Lys Asn Leu Pro Asn Glu Lys Val Leu Pro Lys His Ser 500 505 510 Leu Leu Tyr Glu Tyr Phe Thr Val Tyr Asn Glu Leu Thr Lys Val Lys 515 520 525 Tyr Val Thr Glu Gly Met Arg Lys Pro Ala Phe Leu Ser Gly Glu Gln 530 535 540 Lys Lys Ala Ile Val Asp Leu Leu Phe Lys Thr Asn Arg Lys Val Thr 545 550 555 560 Val Lys Gln Leu Lys Glu Asp Tyr Phe Lys Lys Ile Glu Cys Phe Asp 565 570 575 Ser Val Glu Ile Ser Gly Val Glu Asp Arg Phe Asn Ala Ser Leu Gly 580 585 590 Thr Tyr His Asp Leu Leu Lys Ile Ile Lys Asp Lys Asp Phe Leu Asp 595 600 605 Asn Glu Glu Asn Glu Asp Ile Leu Glu Asp Ile Val Leu Thr Leu Thr 610 615 620 Leu Phe Glu Asp Arg Glu Met Ile Glu Glu Arg Leu Lys Thr Tyr Ala 625 630 635 640 His Leu Phe Asp Asp Lys Val Met Lys Gln Leu Lys Arg Arg Arg Tyr 645 650 655 Thr Gly Trp Gly Arg Leu Ser Arg Lys Leu Ile Asn Gly Ile Arg Asp 660 665 670 Lys Gln Ser Gly Lys Thr Ile Leu Asp Phe Leu Lys Ser Asp Gly Phe 675 680 685 Ala Asn Arg Asn Phe Met Gln Leu Ile His Asp Asp Ser Leu Thr Phe 690 695 700 Lys Glu Asp Ile Gln Lys Ala Gln Val Ser Gly Gln Gly Asp Ser Leu 705 710 715 720 His Glu His Ile Ala Asn Leu Ala Gly Ser Pro Ala Ile Lys Lys Gly 725 730 735 Ile Leu Gln Thr Val Lys Val Val Asp Glu Leu Val Lys Val Met Gly 740 745 750 Arg His Lys Pro Glu Asn Ile Val Ile Glu Met Ala Arg Glu Asn Gln 755 760 765 Thr Thr Gln Lys Gly Gln Lys Asn Ser Arg Glu Arg Met Lys Arg Ile 770 775 780 Glu Glu Gly Ile Lys Glu Leu Gly Ser Gln Ile Leu Lys Glu His Pro 785 790 795 800 Val Glu Asn Thr Gln Leu Gln Asn Glu Lys Leu Tyr Leu Tyr Tyr Leu 805 810 815 Gln Asn Gly Arg Asp Met Tyr Val Asp Gln Glu Leu Asp Ile Asn Arg 820 825 830 Leu Ser Asp Tyr Asp Val Asp Ala Ile Val Pro Gln Ser Phe Leu Lys 835 840 845 Asp Asp Ser Ile Asp Asn Lys Val Leu Thr Arg Ser Asp Lys Asn Arg 850 855 860 Gly Lys Ser Asp Asn Val Pro Ser Glu Glu Val Val Lys Lys Met Lys 865 870 875 880 Asn Tyr Trp Arg Gln Leu Leu Asn Ala Lys Leu Ile Thr Gln Arg Lys 885 890 895 Phe Asp Asp With Thr Lys Ala Glu Arg Gly Gly Leu Ser Glu With Asp 900 905 910 Lys Ala Gly Phe Ile Lys Arg Gln Leu Val Glu Thr Arg Gln Ile Thr 915,920,925 Lys Tyr Asp Ser Arg Met Asn Thr Lys Tyr Asp 930,935,940 Glu Asn Asp Lys Leu Ile Arg Glu Val Lys Val Ile Thr Lys Ser 945 950 955 960 Lys Leu Val Ser Asp Phe Arg Lys Asp Phe Gln Phe Tyr Lys Val Arg 965,970,975 Glu Ile Asn Asn Tyr His Ala His Asp Ala Tyr Leu Asn Ala Val 980,985,990 Val Gly Thr Ala Leu Ile Lys Tyr Pro Lys Leu Glu Ser Glu Phe 995 1000 1005 Val Tyr Gly Asp Tyr Lys Val Tyr Asp Val Arg Lys Met Ile Ala 1010 1015 1020 Lys Ser Glu Gln Glu Ile Gly Lys Ala Thr Ala Lys Tyr Phe Phe 1025 1030 1035 Tyr Ser Asn Ile Met Asn Phe Phe Lys Thr Glu Ile Thr Leu Ala 1040 1045 1050 Asn Gly Glu Ile Arg Lys Arg Pro Leu Ile Glu Thr Asn Gly Glu 1055 1060 1065 Thr Gly Glu Ile Val Trp Asp Lys Gly Arg Asp Phe Ala Thr Val 1070 1075 1080 Arg Lys Val Leu Ser Met Pro Gln Val Asn Ile Val Lys Lys Thr 1085 1090 1095 Glu Val Gln Thr Gly Gly Phe Ser Lys Glu Ser Ile Leu Pro Lys 1100 1105 1110 Arg Asn Ser Asp Lys Leu Ile Ala Arg Lys Lys Asp Trp Asp Pro 1115 1120 1125 Light Light Tyr Gly Gly Phe Asp Ser Pro Thr Val Ala Tyr Ser Val 1130 1135 1140 Leu Val Val Ala Lys Val Glu Lys Gly Lys Ser Lys Lys Leu Lys 1145 1150 1155 Ser Val Lys Glu Leu Leu Gly Ile Thr Ile Met Glu Arg Ser Ser 1160 1165 1170 Phe Glu Lys Asn Pro Ile Asp Phe Leu Glu Ala Lys Gly Tyr Lys 1175 1180 1185 Glu Val Lys Lys Asp Leu Ile Ile Lys Leu Pro Lys Tyr Ser Leu 1190 1195 1200 Phe Glu Leu Glu Asn Gly Arg Lys Arg Met Leu Ala Ser Ala Gly 1205 1210 1215 Glu Leu Gln Lys Gly Asn Glu Leu Ala Leu Pro Ser Lys Tyr Val 1220 1225 1230 Asn Phe Leu Tyr Leu Ala Ser His Tyr Glu Lys Leu Lys Gly Ser 1235 1240 1245 Pro Glu Asp Asn Glu Gln Lys Gln Leu Phe Val Glu Gln His Lys 1250 1255 1260 His Tyr Leu Asp Glu Ile Ile Glu Gln Ile Ser Glu Phe Ser Lys 1265 1270 1275 Arg Val Ile Leu Ala Asp Ala Asn Leu Asp Lys Val Leu Ser Ala 1280 1285 1290 Tyr Asn Lys His Arg Asp Lys Pro Ile Arg Glu Gln Ala Glu Asn 1295 1300 1305 Ile Ile His Leu Phe Thr Leu Thr Asn Leu Gly Ala Pro Ala Ala 1310 1315 1320 Phe Lys Tyr Phe Asp Thr Thr Ile Asp Arg Lys Arg Tyr Thr Ser 1325 1330 1335 Thr Lys Glu Val Leu Asp Ala Thr Leu Ile His Gln Ser Ile Thr 1340 1345 1350 Gly Leu Tyr Glu Thr Arg Ile Asp Leu Ser Gln Leu Gly Gly Asp 1355 1360 1365 <210> 26 <211> 1367 <212> PRT <213> Artificial sequence <220> <223> Cas9 D10A protein <220> <221> Mutation <222> (10)..(10) <223> D10A <400> 26 Asp Lys Lys Tyr Ser Ile Gly Leu Ala Ile Gly Thr Asn Ser Val Gly 1 5 10 15 Trp Ala Val Ile Thr Asp Glu Tyr Lys Val Pro Ser Lys Lys Phe Lys 20 25 30 Val Leu Gly Asn Thr Asp Arg His Ser Ile Lys Lys Asn Leu Ile Gly 35 40 45 Ala Leu Leu Phe Asp Ser Gly Glu Thr Ala Glu Ala Thr Arg Leu Lys 50 55 60 Arg Thr Ala Arg Arg Arg Tyr Thr Arg Arg Lys Asn Arg Ile Cys Tyr.3]] 65 70 75 80 Leu Gln Glu Ile Phe Ser Asn Glu Met Ala Lys Val Asp Asp Ser Phe 85 90 95 Phe His Arg Leu Glu Glu Ser Phe Leu Val Glu Glu Asp Lys Lys His 100 105 110 Glu Arg His Pro Ile Phe Gly Asn Ile Val Asp Glu Val Ala Tyr His 115 120 125 Glu Lys Tyr Pro Thr Ile Tyr His Leu Arg Lys Lys Leu Val Asp Ser 130 135 140 Thr Asp Lys Ala Asp Leu Arg Leu Ile Tyr Leu Ala Leu Ala His Met 145 150 155 160 Ile Lys Phe Arg Gly His Phe Leu Ile Glu Gly Asp Leu Asn Pro Asp 165 170 175 Asn Ser Asp Val Asp Lys Leu Phe Ile Gln Leu Val Gln Thr Tyr Asn 180 185 190 Gln Leu Phe Glu Glu Asn Pro Ile Asn Ala Ser Gly Val Asp Ala Lys 195 200 205 Ala Ile Leu Ser Ala Arg Leu Ser Lys Ser Arg Arg Leu Glu Asn Leu 210 215 220 Ile Ala Gln Leu Pro Gly Glu Lys Lys Asn Gly Leu Phe Gly Asn Leu 225 230 235 240 Ile Ala Leu Ser Leu Gly Leu Thr Pro Asn Phe Lys Ser Asn Phe Asp 245 250 255 Leu Ala Glu Asp Ala Lys Leu Gln Leu Ser Lys Asp Thr Tyr Asp Asp 260 265 270 Asp Leu Asp Asn Leu Leu Ala Gln Ile Gly Asp Gln Tyr Ala Asp Leu 275 280 285 Phe Leu Ala Ala Lys Asn Leu Ser Asp Ala Ile Leu Leu Ser Asp Ile 290 295 300 Leu Arg Val Asn Thr Glu Ile Thr Lys Ala Pro Leu Ser Ala Ser Met 305 310 315 320 Ile Lys Arg Tyr Asp Glu His His Gln Asp Leu Thr Leu Leu Lys Ala 325 330 335 Leu Val Arg Gln Gln Leu Pro Glu Lys Tyr Lys Glu Ile Phe Phe Asp 340 345 350 Gln Ser Lys Asn Gly Tyr Ala Gly Tyr Ile Asp Gly Gly Ala Ser Gln 355 360 365 Glu Glu Phe Tyr Lys Phe Ile Lys Pro Ile Leu Glu Lys Met Asp Gly 370 375 380 Thr Glu Glu Leu Leu Val Lys Leu Asn Arg Glu Asp Leu Leu Arg Lys 385 390 395 400 Gln Arg Thr Phe Asp Asn Gly Ser Ile Pro His Gln Ile His Leu Gly 405 410 415 Glu Leu His Ala Ile Leu Arg Arg Gln Glu Asp Phe Tyr Pro Phe Leu 420 425 430 Lys Asp Asn Arg Glu Lys Ile Glu Lys Ile Leu Thr Phe Arg Ile Pro 435 440 445 Tyr Tyr Val Gly Pro Leu Ala Arg Gly Asn Ser Arg Phe Ala Trp Met 450 455 460 Thr Arg Lys Ser Glu Glu Thr Ile Thr Pro Trp Asn Phe Glu Glu Val 465 470 475 480 Val Asp Lys Gly Ser Ala Gln Ser Phe Ile Glu Arg Met Thr Asn 485,490,495 Phe Asp Lys Asn Leu Pro Asn Glu Lys Val Leu Pro Lys His Ser Leu 500 505 510 Leu Tyr Glu Tyr Phe Thr Val Tyr Asn Glu Leu Thr Lys Val Lys Tyr 515,520,525 Val Thr Glu Gly Met Arg Lys Pro Ala Phe Leu Ser Gly Glu Gln Lys 530 535 540 Lys Ala Ile Val Asp Leu Leu Phe Lys Thr Asn Arg Lys Val Thr Val 545 550 555 560 Lys Gln Leu Lys Glu Asp Tyr Phe Lys Ile Glu Cys Phe Asp Ser 565 570 575 Val Glu Ile Ser Gly Val Glu Asp Arg Phe Asn Ala Ser Leu Gly Thr 580 585 590 Tyr His Asp Leu Leu Lys Ile Ile Lys Asp Lys Asp Phe Leu Asp Asn 595 600 605 Glu Glu Asn Glu Asp Ile Leu Glu Asp Ile Val Leu Thr Leu Thr Leu 610 615 620 Phe Glu Asp Arg Glu Met Ile Glu Glu Arg Leu Lys Thr Tyr Ala His 625 630 635 640 Leu Phe Asp Asp Lys Val Met Lys Gln Leu Lys Arg Arg Arg Tyr Thr 645 650 655 Gly Trp Gly Arg Leu Ser Arg Lys Leu Ile Asn Gly Ile Arg Asp Lys 660 665 670 Gln Ser Gly Lys Thr Ile Leu Asp Phe Leu Lys Ser Asp Gly Phe Ala 675 680 685 Asn Arg Asn Phe Met Gln Leu Ile His Asp Asp Ser Leu Thr Phe Lys 690 695 700 Glu Asp Ile Gln Lys Ala Gln Val Ser Gly Gln Gly Asp Ser Leu His 705 710 715 720 Glu His Ile Ala Asn Leu Ala Gly Ser Pro Ala Ile Lys Lys Gly Ile 725 730 735 Leu Gln Thr Val Lys Val Val Asp Glu Leu Val Lys Val Met Gly Arg 740 745 750 His Lys Pro Glu Asn Ile Val Ile Glu Met Ala Arg Glu Asn Gln Thr 755 760 765 Thr Gln Lys Gly Gln Lys Asn Ser Arg Glu Arg Met Lys Arg Ile Glu 770 775 780 Glu Gly Ile Lys Glu Leu Gly Ser Gln Ile Leu Lys Glu His Pro Val 785 790 795 800 Glu Asn Thr Gln Leu Gln Asn Glu Lys Leu Tyr Leu Tyr Tyr Leu Gln 805 810 815 Asn Gly Arg Asp Met Tyr Val Asp Gln Glu Leu Asp Ile Asn Arg Leu 820 825 830 Ser Asp Tyr Asp Val Asp His Ile Val Pro Gln Ser Phe Leu Lys Asp 835 840 845 Asp Ser Ile Asp Asn Lys Val Leu Thr Arg Ser Asp Lys Asn Arg Gly 850 855 860 Lys Ser Asp Asn Val Pro Ser Glu Glu Val Val Lys Lys Met Lys Asn 865 870 875 880 Tyr Trp Arg Gln Leu Asn Ala Lys Leu And Thr Gln Arg Lys Phe 885,890,895 Asp Asn With Thr Lys Ala Glu Arg Gly Gly Ser Glu With Asp Lys 900 905 910 Only Gly Phe With Lys Arg Gln Leu Val Glu Thr Arg Gln With Thr Lys 915,920,925 Only Gln With Asp Ser Arg With Asn Thr Lys Tyr Asp Glu 930,935,940 Asn Asp Lys With Arg Glu Val Val Lys Val Ile With Thr Lys Ser Ser Lys 945 950 955 960 Leu Val Ser Asp Phe Arg Lys Asp Phe Gln Phe Tyr Lys Val Arg Glu 965,970,975 Ile Asn Asn Tyr His Ala His Asp Ala Tyr Leu Asn Ala Val Val 980,985,990 Gly Thr Ala Leu Ile Lys Tyr Pro Lys Leu Glu Ser Glu Phe Val 995 1000 1005 Tyr Gly Asp Tyr Lys Val Tyr Asp Val Arg Lys Met Ile Ala Lys 1010 1015 1020 Ser Glu Gln Glu Ile Gly Lys Ala Thr Ala Lys Tyr Phe Phe Tyr 1025 1030 1035 Ser Asn Ile Met Asn Phe Phe Lys Thr Glu Ile Thr Leu Ala Asn 1040 1045 1050 Gly Glu Ile Arg Lys Arg Pro Leu Ile Glu Thr Asn Gly Glu Thr 1055 1060 1065 Gly Glu Ile Val Trp Asp Lys Gly Arg Asp Phe Ala Thr Val Arg 1070 1075 1080 Lys Val Leu Ser Met Pro Gln Val Asn Ile Val Lys Lys Thr Glu 1085 1090 1095 Val Gln Thr Gly Gly Phe Ser Lys Glu Ser Ile Leu Pro Lys Arg 1100 1105 1110 Asn Ser Asp Lys Leu Ile Ala Arg Lys Lys Asp Trp Asp Pro Lys 1115 1120 1125 Lys Tyr Gly Gly Phe Asp Ser Pro Thr Val Ala Tyr Ser Val Leu 1130 1135 1140 Val Val Ala Lys Val Glu Lys Gly Lys Ser Lys Lys Leu Lys Ser 1145 1150 1155 Val Lys Glu Leu Leu Gly Ile Thr Ile Met Glu Arg Ser Ser Phe 1160 1165 1170 Glu Lys Asn Pro Ile Asp Phe Leu Glu Ala Lys Gly Tyr Lys Glu 1175 1180 1185 Val Lys Lys Asp Leu Ile Ile Lys Pro Lys Lys Ser Leu Phe 1190 1195 1200 Glu Leu Glu Asn Gly Arg Lys Arg Met Leu Ala Ser Ala Gly Glu 1205 1210 1215 Leu Gln Lys Gly Asn Glu Leu Ala Leu Pro Ser Lys Tyr Val Asn 1220 1225 1230 Phe Leu Tyr Leu Ala Ser His Tyr Glu Lys Leu Lys Gly Ser Pro 1235 1240 1245 Glu Asp Asn Glu Gln Lys Gln Leu Phe Val Glu Gln His Lys His 1250 1255 1260 Tyr Leu Asp Glu Ile Ile Glu Gln Ile Glu Phe Serve Lys Arg 1265 1270 1275 There Is A Lion On Asp Only Asn On The Asp Lys On The Lion On Tyr 1280 1285 1290 Asn Lys His Arg Asp Lys Pro With Arg Glu Gln Ala Glu Asn Ile 1295 1300 1305 Ile His Leu Phe Thr Leu Thr Asn Leu Gly Ala Pro Ala Ala Phe 1310 1315 1320 Lys Tyr Phe Asp Thr Thr Ile Asp Arg Lys Arg Tyr Thr Ser Thr 1325 1330 1335 Lys Glu Val Leu Asp Ala Thr Leu Ile His Gln Ser Ile Thr Gly 1340 1345 1350 Leu Tyr Glu Thr Arg Ile Asp Leu Ser Gln Leu Gly Gly Asp 1355 1360 1365 <210> 27 <211> 4104 <212> DNA <213> Artificial sequence <220> <223> DNA encoding Cas9 D10A protein <220> <221> misc_feature <222> (29)..() <223> A to C <220> <221> misc_feature <222> (29)..(29) <223> A to C <400> 27 atggataaaa agtattctat tggtttagcc atcggcacta attccgttgg atgggctgtc 60 ataaccgatg aatacaaagt accttcaaag aaatttaagg tgttggggaa cacagaccgt 120[[ID=�1]] cattcgatta aaaagaatct tatcggtgcc ctcctattcg atagtggcga aacggcagag 180 gcgactcgcc tgaaacgaac cgctcggaga aggtatacac gtcgcaagaa ccgaatatgt 240 tacttacaag aaatttttag caatgagatg gccaaagttg acgattcttt ctttcaccgt 300 360 aacatagtag atgaggtggc atatcatgaa aagtacccaa cgatttatca cctcagaaaa 420 aagctagttg actcaactga taagcggac ctgaggttaa tctacttggc tcttgcccat 480 atgataaagt tccgtgggca ctttctcatt gagggtgatc taaatccgga caactcggat 540 gtcgacaaac tgttcatcca gttagtacaa acctataatc agttgtttga agagaaccct 600 ataaatgcaa gtggcgtgga tgcgaaggct attcttagcg cccgcctctc taaatcccga 660 cggctgaaa acctgatcgc acaattaccc ggagagaaga aaaatgggtt gttcggtaac 720 cttatagcgc tctcactagg cctgacacca aattttaagt cgaacttcga cttagctgaa 780 gatgccaaat tgcagcttag taggcacag tacgatgacg atctcgacaa tctactggca 840 caaattggag atcagtatgc ggacttattt ttggctgcca aaaaccttag cgatgcaatc 900 ctcctatctg acatactgag agttaatact gagattacca aggcgccgtt atccgcttca 960 atgatcaaaa ggtacgatga acatcaccaa gacttgacac ttctcaaggc cctagtccgt 1020 cagcaactgc ctgagaaata taggaaata ttctttgatc agtcgaaaaa cgggtacgca 1080 ggttatattg acggcggagc gagtcaagag gaattctaca agtttatcaa acccatatta 1140 1200 aagcagcgga ctttcgacaa cggtagcatt ccacatcaaa tccacttagg cgaattgcat 1260 1320 gagaaaatcc taacctttcg cataccttac tatgtgggac ccctggcccg agggaactct 1380 cggttcgcat ggatgacaag aaagtccgaa gaacgatta ctccatggaa ttttgagaa 1440 gttgtcgata aaggtgcgtc agctcaatcg ttcatcgaga ggatgaccaa ctttgacaag 1500 1560 tacaatgaac tcacgaaagt tagtatgtc actgagggca tgcgtaaacc cgcctttcta 1620 1680 gttaagcaat tgaaagga ctactttaag aaaattgaat gcttcgattc tgtcgagatc 1740 tccggggtag aagatcgatt taatgcgtca cttggtacgt atcatgacct cctaaagata 1800 Attaagata Aggactccct ggataacgaa Gagaatgaag Atacttaga Agatatagtg 1860 ttgactctta cccctttga agatcggga atgattgagg aaagactaaa aacatacgct 1920 cacctgttcg acgataggt tatgaacag ttaaagaggc gtcgctatac gggctgggga 1980 cgattgtcgc ggaacttat caacgggata agagacaagc aaagtggtaa aactattctc 2040 gattttctaa agagcgacgg cttcgccaat aggacttta tgcagctgat ccatgatgac 2100 tctttaacct tcaagagga tatacaaag gcacaggtttt ccggacagg ggactcattg 2160 cacgaacata ttgcgaatct tgctggttcg ccagccatca aaaagggcat actccagaca 2220 gtcaaagtag tggatgagct agttaaggtc atgggacgtc aaaccgga aaacattgta 2280 atcgagatgg cacgcgaaaa tcaacgact cagaggggc aaaaaaacag tcgagagcgg 2340 atgagagaa tagagagggg tattaaagaa ctggggcagcc agatcttaa ggagcatccct 2400 gtggaaaata cccaattgca gaacgagaaa ctttacctct attackaca aaatggaagg 2460 gacatgtatg ttgatcagga actggacata aaccgtttat ctgattacga cgtcgatcac 2520 attgtacccc aatccttttt gaaggacgat tcaatcgaca ataaagtgct tacacgctcg 2580 gataagaacc gagggaaaag tgacaatgtt ccaagcgagg aagtcgtaaa gaaaatgaag 2640 aactattggc ggcagctcct aaatgcgaaa ctgataacgc aaagaaagtt cgataactta 2700 actaaagctg agaggggtgg cttgtctgaa cttgacaagg ccggatttat taaacgtcag 2760 ctkgtggaaa cccgccaaat cacaaagcat gttgcacaga tactagattc ccgaatgaat 2820 acgaaatacg acgagaacga taagctgatt cgggaagtca aagtaatcac tttaaaagtca 2880 aaattggtgt cggacttcag aaaggatttt caattctata aagttaggga gataaataac 2940 taccaccatg cgcacgacgc ttatcttaat gccgtcgtag ggaccgcact cattaagaaa 3000 tacccgaagc tagaaagtga gtttgtgtat ggtgattaca aagtttatga cgtccgtaag 3060 atgatcgcga aaagcgaaca ggagataggc aaggctacag ccaaatactt cttttattct 3120 aacattatga atttctttaa gacggaaatc actctggcaa acggagagat acgcaaacga 3180 cctttaattg aaaccaatgg ggagacaggt gaaatcgtat gggataaggg ccgggacttc 3240 gcgacggtga gaaaagtttt gtccatgccc caagtcaaca tagtaagaa aactgaggtg 3300 cagaccggag ggttttcaaa ggaatcgatt cttccaaaaa ggaatagtga taagctcatc 3360 gctcgtaaa aggactggga cccgaaaaag tacggtggct tcgatagccc tacagttgcc 3420 tattctgtcc tagtagtggc aaaagttgag aagggaaaat ccaagaaact gaagtcagtc 3480 aaagaattat tggggataac gattatggag cgctcgtctt ttgaaaagaa ccccatcgac 3540 ttccttgagg cgaaaggtta caaggaagta aaaaaggatc tcataattaa actaccaaag 3600 tatagtctgt ttgagttaga aaatggccga aaacggatgt tggctagcgc cggagagctt 3660 caaaagggga acgaactcgc actaccgtct aaatacgtga atttcctgta tttagcgtcc 3720 cattacgaga agttgaaagg ttcacctgaa gataacgaac agaagcaact ttttgttgag 3780 cagcacaaac attatctcga cgaaatcata gagcaaattt cggaattcag taagagagtc 3840 atcctagctg atgccaatct ggacaaagta ttaagcgcat acaacaagca cagggataaa 3900 cccatacgtg agcaggcgga aaatattatc catttgttta ctcttaccaa cctcggcgct 3960 ccagccgcat tcaagtattt tgacacaacg atagatcgca aacgatacac ttctaccaag 4020 gaggtgctag acgcgacact gattcaccaa tccatcacgg gattatatga aactcggata 4080 gatttgtcac agcttggggg tgac 4104 <210> 28 <211> 128 <212> DNA <213> Artificial sequence <220> <223> The NA scaffold expression cassette (Streptococcus pyogenes) contains a 20-nucleotide programmable sequence, a CRISPR RNA motif (tracrRNA), and an MS2 operator gene motif. <220> <221> misc_feature <222> (1)..(20) <223> 20-nucleotide programmable sequence <220> <221> misc_feature <222> (21)..(96) <223> CRISPR RNA motif (tracrRNA) <220> <221> misc_feature <222> (97)..(98) <223> GC connector <220> <221> misc_feature <222> (99) (100) <223> extend <220> <221> misc_feature <222> (101)..(119) <223> MS2 motif <220> <221> misc_feature <222> (120)...(121) <223> extend <220> <221> misc_feature <222> (122) (128) <223> Termination <400> 28 nnnnnnnnnn nnnnnnnnnn gttttagagc tagaaatagc aagttaaaat aaggctagtc 60 cgttatcaac ttgaaaaagt ggcaccgagt cggtgcgcgc acatgaggat cacccatgtg 120 cttttttt 128 <210> 29 <211> 150 <212> DNA <213> Artificial sequence <220> <223> RNA scaffold containing two MS2 loops (2xMS2) <220> <221> misc_feature <222> (83) (101) <223> MS2 bracket <220> <221> misc_feature <222> (112) (130) <223> MS2 bracket <400> 29 gttttagagc tagaaatagc aagttaaaat aaggctagtc cgttatcaac ttgaaaaagt 60 ggcaccgagt cggtgcggga gcacatgagg atcacccatg tgccacgagc gacatgagga 120 tcacccatgt cgctcgtgttcccttttttt 150 <210> 30 <211> 340 <212> PRT <213> Artificial sequence <220> <223> Effector AID–MCP fusion <220> <221> MISC_FEATURE <222> (1)..(198) <223> AID <220> <22l> MISC_FEATURE <222> (199)..(223) <223> MCP <220> <221> MISC_FEATURE <222> (224)..(34()) <223> UGI <400> 30 Met Asp Ser Leu Leu Met Asn Arg Arg Lys Phe Leu Tyr Gln Phe Lys 1 5 l0 15 Asn Val Arg Trp Ala Lys Gly Arg Arg Glu Thr Tyr Leu Cys Tyr Val 20 25 30 Val Lys Arg Arg Asp Ser Ala Thr Ser Phe Ser Leu Asp Phe Gly Tyr 35 40 45 Leu Arg Asn Lys Asn Gly Cys His Val Glu Leu Leu Phe Leu Arg Tyr 50 55 60 Ile Ser Asp Trp Asp Leu Asp Pro Gly Arg Cys Tyr Arg Val Thr Trp 65 70 75 80 Phe Thr Ser Trp Ser Pro Cys Tyr Asp Cys Ala Arg His Val Ala Asp 85 90 95 Phe Leu Arg Gly Asn Pro Asn Leu Ser Leu Arg Ile Phe Thr Ala Arg 100 105 110 Leu Tyr Phe Cys Glu Asp Arg Lys Ala Glu Pro Glu Gly Leu Arg Arg 115 120 125 Leu His Arg Ala Gly Val Gln Ile Ala Ile Met Thr Phe Lys Asp Tyr 130 135 140 Phe Tyr Cys Trp Asn Thr Phe Val Glu Asn His Glu Arg Thr Phe Lys 145 150 155 160 Ala Trp Glu Gly Leu His Glu Asn Ser Val Arg Leu Ser Arg Gln Leu 165 170 175 Arg Arg Ile Leu Leu Pro Leu Tyr Glu Val Asp Asp Leu Arg Asp Ala 180 185 190 Phe Arg Thr Leu Gly Leu Glu Leu Lys Thr Pro Leu Gly Asp Thr Thr 195 200 205 His Thr Ser Pro Pro Cys Pro Ala Pro Glu Leu Leu Gly Gly Pro Met 210 215 220 Ala Ser Asn Phe Thr Gln Phe Val Leu Val Asp Asn Gly Gly Thr Gly 225 230 235 240 Asp Val Thr Val Ala Pro Ser Asn Phe Ala Asn Gly Ile Ala Glu Trp 245 250 255 Ile Ser Ser Asn Ser Arg Ser Gln Ala Tyr Lys Val Thr Cys Ser Val 260 265 270 Arg Gln Ser Ser Ala Gln Asn Arg Lys Tyr Thr Ile Lys Val Glu Val 275 280 285 Pro Lys Gly Ala Trp Arg Ser Tyr Leu Asn Met Glu Leu Thr Ile Pro 290 295 300 Ile Phe Ala Thr Asn Ser Asp Cys Glu Leu Ile Val Lys Ala Met Gln 305 310 315 320 Gly Leu Leu Lys Asp Gly Asn Pro Ile Pro Ser Ala Ile Ala Ala Asn 325 330 335 Ser Gly Ile Tyr 340 <210> 31[[ID=3〕> <211> 594 <212> DNA <213> Artificial Sequence <220> <223> wtAID cDNA <220> <221> misc_feature <222> (112)..(114) <223> Ser38 Codon <400> 31 atggacagcc tcttgatgaa ccggaggaag tttctttacc aattcaaaaa tgtccgctgg 60 gctaagggtc ggcgtgagac ctacctgtgc tacgtagtga agaggcgtga cagtgctaca 120 tccttttcac tggactttgg ttatcttcgc aataagaacg gctgccacgt ggaattgctc 180 ttcctccgct acatctcgga ctgggaccta gaccctggcc gctgctaccg cgtcacctgg 240 ttcacctcct ggagcccctg ctacgactgt gcccgacatg tggccgactt tctgcgaggg 300 aaccccaacc tcagtctgag gatcttcacc gcgcgcctct acttctgtga ggaccgcaag 360 gctgagcccg aggggctgcg gcggctgcac cgcgccgggg tgcaaatagc catcatgacc 420 ttcaaagatt atttttactg ctggaatact tttgtagaaa accatgaaag aactttcaaa 480 gcctgggaag ggctgcatga aaattcagtt cgtctctcca gacagcttcg gcgcatcctt 540 ttgcccctgt atgaggttga tgacttacga gacgcatttc gtactttggg actt 594 <210> 32 <211> 198 <212> PRT <213> Artificial Sequence <220> <223> wtAID protein <220> <221> MISC_FEATURE <222> (38)..(38) <223> Ser38 <400> 32 Met Asp Ser Leu Leu Met Asn Arg Arg Lys Phe Leu Tyr Gln Phe Lys 1 5 10 15 Asn Val Arg Trp Ala Lys Gly Arg Arg Glu Thr Tyr Leu Cys Tyr Val 20 25 30 Val Lys Arg Arg Asp Ser Ala Thr Ser Phe Ser Leu Asp Phe Gly Tyr 35 40 45 Arg Asn Lys Asn Gly Cys His Val Glu Arg Phe Arg Tyr 50 55 60 Serving Asp Trp Asp Leu Asp Pro Gly Arg Cys Tyr Arg Val Thr Trp 65 70 75 80 Phe Thr Ser Trp Ser Pro Cys Tyr Asp Cys Ala Arg His Val Ala Asp 85 90 95 Phe Leu Arg Gly Asn Pro Asn Leu Ser Leu Arg Ile Phe Thr Ala Arg 100 105 110 Leu Tyr Phe Cys Glu Asp Arg Lys Ala Glu Pro Glu Gly Leu Arg Arg 115 120 125 Leu His Arg Ala Gly Val Gln Ile Ala Ile Met Thr Phe Lys Asp Tyr 130 135 140 Phe Tyr Cys Trp Asn Thr Phe Val Glu Asn His Glu Arg Thr Phe Lys 145 150 155 160 Ala Trp Glu Gly Leu His Glu Asn Ser Val Arg Leu Ser Arg Gln Leu 165 170 175 Arg Arg Ile Leu Leu Pro Leu Tyr Glu Val Asp Asp Leu Arg Asp Ala 180 185 190 Phe Arg Thr Leu Gly Leu 195 <210> 33 <211> 594 <212> DNA <213> Artificial Sequence <220> <223> AID_S38A cDNA <220> <221> misc_feature <222> (112)..(114) [[ID=X]]<223> S38A Mutation <400> 33 atggacagcc tcttgatgaa ccggaggaag tttctttacc aattcaaaaa tgtccgctgg 60 gctaagggtc ggcgtgagac ctacctgtgc tacgtagtga agaggcgtga cgccgctaca 120 tccttttcac tggactttgg ttatcttcgc aataagaacg gctgccacgt ggaattgctc 180 ttcctccgct acatctcgga ctgggaccta gaccctggcc gctgctaccg cgtcacctgg 240 ttcacctcct ggagcccctg ctacgactgt gcccgacatg tggccgactt tctgcgaggg 300 aaccccaacc tcagtctgag gatcttcacc gcgcgcctct acttctgtga ggaccgcaag 360 Note: In the translation, for item , since the original Chinese is "S38A突变", the more accurate English expression in the context of patent text should be "S38A Mutation", so I adjusted it in the translation. If this is not allowed according to strict rules, please let me know and I will restore it to the original form. gctgagcccg aggggctgcg gcggctgcac cgcgccgggg tgcaaatagc catcatgacc 420 ttcaaagatt atttttactg ctggaatact tttgtagaaa accatgaaag aactttcaaa 480 gcctgggaag ggctgcatga aaattcagtt cgtctctcca gacagcttcg gcgcatcctt 540 ttgcccctgt atgaggttga tgacttacga gacgcatttc gtactttggg actt 594 <210> 34 <211> 198 <212> PRT <213> Artificial Sequence <220> <223> AID_S38A protein <220> <221> MISC_FEATURE <222> (38)..(38) <223> S38A mutation <400> 34 Met Asp Ser Leu Leu Met Asn Arg Arg Lys Phe Leu Tyr Gln Phe Lys 1 5 10 15 Asn Val Arg Trp Ala Lys Gly Arg Arg Glu Thr Tyr Leu Cys Tyr Val 20 25 30 Val Lys Arg Arg Asp Ala Ala Thr Ser Phe Ser Leu Asp Phe Gly Tyr 35 40 45 Leu Arg Asn Lys Asn Gly Cys His Val Glu Leu Leu Phe Leu Arg Tyr 50 55 60 Ile Ser Asp Trp Asp Leu Asp Pro Gly Arg Cys Tyr Arg Val Thr Trp 65 70 75 80 Phe Thr Ser Trp Ser Pro Cys Tyr Asp Cys Ala Arg His Val Ala Asp 85 90 95 Phe Leu Arg Gly Asn Pro Asn Leu Ser Leu Arg Ile Phe Thr Ala Arg 100 105 110 Leu Tyr Phe Cys Glu Asp Arg Lys Ala Glu Pro Glu Gly Leu Arg Arg 115 120 125 Leu His Arg Ala Gly Val Gln Ile Ala Ile Met Thr Phe Lys Asp Tyr 130 135 140 Phe Tyr Cys Trp Asn Thr Phe Val Glu Asn His Glu Arg Thr Phe Lys 145 150 155 160 Ala Trp Glu Gly Leu His Glu Asn Ser Val Arg Leu Ser Arg Gln Leu 165 170 175 Arg Arg Ile Leu Leu Pro Leu Tyr Glu Val Asp Asp Leu Arg Asp Ala 180 185 190 Phe Arg Thr Leu Gly Leu 195 <210> 35 <211> 1835 <212> PRT <213> Artificial sequence <220> <223> Protein sequence of the ARNA scaffold-mediated recruitment system nu construct <220> <221> MISC_FEATURE <222> (1)..(7) <223> Nuclear Localization Signal (NLS) <220> <221> MISC_FEATURE <222> (8)..(205) <223> AID <220> <221> MISC_FEATURE <222> (206) (230) <223> connector <220> <221> MISC_FEATURE <222> (231)..(347) <223> MCP <220> <221> MISC_FEATURE <222> (351) (368) <223> T2A peptide <220> <221> MISC_FEATURE <222> (371)..(1737) <223> nCAS9D10A <220> <221> MISC_FEATURE <222> (1742) (1824) <223> UGI <220> <221> MISC_FEATURE <222> (1829) ... (1835) <223> Nuclear Localization Signal (NLS) <400> 35 Pro Lys Lys Lys Arg Lys Val Met Asp Ser Leu Leu Met Asn Arg Arg 1 5 10 15 Lys Phe Leu Tyr Gln Phe Lys Asn Val Arg Trp Ala Lys Gly Arg Arg 20 25 30 Glu Thr Tyr Leu Cys Tyr Val Val Lys Arg Arg Asp Ser Ala Thr Ser 35 40 45 Phe Ser Leu Asp Phe Gly Tyr Leu Arg Asn Lys Asn Gly Cys His Val 50 55 60 Glu Leu Leu Phe Leu Arg Tyr Ile Ser Asp Trp Asp Leu Asp Pro Gly 65 70 75 80 Arg Cys Tyr Arg Val Thr Trp Phe Thr Ser Trp Ser Pro Cys Tyr Asp 85 90 95 Cys Ala Arg His Val Ala Asp Phe Leu Arg Gly Asn Pro Asn Leu Ser 100 105 110 Leu Arg Ile Phe Thr Ala Arg Leu Tyr Phe Cys Glu Asp Arg Lys Ala 115 120 125 Glu Pro Glu Gly Leu Arg Arg Leu His Arg Ala Gly Val Gln Ile Ala 130 135 140 Ile Met Thr Phe Lys Asp Tyr Phe Tyr Cys Trp Asn Thr Phe Val Glu 145 150 155 160 Asn His Glu Arg Thr Phe Lys Ala Trp Glu Gly Leu His Glu Asn Ser 165 170 175 Val Arg Leu Ser Arg Gln Leu Arg Arg Ile Leu Leu Pro Leu Tyr Glu 180 185 190 Val Asp Asp Leu Arg Asp Ala Phe Arg Thr Leu Gly Leu Glu Leu Lys 195 200 205 Thr Pro Leu Gly Asp Thr Thr His Thr Ser Pro Pro Cys Pro Ala Pro 210 215 220 Glu Leu Leu Gly Gly Pro Met Ala Ser Asn Phe Thr Gln Phe Val Leu 225 230 235 240 Val Asp Asn Gly Gly Thr Gly Asp Val Thr Val Ala Pro Ser Asn Phe 245 250 255 Ala Asn Gly Ile Ala Glu Trp Ile Ser Ser Asn Ser Arg Ser Gln Ala 260 265 270 Tyr Lys Val Thr Cys Ser Val Arg Gln Ser Ser Ala Gln Asn Arg Lys 275 280 285 Tyr Thr Ile Lys Val Glu Val Pro Lys Gly Ala Trp Arg Ser Tyr Leu 290 295 300 Asn Met Glu Leu Thr Ile Pro Ile Phe Ala Thr Asn Ser Asp Cys Glu 305 310 315 320 Leu Ile Val Lys Ala Met Gln Gly Leu Leu Lys Asp Gly Asn Pro Ile 325 330 335 Pro Ser Ala Ile Ala Ala Asn Ser Gly Ile Tyr Gly Ser Gly Glu Gly 340 345 350 Arg Gly Ser Leu Leu Thr Cys Gly Asp Val Glu Glu Asn Pro Gly Pro 355 360 365 Gly Thr Asp Lys Lys Tyr Ser Ile Gly Leu Ala Ile Gly Thr Asn Ser 370 375 380 Val Gly Trp Ala Val Ile Thr Asp Glu Tyr Lys Val Pro Ser Lys Lys 385 390 395 400 Phe Lys Val Leu Gly Asn Thr Asp Arg His Ser Ile Lys Lys Asn Leu 405 410 415 Ile Gly Ala Leu Leu Phe Asp Ser Gly Glu Thr Ala Glu Ala Thr Arg 420 425 430 Leu Lys Arg Thr Ala Arg Arg Arg Tyr Thr Arg Arg Lys Asn Arg Ile 435 440 445 Cys Tyr Leu Gln Glu Ile Phe Ser Asn Glu Met Ala Lys Val Asp Asp 450 455 460 Ser Phe Phe His Arg Leu Glu Glu Ser Phe Leu Val Glu Glu Asp Lys 465 470 475 480 Lys His Glu Arg His Pro Ile Phe Gly Asn Ile Val Asp Glu Val Ala 485 490 495 Tyr His Glu Lys Tyr Pro Thr Ile Tyr His Leu Arg Lys Lys Leu Val 500 505 510 Asp Ser Thr Asp Lys Ala Asp Leu Arg Leu Ile Tyr Leu Ala Leu Ala 515 520 525 His Met Ile Lys Phe Arg Gly His Phe Leu Ile Glu Gly Asp Leu Asn 530 535 540 Pro Asp Asn Ser Asp Val Asp Lys Leu Phe Ile Gln Leu Val Gln Thr 545 550 555 560 Tyr Asn Gln Leu Phe Glu Glu Asn Pro Ile Asn Ala Ser Gly Val Asp 565 570 575 Ala Lys Ala Ile Leu Ser Ala Arg Leu Ser Lys Ser Arg Arg Leu Glu 580 585 590 Asn Leu Ile Ala Gln Leu Pro Gly Glu Lys Lys Asn Gly Leu Phe Gly 595 600 605 Asn Leu Ile Ala Leu Ser Leu Gly Leu Thr Pro Asn Phe Lys Ser Asn 610 615 620 Phe Asp Leu Ala Glu Asp Ala Lys Leu Gln Leu Ser Lys Asp Thr Tyr 625 630 635 640 Asp Asp Asp Leu Asp Asn Leu Leu Ala Gln Ile Gly Asp Gln Tyr Ala 645 650 655 Asp Leu Phe Leu Ala Ala Lys Asn Leu Ser Asp Ala Ile Leu Leu Ser 660 665 670 Asp Ile Leu Arg Val Asn Thr Glu Ile Thr Lys Ala Pro Leu Ser Ala 675 680 685 Ser Met Ile Lys Arg Tyr Asp Glu His His Gln Asp Leu Thr Leu Leu 690 695 700 Lys Ala Leu Val Arg Gln Gln Leu Pro Glu Lys Tyr Lys Glu Ile Phe 705 710 715 720 Phe Asp Gln Ser Lys Asn Gly Tyr Ala Gly Tyr Ile Asp Gly Gly Ala 725 730 735 Ser Gln Glu Glu Phe Tyr Lys Phe Ile Lys Pro Ile Leu Glu Lys Met 740 745 750 Asp Gly Thr Glu Glu Leu Leu Val Lys Leu Asn Arg Glu Asp Leu Leu 755 760 765 Arg Lys Gln Arg Thr Phe Asp Asn Gly Ser Ile Pro His Gln Ile His 770 775 780 Leu Gly Glu Leu His Ala Ile Leu Arg Arg Gln Glu Asp Phe Tyr Pro 785 790 795 800 Phe Leu Lys Asp Asn Arg Glu Lys Ile Glu Lys Ile Leu Thr Phe Arg 805 810 815 Ile Pro Tyr Tyr Val Gly Pro Leu Ala Arg Gly Asn Ser Arg Phe Ala 820 825 830 Trp Met Thr Arg Lys Ser Glu Glu Thr Ile Thr Pro Trp Asn Phe Glu 835 840 845 Glu Val Val Asp Lys Gly Ala Ser Ala Gln Ser Phe Ile Glu Arg Met 850 855 860 Thr Asn Phe Asp Lys Asn Leu Pro Asn Glu Lys Val Leu Pro Lys His 865 870 875 880 Ser Leu Leu Tyr Glu Tyr Phe Thr Val Tyr Asn Glu Leu Thr Lys Val 885 890 895 Lys Tyr Val Thr Glu Gly Met Arg Lys Pro Ala Phe Leu Ser Gly Glu 900 905 910 Gln Lys Lys Ala Ile Val Asp Leu Leu Phe Lys Thr Asn Arg Lys Val 915 920 925 Thr Val Lys Gln Leu Lys Glu Asp Tyr Phe Lys Lys Ile Glu Cys Phe 930 935 940 Asp Ser Val Glu Ile Ser Gly Val Glu Asp Arg Phe Asn Ala Ser Leu 945 950 955 960 Gly Thr Tyr His Asp Leu Leu Lys Ile Ile Lys Asp Lys Asp Phe Leu 965 970 975 Asp Asn Glu Glu Asn Glu Asp Ile Leu Glu Asp Ile Val Leu Thr Leu 980 985 990 Thr Leu Phe Glu Asp Arg Glu Met Ile Glu Glu Arg Leu Lys Thr Tyr 995 1000 1005 Ala His Leu Phe Asp Asp Lys Val Met Lys Gln Leu Lys Arg Arg 1010 1015 1020 Arg Tyr Thr Gly Trp Gly Arg Leu Ser Arg Lys Leu Ile Asn Gly 1025 1030 1035 Ile Arg Asp Lys Gln Ser Gly Lys Thr Ile Leu Asp Phe Leu Lys 1040 1045 1050 Ser Asp Gly Phe Ala Asn Arg Asn Phe Met Gln Leu Ile His Asp 1055 1060 1065 Asp Ser Leu Thr Phe Lys Glu Asp Ile Gln Lys Ala Gln Val Ser 1070 1075 1080 Gly Gln Gly Asp Ser Leu His Glu His Ile Ala Asn Leu Ala Gly 1085 1090 1095 Ser Pro Ala Ile Lys Lys Gly Ile Leu Gln Thr Val Lys Val Val 1100 1105 1110 Asp Glu Leu Val Lys Val Met Gly Arg His Lys Pro Glu Asn Ile 1115 1120 1125 Val Ile Glu Met Ala Arg Glu Asn Gln Thr Thr Gln Lys Gly Gln 1130 1135 1140 Lys Asn Ser Arg Glu Arg Met Lys Arg Ile Glu Glu Gly Ile Lys 1145 1150 1155 Glu Leu Gly Ser Gln Ile Leu Lys Glu His Pro Val Glu Asn Thr 1160 1165 1170 Gln Leu Gln Asn Glu Lys Leu Tyr Leu Tyr Tyr Leu Gln Asn Gly 1175 1180 1185 Arg Asp Met Tyr Val Asp Gln Glu Leu Asp Ile Asn Arg Leu Ser 1190 1195 1200 Asp Tyr Asp Val Asp His Ile Val Pro Gln Ser Phe Leu Lys Asp 1205 1210 1215 Asp Ser Ile Asp Asn Lys Val Leu Thr Arg Ser Asp Lys Asn Arg 1220 1225 1230 Gly Lys Ser Asp Asn Val Pro Ser Glu Glu Val Val Lys Lys Met 1235 1240 1245 Lys Asn Tyr Trp Arg Gln Leu Leu Asn Ala Lys Leu Ile Thr Gln 1250 1255 1260 Arg Lys Phe Asp Asn Leu Thr Lys Ala Glu Arg Gly Gly Leu Ser 1265 1270 1275 Glu Leu Asp Lys Ala Gly Phe Ile Lys Arg Gln Leu Val Glu Thr 1280 1285 1290 Arg Gln With Three Lys His Val Ala Gln With Asp Ser Arg Met 1295 1300 1305 Asn Thr Lys Tyr Asp Glu Asn Asp Lys With Arg Glu Val Lys 1310 1315 1320 There Will Be Thr Lys Ser Lys Leu Val Ser Asp Phe Arg Lys Asp 1325 1330 1335 Phe Gln Phe Tyr Lys Val Arg Glu Ile Asn Asn Tyr His Ala 1340 1345 1350 His Asp Ala Tyr Leu Asn Ala Val Val Gly Thr Ala Leu Ile Lys 1355 1360 1365 Lys Tyr Pro Lys Leu Glu Ser Glu Phe Val Tyr Gly Asp Tyr Lys 1370 1375 1380 Val Tyr Asp Val Arg Lys Met Ile Ala Lys Ser Glu Gln Glu Ile 1385 1390 1395 Gly Lys Ala Thr Ala Lys Tyr Phe Phe Tyr Ser Asn Ile Met Asn 1400 1405 1410 Phe Phe Lys Thr Glu Ile Thr Leu Ala Asn Gly Glu Ile Arg Lys 1415 1420 1425 Arg Pro Leu Ile Glu Thr Asn Gly Glu Thr Gly Glu Ile Val Trp 1430 1435 1440 Asp Lys Gly Arg Asp Phe Ala Thr Val Arg Lys Val Leu Ser Met 1445 1450 1455 Pro Gln Val Asn Ile Val Lys Lys Thr Glu Val Gln Thr Gly Gly 1460 1465 1470 Phe Ser Lys Glu Ser Ile Leu Pro Lys Arg Asn Ser Asp Lys Leu 1475 1480 1485 Ile Ala Arg Lys Lys Asp Trp Asp Pro Lys Lys Tyr Gly Gly Phe 1490 1495 1500 Asp Ser Pro Thr Val Ala Tyr Ser Val Leu Val Val Ala Lys Val 1505 1510 1515 Glu Lys Gly Lys Ser Lys Lys Leu Lys Ser Val Lys Glu Leu Leu 1520 1525 1530 Gly Ile Thr Ile Met Glu Arg Ser Ser Phe Glu Lys Asn Pro Ile 1535 1540 1545 Asp Phe Leu Glu Ala Lys Gly Tyr Lys Glu Val Lys Lys Asp Leu 1550 1555 1560 Ile Ile Lys Leu Pro Lys Tyr Ser Leu Phe Glu Leu Glu Asn Gly 1565 1570 1575 Arg Lys Arg Met Leu Ala Ser Ala Gly Glu Leu Gln Lys Gly Asn 1580 1585 1590 Glu Leu Ala Leu Pro Ser Lys Tyr Val Asn Phe Leu Tyr Leu Ala 1595 1600 1605 Ser His Tyr Glu Lys Leu Lys Gly Ser Pro Glu Asp Asn Glu Gln 1610 1615 1620 Lys Gln Leu Phe Val Glu Gln His Lys His Tyr Leu Asp Glu Ile 1625 1630 1635 Ile Glu Gln Ile Ser Glu Phe Ser Lys Arg Val Ile Leu Ala Asp 1640 1645 1650 Ala Asn Leu Asp Lys Val Leu Ser Ala Tyr Asn Lys His Arg Asp 1655 1660 1665 Lys Pro Ile Arg Glu Gln Ala Glu Asn Ile Ile His Leu Phe Thr 1670 1675 1680 Leu Thr Asn Leu Gly Ala Pro Ala Ala Phe Lys Tyr Phe Asp Thr 1685 1690 1695 Thr Ile Asp Arg Lys Arg Tyr Thr Ser Thr Lys Glu Val Leu Asp 1700 1705 1710 Ala Thr Leu Ile His Gln Ser Ile Thr Gly Leu Tyr Glu Thr Arg 1715 1720 1725 Ile Asp Leu Ser Gln Leu Gly Gly Asp Ser Gly Gly Ser Thr Asn 1730 1735 1740 Leu Ser Asp Ile Ile Glu Lys Glu Thr Gly Lys Gln Leu Val Ile 1745 1750 1755 Gln Glu Ser Ile Leu Met Leu Pro Glu Glu Val Glu Glu Val Ile 1760 1765 1770 Gly Asn Lys Pro Glu Ser Asp Ile Leu Val His Thr Ala Tyr Asp 1775 1780 1785 Glu Ser Thr Asp Glu Asn Val Met Leu Leu Thr Ser Asp Ala Pro 1790 1795 1800 Glu Tyr Lys Pro Trp Ala Leu Val Ile Gln Asp Ser Asn Gly Glu 1805 1810 1815 Asn Lys Ile Lys Met Leu Ser Gly Gly Ser Pro Lys Lys Lys Arg 1820 1825 1830 Lys Val 1835 <210> 36 <211> 1941 <212> PRT <213> Artificial sequence <220> <223> Protein sequence of the ARNA scaffold-mediated recruitment system nu.2 construct <220> <221> MISC_FEATURE <222> (1)..(7) <223> Nuclear Localization Signal (NLS) <220> <221> MISC_FEATURE <222> (8)..(205) <223> AID <220> <221> MISC_FEATURE <222> (206) (230) <223> connector <220> <221> MISC_FEATURE <222> (231)..(347) <223> MCP <220> <221> MISC_FEATURE <222> (351) (368) <223> T2A peptide <220> <221> MISC_FEATURE <222> (371) (377) <223> Nuclear Localization Signal (NLS) <220> <221> MISC_FEATURE <222> (378) (1744) <223> nCAS9D10A <220> <221> MISC_FEATURE <222> (1755) (1837) <223> UGI <220> <221> MISC_FEATURE <222> (1848) (1930) <223> UGI <220> <221> MISC_FEATURE <222> (1935) ... (1941) <223> National League for Democracy (NLS) <400> 36 Pro Lys Lys Lys Arg Lys Val Met Asp Ser Leu Leu Met Asn Arg Arg 1 5 10 15 Lys Phe Leu Tyr Gln Phe Lys Asn Val Arg Trp Ala Lys Gly Arg Arg 20 25 30 Glu Thr Tyr Leu Cys Tyr Val Val Lys Arg Arg Ser Ala Thr Ser 35 40 45 Phe Ser Leu Asp Phe Gly Tyr Leu Arg Asn Lys Asn Gly Cys His Val 50 55 60 Glu Leu Phe Leu Arg Tyr Ile Ser Asp Trp Asp Leu Asp Pro Gly 65 70 75 80 Arg Cys Tyr Arg Val Thr Trp Phe Thr Ser Trp Ser Pro Cys Tyr Asp 85 90 95 Cys Ala Arg His Val Ala Asp Phe Leu Arg Gly Asn Pro Asn Leu Ser 100 105 110 Leu Arg Ile Phe Thr Ala Arg Leu Tyr Phe Cys Glu Asp Arg Lys Ala 115 120 125 Glu Pro Glu Gly Leu Arg Arg Leu His Arg Ala Gly Val Gln Ile Ala 130 135 140 With Thr Phe Lys Asp Tyr Phe Tyr Cys Trp Asn Thr Phe Val Glu 145 150 155 160 Asn His Glu Arg Thr Phe Lys Ala Trp Glu Gly Leu His Glu Asn Ser 165 170 175 Val Arg Leu Ser Arg Gln Leu Arg Arg Ile Leu Leu Pro Leu Tyr Glu 180 185 190 Val Asp Asp Leu Arg Asp Ala Phe Arg Thr Leu Gly Leu Glu Leu Lys 195 200 205 Thr Pro Leu Gly Asp Thr Thr His Thr Ser Pro Pro Cys Pro Ala Pro 210 215 220 Glu Leu Leu Gly Gly Pro Met Ala Ser Asn Phe Thr Gln Phe Val Leu 225 230 235 240 Val Asp Asn Gly Gly Thr Gly Asp Val Thr Val Ala Pro Ser Asn Phe 245 250 255 Ala Asn Gly Ile Ala Glu Trp Ile Ser Ser Asn Ser Arg Ser Gln Ala 260 265 270 Tyr Lys Val Thr Cys Ser Val Arg Gln Ser Ser Ala Gln Asn Arg Lys 275 280 285 Tyr Thr Ile Lys Val Glu Val Pro Lys Gly Ala Trp Arg Ser Tyr Leu 290 295 300 Asn Met Glu Leu Thr Ile Pro Ile Phe Ala Thr Asn Ser Asp Cys Glu 305 310 315 320 Leu Ile Val Lys Ala Met Gln Gly Leu Leu Lys Asp Gly Asn Pro Ile 325 330 335 Pro Ser Ala Ile Ala Ala Asn Ser Gly Ile Tyr Gly Ser Gly Glu Gly 340 345 350 Arg Gly Ser Leu Leu Thr Cys Gly Asp Val Glu Glu Asn Pro Gly Pro 355 360 365 Gly Thr Pro Lys Lys Lys Arg Lys Val Asp Lys Lys Tyr Ser Ile Gly 370 375 380 Leu Ala Ile Gly Thr Asn Ser Val Gly Trp Ala Val Ile Thr Asp Glu 385 390 395 400 Tyr Lys Val Pro Ser Lys Lys Phe Lys Val Leu Gly Asn Thr Asp Arg 405 410 415 His Ser Ile Lys Lys Asn Leu Ile Gly Ala Leu Leu Phe Asp Ser Gly 420 425 430 Glu Thr Ala Glu Ala Thr Arg Leu Lys Arg Thr Ala Arg Arg Arg Tyr 435 440 445 Thr Arg Arg Lys Asn Arg Ile Cys Tyr Leu Gln Glu Ile Phe Ser Asn 450 455 460 Glu Met Ala Lys Val Asp Asp Ser Phe Phe His Arg Leu Glu Glu Ser 465 470 475 480 Phe Leu Val Glu Glu Asp Lys Lys His Glu Arg His Pro Ile Phe Gly 485 490 495 Asn Ile Val Asp Glu Val Ala Tyr His Glu Lys Tyr Pro Thr Ile Tyr 500 505 510 His Leu Arg Lys Lys Leu Val Asp Ser Thr Asp Lys Ala Asp Leu Arg 515 520 525 Leu Ile Tyr Leu Ala Leu Ala His Met Ile Lys Phe Arg Gly His Phe 530 535 540 Leu Ile Glu Gly Asp Leu Asn Pro Asp Asn Ser Asp Val Asp Lys Leu 545 550 555 560 Phe Ile Gln Leu Val Gln Thr Tyr Asn Gln Leu Phe Glu Glu Asn Pro 565 570 575 Ile Asn Ala Ser Gly Val Asp Ala Lys Ala Ile Leu Ser Ala Arg Leu 580 585 590 Ser Lys Ser Arg Arg Leu Glu Asn Leu Ile Ala Gln Leu Pro Gly Glu 595 600 605 Lys Lys Asn Gly Leu Phe Gly Asn Leu Ile Ala Leu Ser Leu Gly Leu 610 615 620 Thr Pro Asn Phe Lys Ser Asn Phe Asp Leu Ala Glu Asp Ala Lys Leu 625 630 635 640 Gln Leu Ser Lys Asp Thr Tyr Asp Asp Asp Leu Asp Asn Leu Leu Ala 645 650 655 Gln Ile Gly Asp Gln Tyr Ala Asp Leu Phe Leu Ala Ala Lys Asn Leu 660 665 670 Ser Asp Ala Ile Leu Leu Ser Asp Ile Leu Arg Val Asn Thr Glu Ile 675 680 685 Thr Lys Ala Pro Leu Ser Ala Ser Met Ile Lys Arg Tyr Asp Glu His 690 695 700 His Gln Asp Leu Thr Leu Leu Lys Ala Leu Val Arg Gln Gln Leu Pro 705 710 715 720 Glu Lys Tyr Lys Glu Ile Phe Phe Asp Gln Ser Lys Asn Gly Tyr Ala 725 730 735 Gly Tyr Ile Asp Gly Gly Ala Ser Gln Glu Glu Phe Tyr Lys Phe Ile 740 745 750 Lys Pro Ile Leu Glu Lys Met Asp Gly Thr Glu Glu Leu Leu Val Lys 755 760 765 Leu Asn Arg Glu Asp Leu Leu Arg Lys Gln Arg Thr Phe Asp Asn Gly 770 775 780 Ser Ile Pro His Gln Ile His Leu Gly Glu Leu His Ala Ile Leu Arg 785 790 795 800 Arg Gln Glu Asp Phe Tyr Pro Phe Leu Lys Asp Asn Arg Glu Lys Ile 805 810 815 Glu Lys Ile Leu Thr Phe Arg Ile Pro Tyr Tyr Val Gly Pro Leu Ala 820 825 830 Arg Gly Asn Ser Arg Phe Ala Trp Met Thr Arg Lys Ser Glu Glu Thr 835 840 845 Ile Thr Pro Trp Asn Phe Glu Glu Val Val Asp Lys Gly Ala Ser Ala 850 855 860 Gln Ser Phe Ile Glu Arg Met Thr Asn Phe Asp Lys Asn Leu Pro Asn 865 870 875 880 Glu Lys Val Leu Pro Lys His Ser Leu Leu Tyr Glu Tyr Phe Thr Val 885 890 895 Tyr Asn Glu Leu Thr Lys Val Lys Tyr Val Thr Glu Gly Met Arg Lys 900 905 910 Pro Ala Phe Leu Ser Gly Glu Gln Lys Lys Ala Ile Val Asp Leu Leu 915 920 925 Phe Lys Thr Asn Arg Lys Val Thr Val Lys Gln Leu Lys Glu Asp Tyr 930 935 940 Phe Lys Lys Ile Glu Cys Phe Asp Ser Val Glu Ile Ser Gly Val Glu 945 950 955 960 Asp Arg Phe Asn Ala Ser Leu Gly Thr Tyr His Asp Leu Leu Lys Ile 965 970 975 Ile Lys Asp Lys Asp Phe Leu Asp Asn Glu Glu Asn Glu Asp Ile Leu 980 985 990 Glu Asp Ile Val Leu Thr Leu Thr Leu Phe Glu Asp Arg Glu Met Ile 995 1000 1005 Glu Glu Arg Leu Lys Thr Tyr Ala His Leu Phe Asp Asp Lys Val 1010 1015 1020 Met Lys Gln Leu Lys Arg Arg Arg Tyr Thr Gly Trp Gly Arg Leu 1025 1030 1035 Ser Arg Lys Leu Ile Asn Gly Ile Arg Asp Lys Gln Ser Gly Lys 1040 1045 1050 Thr Ile Leu Asp Phe Leu Lys Ser Asp Gly Phe Ala Asn Arg Asn 1055 1060 1065 Phe Met Gln Leu Ile His Asp Asp Ser Leu Thr Phe Lys Glu Asp 1070 1075 1080 Ile Gln Lys Ala Gln Val Ser Gly Gln Gly Asp Ser Leu His Glu 1085 1090 1095 His Ile Ala Asn Leu Ala Gly Ser Pro Ala Ile Lys Lys Gly Ile 1100 1105 1110 Leu Gln Thr Val Lys Val Val Asp Glu Leu Val Lys Val Met Gly 1115 1120 1125 Arg His Lys Pro Glu Asn Ile Val Ile Glu Met Ala Arg Glu Asn 1130 1135 1140 Gln Thr Thr Gln Lys Gly Gln Lys Asn Ser Arg Glu Arg Met Lys 1145 1150 1155 Arg Ile Glu Glu Gly Ile Lys Glu Leu Gly Ser Gln Ile Leu Lys 1160 1165 1170 Glu His Pro Val Glu Asn Thr Gln Leu Gln Asn Glu Lys Leu Tyr 1175 1180 1185 Leu Tyr Tyr Leu Gln Asn Gly Arg Asp Met Tyr Val Asp Gln Glu 1190 1195 1200 Leu Asp Ile Asn Arg Leu Ser Asp Tyr Asp Val Asp His Ile Val 1205 1210 1215 Pro Gln Ser Phe Leu Lys Asp Asp Ser Ile Asp Asn Lys Val Leu 1220 1225 1230 Thr Arg Ser Asp Lys Asn Arg Gly Lys Ser Asp Asn Val Pro Ser 1235 1240 1245 Glu Glu Val Val Lys Lys Met Lys Asn Tyr Trp Arg Gln Leu Leu 1250 1255 1260 Asn Ala Lys Leu Ile Thr Gln Arg Lys Phe Asp Asn Leu Thr Lys 1265 1270 1275 Ala Glu Arg Gly Gly Leu Ser Glu Leu Asp Lys Ala Gly Phe Ile 1280 1285 1290 Lys Arg Gln Leu Val Glu Thr Arg Gln Ile Thr Lys His Val Ala 1295 1300 1305 Gln Ile Leu Asp Ser Arg Met Asn Thr Lys Tyr Asp Glu Asn Asp 1310 1315 1320 Lys Leu Ile Arg Glu Val Lys Val Ile Thr Leu Lys Ser Lys Leu 1325 1330 1335 Val Ser Asp Phe Arg Lys Asp Phe Gln Phe Tyr Lys Val Arg Glu 1340 1345 1350 Ile Asn Asn Tyr His His Ala His Asp Ala Tyr Leu Asn Ala Val 1355 1360 1365 Val Gly Thr Ala Leu Ile Lys Lys Tyr Pro Lys Leu Glu Ser Glu 1370 1375 1380 Phe Val Tyr Gly Asp Tyr Lys Val Tyr Asp Val Arg Lys Met Ile 1385 1390 1395 Ala Lys Ser Glu Gln Glu Ile Gly Lys Ala Thr Ala Lys Tyr Phe 1400 1405 1410 Phe Tyr Ser Asn Ile Met Asn Phe Phe Lys Thr Glu Ile Thr Leu 1415 1420 1425 Ala Asn Gly Glu Ile Arg Lys Arg Pro Leu Ile Glu Thr Asn Gly 1430 1435 1440 Glu Thr Gly Glu Ile Val Trp Asp Lys Gly Arg Asp Phe Ala Thr 1445 1450 1455 Val Arg Lys Val Leu Ser Met Pro Gln Val Asn Ile Val Lys Lys 1460 1465 1470 Thr Glu Val Gln Thr Gly Gly Phe Ser Lys Glu Ser Ile Leu Pro 1475 1480 1485 Lys Arg Asn Ser Asp Lys Leu Ile Ala Arg Lys Lys Asp Trp Asp 1490 1495 1500 Pro Lys Lys Tyr Gly Gly Phe Asp Ser Pro Thr Val Ala Tyr Ser 1505 1510 1515 Val Leu Val Val Ala Lys Val Glu Lys Gly Lys Ser Lys Lys Leu 1520 1525 1530 Lys Ser Val Lys Glu Leu Leu Gly Ile Thr Ile Met Glu Arg Ser 1535 1540 1545 Ser Phe Glu Lys Asn Pro Ile Asp Phe Leu Glu Ala Lys Gly Tyr 1550 1555 1560 Lys Glu Val Lys Lys Asp Leu Ile Ile Lys Leu Pro Lys Tyr Ser 1565 1570 1575 Leu Phe Glu Leu Glu Asn Gly Arg Lys Arg Met Leu Ala Ser Ala 1580 1585 1590 Gly Glu Leu Gln Lys Gly Asn Glu Leu Ala Leu Pro Ser Lys Tyr 1595 1600 1605 Val Asn Phe Leu Tyr Leu Ala Ser His Tyr Glu Lys Leu Lys Gly 1610 1615 1620 Ser Pro Glu Asp Asn Glu Gln Lys Gln Leu Phe Val Glu Gln His 1625 1630 1635 Lys His Tyr Leu Asp Glu Ile Ile Glu Gln Ile Ser Glu Phe Ser 1640 1645 1650 Lys Arg Val Ile Leu Ala Asp Ala Asn Leu Asp Lys Val Leu Ser 1655 1660 1665 Ala Tyr Asn Lys His Arg Asp Lys Pro Ile Arg Glu Gln Ala Glu 1670 1675 1680 Asn Ile Ile His Leu Phe Thr Leu Thr Asn Leu Gly Ala Pro Ala 1685 1690 1695 Ala Phe Lys Tyr Phe Asp Thr Thr Ile Asp Arg Lys Arg Tyr Thr 1700 1705 1710 Ser Thr Lys Glu Val Leu Asp Ala Thr Leu Ile His Gln Ser Ile 1715 1720 1725 Thr Gly Leu Tyr Glu Thr Arg Ile Asp Leu Ser Gln Leu Gly Gly 1730 1735 1740 Asp Ser Gly Gly Ser Gly Gly Ser Gly Gly Ser Thr Asn Leu Ser 1745 1750 1755 Asp Ile Ile Glu Lys Glu Thr Gly Lys Gln Leu Val Ile Gln Glu 1760 1765 1770 Ser Ile Leu Met Leu Pro Glu Glu Val Glu Glu Val Ile Gly Asn 1775 1780 1785 Lys Pro Glu Ser Asp Ile Leu Val His Thr Ala Tyr Asp Glu Ser 1790 1795 1800 Thr Asp Glu Asn Val Met Leu Leu Thr Ser Asp Ala Pro Glu Tyr 1805 1810 1815 Lys Pro Trp Ala Leu Val Ile Gln Asp Ser Asn Gly Glu Asn Lys 1820 1825 1830 Ile Lys Met Leu Ser Gly Gly Ser Gly Gly Ser Gly Gly Ser Thr 1835 1840 1845 Asn Leu Ser Asp Ile Ile Glu Lys Glu Thr Gly Lys Gln Leu Val 1850 1855 1860 Ile Gln Glu Ser Ile Leu Met Leu Pro Glu Glu Val Glu Glu Val 1865 1870 1875 Ile Gly Asn Lys Pro Glu Ser Asp Ile Leu Val His Thr Ala Tyr 1880 1885 1890 Asp Glu Ser Thr Asp Glu Asn Val Met Leu Leu Thr Ser Asp Ala 1895 1900 1905 Pro Glu Tyr Lys Pro Trp Ala Leu Val Ile Gln Asp Ser Asn Gly 1910 1915 1920 Glu Asn Lys Ile Lys Met Leu Ser Gly Gly Ser Pro Lys Lys Lys 1925 1930 1935 Arg Lys Val 1940 <210> 37 <211> 1866 <212> PRT <213> Artificial sequence <220> <223> Protein sequences of RNA scaffold-mediated recruitment systems <220> <221> MISC_FEATURE <222> (1)..(7) <223> Nuclear Localization Signal (NLS) <220> <221> MISC_FEATURE <222> (8)..(235) <223> AID <220> <221> MISC_FEATURE <222> (236) (261) <223> connector <220> <221> MISC_FEATURE <222> (262) (378) <223> MCP <220> <221> MISC_FEATURE <222> (382) (399) <223> T2A peptide <220> <221> MISC_FEATURE <222> (402)..(1768) <223> nCAS9D10A <220> <221> MISC_FEATURE <222> (1772) ... (1855) <223> UGI <220> <221> MISC_FEATURE <222> (1860) ... (1866) <223> Nuclear Localization Signal (NLS) <400> 37 Pro Lys Lys Lys Arg Lys Val Met Ser Ser Ser Glu Thr Gly Pro Val Ala 1 5 10 15 Val Asp Pro Thr Leu Arg Arg Arg Ile Glu Pro His Glu Phe Glu Val 20 25 30 Phe Phe Asp Pro Arg Glu Leu Arg Lys Glu Thr Cys Leu Leu Tyr Glu 35 40 45 Ile Asn Trp Gly Gly Arg His Ser Ile Trp Arg His Thr Ser Gln Asn 50 55 60 Thr Asn Lys His Val Glu Val Asn Phe Ile Glu Lys Phe Thr Thr Glu 65 70 75 80 Arg Tyr Phe Cys Pro Asn Thr Arg Cys Ser Ile Thr Trp Phe Leu Ser 85 90 95 Trp Ser Pro Cys Gly Glu Cys Ser Arg Ala Ile Thr Glu Phe Leu Ser 100 105 110 Arg Tyr Pro His Val Thr Leu Phe Ile Tyr Ile Ala Arg Leu Tyr His 115 120 125 His Ala Asp Pro Arg Asn Arg Gln Gly Leu Arg Asp Leu Ile Ser Ser 130 135 140 Gly Val Thr Ile Gln Ile Met Thr Glu Gln Glu Ser Gly Tyr Cys Trp 145 150 155 160 Arg Asn Phe Val Asn Tyr Ser Pro Ser Asn Glu Ala His Trp Pro Arg 165 170 175 Tyr Pro His Leu Trp Val Arg Leu Tyr Val Leu Glu Leu Tyr Cys Ile 180 185 190 Ile Leu Gly Leu Pro Pro Cys Leu Asn Ile Leu Arg Arg Lys Gln Pro 195 200 205 Gln Leu Thr Phe Phe Thr Ile Ala Leu Gln Ser Cys His Tyr Gln Arg 210 215 220 Leu Pro Pro His Ile Leu Trp Ala Thr Gly Leu Lys Glu Leu Lys Thr 225 230 235 240 Pro Leu Gly Asp Thr Thr His Thr Ser Pro Pro Cys Pro Ala Pro Glu 245 250 255 Leu Leu Gly Gly Pro Met Ala Ser Asn Phe Thr Gln Phe Val Leu Val 260 265 270 Asp Asn Gly Gly Thr Gly Asp Val Thr Val Ala Pro Ser Asn Phe Ala 275 280 285 Asn Gly Ile Ala Glu Trp Ile Ser Ser Asn Ser Arg Ser Gln Ala Tyr 290 295 300 Lys Val Thr Cys Ser Val Arg Gln Ser Ser Ala Gln Asn Arg Lys Tyr 305 310 315 320 Thr Ile Lys Val Glu Val Pro Lys Gly Ala Trp Arg Ser Tyr Leu Asn 325 330 335 Met Glu Leu Thr Ile Pro Ile Phe Ala Thr Asn Ser Asp Cys Glu Leu 340 345 350 Ile Val Lys Ala Met Gln Gly Leu Leu Lys Asp Gly Asn Pro Ile Pro 355 360 365 Ser Ala Ile Ala Ala Asn Ser Gly Ile Tyr Gly Ser Gly Glu Gly Arg 370 375 380 Gly Ser Leu Leu Thr Cys Gly Asp Val Glu Glu Asn Pro Gly Pro Gly 385 390 395 400 Thr Asp Lys Lys Tyr Ser Ile Gly Leu Ala Ile Gly Thr Asn Ser Val 405 410 415 Gly Trp Ala Val Ile Thr Asp Glu Tyr Lys Val Pro Ser Lys Lys Phe 420 425 430 Lys Val Leu Gly Asn Thr Asp Arg His Ser Ile Lys Lys Asn Leu Ile 435 440 445 Gly Ala Leu Leu Phe Asp Ser Gly Glu Thr Ala Glu Ala Thr Arg Leu 450 455 460 Lys Arg Thr Ala Arg Arg Arg Tyr Thr Arg Arg Lys Asn Arg Ile Cys 465 470 475 480 Tyr Leu Gln Glu Ile Phe Ser Asn Glu Met Ala Lys Val Asp Asp Ser 485 490 495 Phe Phe His Arg Leu Glu Glu Ser Phe Leu Val Glu Glu Asp Lys Lys 500 505 510 His Glu Arg His Pro Ile Phe Gly Asn Ile Val Asp Glu Val Ala Tyr 515 520 525 His Glu Lys Tyr Pro Thr Ile Tyr His Leu Arg Lys Lys Leu Val Asp 530 535 540 Ser Thr Asp Lys Ala Asp Leu Arg Leu Ile Tyr Leu Ala Leu Ala His 545 550 555 560 Met Ile Lys Phe Arg Gly His Phe Leu Ile Glu Gly Asp Leu Asn Pro 565 570 575 Asp Asn Ser Asp Val Asp Lys Leu Phe Ile Gln Leu Val Gln Thr Tyr 580 585 590 Asn Gln Leu Phe Glu Glu Asn Pro Ile Asn Ala Ser Gly Val Asp Ala 595 600 605 Lys Ala Ile Leu Ser Ala Arg Leu Ser Lys Ser Arg Arg Leu Glu Asn 610 615 620 Leu Ile Ala Gln Leu Pro Gly Glu Lys Lys Asn Gly Leu Phe Gly Asn 625 630 635 640 Leu Ile Ala Leu Ser Leu Gly Leu Thr Pro Asn Phe Lys Ser Asn Phe 645 650 655 Asp Leu Ala Glu Asp Ala Lys Leu Gln Leu Ser Lys Asp Thr Tyr Asp 660 665 670 Asp Asp Leu Asp Asn Leu Leu Ala Gln Ile Gly Asp Gln Tyr Ala Asp 675 680 685 Leu Phe Leu Ala Ala Lys Asn Leu Ser Asp Ala Ile Leu Leu Ser Asp 690 695 700 Ile Leu Arg Val Asn Thr Glu Ile Thr Lys Ala Pro Leu Ser Ala Ser 705 710 715 720 Met Ile Lys Arg Tyr Asp Glu His His Gln Asp Leu Thr Leu Leu Lys 725 730 735 Ala Leu Val Arg Gln Gln Leu Pro Glu Lys Tyr Lys Glu Ile Phe Phe 740 745 750 Asp Gln Ser Lys Asn Gly Tyr Ala Gly Tyr Ile Asp Gly Gly Ala Ser 755 760 765 Gln Glu Glu Phe Tyr Lys Phe Ile Lys Pro Ile Leu Glu Lys Met Asp 770 775 780 Gly Thr Glu Glu Leu Leu Val Lys Leu Asn Arg Glu Asp Leu Leu Arg 785 790 795 800 Lys Gln Arg Thr Phe Asp Asn Gly Ser Ile Pro His Gln Ile His Leu 805 810 815 Gly Glu Leu His Ala Ile Leu Arg Arg Gln Glu Asp Phe Tyr Pro Phe 820 825 830 Leu Lys Asp Asn Arg Glu Lys Ile Glu Lys Ile Leu Thr Phe Arg Ile 835 840 845 Pro Tyr Tyr Val Gly Pro Leu Ala Arg Gly Asn Ser Arg Phe Ala Trp 850 855 860 Met Thr Arg Lys Ser Glu Glu Thr Ile Thr Pro Trp Asn Phe Glu Glu 865 870 875 880 Val Val Asp Lys Gly Ala Ser Ala Gln Ser Phe Ile Glu Arg Met Thr 885 890 895 Asn Phe Asp Lys Asn Leu Pro Asn Glu Lys Val Leu Pro Lys His Ser 900 905 910 Leu Leu Tyr Glu Tyr Phe Thr Val Tyr Asn Glu Leu Thr Lys Val Lys 915 920 925 Tyr Val Thr Glu Gly Met Arg Lys Pro Ala Phe Leu Ser Gly Glu Gln 930 935 940 Lys Lys Ala Ile Val Asp Leu Leu Phe Lys Thr Asn Arg Lys Val Thr 945 950 955 960 Val Lys Gln Leu Lys Glu Asp Tyr Phe Lys Lys Ile Glu Cys Phe Asp 965 970 975 Ser Val Glu Ile Ser Gly Val Glu Asp Arg Phe Asn Ala Ser Leu Gly 980 985 990 Thr Tyr His Asp Leu Leu Lys Ile Ile Lys Asp Lys Asp Phe Leu Asp 995 1000 1005 Asn Glu Glu Asn Glu Asp Ile Leu Glu Asp Ile Val Leu Thr Leu 1010 1015 1020 Thr Leu Phe Glu Asp Arg Glu Met Ile Glu Glu Arg Leu Lys Thr 1025 1030 1035 Tyr Ala His Leu Phe Asp Asp Lys Val Met Lys Gln Leu Lys Arg 1040 1045 1050 Arg Arg Tyr Thr Gly Trp Gly Arg Leu Ser Arg Lys Leu Ile Asn 1055 1060 1065 Gly Ile Arg Asp Lys Gln Ser Gly Lys Thr Ile Leu Asp Phe Leu 1070 1075 1080 Lys Ser Asp Gly Phe Ala Asn Arg Asn Phe Met Gln Leu Ile His 1085 1090 1095 Asp Asp Ser Leu Thr Phe Lys Glu Asp Ile Gln Lys Ala Gln Val 1100 1105 1110 Ser Gly Gln Gly Asp Ser Leu His Glu His Ile Ala Asn Leu Ala 1115 1120 1125 Gly Ser Pro Ala Ile Lys Lys Gly Ile Leu Gln Thr Val Lys Val 1130 1135 1140 Val Asp Glu Leu Val Lys Val Met Gly Arg His Lys Pro Glu Asn 1145 1150 1155 Ile Val Ile Glu Met Ala Arg Glu Asn Gln Thr Thr Gln Lys Gly 1160 1165 1170 Gln Lys Asn Ser Arg Glu Arg Met Lys Arg Ile Glu Glu Gly Ile 1175 1180 1185 Lys Glu Leu Gly Ser Gln Ile Leu Lys Glu His Pro Val Glu Asn 1190 1195 1200 Thr Gln Leu Gln Asn Glu Lys Leu Tyr Leu Tyr Tyr Leu Gln Asn 1205 1210 1215 Gly Arg Asp Met Tyr Val Asp Gln Glu Leu Asp Ile Asn Arg Leu 1220 1225 1230 Ser Asp Tyr Asp Val Asp His Ile Val Pro Gln Ser Phe Leu Lys 1235 1240 1245 Asp Asp Ser Ile Asp Asn Lys Val Leu Thr Arg Ser Asp Lys Asn 1250 1255 1260 Arg Gly Lys Ser Asp Asn Val Pro Ser Glu Glu Val Val Lys Lys 1265 1270 1275 Met Lys Asn Tyr Trp Arg Gln Leu Leu Asn Ala Lys Leu Ile Thr 1280 1285 1290 Gln Arg Lys Phe Asp Asn Leu Thr Lys Ala Glu Arg Gly Gly Leu 1295 1300 1305 Ser Glu Leu Asp Lys Ala Gly Phe Ile Lys Arg Gln Leu Val Glu 1310 1315 1320 Three Arg Gln With Three Lys His Val Only Gln With Asp Ser Arg 1325 1330 1335 Met Asn Thr Lys Tyr Asp Glu Asn Asp Lys Leu Ile Arg Glu Val 1340 1345 1350 Lys Val Ile Thr Leu Lys Ser Lys Leu Val Ser Asp Phe Arg Lys 1355 1360 1365 Asp Phe Gln Phe Tyr Lys Val Arg Glu Ile Asn Asn Tyr His His 1370 1375 1380 Only His Asp Only Tyr Leu Asn Only Val Val Gly Thr Only Leu Ile 1385 1390 1395 Lys Lys Tyr Pro Lys Leu Glu Ser Glu Phe Val Tyr Gly Asp Tyr 1400 1405 1410 Lys Val Tyr Asp Val Arg Lys Met Ile Ala Lys Ser Glu Gln Glu 1415 1420 1425 Ile Gly Lys Ala Thr Ala Lys Tyr Phe Phe Tyr Ser Asn Ile Met 1430 1435 1440 Asn Phe Phe Lys Thr Glu Ile Thr Leu Ala Asn Gly Glu Ile Arg 1445 1450 1455 Lys Arg Pro Leu Ile Glu Thr Asn Gly Glu Thr Gly Glu Ile Val 1460 1465 1470 Trp Asp Lys Gly Arg Asp Phe Ala Thr Val Arg Lys Val Leu Ser 1475 1480 1485 Met Pro Gln Val Asn Ile Val Lys Lys Thr Glu Val Gln Thr Gly 1490 1495 1500 Gly Phe Ser Lys Glu Ser Ile Leu Pro Lys Arg Asn Ser Asp Lys 1505 1510 1515 Leu Ile Ala Arg Lys Lys Asp Trp Asp Pro Lys Lys Tyr Gly Gly 1520 1525 1530 Phe Asp Ser Pro Thr Val Ala Tyr Ser Val Leu Val Val Ala Lys 1535 1540 1545 Val Glu Lys Gly Lys Ser Lys Lys Leu Lys Ser Val Lys Glu Leu 1550 1555 1560 Leu Gly Ile Thr Ile Met Glu Arg Ser Ser Phe Glu Lys Asn Pro 1565 1570 1575 Ile Asp Phe Leu Glu Ala Lys Gly Tyr Glu Val Lys Asp 1580 1585 1590 Leu Ile Lys Leu Pro Lys Ser Leu Phe Glu Leu Glu Asn 1595 1600 1605 Gly Arg Lys Arg Met Leu Ser Ala Gly Glu Leu Gln Lys Gly 1610 1615 1620 Asn Glu Leu Ala Leu Pro Ser Lys Tyr Val Asn Phe Leu Tyr Leu 1625 1630 1635 Ala Ser His Tyr Glu Lys Leu Lys Gly Ser Pro Glu Asp Asn Glu 1640 1645 1650 Gln Lys Gln Leu Phe Val Glu Gln His Lys His Tyr Leu Asp Glu 1655 1660 1665 Glu Gln Synthesis of Glu Phe Synthesis of Lys Arg Val Leu Wing 1670 1675 1680 Asp Ala Asn Leu Asp Lys Val Leu Ser Ala Tyr Asn Lys His Arg 1685 1690 1695 Asp Lys Pro With Arg Glu Gln Ala Glu Asn With Leu Phe 1700 1705 1710 Thr Leu Thr Asn Leu Gly Ala Pro Ala Ala Phe Lys Tyr Phe Asp 1715 1720 1725 Thr Thr Ile Asp Arg Lys Arg Tyr Thr Ser Thr Lys Glu Val Leu 1730 1735 1740 Asp Ala Thr Leu Ile His Gln Ser Ile Thr Gly Leu Tyr Glu Thr 1745 1750 1755 Arg Ile Asp Leu Ser Gln Leu Gly Gly Asp Ser Gly Gly Ser Thr 1760 1765 1770 Asn Leu Ser Asp Ile Ile Glu Lys Glu Thr Gly Lys Gln Leu Val 1775 1780 1785 Ile Gln Glu Ser Ile Leu Met Leu Pro Glu Glu Val Glu Glu Val 1790 1795 1800 Ile Gly Asn Lys Pro Glu Ser Asp Ile Leu Val His Thr Ala Tyr 1805 1810 1815 Asp Glu Ser Thr Asp Glu Asn Val Met Leu Leu Thr Ser Asp Ala 1820 1825 1830 Pro Glu Tyr Lys Pro Trp Ala Leu Val Ile Gln Asp Ser Asn Gly 1835 1840 1845 Glu Asn Lys Ile Lys Met Leu Ser Gly Gly Ser Pro Lys Lys Lys 1850 1855 1860 Arg Lys Val 1865 <210> 38 <211> 128 <212> RNA <213> Artificial sequence <220> <223> Sequence of gRNA_MS2 construct <220> <221> misc_feature <222> (1)..(20) <223> N represents a customizable target sequence. <220> <221> misc_feature <222> (21)..(96) <223> gRNA scaffold <220> <221> misc_feature <222> (101)..(119) <223> MS2 aptamer <400> 38 nnnnnnnnnn nnnnnnnnnn guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcgcgc acaugaggau cacccaug 120 cuuuuuuu 128 <210> 39 <211> 168 <212> DNA <213> Artificial sequence <220> <223> Sequence of gRNA_2xMS2 construct <220> <221> misc_feature <222> (1)..(20) <223> N represents a customizable target sequence. <220> <221> misc_feature <222> (21)..(96) <223> gRNA scaffold <220> <221> misc_feature <222> (103) (121) <223> MS2 aptamers <220> <221> misc_feature <222> (132) (150) <223> MS2 aptamers <220> <221> misc_feature <222> (164) (170) <223> Termination <400> 39 nnnnnnnnnn nnnnnnnngu uuuagagcua gaaauagcaa guuaaaauaa ggcuaguccg 60 uuaucaacuu gaaaaagugg caccgagucg gugcgggagc acaugaggau cacccaug 120 ccacgagcga caugaggauc acccaugucg ctcgtgttcc cuuuuuuu 168 <210> 40 <211> 103 <212> DNA <213> Artificial sequence <220> <223> S. pyrogenes sgRNA <220> <221> misc_feature <222> (1)..(20) <223> N represents a 20 nt target-specific interval sequence. <220> <221> misc_feature <222> (21)..(96) <223> constant sgRNA sequence <220> <221> misc_feature <222> (97) (103) <223> Termination <400> 40 nnnnnnnnnn nnnnnnnnnn gttttagagc tagaaatagc aagttaaaat aaggctagtc 60 cgttatcaac ttgaaaaagt ggcaccgagt cggtgctttt ttt 103 <210> 41 <211> 170 <212> DNA <213> Artificial sequence <220> <223> 2xMS2_3' <220> <221> misc_feature <222> (1)..(20) <223> N represents a 20 nt target-specific spacer sequence. <220> <221> misc_feature <222> (21)..(96) <223> constant sgRNA sequence <220> <221> misc_feature <222> (97)..(102) <223> extend <220> <221> misc_feature <222> (103) (121) <223> MS2 (C5 variant) <220> <221> misc_feature <222> (122) (131) <223> extend <220> <221> misc_feature <222> (132) (150) <223> MS2 (C5 variant) <220> <221> misc_feature <222> (151) (163) <223> extend <220> <221> misc_feature <222> (164) (170) <223> Termination <400> 41 nnnnnnnnnn nnnnnnnnnn gttttagagc tagaaatagc aagttaaaat aaggctagtc 60 cgttatcaac ttgaaaaagt ggcaccgagt cggtgcggga gcacatgagg atcacccatg 120 tgccacgagc gacatgagga tcacccatgt cgctcgtgtt cccttttttt 170 <210> 42 <211> 133 <212> DNA <213> Artificial sequence <220> <223> 1xMS2_TL <220> <221> misc_feature <222> (1)..(20) <223> N represents a 20 nt target-specific spacer sequence. <220> <221> misc_feature <222> (21)..(32) <223> constant sgRNA sequence <220> <221> misc_feature <222> (33) (37) <223> extend <220> <221> misc_feature <222> (38) (56) <223> MS2 (C5 variant) aptamer <220> <221> misc_feature <222> (57) (66) <223> extend <220> <221> misc_feature <222> (67) (126) <223> constant sgRNA sequence <220> <221> misc_feature <222> (127) (133) <223> Termination <400> 42 nnnnnnnnnn nnnnnnnnnn gttttagagc taggccaaca tgaggatcac ccatgtctgc 60 agggcctagc aagttaaaat aaggctagtc cgttatcaac ttgaaaaagt ggcaccgagt 120 cggtgctttt ttt 133 <210> 43 <211> 128 <212> DNA <213> Artificial sequence <220> <223> 1xMS2_3' <220> <221> misc_feature <222> (1)..(20) <223> N represents a 20 nt target-specific spacer sequence. <220> <221> misc_feature <222> (21)..(96) <223> constant sgRNA sequence <220> <221> misc_feature <222> (97) (100) <223> extend <220> <221> misc_feature <222> (101)..(119) <223> MS2 (C5 variant) aptamer <220> <221> misc_feature <222> (120)...(121) <223> extend <220> <221> misc_feature <222> (122) (128) <223> Termination <400> 43 nnnnnnnnnn nnnnnnnnnn gttttagagc tagaaatagc aagttaaaat aaggctagtc 60 cgttatcaac ttgaaaaagt ggcaccgagt cggtgcgcgc acatgaggat cacccatgtg 120 cttttttt 128 <210> 44 <211> 117 <212> DNA <213> Artificial sequence <220> <223> 7bp extension_US <220> <221> misc_feature <222> (1)..(20) <223> N represents a 20 nt target-specific spacer sequence. <220> <221> misc_feature <222> (21)..(32) <223> constant sgRNA sequence <220> <221> misc_feature <222> (33)..(39) <223> Extended repeats: anti-repetition sequences <220> <221> misc_feature <222> (40) (43) <223> constant sgRNA sequence <220> <221> misc_feature <222> (44) (50) <223> Extended repeats: anti-repetition sequences <220> <221> misc_feature <222> (51)..(110) <223> constant sgRNA sequence <220> <221> misc_feature <222> (111)...(117) <223> Termination <400> 44 nnnnnnnnnn nnnnnnnnnn gttttagagc tatgctgttg aaaaacagca tagcaagtta 60 aaataaggct agtccgttat caacttgaaa aagtggcacc gagtcggtgc ttttttt 117 <210> 45 <211> 184 <212> DNA <213> Artificial sequence <220> <223> 2xMS2_3' 7bp_extended_US <220> <221> misc_feature <222> (1)..(20) <223> N represents a 20 nt target-specific spacer sequence. <220> <221> misc_feature <222> (21)..(32) <223> constant sgRNA sequence <220> <221> misc_feature <222> (33)..(39) <223> Extended repeats: anti-repetition sequences <220> <221> misc_feature <222> (40) (43) <223> constant sgRNA sequence <220> <221> misc_feature <222> (44) (50) <223> Extended repeats: anti-repetition sequences <220> <221> misc_feature <222> (51)..(110) <223> constant sgRNA sequence <220> <221> misc_feature <222> (111)..(116) <223> extend <220> <221> misc_feature <222> (117) (135) <223> MS2 (C5 variant) <220> <221> misc_feature <222> (136) (145) <223> extend <220> <221> misc_feature <222> (146) (164) <223> MS2 (C5 variant) aptamer <220> <221> misc_feature <222> (165) (177) <223> extend <220> <221> misc_feature <222> (178) (184) <223> Termination <400> 45 nnnnnnnnnn nnnnnnnnnn gttttagagc tatgctgttg aaaaacagca tagcaagtta 60 aaataaggct agtccgttat caacttgaaa aagtggcacc gagtcggtgc gggagcacat 120 gaggatcacc catgtgccac gagcgacatg aggatcaccc atgtcgctcg tgttcccttt 180 tttt 184 <210> 46 <211> 147 <212> DNA <213> Artificial sequence <220> <223> 1xMS2-TL _7bp Extension_US <220> <221> misc_feature <222> (1)..(20) <223> N represents a 20 nt target-specific spacer sequence. <220> <221> misc_feature <222> (21)..(32) <223> constant sgRNA sequence <220> <221> misc_feature <222> (33)..(39) <223> Extended repeats: anti-repetition sequences <220> <221> misc_feature <222> (40) (44) <223> extend <220> <221> misc_feature <222> (45) (63) <223> MS2 (C5 variant) aptamer <220> <221> misc_feature <222> (64) (73) <223> extend <220> <221> misc_feature <222> (74) (80) <223> Extended repeats: anti-repetition sequences <220> <221> misc_feature <222> (81)..(140) <223> constant sgRNA sequence <220> <221> misc_feature <222> (141) (147) <223> constant sgRNA sequence <220> <221> misc_feature <222> (141) (147) <223> Termination <400> 46 nnnnnnnnnn nnnnnnnnnn gttttagagc tatgctgttg gccaacatga ggatcaccca 60 tgtctgcagg gccaacagca tagcaagtta aaataaggct agtccgttat caacttgaaa 120 aagtggcacc gagtcggtgc ttttttt 147 <210> 47 <211> 132 <212> DNA <213> Artificial sequence <220> <223> 1xMS2-3'_2bp extension_US <220> <221> misc_feature <222> (1)..(20) <223> N represents a 20 nt target-specific spacer sequence. <220> <221> misc_feature <222> (21)..(32) <223> constant sgRNA sequence <220> <221> misc_feature <222> (33)..(34) <223> Extended repeats: anti-repetition sequences <220> <221> misc_feature <222> (35) (38) <223> constant sgRNA sequence <220> <221> misc_feature <222> (39) (40) <223> Extended repeats: anti-repetition sequences <220> <221> misc_feature <222> (41) (100) <223> constant sgRNA sequence <220> <221> misc_feature <222> (101)..(104) <223> extend <220> <221> misc_feature <222> (105) (123) <223> MS2 (C5 variant) aptamer <220> <221> misc_feature <222> (124) (126) <223> extend <220> <221> misc_feature <222> (127) (132) <223> Termination <400> 47 nnnnnnnnnn nnnnnnnnnn gttttagagc tatggaaaca tagcaagtta aaataaggct 60 agtccgttat caacttgaaa aagtggcacc gagtcggtgc gcgcacatga ggatcaccca 120 tgtgcttttt tt 132 <210> 48 <211> 138 <212> DNA <213> Artificial sequence <220> <223> 1xMS2-3'_5bp-Extension_US <220> <221> misc_feature <222> (1)..(20) <223> N represents a 20 nt target-specific spacer sequence. <220> <221> misc_feature <222> (21)..(32) <223> constant sgRNA sequence <220> <221> misc_feature <222> (33) (37) <223> Extended repeats: anti-repetition sequences <220> <221> misc_feature <222> (38) (41) <223> constant sgRNA sequence <220> <221> misc_feature <222> (42) (46) <223> Extended repeats: anti-repetition sequences <220> <221> misc_feature <222> (47) (106) <223> constant sgRNA sequence <220> <221> misc_feature <222> (107) (110) <223> extend <220> <221> misc_feature <222> (111)..(129) <223> MS2 (C5 variant) aptamer <220> <221> misc_feature <222> (130) (131) <223> extend <220> <221> misc_feature <222> (132) (138) <223> Termination <400> 48 nnnnnnnnnn nnnnnnnnnn gttttagagc tatgctggaa acagcatagc aagttaaaat 60 aaggctagtc cgttatcaac ttgaaaaagt ggcaccgagt cggtgcgcgc acatgaggat 120 cacccatgtg cttttttt 138 <210> 49 <211> 142 <212> DNA <213> Artificial sequence <220> <223> 1xMS2-3'_7bp-Extension_US <220> <221> misc_feature <222> (1)..(20) <223> N denotes the 20 nt target specific spacer sequence <220> <221> misc_feature <222> (21)..(32) <223> constant sgRNA sequence <220> <221> misc_feature <222> (33)..(39) <223> Extended repeats: anti-repetition sequences <220> <221> misc_feature <222> (40) (43) <223> constant sgRNA sequence <220> <221> misc_feature <222> (44) (50) <223> Extended repeats: anti-repetition sequences <220> <221> misc_feature <222> (51)..(110) <223> constant sgRNA sequence <220> <221> misc_feature <222> (111) (114) <223> extensions <220> <221> misc_feature <222> (111)..(133) <223> MS2 (C5 variant) aptamer <220> <221> misc_feature <222> (134) (135) <223> extend <220> <221> misc_feature <222> (136) (142) <223> extend <220> <221> misc_feature <222> (136) (142) <223> Termination <400> 49 nnnnnnnnnn nnnnnnnnnn gttttagagc tatgctgttg aaaaacagca tagcaagtta 60 aaataaggct agtccgttat caacttgaaa aagtggcacc gagtcggtgc gcgcacatga 120 ggatcaccca tgtgcttttt tt 142 <210> 50 <211> 148 <212> DNA <213> Artificial sequence <220> <223> 1xMS2-3'_10bp-Extension_US <220> <221> misc_feature <222> (1)..(20) <223> N represents a 20 nt target-specific spacer sequence. <220> <221> misc_feature <222> (21)..(32) <223> constant sgRNA sequence <220> <221> misc_feature <222> (33)..(42) <223> Extended repeats: anti-repetition sequences <220> <221> misc_feature <222> (43) (46) <223> constant sgRNA sequence <220> <221> misc_feature <222> (47) (56) <223> Extended repeats: anti-repetition sequences <220> <221> misc_feature <222> (57) (116) <223> constant sgRNA sequence <220> <221> misc_feature <222> (117) (120) <223> extend <220> <221> misc_feature <222> (121) (139) <223> MS2 (C5 variant) aptamer <220> <221> misc_feature <222> (140) (141) <223> extend <220> <221> misc_feature <222> (142) (148) <223> Termination <400> 50 nnnnnnnnnn nnnnnnnnnn gttttagagc tatgctgttt tggaaacaaa acagcatagc 60 aagttaaaat aaggctagtc cgttatcaac ttgaaaaagt ggcaccgagt cggtgcgcgc 120 acatgaggat cacccatgtg cttttttt 148 <210> 51 <211> 147 <212> DNA <213> Artificial sequence <220> <223> 1xMS2-SL2_7bp-Extended_US <220> <221> misc_feature <222> (1)..(20) <223> N represents a 20 nt target-specific spacer sequence. <220> <221> misc_feature <222> (21)..(32) <223> constant sgRNA sequence <220> <221> misc_feature <222> (33)..(39) <223> Extended repeats: anti-repetition sequences <220> <221> misc_feature <222> (40) (43) <223> constant sgRNA sequence <220> <221> misc_feature <222> (44) (50) <223> The extended repeat: anti-repetition sequence <220> <221> misc_feature <222> (51)..(86) <223> constant sgRNA sequence <220> <221> misc_feature <222> (87) (91) <223> extend <220> <221> misc_feature <222> (92)..(110) <223> MS2 (C5 variant) aptamer <220> <221> misc_feature <222> (111) (120) <223> extend <220> <221> misc_feature <222> (121) (140) <223> constant sgRNA sequence <220> <221> misc_feature <222> (141) (147) <223> Termination <400> 51 nnnnnnnnnn nnnnnnnnnn gttttagagc tatgctgttg aaaaacagca tagcaagtta 60 aaataaggct agtccgttat caacttggcc aacatgagga tcacccatgt ctgcagggcc 120 aagtggcacc gagtcggtgc ttttttt 147 <210> 52 <211> 167 <212> DNA <213> Artificial sequence <220> <223> 2xMS2_C5-SL2_f6-3'_7bp Extension-US <220> <221> misc_feature <222> (1)..(20) <223> N represents a 20 nt target-specific spacer sequence. <220> <221> misc_feature <222> (21)..(32) <223> constant sgRNA sequence <220> <221> misc_feature <222> (33)..(39) <223> Extended repeats: anti-repetition sequences <220> <221> misc_feature <222> (40) (43) <223> constant sgRNA sequence <220> <221> misc_feature <222> (44) (50) <223> Extended repeats: anti-repetition sequences <220> <221> misc_feature <222> (51)..(86) <223> constant sgRNA sequence <220> <221> misc_feature <222> (87) (91) <223> extend <220> <221> misc_feature <222> (92)..(110) <223> f6 Fitt <220> <221> misc_feature <222> (111) (120) <223> extend <220> <221> misc_feature <222> (121) (140) <223> constant sgRNA sequence <220> <221> misc_feature <222> (141) (144) <223> extend <220> <221> misc_feature <222> (145) (158) <223> f6 Fitt <220> <221> misc_feature <222> (159) (160) <223> extend <220> <221> misc_feature <222> (161) (167) <223> Termination <400> 52 nnnnnnnnnn nnnnnnnnnn gttttagagc tatgctgttg aaaaacagca tagcaagtta 60 aaataaggct agtccgttat caacttggcc aacatgagga tcacccatgt ctgcagggcc 120 aagtggcacc gagtcggtgc gcgcccacag tcactggggc ttttttt 167 <210> 53 <211> 20 <212> DNA <213> Artificial sequence <220> <223> sgRNa target site sequence site 2 <400> 53 gaacacaaag catagactgc 20 <210> 54 <211> 20 <212> DNA <213> Artificial sequence <220> <223> sgRNa target site sequence site 3 <400> 54 ggcccagact gagcacgtga 20 <210> 55 <211> 20 <212> DNA <213> Artificial sequence <220> <223> sgRNa target site sequence CTNNB1 <400> 55 ctggactctg gaatccattc 20 <210> 56 <211> 20 <212> DNA <213> Artificial sequence <220> <223> sgRNa target site sequence EGFR <400> 56 atcacgcagc tcatgccctt 20 <210> 57 <211> 20 <212> DNA <213> Artificial sequence <220> <223> sgRNa target site sequence PCSK9 <400> 57 caggttccac gggatgctct 20 <210> 58 <211> 20 <212> DNA <213> Artificial sequence <220> <223> sgRNa target site sequence FANCF <400> 58 ggaatccctt ctgcagcacc 20 <210> 59 <211> 20 <212> DNA <213> Artificial sequence <220> <223> sgRNa target site sequence TRAC <400> 59 ttcgtatctg taaaaccaag 20 <210> 60 <211> 20 <212> DNA <213> Artificial sequence <220> <223> sgRNa target site sequence B2M <400> 60 cttaccccac ttaactatct 20 <210> 61 <211> 20 <212> DNA <213> Artificial sequence <220> <223> sgRNA target site sequence CR0118_PDCD1 <400> 61 cagttccaaa ccctggtggt 20 <210> 62 <211> 20 <212> DNA <213> Artificial sequence <220> <223> sgRNA target site sequence CR0107_PDCD1 <400> 62 gggggttcca gggcctgtct 20 <210> 63 <211> 20 <212> DNA <213> Artificial sequence <220> <223> sgRNA target site sequence CR0057-TRAC_EX3 <400> 63 ttcgtatctg taaaaccaag 20 <210> 64 <211> 20 <212> DNA <213> Artificial sequence <220> <223> sgRNA target site sequence CR0151_CD2 <400> 64 gttcagccaa aacctcccca 20 <210> 65 <211> 20 <212> DNA <213> Artificial sequence <220> <223> sgRNA target site sequence CR0121_PDCD1 <400> 65 ggagtctgag agatggagag 20 <210> 66 <211> 20 <212> DNA <213> Artificial sequence <220> <223> sgRNA target site sequence CR0165_CIITA <400> 66 cagctcacag tgtgccacca 20 <210> 67 <211> 20 <212> DNA <213> Artificial sequence <220> <223> sgRNA target site sequence TRAC_22550571 <400> 67 ttcaaaacct gtcagtgatt 20 <210> 68 <211> 20 <212> DNA <213> Artificial sequence <220> <223> sgRNA target site sequence PDCD1_241852953 <400> 68 gggggttcca gggcctgtct 20 <210> 69 <211> 20 <212> DNA <213> Artificial sequence <220> <223> sgRNA target site sequence CTNNB1 <400> 69 ctggactctg gaatccattc 20 <210> 70 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Site 2 Forward Primer <400> 70 tggcccttca agttactgca 20 <210> 71 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Site 2 reverse primer <400> 71 agcacatgac agttaaggtt tgt 23 <210> 72 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Site 3 forward primer <400> 72 aaacgcccat gcaattagtc 20 <210> 73 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Site 3 reverse primer <400> 73 agcccctgtc taggaaaagc 20 <210> 74 <211> twenty four <212> DNA <213> Artificial sequence <220> <223> CTNNB1 forward primer <400> 74 caatgggtca tatcacagat tctt 24 <210> 75 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> CTNNB1 reverse primer <400> 75 ccagctactt gttcttgagt gaa 23 <210> 76 <211> 20 <212> DNA <213> Artificial sequence <220> <223> EGFR forward primer <400> 76 tcatgcgtct tcacctggaa 20 <210> 77 <211> 20 <212> DNA <213> Artificial sequence <220> <223> EGFR reverse primer <400> 77 cgcacacaca tatccccatg 20 <210> 78 <211> 20 <212> DNA <213> Artificial sequence <220> <223> PCSK9 forward primer <400> 78 cactagcagg gacaaggtgg 20 <210> 79 <211> 20 <212> DNA <213> Artificial sequence <220> <223> PCSK9 reverse primer <400> 79 attcagctca gatggggtgg 20 <210> 80 <211> 20 <212> DNA <213> Artificial sequence <220> <223> FANCF forward primer <400> 80 cgctgggaga ttgacatgca 20 <210> 81 <211> 20 <212> DNA <213> Artificial sequence <220> <223> FANCF reverse primer <400> 81 ctcttgcctc cactggttgt 20 <210> 82 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRAC forward primer <400> 82 acctacccca tccccagaag 20 <210> 83 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRAC reverse primer <400> 83 tccctaaacc ccactcccag 20 <210> 84 <211> twenty one <212> DNA <213> Artificial sequence <220> <223> B2M forward primers <400> 84 tgggtttcat ccatccgaca t 21 <210> 85 <211> twenty one <212> DNA <213> Artificial sequence <220> <223> B2M reverse primer <400> 85 atgggatggg actcattcag g 21 <210> 86 <211> 42 <212> DNA <213> Artificial sequence <220> <223> Synthesized crRNA sequence <220> <221> misc_feature <222> (1)..(20) <223> N represents a 20 nt target-specific spacer sequence. <220> <221> misc_feature <222> (1)..(1) <223> 2'OMe (2'-O-methyl nucleotide) and thiophosphate-modified residues <220> <221> misc_feature <222> (2)..(2) <223> 2'OMe (2'-O-methyl nucleotide) and thiophosphate-modified residues <400> 86 nnnnnnnnnn nnnnnnnnnn guuuuagagc uaugcuguuu ug 42 <210> 87 <211> 99 <212> DNA <213> Artificial sequence <220> <223> Synthesize 1xMS2 tracrRNA-aptamer sequence <220> <221> misc_feature <222> (97)..(98) <223> 2'OMe (m) and thiophosphate (*) modified residues <220> <221> misc_feature <222> (97)..(98) <223> 2'OMe and thiophosphate modified residues <400> 87 aacagcauag caaguuaaaa uaaggcuagu ccguuaucaa cuugaaaaag uggcaccgag 60 ucggugcgcg cacaugagga ucaccccaugu gcuuuuuuu 99 <210> 88 <211> 141 <212> DNA <213> Artificial sequence <220> <223> Synthesize 2xMS2 tracrRNA-aptamer sequence <220> <221> misc_feature <222> (139) (140) <223> 2'OMe and phosphorothioate modified residue <220> <221> misc_feature <222> (139) (140) <223> 2'OMe and thiophosphate modified residues <400> 88 aacagcauag caaguuaaaa uaaggcuagu ccguuaucaa cuugaaaaag uggcaccgag 60 ucggugcggg agcacaugag gaucacccau gugccacgag cgacaugagg aucaccaug 120 ucgcucgugu ucccuuuuuu u 141 <210> 89 <211> 167 <212> DNA <213> Artificial sequence <220> <223> Lentiviral sgRNA sequence <220> <221> misc_feature <222> (1)..(20) <223> n's represents a 20-base target-specific sequence. <220> <221> misc_feature <222> (1)..(2) <223> Residues 1 and 2 are modified with thiophosphate. <220> <221> misc_feature <222> (1)..(2) <223> r-bases 1 and 2 are 2'OME modified. <220> <221> misc_feature <222> (168) (169) <223> Residues 168 and 169 are modified with thiophosphate. <220> <221> misc_feature <222> (168) (169) <223> Residues 168 and 169 are 2'OME modified. <400> 89 nnnnnnnnnn nnnnnnnnnn guuuuagagc uagaaauagc aaguuaaaau aaggcuaguc 60 cguuaucaac uugaaaaagu ggcaccgagu cggugcggga gcacaugagg aucacccaug 120 ugccacgagc gacaugagga ucaccccaugu cgcucguguu cccuuuu 167 <210> 90 <211> 20 <212> DNA <213> Artificial sequence <220> <223> B2M_1 sgRNA guide sequence <400> 90 cacagcccaa gatagttaag 20 <210> 91 <211> 20 <212> DNA <213> Artificial sequence <220> <223> B2M_2 sgRNA guide sequence <400> 91 acagcccaag atagttaagt 20 <210> 92 <211> 20 <212> DNA <213> Artificial sequence <220> <223> B2M_3 sgRNA guide sequence <400> 92 ttaccccact taactatctt 20 <210> 93 <211> 20 <212> DNA <213> Artificial sequence <220> <223> B2M_4 sgRNA guide sequence <400> 93 cttaccccac ttaactatct 20 <210> 94 <211> 20 <212> DNA <213> Artificial sequence <220> <223> B2M_5 sgRNA guide sequence <400> 94 actcacgctg gatagcctcc 20 <210> 95 <211> 20 <212> DNA <213> Artificial sequence <220> <223> B2M_6 sgRNA guide sequence <400> 95 ttggagtacc tgaggaatat 20 <210> 96 <211> 20 <212> DNA <213> Artificial sequence <220> <223> B2M_7 sgRNA guide sequence <400> 96 tcgatctatg aaaaagacag 20 <210> 97 <211> 20 <212> DNA <213> Artificial sequence <220> <223> B2M_8 sgRNA guide sequence <400> 97 aacctgaaaa gaaaagaaaa 20 <210> 98 <211> 20 <212> DNA <213> Artificial sequence <220> <223> CD52_1 sgRNA guide sequence <400> 98 gtacaggtaa gagcaacgcc 20 <210> 99 <211> 20 <212> DNA <213> Artificial sequence <220> <223> CD52_2 sgRNA guide sequence <400> 99 ctcctcctac agatacaaac 20 <210> 100 <211> 20 <212> DNA <213> Artificial sequence <220> <223> CD52_3 sgRNA guide sequence <400> 100 cagatacaaa ctggactctc 20 <210> 101 <211> 20 <212> DNA <213> Artificial sequence <220> <223> CD52_4 sgRNA guide sequence <400> 101 ctcttacctg taccataacc 20 <210> 102 <211> 20 <212> DNA <213> Artificial sequence <220> <223> CD52_5 sgRNA guide sequence <400> 102 gtatctgtag gaggagaagt 20 <210> 103 <211> 20 <212> DNA <213> Artificial sequence <220> <223> CD52_6 sgRNA guide sequence <400> 103 tgtatctgta ggaggagaag 20 <210> 104 <211> 20 <212> DNA <213> Artificial sequence <220> <223> CD52_7 sgRNA guide sequence <400> 104 gtccagtttg tatctgtagg 20 <210> 105 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRAC_1 sgRNA guide sequence <400> 105 aacaaatgtg tcacaaagta 20 <210> 106 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRAC_2 sgRNA guide sequence <400> 106 cttcttcccc agcccaggta 20 <210> 107 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRAC_3 sgRNA guide sequence <400> 107 ttcttcccca gcccaggtaa 20 <210> 108 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRAC_4 sgRNA guide sequence <400> 108 agcccaggta agggcagctt 20 <210> 109 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRAC_5 sgRNA guide sequence <400> 109 tttcaaaacc tgtcagtgat 20 <210> 110 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRAC_6 sgRNA guide sequence <400> 110 ttcaaaacct gtcagtgatt 20 <210> 111 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRAC_7 sgRNA guide sequence <400> 111 ccgaatcctc ctcctgaaag 20 <210> 112 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRAC_8 sgRNA guide sequence <400> 112 cttacctggg ctggggaaga 20 <210> 113 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRAC_9 sgRNA guide sequence <400> 113 ttcgtatctg taaaaccaag 20 <210> 114 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRBC1 / 2_1 sgRNA guide sequence <400> 114 ccacacccaa aaggccacac 20 <210> 115 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRBC1 / 2_2 sgRNA guide sequence <400> 115 cccaccagct cagctccacg 20 <210> 116 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRBC1 / 2_3 sgRNA guide sequence <400> 116 cgctgtcaag tccagttcta 20 <210> 117 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRBC1 / 2_4 sgRNA guide sequence <400> 117 gctgtcaagt ccagttctac 20 <210> 118 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRBC1 / 2_5 sgRNA guide sequence <400> 118 agtccagttc tacgggctct 20 <210> 119 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRBC1 / 2_6 sgRNA guide sequence <400> 119 cacccagatc gtcagcgccg 20 <210> 120 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRBC1 / 2_7 sgRNA guide sequence <400> 120 acctgctcta ccccaggcct 20 <210> 121 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRBC1 / 2_8 sgRNA guide sequence <400> 121 ccactcacct gctctacccc 20 <210> 122 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRBC1_1 sgRNA guide sequence <400> 122 cacggacccg cagcccctca 20 <210> 123 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRBC1_2 sgRNA guide sequence <400> 123 gcgggggttc tgccagaagg 20 <210> 124 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRBC1_3 sgRNA guide sequence <400> 124 gttgcggggg ttctgccaga 20 <210> 125 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRBC1_4 sgRNA guide sequence <400> 125 atgacgagtg gacccaggat 20 <210> 126 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRBC1_5 sgRNA guide sequence <400> 126 tgacgagtgg acccaggata 20 <210> 127 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRBC1_6 sgRNA guide sequence <400> 127 acctgctcta ccccaggcct 20 <210> 128 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRBC1_7 sgRNA guide sequence <400> 128 ccaacagtgt cctaccagca 20 <210> 129 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRBC1_8 sgRNA guide sequence <400> 129 caacagtgtc ctaccagcaa 20 <210> 130 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRBC1_9 sgRNA guide sequence <400> 130 aacagtgtcc taccagcaag 20 <210> 131 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRBC1_10 sgRNA guide sequence <400> 131 gtctgaaaga aagcagggag 20 <210> 132 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRBC1_11 sgRNA guide sequence <400> 132 ccacagtctg aaagaaagca 20 <210> 133 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRBC1_12 sgRNA guide sequence <400> 133 gccacagtct gaaagaaagc 20 <210> 134 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRBC1_13 sgRNA guide sequence <400> 134 gacactgttg gcacggagga 20 <210> 135 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRBC1_14 sgRNA guide sequence <400> 135 gtaggacact gttggcacgg 20 <210> 136 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRBC1_15 sgRNA guide sequence <400> 136 taccatggcc atcaacacaa 20 <210> 137 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRBC1_16 sgRNA guide sequence <400> 137 ttaccatggc catcaacaca 20 <210> 138 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRBC2_1 sgRNA guide sequence <400> 138 ccagctcagc tccacgtggt 20 <210> 139 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRBC2_2 sgRNA guide sequence <400> 139 cacagacccg cagcccctca 20 <210> 140 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRBC2_3 sgRNA guide sequence <400> 140 gcgggggttc tgccagaagg 20 <210> 141 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRBC2_4 sgRNA guide sequence <400> 141 gttgcggggg ttctgccaga 20 <210> 142 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRBC2_5 sgRNA guide sequence <400> 142 atgacgagtg gacccaggat 20 <210> 143 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRBC2_6 sgRNA guide sequence <400> 143 tgacgagtgg acccaggata 20 <210> 144 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRBC2_7 sgRNA guide sequence <400> 144 acctgctcta ccccaggcct 20 <210> 145 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRBC2_8 sgRNA guide sequence <400> 145 tcaacagagt cttaccagca 20 <210> 146 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRBC2_9 sgRNA guide sequence <400> 146 caacagagtc ttaccagcaa 20 <210> 147 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRBC2_10 sgRNA guide sequence <400> 147 aacagagtct taccagcaag 20 <210> 148 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRBC2_11 sgRNA guide sequence <400> 148 cacagtctga aagaaaacag 20 <210> 149 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRBC2_12 sgRNA guide sequence <400> 149 ccacagtctg aaagaaaaca 20 <210> 150 <211> 20 <212> DNA <213> Artificial sequence <220> <223> TRBC2_13 sgRNA guide sequence <400> 150 gccacagtct gaaagaaaac 20 <210> 151 <211> 20 <212> DNA <213> Artificial sequence <220> <223> PDCD1_1 sgRNA guide sequence <400> 151 tccaggcatg cagatcccac 20 <210> 152 <211> 20 <212> DNA <213> Artificial sequence <220> <223> PDCD1_2 sgRNA guide sequence <400> 152 tgcagatccc acaggcgccc 20 <210> 153 <211> 20 <212> DNA <213> Artificial sequence <220> <223> PDCD1_3 sgRNA guide sequence <400> 153 cgactggcca gggcgcctgt 20 <210> 154 <211> 20 <212> DNA <213> Artificial sequence <220> <223> PDCD1_4 sgRNA guide sequence <400> 154 acgactggcc agggcgcctg 20 <210> 155 <211> 20 <212> DNA <213> Artificial sequence <220> <223> PDCD1_5 sgRNA guide sequence <400> 155 accgcccaga cgactggcca 20 <210> 156 <211> 20 <212> DNA <213> Artificial sequence <220> <223> PDCD1_6 sgRNA guide sequence <400> 156 caccgcccag acgactggcc 20 <210> 157 <211> 20 <212> DNA <213> Artificial sequence <220> <223> PDCD1_7 sgRNA guide sequence <400> 157 tgtagcaccg cccagacgac 20 <210> 158 <211> 20 <212> DNA <213> Artificial sequence <220> <223> PDCD1_8 sgRNA guide sequence <400> 158 gggcggtgct acaactgggc 20 <210> 159 <211> 20 <212> DNA <213> Artificial sequence <220> <223> PDCD1_9 sgRNA guide sequence <400> 159 cggtgctaca actgggctgg 20 <210> 160 <211> 20 <212> DNA <213> Artificial sequence <220> <223> PDCD1_10 sgRNA guide sequence <400> 160 ctacaactgg gctggcggcc 20 <210> 161 <211> 20 <212> DNA <213> Artificial sequence <220> <223> PDCD1_11 sgRNA guide sequence <400> 161 cacctaccta agaaccatcc 20 <210> 162 <211> 20 <212> DNA <213> Artificial sequence <220> <223> =PDCD1_12 sgRNA guide sequence <400> 162 ggggttccag ggcctgtctg 20 <210> 163 <211> 20 <212> DNA <213> Artificial sequence <220> <223> PDCD1_13 sgRNA guide sequence <400> 163 gggggttcca gggcctgtct 20 <210> 164 <211> 20 <212> DNA <213> Artificial sequence <220> <223> PDCD1_14 sgRNA guide sequence <400> 164 ggggggttcc agggcctgtc 20 <210> 165 <211> 20 <212> DNA <213> Artificial sequence <220> <223> PDCD1_15 sgRNA guide sequence <400> 165 cagcaaccag acggacaagc 20 <210> 166 <211> 20 <212> DNA <213> Artificial sequence <220> <223> PDCD1_16 sgRNA guide sequence <400> 166 cccgaggacc gcagccagcc 20 <210> 167 <211> 20 <212> DNA <213> Artificial sequence <220> <223> PDCD1_17 sgRNA guide sequence <400> 167 ggaccgcagc cagcccggcc 20 <210> 168 <211> 20 <212> DNA <213> Artificial sequence <220> <223> PDCD1_18 sgRNA guide sequence <400> 168 cgtgtcacac aactgcccaa 20 <210> 169 <211> 20 <212> DNA <213> Artificial sequence <220> <223> PDCD1_19 sgRNA guide sequence <400> 169 gtgtcacaca actgcccaac 20 <210> 170 <211> 20 <212> DNA <213> Artificial sequence <220> <223> PDCD1_20 sgRNA guide sequence <400> 170 cgcagatcaa agagagcctg 20 <210> 171 <211> 20 <212> DNA <213> Artificial sequence <220> <223> PDCD1_21 sgRNA guide sequence <400> 171 gcagatcaaa gagagcctgc 20 <210> 172 <211> 20 <212> DNA <213> Artificial sequence <220> <223> PDCD1_22 sgRNA guide sequence <400> 172 agccggccag ttccaaaccc 20 <210> 173 <211> 20 <212> DNA <213> Artificial sequence <220> <223> PDCD1_23 sgRNA guide sequence <400> 173 cggccagttc caaaccctgg 20 <210> 174 <211> 20 <212> DNA <213> Artificial sequence <220> <223> PDCD1_24 sgRNA guide sequence <400> 174 cagttccaaa ccctggtggt 20 <210> 175 <211> 20 <212> DNA <213> Artificial sequence <220> <223> PDCD1_25 sgRNA guide sequence <400> 175 ggacccagac tagcagcacc 20 <210> 176 <211> 20 <212> DNA <213> Artificial sequence <220> <223> PDCD1_26 sgRNA guide sequence <400> 176 cacctaccta agaaccatcc 20 <210> 177 <211> 20 <212> DNA <213> Artificial sequence <220> <223> PDCD1_27 sgRNA guide sequence <400> 177 ggagtctgag agatggagag 20 <210> 178 <211> 20 <212> DNA <213> Artificial sequence <220> <223> PDCD1_28 sgRNA guide sequence <400> 178 tctggaaggg cacaaaggtc 20 <210> 179 <211> 20 <212> DNA <213> Artificial sequence <220> <223> PDCD1_29 sgRNA guide sequence <400> 179 ttctctctgg aagggcacaa 20 <210> 180 <211> 20 <212> DNA <213> Artificial sequence <220> <223> PDCD1_30 sgRNA guide sequence <400> 180 tgacgttacc tcgtgcggcc 20 <210> 181 <211> 20 <212> DNA <213> Artificial sequence <220> <223> PDCD1_31 sgRNA guide sequence <400> 181 tccctgcaga gaaacacact 20 <210> 182 <211> 20 <212> DNA <213> Artificial sequence <220> <223> PDCD1_32 sgRNA guide sequence <400> 182 gagactcacc aggggctggc 20 <210> 183 <211> 20 <212> DNA <213> Artificial sequence <220> <223> PDCD1_33 sgRNA guide sequence <400> 183 tctttgagga gaaagggaga 20 <210> 184 <211> 20 <212> DNA <213> Artificial sequence <220> <223> PDCD1_34 sgRNA guide sequence <400> 184 ttctttgagg agaaagggag 20 <210> 185 <211> 99 <212> RNA <213> Artificial sequence <220> <223> 1x MS2_3 tracrRNA (F-5) <220> <221> misc_feature <222> (77)..(77) <223> x = 2AdP (2-aminopurine) <220> <221> misc_feature <222> (97)..(97) <223> U is modified with 2'OME and thiophosphate. <220> <221> misc_feature <222> (98)..(98) <223> U is modified with 2'OME and thiophosphate. <400> 185 aacagcauag caaguuaaaa uaaggcuagu ccguuaucaa cuugaaaaag uggcaccgag 60 ucggugcgcg gcccggngga ucaccacggg ccuuuuuuu 99 <210> 186 <211> 1981 <212> PRT <213> Artificial sequence <220> <223> Protein sequences of RNA scaffold-mediated recruitment systems (2) xUGI) <220> <221> MISC_FEATURE <222> (1)..(7) <223> Nuclear Localization Signal (NLS) <220> <221> MISC_FEATURE <222> (8)..(235) <223> APOBEC1 <220> <221> MISC_FEATURE <222> (236) (261) <223> connector <220> <221> MISC_FEATURE <222> (262) (378) <223> MCP <220> <221> MISC_FEATURE <222> (382) (399) <223> T2A peptide <220> <221> MISC_FEATURE <222> (402)..(408) <223> Nuclear localization signal (NLS) <220> <221> MISC_FEATURE <222> (409)..(1775) <223> nCAS9D10A <220> <221> MISC_FEATURE <222> (1786)..(1868) <223> UGI <220> <221> MISC_FEATURE <222> (1879)..(1961) <223> UGI <220> <221> MISC_FEATURE <222> (1976)..(1982) <223> Nuclear localization signal (NLS) <400> 186 Lys Lys Lys Arg Lys Val Met Ser Ser Glu Thr Gly Pro Val Ala Val 1 5 10 15 Asp Pro Thr Leu Arg Arg Arg Ile Glu Pro His Glu Phe Glu Val Phe 20 25 30 Phe Asp Pro Arg Glu Leu Arg Lys Glu Thr Cys Leu Leu Tyr Glu Ile 35 40 45 Asn Trp Gly Gly Arg His Ser Ile Trp Arg His Thr Ser Gln Asn Thr 50 55 60 Asn Lys His Val Glu Val Asn Phe Ile Glu Lys Phe Thr Thr Glu Arg 65 70 75 80 Tyr Phe Cys Pro Asn Thr Arg Cys Ser Ile Thr Trp Phe Leu Ser Trp 85 90 95 Ser Pro Cys Gly Glu Cys Ser Arg Ala Ile Thr Glu Phe Leu Ser Arg 100 105 110 Tyr Pro His Val Thr Leu Phe Ile Tyr Ile Ala Arg Leu Tyr His His 115 120 125 Ala Asp Pro Arg Asn Arg Gln Gly Leu Arg Asp Leu Ile Ser Ser Gly 130 135 140 Val Thr Ile Gln Ile Met Thr Glu Gln Glu Ser Gly Tyr Cys Trp Arg 145 150 155 160 Asn Phe Val Asn Tyr Ser Pro Ser Asn Glu Ala His Trp Pro Arg Tyr 165 170 175 Pro His Leu Trp Val Arg Leu Tyr Val Leu Glu Leu Tyr Cys Ile Ile 180 185 190 Leu Gly Leu Pro Pro Cys Leu Asn Ile Leu Arg Arg Lys Gln Pro Gln 195 200 205 Leu Thr Phe Phe Thr Ile Ala Leu Gln Ser Cys His Tyr Gln Arg Leu 210 215 220 Pro Pro His Ile Leu Trp Ala Thr Gly Leu Lys Glu Leu Lys Thr Pro 225 230 235 240 Leu Gly Asp Thr Thr His Thr Ser Pro Pro Cys Pro Ala Pro Glu Leu 245 250 255 Leu Gly Gly Pro Met Ala Ser Asn Phe Thr Gln Phe Val Leu Val Asp 260 265 270 Asn Gly Gly Thr Gly Asp Val Thr Val Ala Pro Ser Asn Phe Ala Asn 275 280 285 Gly Ile Ala Glu Trp Ile Ser Ser Asn Ser Arg Ser Gln Ala Tyr Lys 290 295 300 Val Thr Cys Ser Val Arg Gln Ser Ser Ala Gln Asn Arg Lys Tyr Thr 305 310 315 320 Ile Lys Val Glu Val Pro Lys Gly Ala Trp Arg Ser Tyr Leu Asn Met 325 330 335 Glu Leu Thr Ile Pro Ile Phe Ala Thr Asn Ser Asp Cys Glu Leu Ile 340 345 350 Val Lys Ala Met Gln Gly Leu Leu Lys Asp Gly Asn Pro Ile Pro Ser 355 360 365 Ala Ile Ala Ala Asn Ser Gly Ile Tyr Gly Ser Gly Glu Gly Arg Gly 370 375 380 Ser Leu Leu Thr Cys Gly Asp Val Glu Glu Asn Pro Gly Pro Gly Thr 385 390 395 400 Pro Lys Lys Lys Arg Lys Val Asp Lys Lys Tyr Ser Ile Gly Leu Ala 405 410 415 Ile Gly Thr Asn Ser Val Gly Trp Ala Val Ile Thr Asp Glu Tyr Lys 420 425 430 Val Pro Ser Lys Lys Phe Lys Val Leu Gly Asn Thr Asp Arg His Ser 435 440 445 Ile Lys Lys Asn Leu Ile Gly Ala Leu Leu Phe Asp Ser Gly Glu Thr 450 455 460 Ala Glu Ala Thr Arg Leu Lys Arg Thr Ala Arg Arg Arg Tyr Thr Arg 465 470 475 480 Arg Lys Asn Arg Ile Cys Tyr Leu Gln Glu Ile Phe Ser Asn Glu Met 485 490 495 Ala Lys Val Asp Asp Ser Phe Phe His Arg Leu Glu Glu Ser Phe Leu 500 505 510 Val Glu Glu Asp Lys Lys His Glu Arg His Pro Ile Phe Gly Asn Ile 515 520 525 Val Asp Glu Val Ala Tyr His Glu Lys Tyr Pro Thr Ile Tyr His Leu 530 535 540 Arg Lys Lys Leu Val Asp Ser Thr Asp Lys Ala Asp Leu Arg Leu Ile 545 550 555 560 Tyr Leu Ala Leu Ala His Met Ile Lys Phe Arg Gly His Phe Leu Ile 565 570 575 Glu Gly Asp Leu Asn Pro Asp Asn Ser Asp Val Asp Lys Leu Phe Ile 580 585 590 Gln Leu Val Gln Thr Tyr Asn Gln Leu Phe Glu Glu Asn Pro Ile Asn 595 600 605 Ala Ser Gly Val Asp Ala Lys Ala Ile Leu Ser Ala Arg Leu Ser Lys 610 615 620 Ser Arg Arg Leu Glu Asn Leu Ile Ala Gln Leu Pro Gly Glu Lys Lys 625 630 635 640 Asn Gly Leu Phe Gly Asn Leu Ile Ala Leu Ser Leu Gly Leu Thr Pro 645 650 655 Asn Phe Lys Ser Asn Phe Asp Leu Ala Glu Asp Ala Lys Leu Gln Leu 660 665 670 Ser Lys Asp Thr Tyr Asp Asp Asp Leu Asp Asn Leu Leu Ala Gln Ile 675 680 685 Gly Asp Gln Tyr Ala Asp Leu Phe Leu Ala Ala Lys Asn Leu Ser Asp 690 695 700 Ala Ile Leu Leu Ser Asp Ile Leu Arg Val Asn Thr Glu Ile Thr Lys 705 710 715 720 Ala Pro Leu Ser Ala Ser Met Ile Lys Arg Tyr Asp Glu His His Gln 725 730 735 Asp Leu Thr Leu Leu Lys Ala Leu Val Arg Gln Gln Leu Pro Glu Lys 740 745 750 Tyr Lys Glu Ile Phe Phe Asp Gln Ser Lys Asn Gly Tyr Ala Gly Tyr 755 760 765 Ile Asp Gly Gly Ala Ser Gln Glu Glu Phe Tyr Lys Phe Ile Lys Pro 770 775 780 Ile Leu Glu Lys Met Asp Gly Thr Glu Glu Leu Leu Val Lys Leu Asn 785 790 795 800 Arg Glu Asp Leu Leu Arg Lys Gln Arg Thr Phe Asp Asn Gly Ser Ile 805 810 815 Pro His Gln Ile His Leu Gly Glu Leu His Ala Ile Leu Arg Arg Gln 820 825 830 Glu Asp Phe Tyr Pro Phe Leu Lys Asp Asn Arg Glu Lys Ile Glu Lys 835 840 845 Ile Leu Thr Phe Arg Ile Pro Tyr Tyr Val Gly Pro Leu Ala Arg Gly 850 855 860 Asn Ser Arg Phe Ala Trp Met Thr Arg Lys Ser Glu Glu Thr Ile Thr 865 870 875 880 Pro Trp Asn Phe Glu Glu Val Val Asp Lys Gly Ala Ser Ala Gln Ser 885 890 895 Phe Ile Glu Arg Met Thr Asn Phe Asp Lys Asn Leu Pro Asn Glu Lys 900 905 910 Val Leu Pro Lys His Ser Leu Leu Tyr Glu Tyr Phe Thr Val Tyr Asn 915 920 925 Glu Leu Thr Lys Val Lys Tyr Val Thr Glu Gly Met Arg Lys Pro Ala 930 935 940 Phe Leu Ser Gly Glu Gln Lys Lys Ala Ile Val Asp Leu Leu Phe Lys 945 950 955 960 Thr Asn Arg Lys Val Thr Val Lys Gln Leu Lys Glu Asp Tyr Phe Lys 965 970 975 Lys Ile Glu Cys Phe Asp Ser Val Glu Ile Ser Gly Val Glu Asp Arg 980 985 990 Phe Asn Ala Ser Leu Gly Thr Tyr His Asp Leu Leu Lys Ile Ile Lys 995 1000 1005 Asp Lys Asp Phe Leu Asp Asn Glu Glu Asn Glu Asp Ile Leu Glu 1010 1015 1020 Asp Ile Val Leu Thr Leu Thr Leu Phe Glu Asp Arg Glu Met Ile 1025 1030 1035 Glu Glu Arg Leu Lys Thr Tyr Ala His Leu Phe Asp Asp Lys Val 1040 1045 1050 Met Lys Gln Leu Lys Arg Arg Arg Tyr Thr Gly Trp Gly Arg Leu 1055 1060 1065 Ser Arg Lys Leu Ile Asn Gly Ile Arg Asp Lys Gln Ser Gly Lys 1070 1075 1080 Thr Ile Leu Asp Phe Leu Lys Ser Asp Gly Phe Ala Asn Arg Asn 1085 1090 1095 Phe Met Gln Leu Ile His Asp Asp Ser Leu Thr Phe Lys Glu Asp 1100 1105 1110 Ile Gln Lys Ala Gln Val Ser Gly Gln Gly Asp Ser Leu His Glu 1115 1120 1125 His Ile Ala Asn Leu Ala Gly Ser Pro Ala Ile Lys Lys Gly Ile 1130 1135 1140 Leu Gln Thr Val Lys Val Val Asp Glu Leu Val Lys Val Met Gly 1145 1150 1155 Arg His Lys Pro Glu Asn Ile Val Ile Glu Met Ala Arg Glu Asn 1160 1165 1170 Gln Thr Thr Gln Lys Gly Gln Lys Asn Ser Arg Glu Arg Met Lys 1175 1180 1185 Arg Ile Glu Glu Gly Ile Lys Glu Leu Gly Ser Gln Ile Leu Lys 1190 1195 1200 Glu His Pro Val Glu Asn Thr Gln Leu Gln Asn Glu Lys Leu Tyr 1205 1210 1215 Leu Tyr Tyr Leu Gln Asn Gly Arg Asp Met Tyr Val Asp Gln Glu 1220 1225 1230 Leu Asp Ile Asn Arg Leu Ser Asp Tyr Asp Val Asp His Ile Val 1235 1240 1245 Pro Gln Ser Phe Leu Lys Asp Asp Ser Ile Asp Asn Lys Val Leu 1250 1255 1260 Thr Arg Ser Asp Lys Asn Arg Gly Lys Ser Asp Asn Val Pro Ser 1265 1270 1275 Glu Glu Val Val Lys Lys Met Lys Asn Tyr Trp Arg Gln Leu Leu 1280 1285 1290 Asn Ala Lys Leu Ile Thr Gln Arg Lys Phe Asp Asn Leu Thr Lys 1295 1300 1305 Ala Glu Arg Gly Gly Leu Ser Glu Leu Asp Lys Ala Gly Phe Ile 1310 1315 1320 Lys His Val Ala 1325 1330 1335 Gln Ile Leu Asp Ser Arg Met Thr Lys Tyr Asp Glu Asp 1340 1345 1350 Lys Leu Ile Arg Glu Val Lys Val Ile Thr Leu Lys Ser Lys Leu 1355 1360 1365 Val Ser Asp Phe Arg Lys Asp Phe Gln Phe Tyr Lys Val Arg Glu 1370 1375 1380 Ile Asn Asn Tyr His Ala His Asp Ala Tyr Leu Asn Ala Val 1385 1390 1395 Val Gly Thr Ala Leu Ile Lys Tyr Pro Lys Leu Glu Ser Glu 1400 1405 1410 Phe Val Tyr Gly Asp Tyr Lys Val Tyr Asp Val Arg Lys Met Ile 1415 1420 1425 Only Lys Ser Glu Gln Glu Ile Gly Lys Ala Thr Ala Lys Tyr Phe 1430 1435 1440 Phe Tyr Ser Asn Ile Met Asn Phe Phe Lys Thr Glu Ile Thr Leu 1445 1450 1455 Ala Asn Gly Glu Ile Arg Lys Arg Pro Leu Ile Glu Thr Asn Gly 1460 1465 1470 Glu Thr Gly Glu Ile Val Trp Asp Lys Gly Arg Asp Phe Ala Thr 1475 1480 1485 Val Arg Lys Val Leu Ser Met Pro Gln Val Asn Ile Val Lys Lys 1490 1495 1500 Thr Glu Val Gln Thr Gly Gly Phe Ser Lys Glu Ser Ile Leu Pro 1505 1510 1515 Lys Arg Asn Ser Asp Lys Leu Ile Ala Arg Lys Lys Asp Trp Asp 1520 1525 1530 Pro Lys Lys Tyr Gly Gly Phe Asp Ser Pro Thr Val Ala Tyr Ser 1535 1540 1545 Val Leu Val Val Ala Lys Val Glu Lys Gly Lys Ser Lys Lys Leu 1550 1555 1560 Lys Ser Val Lys Glu Leu Leu Gly Ile Thr Ile Met Glu Arg Ser 1565 1570 1575 Ser Phe Glu Lys Asn Pro Ile Asp Phe Leu Glu Ala Lys Gly Tyr 1580 1585 1590 Lys Glu Val Lys Lys Asp Leu Ile Ile Lys Leu Pro Lys Tyr Ser 1595 1600 1605 Leu Phe Glu Leu Glu Asn Gly Arg Lys Arg Met Leu Ala Ser Ala 1610 1615 1620 Gly Glu Leu Gln Lys Gly Asn Glu Leu Ala Leu Pro Ser Lys Tyr 1625 1630 1635 Val Asn Phe Leu Tyr Leu Ala Ser His Tyr Glu Lys Leu Lys Gly 1640 1645 1650 Ser Pro Glu Asp Asn Glu Gln Lys Gln Leu Phe Val Glu Gln His 1655 1660 1665 Lys His Tyr Leu Asp Glu Ile Ile Glu Gln Ile Ser Glu Phe Ser 1670 1675 1680 Lys Arg Val Ile Leu Ala Asp Ala Asn Leu Asp Lys Val Leu Ser 1685 1690 1695 Ala Tyr Asn Lys His Arg Asp Lys Pro Ile Arg Glu Gln Ala Glu 1700 1705 1710 Asn Ile Ile His Leu Phe Thr Leu Thr Asn Leu Gly Ala Pro Ala 1715 1720 1725 Ala Phe Lys Tyr Phe Asp Thr Thr Ile Asp Arg Lys Arg Tyr Thr 1730 1735 1740 Ser Thr Lys Glu Val Leu Asp Ala Thr Leu Ile His Gln Ser Ile 1745 1750 1755 Thr Gly Leu Tyr Glu Thr Arg Ile Asp Leu Ser Gln Leu Gly Gly 1760 1765 1770 Asp Ser Gly Gly Ser Gly Gly Ser Gly Gly Ser Thr Asn Leu Ser 1775 1780 1785 Asp Ile Ile Glu Lys Glu Thr Gly Lys Gln Leu Val Ile Gln Glu 1790 1795 1800 Ser Ile Leu Met Leu Pro Glu Glu Val Glu Glu Val Ile Gly Asn 1805 1810 1815 Lys Pro Glu Ser Asp Ile Leu Val His Thr Ala Tyr Asp Glu Ser 1820 1825 1830 Thr Asp Glu Asn Val Met Leu Leu Thr Ser Asp Ala Pro Glu Tyr 1835 1840 1845 Lys Pro Trp Ala Leu Val Ile Gln Asp Ser Asn Gly Glu Asn Lys 1850 1855 1860 Ile Lys Met Leu Ser Gly Gly Ser Gly Gly Ser Gly Gly Ser Thr 1865 1870 1875 Asn Leu Ser Asp Ile Ile Glu Lys Glu Thr Gly Lys Gln Leu Val 1880 1885 1890 Ile Gln Glu Ser Ile Leu Met Leu Pro Glu Glu Val Glu Glu Val 1895 1900 1905 Ile Gly Asn Lys Pro Glu Ser Asp Ile Leu Val His Thr Ala Tyr 1910 1915 1920 Asp Glu Ser Thr Asp Glu Asn Val Met Leu Leu Thr Ser Asp Ala 1925 1930 1935 Pro Glu Tyr Lys Pro Trp Ala Leu Val Ile Gln Asp Ser Asn Gly 1940 1945 1950 Glu Asn Lys Ile Lys...

Claims

1. An RNA scaffold comprising: (a) tracrRNA; (b) An RNA motif containing an extended sequence, wherein the RNA motif is an aptamer that recruits effector modules, and wherein the effector is a non-nuclease effector, and wherein, The extended sequence of the RNA motif comprises 2 to 24 nucleotides, and the total length of the RNA motif is 23 to 45 nucleotides; and (c) crRNA containing a guide RNA sequence, and The RNA scaffold therein forms a single RNA molecule.

2. The RNA scaffold according to claim 1, wherein, The RNA scaffold contains one or more modifications.

3. The RNA scaffold according to claim 1 or claim 2, wherein, The RNA motif is attached to the 3' end of the tracrRNA via a linker.

4. The RNA scaffold according to claim 3, wherein, The adapter is a single-stranded RNA or a chemically linked one.

5. The RNA scaffold according to claim 1, wherein, The tracrRNA is fused with the crRNA containing the guide RNA sequence.

6. The RNA scaffold according to claim 1, wherein, The tracrRNA hybridizes with the crRNA through repeat-anti-repeat regions.

7. The RNA scaffold according to claim 6, wherein, The repeating-anti-repetition region is extended.

8. The RNA scaffold according to claim 7, wherein, The repeat-anti-repeat region includes an upper stem that is extended and contains a total length of 20 to 26 nucleotides.

9. The RNA scaffold according to claim 8, wherein, The upper stem of the repeat-anti-repeat region contains a total length of 22 nucleotides.

10. The RNA scaffold according to claim 1, wherein, The RNA scaffold contains one or more RNA motifs.

11. The RNA scaffold according to claim 10, wherein, The one or more RNA motifs contain one or more modifications.

12. The RNA scaffold according to claim 11, wherein, The one or more modifications are located at the 5' end and / or 3' end of one or more RNA motifs.

13. The RNA scaffold according to claim 11, wherein, The one or more modifications involve replacing the A base at position 10 with 2-aminopurine (2AP).

14. The RNA scaffold according to claim 13, wherein, 2-Aminopurine (2AP) is 2'-deoxy-2-aminopurine or 2'-ribose-2-aminopurine.

15. The RNA scaffold according to claim 2, wherein, The one or more modifications target the backbone and / or sugar portion of the RNA scaffold.

16. The RNA scaffold according to claim 1, wherein, The extended sequence of the RNA motif is a double-stranded extension.

17. The RNA scaffold according to claim 1, wherein, The total length of the RNA motif is 23, 29, 35, or 45 nucleotides.

18. The RNA scaffold according to claim 4, wherein, Single-stranded RNA adapters contain 1 to 10 nucleotides, preferably 2 to 6 nucleotides.

19. The RNA scaffold according to claim 10, wherein, One or more RNA motifs bind to aptamer-binding molecules.

20. The RNA scaffold according to claim 10, wherein, The one or more RNA motifs are aptamers selected from the group consisting of MS2, Ku, PP7, SfMu, and Sm7.

21. The RNA scaffold according to claim 20, wherein, The MS2 aptamer binds to the MCP protein.

22. The RNA scaffold according to claim 20, wherein, The MS2 aptamer is wild-type MS2, mutant MS2, or a variant thereof.

23. The RNA scaffold according to claim 22, wherein, The mutant MS2 is a C-5 or F-5 mutant.

24. The RNA scaffold according to claim 1, wherein, The effector submodule comprises: (i) an RNA-binding domain capable of binding to the RNA motif, and (ii) an effector domain.

25. The RNA scaffold according to claim 24, wherein, The effector domains are selected from the group consisting of reporter molecules, tags, molecules, proteins, microparticles, and nanoparticles.

26. The RNA scaffold according to claim 24, wherein, The effector domain is a DNA-modifying enzyme.

27. The RNA scaffold according to claim 26, wherein, The DNA modifying enzyme is selected from the group consisting of AID, CDA, APOBEC1, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D, APOBEC3F or other APOBEC family enzymes, ADA, ADAR family enzymes and tRNA adenosine deaminase.

28. The RNA scaffold according to claim 1, wherein, The RNA scaffold consists of sequences corresponding to any one of SEQ ID NO: 40 to SEQ ID NO:

52.

29. The RNA scaffold according to claim 1, wherein, The RNA motif consists of a sequence selected from any one of SEQ ID NO: 21 to SEQ ID NO: 24.

Citation Information

Patent Citations

  • Fusion protein constructs

    US20100063258A1

  • Crispr-CAS component systems, methods and compositions for sequence manipulation

    US20140179006A1

  • Crispr / CAS systems for genomic modification and gene modulation

    US20140273226A1

  • Crispr-based genome modification and regulation

    US20140273233A1

  • Synthetic polynucleotides

    US3687808A