Fusion protein capable of generating point mutation in cells, preparation and use thereof

By designing fusion proteins, using the fusion protein of the Cas enzyme with missing nuclease activity and the cytosine deaminase AID, combined with sgRNA guidance, efficient and targeted gene mutations are achieved in vivo, solving the problems of low mutation frequency and strong randomness in existing technologies, and meeting the needs of experimental screening of functional mutants.

CN107522787BActive Publication Date: 2025-09-16SHANGHAI INST OF BIOLOGICAL SCI CHINESE ACAD OF SCI

Patent Information

Application Number
CN201710451424.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2016-06-15
Filing Date
2017-06-15
Publication Date
2025-09-16
Estimated Expiration
2037-06-15

AI Technical Summary

Technical Problem

Existing technologies make it difficult to efficiently and targetedly introduce gene mutation frequencies in vivo, especially in immune system B cells, and cannot meet the needs of experimental screening of functional mutants.

Method used

A fusion protein is designed, which contains a Cas enzyme with missing nuclease activity and a cytosine deaminase AID. It is guided to a specific DNA site by sgRNA to deaminize cytosine and randomly mutate it to other bases to achieve site-directed mutagenesis.

Benefits of technology

It increases the frequency of gene mutations, achieves efficient and targeted point mutations, and meets the needs of experimental screening of functional mutants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure HDA0001322562080000011
    Figure HDA0001322562080000011
  • Figure HDA0001322562080000012
    Figure HDA0001322562080000012
  • Figure HDA0001322562080000021
    Figure HDA0001322562080000021
Patent Text Reader

Abstract

The present invention relates to a fusion protein that produces point mutations in cells, its preparation and use. Specifically, the fusion protein provided by the present invention contains a Cas enzyme that lacks cytosine deaminase and nuclease activity and retains helicase activity, or is formed by a Cas enzyme that lacks cytosine deaminase and nuclease activity and retains helicase activity. The present invention also relates to the coding sequence of the fusion protein, a polynucleotide sequence containing the coding sequence, a nucleic acid construct containing the polynucleotide sequence, a corresponding host cell, a method for producing point mutations in a cell, and a kit, etc. By using the present invention, it is possible to achieve site-directed mutagenesis while obtaining high mutation efficiency and a variety of mutation combinations in a specific gene region.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a fusion protein capable of generating point mutations in cells, and its preparation and use. Background Art

[0002] There is a close relationship between genotype and phenotype. In nature, spontaneous mutations can cause changes in genotype, resulting in a variety of phenotypes. In the laboratory, mutations are still used to diversify genes and produce a variety of phenotypes, thereby screening functional mutants, studying the relationship between genes and functions, and obtaining more functional proteins. In nature, the frequency of spontaneous mutations is extremely low. Among common organisms, the spontaneous mutation rate of the human genome is 5.0×10 -10 The spontaneous mutation rate of the mouse genome is 1.8×10 -10 The spontaneous mutation rate of the Escherichia coli genome is 5.4×10 -10 The spontaneous mutation rate of HIV is 3×10 -5 As the size of an organism's genome decreases, the frequency of spontaneous mutation increases〔Holmes E C. The comparative genomics of viral emergence[J]. Proceedings of the National Academy of Sciences, 2010, 107(4): 1742-1746〕. However, this low frequency of gene mutation cannot produce a sufficient number of phenotypes to study the relationship between genes, phenotypes, and functions.

[0003] In order to increase the frequency of gene mutation, the existing methods in the laboratory are mainly divided into in vivo mutation method and in vitro mutation method. In vivo point mutation method: 1. Physical method: ultraviolet radiation, the mutation frequency is 1×10 -10 〔Packer MS, Liu DR. Methods for the directed evolution of proteins[J]. Nature Reviews Genetics, 2015〕. 2. Chemical method: ENU is an alkylating agent that transfers ethyl groups to oxygen and nitrogen atoms in DNA, causing mismatches, base substitutions, or deletions, with a mutation frequency of 1-1.5×10 -5〔FILBY.ZEBRAFISH:METHODS ANDPROTOCOLS.METHODS IN MOLECULAR BIOLOGY-By GJLieschke,AC Oates andK.Kawakami.[J].Journal of Fish Biology,2010,76(7):1874-1876〕. Although ENU is easily available, it is very sensitive to light, heat, and pH, which limits its application. Both methods can change the mutation frequency by dosage, but the point mutations induced are random, the mutation frequency is low, the mutation pattern is uneven, and it is harmful to the organism〔Guénet JL.Chemical mutagenesis of the mouse genome:an overview[J].Genetica,2004,122(1):9-24〕. 3. Biological methods: Transposons are basic units on chromosomal DNA that can replicate and shift autonomously. They can cause insertional mutations, leading to gene knockout and gene activation through gene insertion. Different insertion sites can be selected by choosing different vectors. However, their mutation rate is lower than that of ENU, and only 3×10 -5 Insertion event, and the host needs to express transposase at the same time to complete transposition〔Kitada K, Ishishita S, Tosaka K, et al. Transposon-tagged mutagenesis in the rat. [J]. Nature Methods, 2007, 4(2): 131-133〕.

[0004] In the immune system, germinal center B cells can produce diverse antibodies through somatic hypermutation to resist pathogen invasion [Odegard VH, Schatz DG. Targeting of somatic hypermutation. [J]. Nature Reviews Immunology, 2006, 6(8): 573-583]. Somatic hypermutation refers to non-templated point mutations in the variable regions of immunoglobulin heavy and light chains, which are related to B cell affinity maturation [Odegard VH et al., supra]. The enzyme that mediates this process is activation-induced cytosine deaminase (AID). AID is a cytosine deaminase that belongs to the APOBEC family, a family of RNA editing enzymes. It has a nuclear localization signal at the N-terminus and a nuclear export signal at the C-terminus. Its catalytic domain is shared by the APOBEC family [Zhenming X, Hong Z, Pone EJ, et al. Immunoglobulin class-switch DNA recombination: induction, targeting and beyond. [J]. Nature Reviews Immunology, 2012, 12(7): 517-31]. It is generally believed that the N-terminal structure is required for SHM. AID expression is restricted to B cells in the germinal center. Its point mutation function is conditional, requiring it to act on single-stranded DNA and exhibiting sequence preference. Its hotspot domain is RGYW [Kiyotsugu Y, Il-Mi O, Tomonori E, et al. AID Enzyme-Induced Hypermutation in an Actively Transcribed Gene in Fibroblasts [J]. Science, 2002, 296(5575):2033-2036]. R represents A / G, Y represents C / T, and W represents A / T. This indicates that AID's function is related to the primary structure of DNA. First, it deaminates cytosine on single-stranded DNA to U, forming a UG mismatch. If the UG is not repaired, a CT GA transition mutation will occur during DNA replication. Furthermore, the U can be removed by UNG (uracil DNA glycosidase), creating an apyrimidine site, allowing the random incorporation of four bases [Odegard VH et al., supra]. The point mutations generated by the above process are of great significance for the high-frequency mutation of somatic cells and can produce diverse antibodies. However, the frequency of point mutations caused in vivo is 1×10 -4 -1×10 -3, and the sites are random〔Masatoshi A,Nesreen H,Andre S,et al.Accumulation of the FACT complex,as well as histone H3.3,serves as atarget marker for somatic hypermutation.[J].Proceedings of the NationalAcademy of Sciences of the United States of America,2013,110(19):7784-7789〕, which still cannot meet the needs of experimental screening of mutants. Summary of the Invention

[0005] The first aspect of the present invention provides a fusion protein comprising a Cas enzyme that lacks cytosine deaminase and nuclease activities but retains helicase activity.

[0006] In one or more embodiments, the fusion protein is formed by a Cas enzyme that lacks cytosine deaminase and nuclease activities but retains helicase activity.

[0007] In one or more embodiments, the Cas enzyme is selected from the group consisting of: Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, homologs thereof, or modified forms thereof.

[0008] In one or more embodiments, the nuclease activity of the Cas enzyme is partially deleted, so that the Cas enzyme can only cause single-strand breaks of DNA; or the nuclease activity of the Cas enzyme is completely deleted, which can cause double-strand breaks of DNA.

[0009] In one or more embodiments, the Cas enzyme is a Cas9 enzyme selected from: Cas9 from Streptococcus pyogenes (SpCas9), Cas9 from Staphylococcus aureus (SaCas9), and Cas9 from Streptococcus thermophilus (St1Cas9).

[0010] In one or more embodiments, the Cas enzyme is a Cas9 enzyme, and the two endonuclease catalytic domains RuvC1 and / or HNH of the enzyme are mutated, resulting in the loss of nuclease activity of the enzyme and retention of helicase activity.

[0011] In one or more embodiments, both RuvC1 and HNH of the Cas9 enzyme are mutated, resulting in the loss of nuclease activity but retention of helicase activity.

[0012] In one or more embodiments, the 10th amino acid asparagine of the Cas9 enzyme is mutated to alanine or other amino acids, and the 841st amino acid histidine is mutated to alanine or other amino acids.

[0013] In one or more embodiments, the amino acid sequence of the Cas9 enzyme is as shown in SEQ ID NO:2, amino acid residues 42-1452, or as shown in SEQ ID NO:72, amino acid residues 42-1419.

[0014] In one or more embodiments, the CDase is a full-length CDase or a fragment thereof, wherein the fragment comprises at least the NLS domain, the catalytic domain, and the APOBEC-like domain of CDase.

[0015] In one or more embodiments, the cytosine deaminase has substitution mutations at amino acid residues 10, 82, and 156.

[0016] In one or more embodiments, the substitution mutations are K10E, T82I, and E156G.

[0017] In one or more embodiments, the fragment comprises at least amino acid residues 9 to 182 of AID, for example, at least amino acid residues 1 to 182 of AID.

[0018] In one or more embodiments, the amino acid sequence of the CDase is as shown in amino acids 1457-1654 of SEQ ID NO: 2, or as shown in amino acid residues 1447-1629 of SEQ ID NO: 68.

[0019] In one or more embodiments, the fragment comprises at least amino acid residues 1465-1638 of SEQ ID NO: 2, for example, at least amino acid residues 1457-1638 of SEQ ID NO: 2.

[0020] In one or more embodiments, the fragment consists of amino acid residues 1-182, consists of amino acid residues 1-186, or consists of amino acid residues 1-190.

[0021] In one or more embodiments, the fusion protein further comprises one or more of the following sequences: a linker, a nuclear localization sequence, and amino acid residues or amino acid sequences introduced to construct a fusion protein, promote the expression of a recombinant protein, obtain a recombinant protein that is automatically secreted outside the host cell, or facilitate the purification of a recombinant protein.

[0022] In one or more embodiments, the amino acid sequence of the fusion protein is as shown in SEQ ID NO:2, 4, 66, 68, 70 or 72, or as shown in amino acids 26-1654 of SEQ ID NO:2, or as shown in amino acids 26-1638 of SEQ ID NO:4, or as shown in amino acids 26-1629 of SEQ ID NO:68, or as shown in amino acids 26-1629 of SEQ ID NO:70, or as shown in amino acids 26-1638 of SEQ ID NO:72.

[0023] A second aspect of the present invention provides a polynucleotide sequence selected from:

[0024] (1) a polynucleotide sequence encoding the fusion protein described in the first aspect of this invention; and

[0025] (2) A complementary sequence of the sequence described in (1).

[0026] The third aspect of the present invention provides a nucleic acid construct comprising the polynucleotide sequence described in the second aspect of this invention.

[0027] In one or more embodiments, the nucleic acid construct is an expression vector for expressing the fusion protein described herein in a host cell.

[0028] The fourth aspect of the present invention provides a host cell, wherein the host cell contains the fusion protein, its coding sequence or nucleic acid construct described herein.

[0029] The fifth aspect of the present invention provides a method for generating a point mutation in a cell, comprising the step of expressing the fusion protein and sgRNA described herein in the cell.

[0030] In one or more embodiments, the method includes the steps of transferring the fusion protein or its expression vector and the sgRNA or its expression vector described herein into the cell, and then screening to obtain the desired mutant nucleic acid sequence.

[0031] In one or more embodiments, the sgRNA includes a target binding region and a Cas protein recognition region, wherein the target binding region can specifically bind to the nucleic acid sequence to be mutated, and the Cas protein recognition region can be recognized and bound by the Cas enzyme in the fusion protein.

[0032] In one or more embodiments, the target binding region of the sgRNA specifically binds to the template strand of the nucleic acid sequence to be mutated, and the region opposite the sgRNA binding region on the template strand is adjacent to the protospacer sequence motif recognized by the Cas protein, or is separated by no more than 10 bases.

[0033] In one or more embodiments, the gene to be mutated encodes a functional protein.

[0034] In one or more embodiments, the functional proteins include proteins involved in the occurrence, development and metastasis of diseases, proteins involved in cell differentiation, proliferation and apoptosis, proteins involved in metabolism, development-related proteins, various drug targets, etc.

[0035] In one or more embodiments, the functional protein is selected from the group consisting of antibodies, enzymes, lipoproteins, hormone-like proteins, transport and storage proteins, motor proteins, receptor proteins, and membrane proteins.

[0036] A sixth aspect of the present invention provides a kit comprising the fusion protein, polynucleotide sequence or nucleic acid construct described herein.

[0037] A seventh aspect of the present invention provides the use of the fusion protein, polynucleotide sequence or nucleic acid construct described herein in generating point mutations in cells, or in preparing a composition or kit for generating point mutations in cells. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 : A and C show PCR-amplified AID (lane 1) and AIDX fragments (lane 1), respectively; B shows an agarose gel image of the pEntr11-dCas9-AID plasmid, where lane 1 is the pEntr11 empty plasmid, lane 2 is the pEntr11-dCas9 plasmid, and lanes 3-7 are the pEntr11-dCas9-AID plasmids; D shows the PCR results of the pEntr11-dCas9-AIDX plasmid in culture, where the amplified fragment is AIDX. Lanes 1-5 in D represent five different positive clones, respectively; lane 6 is the empty plasmid as a negative control.

[0039] Figure 2: A, Lanes 1 and 2 are dCas9-AID and dCas9-AIDX fragments amplified by PCR, respectively; B, enzyme digestion of MO91 empty plasmid, where lane 1 is single digestion with BglⅡ, lane 2 is MO91 empty plasmid, and lane 3 is double digestion with BglⅡ and XhoⅠ; C, PCR results of MO91-dCas9-AIDX plasmid in bacterial solution, the amplified fragment is AIDX; D, PCR results of MO91-dCas9-AID plasmid in bacterial solution, the amplified fragment is AID.

[0040] Figure 3 : A, lane 1 is the 3*flag+NLS fragment amplified by PCR, lanes 2 and 3 are BglⅡ single-enzyme digested MO91-dCas9-AID plasmid and MO91-dCas9-AIDX plasmid, respectively, and lane 4 is the MO91-dCas9-AID plasmid control; B, lanes 1-4 are MO91-dCas9 (3*flag, NLS)-AID plasmid, lane 5 is MO91-dCas9-AID plasmid, and lanes 6-9 are MO91-dCas9 (3*flag, NLS)-AIDX plasmid.

[0041] Figure 4 : Sequence of the EGFP reporter, with the stop codon in bold. The designed sgRNA is indicated by an arrow.

[0042] Figure 5 : Schematic diagram of the reporter plasmid.

[0043] Figure 6 Flow cytometry analysis of reporter cell lines. The three curves (left to right) represent the Thy1.1 expression levels of unstained control, reporter-negative cells, and reporter-positive cells, respectively.

[0044] Figure 7 : Comparison of dCas9-AID, dCas9-AIDX, AID and AIDX point mutation efficiency in reporter cells.

[0045] Figure 8 Optimizing dCas9-AID point mutation efficiency in reporter cells. A, dCas9-AID induces GFP expression; B, Schematic diagram of different AID variants and their efficiency in inducing point mutations; C, dCas9-AIDX-induced point mutations require the cytosine deaminase activity of AID.

[0046] Figure 9 : Frequency distribution of point mutations caused by dCas9-AIDX and AID on EGFP and cMyc genes.

[0047] Figure 10dCas9-AIDX randomly mutates C and G bases into three other bases. A, Statistics of base mutation types; B, Mechanism of dCas9-AIDX-induced point mutations.

[0048] Figure 11 :UGI increases the base substitution frequency of the dCas9-AIDX system, reveals the action trajectory of dCas9-AIDX on genes, and makes the direction of base mutation more unified.

[0049] Figure 12 : dCas9-AIDX can not only act on exogenous genes, but also on endogenous genes.

[0050] Figure 13 : Structural and functional domains of AID.

[0051] Figure 14 : Experimental process (a) and results (bd) of applying dCas9-AIDX to screen for Gleevec resistance of K562 BCR-ABL gene.

[0052] Figure 15 :TAM (targeted cytosine deaminase AID-mediated gene mutagenesis technology) mutated the amino acids in the variable region of anti-HEL-IgG1.

[0053] Figure 16 : TAM induces base mutations in the variable region of anti-HEL-IgG1 (upper panel) and reproducibly induces base mutations in IgG1 CDR (lower panel).

[0054] Figure 17 : The affinity of the mutated antibody to HEL was enhanced by more than 10 times.

[0055] Figure 18 Figure 3: Expression results of nCas9-AIDX in bacteria. The boxed bands represent the nCas9-AIDX fusion protein.

[0056] Figure 19 Functional testing results of different fusion proteins. For each data set, the three columns from left to right represent the results of MO91-AIDX-XTEN-dCas9, MO91-dCas9-XTEN-AIDX, and MO91-dCas9-AIDX, respectively.

[0057] Figure 20 Functional testing results of different fusion proteins. For each data set, the three columns from left to right represent the results of MO91-dCas9-AIDX, MO91-dCas9-XTEN-AIDX(K10E T82I E156G), and MO91-dCas9-XTEN-AIDX, respectively.

[0058] Figure 21 : Functional verification results of nCas9-AIDX fusion protein. DETAILED DESCRIPTION

[0059] This paper involves fusion proteins of nuclease-deficient Cas proteins with the cytosine deaminase AID or its mutants. Under the guidance of sgRNA, the fusion protein is recruited to specific DNA sequences, where AID or its mutants deaminate cytosine to produce uracil. This uracil is then randomly mutated into other bases during the DNA repair process, achieving both site-directed mutagenesis and high mutation efficiency.

[0060] For details about Cas / sgRNA, in addition to those described below, please refer to CN 201380049665.5 and CN201380072752.2, the entire contents of which are incorporated herein by reference.

[0061] Cas proteins

[0062] CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) is a gene-editing system that bacteria use to defend against viral invasion and evade the mammalian immune response. This system has been modified and optimized and is now widely used in in vitro biochemical reactions, as well as in gene editing in cells and individuals.

[0063] Typically, a complex formed by a Cas protein with endonuclease activity and its specifically recognized sgRNA complements the template strand in the target DNA through the sgRNA's pairing region, and Cas cuts the double-stranded DNA at a specific position. It should be understood that herein, "Cas protein" and "Cas enzyme" are used interchangeably.

[0064] This article utilizes the above-mentioned characteristics of Cas / sgRNA, that is, the specific binding of sgRNA to the target is utilized to locate Cas to the desired position, at which position the cytosine is deaminated by AID or its mutant in the fusion protein. The Cas protein suitable for the present invention that partially or completely lacks nuclease activity, especially partially or completely lacks endonuclease activity but retains helicase activity, can be derived from various Cas proteins and variants thereof known in the art, including but not limited to Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, homologs thereof, or modified forms thereof.

[0065] In some embodiments, a Cas9 enzyme lacking nuclease activity and a single-stranded sgRNA that is specifically recognized by the enzyme are used. The Cas9 enzyme can be a Cas9 enzyme from different species, including but not limited to Cas9 from Streptococcus pyogenes (SpCas9), Cas9 from Staphylococcus aureus (SaCas9), and Cas9 from Streptococcus thermophilus (St1Cas9). Various variants of the Cas9 enzyme can be used, as long as the Cas9 enzyme can specifically recognize its sgRNA and lacks nuclease activity.

[0066] Cas proteins lacking nuclease activity can be prepared by methods well known in the art, including but not limited to deleting the entire catalytic domain of the nuclease endonuclease in the Cas protein or mutating one or several amino acids in the domain, thereby producing Cas proteins lacking nuclease activity. The mutation can be the deletion or substitution of one or more (e.g., more than 2, more than 3, more than 4, more than 5, more than 10, to the entire catalytic domain) amino acid residues, or the insertion of one or more new amino acid residues (e.g., more than 1, more than 2, more than 3, more than 4, more than 5, more than 10, or 1 to 10, 1 to 15). Conventional methods in the art can be used to delete the above-mentioned domains or mutate amino acid residues, and to detect whether the mutated Cas protein still has nuclease activity. For example, for Cas9, its two endonuclease catalytic domains RuvC1 and HNH can be mutated separately, for example, the 10th amino acid (located in the RuvC1 domain) asparagine of the enzyme is mutated to alanine or other amino acids, and the 841st amino acid (located in the HNH domain) histidine is mutated to alanine or other amino acids. These two mutations cause Cas9 to lose endonuclease activity. Preferably, the Cas enzyme has no nuclease activity at all. In one or more embodiments, the amino acid sequence of the Cas9 enzyme without nuclease activity used herein is shown in SEQ ID NO: 2 42-1452. In other embodiments, the Cas enzyme used herein partially lacks nuclease activity, that is, the Cas enzyme can cause DNA single-strand breaks. Representative examples of such Cas enzymes can be shown in SEQ ID NO: 72 amino acid residues 42-1419.

[0067] The Cas / sgRNA complex requires a protospacer adjacent motif (PAM) on the non-template strand (3' to 5') of DNA for proper function. Different Cas enzymes have different PAMs. For example, the PAM for SpCas9 is typically NGG; the PAM for SaCas9 is typically NNGRR; and the PAM for St1Cas9 is typically NNAGAA, where N is A, C, T, or G, and R is G or A.

[0068] In certain preferred embodiments, the PAM for SaCas9 enzyme is NNGRRT. In certain preferred embodiments, the PAM for SpCas9 is TGG.

[0069] sgRNA

[0070] sgRNA typically consists of two parts: a target binding region and a Cas protein recognition region. The target binding region and the Cas protein recognition region are usually connected in a 5' to 3' direction.

[0071] The length of the target binding region is typically 15 to 25 bases, more typically 18 to 22 bases, such as 20 bases. The target binding region specifically binds to the template strand of the DNA, thereby recruiting the fusion protein to a predetermined site. Typically, the contralateral region of the sgRNA binding region on the DNA template strand is adjacent to the PAM, or separated by several bases (e.g., within 10, or within 8, or within 5). Therefore, when designing sgRNA, the PAM of the enzyme is usually determined based on the Cas enzyme used, and then a site that can serve as the PAM is searched on the non-template strand of the DNA. After that, a fragment of 15 to 25 bases, more typically 18 to 22 bases, downstream of the PAM site of the non-template strand (3' to 5') is used as the sequence of the target binding region of the sgRNA.

[0072] The Cas protein recognition region of the sgRNA is determined according to the Cas protein used, which is known to those skilled in the art.

[0073] Therefore, the sequence of the target binding region of the sgRNA herein is a fragment of 15 to 25 bases, more usually 18 to 22 bases, downstream of the DNA chain containing the PAM site recognized by the selected Cas enzyme, immediately adjacent to the PAM site or separated from the PAM site by less than 10 bases (e.g., less than 8, less than 5, etc.); its Cas protein recognition region is specifically recognized by the selected Cas enzyme.

[0074] sgRNA can be prepared using conventional methods in the art, for example, using conventional chemical synthesis methods. sgRNA can also be transferred into cells via an expression vector and expressed in the cells. sgRNA expression vectors can be constructed using methods well known in the art.

[0075] Activation-induced cytosine deaminase (AID)

[0076] AID is a cytosine deaminase that belongs to the APOBEC family, a family of RNA editing enzymes: there is a nuclear localization signal at the N-terminus and a nuclear export signal at the C-terminus, and its catalytic domain is shared by the APOBEC family. It is generally believed that the N-terminal structure is necessary for somatic hypermutation (SHM). The function of AID is to deaminate cytosine, converting cytosine into uracil, and subsequent DNA repair can convert uracil into other bases. It should be understood that cytosine deaminases known in the art or fragments or mutants thereof that retain the biological activity of deaminating cytosine and converting cytosine into uracil can be used herein.

[0077] like Figure 14The structural and functional domains of AID are shown. Amino acids 9-26 form the nuclear localization (NLS) domain, with amino acids 13-26 in particular involved in DNA binding. Amino acids 56-94 form the catalytic domain, amino acids 109-182 form the APOBEC-like domain, amino acids 193-198 form the nuclear export (NES) domain, amino acids 39-42 interact with catenin-like protein 1 (CTNNBL1), and amino acids 113-123 form the hotspot recognition loop.

[0078] The full-length sequence of AID (as shown in amino acids 1457-1654 of SEQ ID NO: 2) can be used herein, or fragments of AID can be used. Preferably, the fragment includes at least the NLS domain, the catalytic domain, and the APOBEC-like domain. Therefore, in certain embodiments, the fragment comprises at least amino acid residues 9-182 of AID (i.e., amino acid residues 1465-1638 of SEQ ID NO: 2). In other embodiments, the fragment comprises at least amino acid residues 1-182 of AID (i.e., amino acid residues 1457-1638 of SEQ ID NO: 2). For example, in certain embodiments, the AID fragment used herein consists of amino acid residues 1-182, consists of amino acid residues 1-186, or consists of amino acid residues 1-190. Therefore, in certain embodiments, the AID fragment used herein consists of amino acid residues 1457-1638 of SEQ ID NO: 2, amino acid residues 1457-1642 of SEQ ID NO: 2, or amino acid residues 1457-1646 of SEQ ID NO: 2.

[0079] Variants of AID that retain their cytosine deaminase activity may also be used herein. For example, such variants may have 1-10, such as 1-8, 1-5, or 1-3 amino acid variations, including amino acid deletions, substitutions, and mutations, relative to the wild-type sequence of AID. Preferably, these amino acid variations do not occur within the aforementioned NLS domain, catalytic domain, and APOBEC-like domain, or, even if they do occur within these domains, do not affect the original biological functions of these domains. For example, preferably, these variations do not occur at positions 24, 27, 38, 56, 58, 87, 90, 112, or 140 of the AID amino acid sequence. In certain embodiments, these variations also do not occur within amino acids 39-42 or 113-123. Thus, for example, variations may occur within amino acids 1-8, 28-37, 43-55, and / or 183-198. In certain embodiments, variations occur at positions 10, 82, and 156. For example, substitution mutations occur at positions 10, 82, and 156, and such substitution mutations may be K10E, T82I, and E156G. In these embodiments, the amino acid sequence of the exemplary AID mutant comprises the amino acid sequence shown at positions 1447-1629 of SEQ ID NO: 68, or consists of the amino acid residues shown at positions 1447-1629 of SEQ ID NO: 68.

[0080] Fusion protein

[0081] Provided herein is a fusion protein containing a Cas enzyme and an AID. In the fusion protein herein, the Cas enzyme is generally at the N-terminus of the fusion protein amino acid sequence, and the AID is at the C-terminus. In certain embodiments, provided herein is a fusion protein mainly formed by a Cas enzyme and an AID. It should be understood that the fusion protein "mainly formed by..." or similar expressions described herein does not mean that the fusion protein only includes the Cas enzyme and the AID. This limitation should be understood as the fusion protein may only include the Cas enzyme and the AID, or may also contain other parts that do not affect the targeting effect of the Cas enzyme in the fusion protein and the function of the AID mutant target sequence, including but not limited to various linker sequences, nuclear localization sequences, and amino acid sequences introduced into the fusion protein as described below due to gene cloning operations, and / or in order to construct a fusion protein, promote the expression of a recombinant protein, obtain a recombinant protein that is automatically secreted outside the host cell, or facilitate the detection and / or purification of the recombinant protein.

[0082] The Cas enzyme can be fused to the AID via a linker. The linker can be a peptide of 3 to 25 residues, such as a peptide of 3 to 15, 5 to 15, or 10 to 20 residues. Suitable examples of peptide linkers are well known in the art. Typically, the linker contains one or more repeated motifs, which typically contain Gly and / or Ser. For example, the motif can be SGGS, GSSGS, GGGS, GGGGS, SSSSG, GSGSA, and GGSGG. Preferably, the motifs are adjacent in the linker sequence, with no amino acid residues inserted between the repeats. The linker sequence can consist of 1, 2, 3, 4, or 5 repeated motifs. In certain embodiments, the linker sequence is a polyglycine linker sequence. The number of glycine residues in the linker sequence is not particularly limited and is typically 2 to 20, such as 2 to 15, 2 to 10, or 2 to 8. In addition to glycine and serine, the linker may also contain other known amino acid residues, such as alanine (A), leucine (L), threonine (T), glutamic acid (E), phenylalanine (F), arginine (R), glutamine (Q), etc. In certain embodiments, the linker sequence is XTEN, and its amino acid sequence is shown in amino acid residues 183-198 of SEQ ID NO: 66.

[0083] As an example, the linker can be composed of the following amino acid sequence: G(SGGGG)2SGGGLGSTEF (SEQ ID NO:21), RSTSGLGGGS(GGGGS)2G (SEQ ID NO:22), QLTSGLGGGS(GGGGS)2G (SEQ ID NO:23), GGGS (SEQ ID NO:24), GGGGS (SEQ ID NO:25), SSSSG (SEQ ID NO:26), GSGSA (SEQ ID NO:27), GGSGGGGGGSGGGGSGGGGS (SEQ ID NO:28), SSSSGSSSSGSSSSG (SEQ ID NO:29), GSGSAGSGSAGSGSA (SEQ ID NO:30), GGSGGGGSGGGGSGG (SEQ ID NO:31), amino acid residues 1420-1456 of SEQ ID NO:72, etc.

[0084] It should be understood that during gene cloning, it is often necessary to design appropriate restriction sites, which inevitably introduces one or more irrelevant residues at the termini of the expressed amino acid sequence, but this does not affect the activity of the target sequence. To construct fusion proteins, promote expression of recombinant proteins, obtain recombinant proteins that are automatically secreted outside the host cell, or facilitate purification of recombinant proteins, it is often necessary to add certain amino acids to the N-terminus, C-terminus, or other suitable regions within the recombinant protein. For example, these include, but are not limited to, suitable linker peptides, signal peptides, leader peptides, terminal extensions, etc. Therefore, the amino or carboxyl termini of the fusion proteins herein may also contain one or more polypeptide fragments as protein tags. Any suitable tag may be used herein. For example, the tag may include FLAG (DYKDDDDK, SEQ ID NO: 32), HA, HA1, c-Myc, Poly-His, Poly-Arg, Strep-TagII, AU1, EE, T7, 4A6, ε, B, gE, and Ty1. These tags can be used to purify proteins.

[0085] The fusion protein herein may also contain a nuclear localization sequence (NLS). Nuclear localization sequences of various sources and various amino acid compositions well known in the art can be used. Such nuclear localization sequences include, but are not limited to, the NLS of the SV40 virus large T antigen, which has the amino acid sequence PKKKRKV (SEQ ID NO: 33); NLSs from nucleoplasmin, for example, the bipartite NLS of nucleoplasmin having the sequence KRPAATKKAGQAKKKK (SEQ ID NO: 34); NLSs from c-myc, which have the amino acid sequence PAAKRVKLD (SEQ ID NO: 35) or RQRRNELKRSP (SEQ ID NO: 36); NLSs from hRNPA1M9, which have the sequence NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 37); the sequence RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 38) from the IBB domain of importin-α; the sequences VSRKRPRP (SEQ ID NO: 39) and PPKKARED (SEQ ID NO: 40) of myoma T protein. ID NO:40); the sequence of mouse c-ablIV SALIKKKKKMAP (SEQ ID NO:41); the sequences of influenza virus NS1 DRLRR (SEQ ID NO:42) and PKQKKRK (SEQ ID NO:43); the sequence of hepatitis virus delta antigen RKLKKKIKKL (SEQ ID NO:44); the sequence of mouse Mx1 protein REKKKFLKRR (SEQ ID NO:45); the sequence of human poly (ADP-ribose) polymerase KRKGDEVDGVDEVAKKKSKK (SEQ ID NO:46); and the sequence of steroid hormone receptor (human) glucocorticoid RKCLQAGMNLEARKTKK (SEQ ID NO:47); etc. In certain specific embodiments, the sequence shown in amino acid residues 26-33 of SEQ ID NO:2 is used herein as the NLS. The NLS may be located at the N-terminus or C-terminus of the fusion protein; or it may be located in the fusion protein sequence, for example, at the N-terminus and / or C-terminus of the Cas9 enzyme in the fusion protein, or at the N-terminus and / or C-terminus of the AID in the fusion protein.

[0086] The accumulation of the fusion protein of the present invention in the nucleus can be detected by any suitable technology.For example, a detection marker can be fused to the Cas enzyme so that the position of the fusion protein in the cell can be visualized when combined with a means for detecting the position of the nucleus (for example, a dye specific for the nucleus, such as DAPI). In certain embodiments, 3*flag is used herein as a marker, and the peptide sequence can be shown as SEQ ID NO:2, amino acid residues 1-23. It should be understood that, generally, when a marker sequence is present, the marker sequence is generally at the N-terminus of the fusion protein. The marker sequence can be directly connected to the NLS or connected by an appropriate linker sequence. The NLS sequence can be directly connected to the Cas enzyme or AID or connected to the Cas enzyme or AID by an appropriate linker sequence.

[0087] Therefore, in certain embodiments, the fusion protein herein consists of a Cas enzyme and an AID. In other embodiments, the fusion protein herein consists of a Cas enzyme connected to an AID via a linker. In certain embodiments, the fusion protein herein consists of an NLS, a Cas enzyme, an AID, and an optional linker sequence between the Cas enzyme and the AID. In certain specific embodiments, the Cas enzyme in the fusion protein is the Cas9 enzyme described above. In certain specific embodiments, the amino acid sequence of AID in the fusion protein is as shown in amino acid residues 1457-1654 of SEQ ID NO: 2. In other specific embodiments, the amino acid sequence of AID in the fusion protein is as shown in amino acid residues 1457-1646 of SEQ ID NO: 4. In other specific embodiments, the amino acid sequence of AID in the fusion protein is as shown in amino acid residues 1447-1629 of SEQ ID NO: 68.

[0088] In certain embodiments, the amino acid sequence of the fusion protein herein is as shown in SEQ ID NO:2, 4, 66, 68, 70 or 72, or as shown in amino acids 26-1654 of SEQ ID NO:2, or as shown in amino acids 26-1638 of SEQ ID NO:4, or as shown in amino acids 26-1629 of SEQ ID NO:68, or as shown in amino acids 26-1629 of SEQ ID NO:70, or as shown in amino acids 26-1638 of SEQ ID NO:72.

[0089] Polynucleotide sequences, hosts, and protein expression

[0090] The present invention includes polynucleotide sequences encoding the fusion proteins herein. The polynucleotides herein can be in the form of DNA or RNA. DNA forms include cDNA, genomic DNA, or synthetic DNA. DNA can be single-stranded or double-stranded. DNA can be the coding strand or the non-coding strand.

[0091] The nucleotide sequences described herein can generally be obtained by PCR amplification. Specifically, primers can be designed based on the nucleotide sequences disclosed herein, particularly the open reading frame sequences, and amplified using a commercially available cDNA library or a cDNA library prepared by conventional methods known to those skilled in the art as a template to obtain the relevant sequence. When the sequence is long, two or more PCR amplifications are often required, and then the fragments amplified from each amplification are spliced ​​together in the correct order. For example, in certain embodiments, the polynucleotide sequence encoding the fusion protein described herein is as shown in SEQ ID NO: 1, 3, 65, 67, 79, or 71, or as shown in bases 73-4965 of SEQ ID NO: 1, or as shown in bases 73-4917 of SEQ ID NO: 3, or as shown in bases 76-4890 of SEQ ID NO: 67, or as shown in bases 76-4890 of SEQ ID NO: 70, or as shown in bases 76-4917 of SEQ ID NO: 72.

[0092] This article also includes nucleic acid constructs comprising the polynucleotides. The nucleic acid constructs contain the coding sequence of the fusion protein described herein, and one or more regulatory sequences operably linked to these sequences. The coding sequence of the fusion protein described in the present invention can be manipulated in a variety of ways to ensure expression of the protein. Before the nucleic acid construct is inserted into a vector, the nucleic acid construct can be manipulated according to the different or required expression vectors. The technology of using recombinant DNA methods to alter polynucleotide sequences is known in the art.

[0093] The regulatory sequence may be a suitable promoter sequence. The promoter sequence is generally operably linked to the coding sequence of the protein to be expressed. The promoter may be any nucleotide sequence that exhibits transcriptional activity in the selected host cell, including mutant, truncated, and hybrid promoters, and may be obtained from a gene encoding an extracellular or intracellular polypeptide that is homologous or heterologous to the host cell.

[0094] The regulatory sequence may also be a suitable transcription terminator sequence, a sequence recognized by the host cell to terminate transcription. The terminator sequence is operably linked to the 3' end of the nucleotide sequence encoding the polypeptide. Any terminator that is functional in the selected host cell may be used in the present invention.

[0095] The regulatory sequence may also be a suitable leader sequence, a non-translated region of an mRNA that is important for translation by the host cell. The leader sequence is operably linked to the 5' end of the nucleotide sequence encoding the polypeptide. Any terminator that is functional in the selected host cell may be used in the present invention.

[0096] In certain embodiments, the nucleic acid construct is a vector. For example, the polynucleotide sequences described herein can be inserted into a recombinant expression vector. The term "recombinant expression vector" refers to bacterial plasmids, bacteriophages, yeast plasmids, plant cell viruses, mammalian cell viruses such as adenoviruses, retroviruses, or other vectors known in the art. Any plasmid or vector can be used as long as it is replicable and stable in the host. An important feature of an expression vector is that it typically contains an origin of replication, a promoter, a marker gene, and translational control elements. An expression vector may also include a ribosome binding site for translation initiation and a transcription terminator. The polynucleotide sequences described herein can be operably linked to an appropriate promoter in the expression vector to direct mRNA synthesis via the promoter. Representative examples of such promoters include the lac or trp promoters of Escherichia coli; the λ phage PL promoter; eukaryotic promoters including the CMV immediate early promoter, the HSV thymidine kinase promoter, the early and late SV40 promoters, retroviral LTRs, and other known promoters that control gene expression in prokaryotic or eukaryotic cells or their viruses. Marker genes can be used to provide phenotypic traits for selection of transformed host cells, including but not limited to dihydrofolate reductase, neomycin resistance, and green fluorescent protein (GFP) for eukaryotic cell culture, or tetracycline or ampicillin resistance for Escherichia coli. When the polynucleotides described herein are expressed in higher eukaryotic cells, transcription will be enhanced if an enhancer sequence is inserted into the vector. Enhancers are cis-acting elements of DNA, usually about 10 to 300 base pairs, that act on a promoter to increase gene transcription.

[0097] Those skilled in the art will readily appreciate how to select appropriate vectors, promoters, enhancers, and host cells. Methods well known to those skilled in the art can be used to construct expression vectors containing the polynucleotide sequences described herein and appropriate transcriptional / translational control signals. These methods include in vitro recombinant DNA techniques, DNA synthesis techniques, in vivo recombination techniques, and the like.

[0098] The vectors described herein can be transformed into appropriate host cells to enable them to express the fusion proteins described herein. The host cell can be a prokaryotic cell, such as a bacterial cell; or a lower eukaryotic cell, such as a yeast cell; a filamentous fungal cell, or a higher eukaryotic cell, such as a mammalian cell. The host cell can also be a plant cell. Representative examples of host cells include: Escherichia coli; Streptomyces; bacterial cells of Salmonella typhimurium; fungal cells such as yeast, filamentous fungi; plant cells; insect cells of Drosophila S2 or Sf9; animal cells such as CHO, COS, 293 cells, or Bowes melanoma cells. In addition to cells used to express fusion proteins, other cells containing the polynucleotide sequences or vectors described herein and sgRNA or its expression vector, such as cells used to prepare point mutant proteins, are also within the scope of the host cells described herein.

[0099] Transformation of host cells with recombinant DNA can be performed using conventional techniques well known to those skilled in the art. When the host is a prokaryotic organism such as Escherichia coli, competent cells capable of absorbing DNA can be harvested after the exponential growth phase and treated with CaCl2, using procedures well known in the art. Another method is to use MgCl2. If desired, transformation can also be performed using electroporation. When the host is a eukaryotic organism, the following DNA transfection methods can be used: calcium phosphate coprecipitation, conventional mechanical methods such as microinjection, electroporation, liposome packaging, etc.

[0100] After transforming the host cell, the transformant obtained can be cultured using conventional methods to allow it to express the fusion protein described herein. Depending on the host cell used, the culture medium used in the culture can be selected from various conventional culture media. Various separation methods known in the art can be used to separate and purify the recombinant fusion protein herein. These methods are well known to those skilled in the art and include, but are not limited to, conventional renaturation treatment, treatment with a protein precipitant (salting out method), centrifugation, osmotic breaking, ultra-treatment, ultracentrifugation, molecular sieve chromatography (gel filtration), adsorption chromatography, ion exchange chromatography, high performance liquid chromatography (HPLC) and other various liquid chromatography techniques and combinations of these methods.

[0101] Therefore, the present invention also includes host cells containing the fusion protein described herein, its coding sequence or expression vector and optionally sgRNA or its expression vector. Such host cells can constitutively express the fusion protein described herein, or can express the fusion protein described herein under certain induction conditions. Methods for making host cells express the fusion protein of the present invention constitutively or under induction conditions are well known in the art. For example, in certain embodiments, an inducible promoter is used to construct the expression vector of the present invention, thereby achieving inducible expression of the fusion protein.

[0102] Composition, kit

[0103] The fusion protein, its coding sequence or expression vector, and / or sgRNA, its coding sequence or expression vector herein can be provided in the form of a composition. For example, the composition can contain the fusion protein and sgRNA or an expression vector of sgRNA herein, or can contain the expression vector of the fusion protein and sgRNA or an expression vector of sgRNA herein. In the composition, the fusion protein or its expression vector, or the sgRNA or its expression vector can be provided in the form of a mixture, or can be packaged separately. The composition can be in the form of a solution or in a lyophilized form.

[0104] The composition can be provided in a kit. Therefore, a kit containing the composition described herein is provided herein. Alternatively, a kit is also provided herein, which contains the fusion protein and sgRNA or sgRNA expression vector herein, or contains the expression vector of the fusion protein and sgRNA or sgRNA expression vector herein. In the kit, the fusion protein or its expression vector, or the sgRNA or its expression vector can be packaged independently or provided in the form of a mixture. The kit can also include, for example, reagents for transferring the fusion protein or its expression vector and / or sgRNA or its expression vector into cells, and instructions for guiding technicians to perform the transfer. Alternatively, the kit can also include instructions for guiding technicians to use the components contained in the kit to implement the various methods and uses described herein. The kit also includes other reagents, such as reagents for PCR, etc.

[0105] Methods and uses

[0106] The third aspect of this invention provides a method for generating a point mutation in a cell, the method comprising the step of expressing the fusion protein and sgRNA described herein in the cell. In certain embodiments, the fusion protein of the present invention or its expression vector and the sgRNA or its expression vector are transferred into the cell. In the case where the cell constitutively expresses the fusion protein described herein, only the corresponding sgRNA or its expression vector can be transferred into the cell. In the case where the cell inducibly expresses the fusion protein described herein, after the sgRNA is transferred, the cell can also be incubated with an inducer, or the cell can be subjected to corresponding induction measures (such as light). Conventional transfection methods can be used to transfer the fusion protein or its expression vector and / or sgRNA or its expression vector into the cell. For example, in certain embodiments, during transfection, a plasmid DNA-liposome complex is first prepared, and then the plasmid DNA-liposome complex and the corresponding sgRNA are co-transfected into the cell. After obtaining cells that have produced point mutations, the cells can be cultured under conditions suitable for the cell growth and expression of the desired protein, and the mutants produced can be separated and analyzed by various conventional methods (such as high-throughput methods).

[0107] Therefore, the methods described herein for generating point mutations in cells can also be used to generate mutant libraries, which can then be isolated and screened using conventional techniques to obtain mutants with desired biological functions. Therefore, the present invention also provides a method for constructing a mutant library, comprising the step of expressing the fusion protein described herein and the sgRNA in the cells.

[0108] One or more sgRNAs can be designed for the same mutation site. When designing multiple sgRNAs, the sgRNAs have different target binding regions but share the same Cas protein recognition region. These one or more sgRNAs can then be introduced into cells along with the corresponding fusion protein.

[0109] The cells can be any cell of interest, including prokaryotic and eukaryotic cells, such as plant cells, animal cells, microbial cells, and the like. Animal cells are particularly preferred, such as mammalian cells and rodent cells, including those from humans, horses, cows, sheep, mice, rabbits, and the like. Microbial cells include cells from various microbial species known in the art, especially those with medical research value and production value (e.g., production of fuels such as ethanol, protein production, and oils such as DHA). The cells can also be cells from various organ sources, such as cells from human liver, kidney, skin, and the like. The cells can also be various mature cell lines currently available on the market, such as 293 cells and COS cells. In certain embodiments, the cells are cells from healthy individuals; in other embodiments, the cells are cells from diseased tissues of diseased individuals, such as cells from inflammatory tissues, tumor cells, induced pluripotent stem cells, and the like. The cells can also be cells that have been genetically engineered to have a specific function (e.g., production of a protein of interest) or to produce a phenotype of interest. In other words, the gene or nucleic acid sequence to be mutated can be a naturally occurring (endogenous) gene or nucleic acid sequence in the cell, or it can be an exogenously introduced (exogenous) gene or nucleic acid sequence. The exogenously introduced gene or nucleic acid sequence can be integrated into the cell's genomic sequence, or it can be independently expressed outside the genome and stably expressed.

[0110] For different cells, expression vectors expressing the fusion protein and sgRNA described herein can be designed using known techniques to make these expression vectors suitable for expression in the cell. For example, a promoter and other relevant regulatory sequences that facilitate expression in the cell can be provided in the expression vector. These can be selected and implemented by technicians based on actual circumstances.

[0111] The nucleic acid sequence in which point mutations are desired can be any nucleic acid sequence of interest, such as a gene sequence, particularly genes or nucleic acid sequences associated with various diseases, the production of various proteins of interest, or various biological functions of interest. Such genes or nucleic acid sequences of interest include, but are not limited to, nucleic acid sequences encoding various functional proteins. As used herein, a functional protein refers to a protein capable of performing a physiological function in an organism, including catalytic proteins, transport proteins, immune proteins, and regulatory proteins. In certain embodiments, such functional proteins include, but are not limited to, proteins involved in the development, progression, and metastasis of diseases, proteins involved in cell differentiation, proliferation, and apoptosis, proteins involved in metabolism, proteins related to development, and various drug targets. For example, a functional protein can be an antibody, enzyme, lipoprotein, hormone protein, transport and storage protein, motor protein, receptor protein, membrane protein, and the like. Therefore, the fusion proteins, polynucleotides, nucleic acid constructs, cells, and methods described herein can be used to construct mutant libraries, which can then be screened to obtain proteins with new or enhanced functions, such as antibodies, enzymes, or other functional proteins.

[0112] The method described herein can be used to generate random mutations on a nucleic acid sequence of interest, or to generate mutations at specific sites in a nucleic acid sequence of interest. For the former, the PAM site on the template chain can be found according to the Cas enzyme used, and a fragment of 15 to 25 bases, more usually 18 to 22 bases, immediately downstream of the PAM site or separated from the PAM site by less than 10 (such as less than 8, less than 5, or less than 3) bases is used as the target recognition region of the sgRNA to design the sgRNA recognized by the Cas enzyme. For the latter, a site that can serve as a PAM can be found near the specific site, and a Cas enzyme that can recognize the PAM can be selected based on the PAM, and the fusion protein of the present invention containing the Cas enzyme and the corresponding sgRNA can be designed and prepared as described herein.

[0113] The methods herein can be either in vitro or in vivo. When implemented in vivo, the fusion protein or its expression vector and sgRNA or its expression vector herein can be transferred into the body of the experimental subject, such as the corresponding tissue cells, by means well known in the art, and the functional variants of interest can be screened by observing the phenotypic changes of the animal. It should be understood that in in vivo experiments, the experimental subjects can be various non-human animals, especially various non-human model organisms commonly used in the art. In vivo experiments should also meet ethical requirements.

[0114] The present invention will be described below in the form of specific examples. It should be understood that these examples are merely illustrative and do not limit the scope of the present invention. The experimental methods in the following examples, for which specific conditions are not specified, are generally performed according to conventional conditions such as those described in Sambrook & Russell's Molecular Cloning: A Laboratory Manual (Molecular Cloning Laboratory Manual, 3rd Edition), or according to the conditions recommended by the manufacturer. Unless otherwise defined, all professional and scientific terms used herein have the same meaning as those familiar to those skilled in the art. In addition, any methods and materials similar or equivalent to those described herein can be applied to the present invention. The preferred embodiments and materials described herein are for demonstration purposes only.

[0115] Example 1: Construction of pEntr11-dCas9-AID plasmid and pEntr11-dCas9-AIDX plasmid

[0116] 1. Using cDNA reverse transcribed from RNA of A20 cell line (purchased from the cell bank of Type Culture Collection Committee of Chinese Academy of Sciences) as template, the full-length sequence of AID and the AIDX fragment (truncated from amino acid residue 183) were amplified using primers shown in SEQ ID NOs: 5 and 6 and primers shown in SEQ ID NOs: 5 and 7, respectively (see Figure 1 , A and C);

[0117] 2. Construction of pEntr11-dCas9-TET1CD plasmid:

[0118] (1) Amplify the dCas9 target gene fragment from the dCas9 plasmid (Addgene) using PCR;

[0119] (2) The dCas9 target gene fragment and pEntr11 plasmid (Invitrogen) were digested with restriction endonucleases BamHI and NcoⅠ, and the fragments were recovered;

[0120] (3) Ligate the digested dCas9 fragment and the pEntr11 vector, and then transform the ligation product into TOP10 competent cells;

[0121] (4) Select positive clones, extract plasmids and send them for sequencing verification, thus completing the construction of pEntr11-dCas9 plasmid;

[0122] (5) Amplify the TET1CD target gene fragment using PCR;

[0123] (6) The pEntr11-dCas9 plasmid was digested with restriction endonucleases BamHI and XhoⅠ, and the fragments were recovered;

[0124] (7) TET1CD was cloned into the pEntr11-dCas9 plasmid using the Gibson Assembly method, thus completing the construction of the pEntr11-dCas9-TET1CD plasmid;

[0125] 3. Use restriction endonucleases BamHⅠ and XhoⅠ to digest the pEntr11-dCas9-TET1CD plasmid, AID, and AIDX fragments, and then recover the pEntr11-dCas9 vector and AID and AIDX fragments;

[0126] 4. Ligate the AID and AIDX fragments after enzyme digestion to the pEntr11-dCas9 vector respectively, and then transform the ligation products into TOP10 competent cells;

[0127] 5. Select positive clones, extract plasmids and send them for sequencing verification. This completes the construction of pEntr11-dCas9-AID and pEntr11-dCas9-AIDX plasmids ( Figure 1 , B and D).

[0128] Example 2: Construction of MO91-dCas9-AID plasmid and MO91-dCas9-AIDX plasmid

[0129] 1. Amplify the dCas9-AID fragment and dCas9-AIDX fragment from the pEntr11-dCas9-AID plasmid and the pEntr11-dCas9-AIDX plasmid using the primers shown in SEQ ID NO: 8 and 9 ( Figure 2 , A);

[0130] 2. Use restriction endonucleases BglⅡ and XhoⅠ to digest MO91 plasmid (Addgene Plasmid#19755) and AID and AIDX fragments, and then recover the vector, AID fragment and AIDX fragment ( Figure 2 , B);

[0131] 3. Ligate the AID and AIDX fragments after enzyme digestion to the MO91 vector respectively, and then transform the ligation products into Stbl3 competent cells;

[0132] 4. Select positive clones, extract plasmids and send them for sequencing verification. This completes the construction of MO91-dCas9-AID and MO91-dCas9-AIDX plasmids ( Figure 2 , C and D).

[0133] Example 3: Construction of MO91-dCas9 (3*flag, NLS)-AID plasmid and MO91-dCas9 (3*flag, NLS)-AIDX plasmid

[0134] Using pCW-Cas9 plasmid (Wuhan Miaoling Biotechnology Co., Ltd.) as a template, primers were designed to PCR amplify the 3*flag+NLS fragment, and the 3*flag+NLS fragment was cloned into the dCas9 N-terminus of MO91-dCas9-AID plasmid and MO91-dCas9-AIDX plasmid using the Gibson Assembly method, respectively, to construct MO91-dCas9(3*flag,NLS)-AID plasmid and MO91-dCas9(3*flag,NLS)-AIDX plasmid ( Figure 3 ).

[0135] Example 4: Establishment of an effective reporter system to indicate the efficiency of AID point mutations

[0136] The level of point mutations caused at the genomic level needs to be detected by a simple and intuitive method. The present invention mainly uses flow cytometry to indirectly detect the level of point mutations at the protein level. The stop codon (TAG) is artificially inserted into the EGFP gene, and EGFP cannot be expressed normally. When the fusion protein herein acts on the stop codon in the EGFP gene, the stop codon is point mutated, so that the EGFP gene mutation is expressed normally. Therefore, the higher the EGFP expression level, the higher the efficiency of the point mutation.

[0137] In this example, the EGFP gene containing a stop codon (sequence as Figure 4 The plasmid was inserted into the MO405-thy1.1 plasmid (Addgene) and MSCV was used to initiate gene expression. The plasmid was used to infect 293T cells, specifically comprising:

[0138] 1. Plate 293T cells and ensure that the cell density reaches 90% when encapsulating the virus;

[0139] 2. Encapsulate the virus after 24 hours. The encapsulation method is the same as transfection.

[0140] 3. Change the solution 24 hours after poisoning;

[0141] 4. 24 hours after the poison was packaged, collect the poison for the first time, add 1ug / ml of polybrene, 800g, 90min, and change the solution after 6-8h;

[0142] 5. 48 hours after the poison was packaged, the poison was collected for the second time, 1 μg / ml of polybrene was added, 800g, 90min, and the solution was changed after 6-8 hours;

[0143] 6. After the cells have grown to a sufficient number, flow cytometry staining (PE-thy1.1) is performed and thy1.1 positive cells are sorted as reporter cells. Figure 6 The schematic diagram of the reporter cell model is shown in Figure 5 middle.

[0144] Example 5: Preparation of sgRNA

[0145] 1. Find a 20-bp target sequence. If the starting base of the 20-bp target sequence is not G, add a G to its 5' end to enable efficient transcription by the RNA polymerase III U6 promoter. Note that the target sequence must not contain XhoI or NheI recognition sites.

[0146] 2. Clone the sgRNA into pLX (Addgene 50662) to obtain pLX sgRNA. The following four primers are required, of which R1 and F2 are specific for sgRNA:

[0147] F1:AAACTCGAGTGTACAAAAAAGCAGGCTTTAAAG(SEQ ID NO:10)

[0148] R1:rc(GN 19 )GGTGTTTCGTCCTTTCC(SEQ ID NO:11)

[0149] F2:GN 19 GTTTTAGAGCTAGAAATAGCAA(SEQ ID NO:12)

[0150] R2: AAAGCTAGCTAATGCCAACTTTGTACAAGAAAGCTG (SEQ ID NO: 13)

[0151] Among them, GN 19 = new target sequence, rc(GN 19 ) = reverse complement of the new target sequence.

[0152] 3. Amplify pLX sgRNA using F1+R1 and F2+R2 respectively;

[0153] 4. Gel-purify the products obtained from the two amplifications, combine them, and use them for the third PCR on F1+R2;

[0154] 5. Digest the product obtained from the PCR performed in step 4 with NheI and XhoI; and

[0155] 6. Ligation and transformation to prepare the expression vector of sgRNA.

[0156] The base sequences of the target binding regions of the four sgRNAs are as follows:

[0157] GCATGCCCGAAGGCTACGTCC (SEQ ID NO: 14);

[0158] GCAACTAGTATACCCGCGCCG(SEQ ID NO:15);

[0159] GCCTCGAACTTCACCTCGGCG(SEQ ID NO:16);

[0160] GTCAGCTCGATGCGGTTCACC (SEQ ID NO: 17).

[0161] Example 6: CRISPR-Cas9 improves AID point mutation efficiency

[0162] The reporter cells constructed in Example 4 were cultured to 70-90% confluence and then transfected. During transfection, a plasmid DNA-liposome complex was first prepared, including four times the amount of 2000 reagent diluted in In the culture medium, MO91-dCas9 (3*flag, NLS)-AID plasmid or MO91-dCas9 (3*flag, NLS)-AIDX plasmid was diluted in culture medium, and then add the diluted plasmid to the diluted 2000 reagent (1:1) and incubated for 30 minutes. The reporter cells constructed in Example 4 were then co-transfected with the plasmid DNA-liposome complex and the four sgRNAs targeting the EGFP stop codon prepared in Example 5. As a control, the reporter cells constructed in Example 4 were transfected with the plasmid DNA-liposome complex alone. The cells were cultured with 2 μg / ml of puromycin and 20 μg / ml of blasticidin, screened for 3 days, and EGFP expression levels were analyzed by flow cytometry on days 4 and 7 after transfection.

[0163] The results are as follows Figure 7 As shown, the %EGFP+ of AID and AIDX were 0.14% and 0.30%, respectively, while the %EGFP+ of dCas9-AID+sgRNA and dCas9-AIDX+sgRNA were 2.14% and 4.36%, respectively.

[0164] The results showed that fusing AID or AIDX with dCas9, under the guidance of sgRNA, can confine the point mutation function of AID to specific sites under the targeting action of sgRNA, while increasing its action concentration and improving its mutation efficiency.

[0165] Example 7: CRISPR-Cas9 improves AID point mutation efficiency and optimization

[0166] The same method as in Example 6 was used to co-transfect sgRNA and dCas9-AID expression vectors into the reporter cells constructed in Example 4. The sgRNAs were divided into two groups. One group was a control sgRNA targeting AAVS1, and its target binding regions were as follows: GATTCCCAGGGCCGGTTAATG (SEQ ID NO: 18); GTCCCCTCCACCCCACAGTG (SEQ ID NO: 19); and GGGGCCACTAGGGACAGGAT (SEQ ID NO: 20). The other group was an sgRNA group targeting EGFP (SEQ ID NOs: 14-17). At the same time, a control group was set up in which AID was transfected alone into the reporter cells. The expression vector for the control sgRNA was constructed as described in Example 5.

[0167] On the 8th day after transfection, FACS analysis showed that the EGFP%+ in the AID group was only 0.13%, while that in the dCas9-AID+sgRNA group reached 2.1% ( Figure 8 , A), EGFP%+ increased by 16 times. To further optimize the efficiency of the dCas9-AID system, dCas9 was fused with different AID mutants: AID-FL (full length), AID-CD (containing only the catalytic domain), P182X (truncated from amino acid residue 183), R186X (truncated from amino acid residue 187), and R190X (truncated from amino acid residue 191). Each dCas9-AID expression vector and sgRNA were co-transfected in reporter cells, among which dCas9-R186X had the highest efficiency ( Figure 8 , B and C). Therefore, dCas9-R186X was used to carry out the experiments of Examples 8-13. In these examples, dCas9-R186X is referred to as dCas9-AIDX.

[0168] In order to prove that the fusion of AID and dCas9 in the dCas9-AID system is indeed the basis for the base substitution function of the entire system, Cas9, dCas9, the functional mutant of dCas9-AIDX [R186X (E58Q)], dCas9-AIDX and sgRNA were co-transfected into reporter cells. Only the dcas9-AIDX and sgRNA groups showed EGFP%+, while the other groups were all 0( Figure 8 , C). This proves that it is indeed the fusion of AID and dCas9 that gives the entire system the base substitution function.

[0169] Example 8: CRISPR-Cas9 confines AID point mutation function to the sgRNA targeting site

[0170] To investigate whether CRISPR-Cas9 can localize the AID point mutation function to the sgRNA targeting site, PCR was performed on EGFP containing a stop codon using the genomic DNA of the reporter system constructed in Example 4 as a template to construct a library, and Miseq sequencing was performed using cMyc as a control gene. Figure 9 As shown. From the sequencing results of the reporter cells, it can be seen that although Miseq has a high sequencing throughput and filters out low-quality reads, there is still a sequencing base mutation frequency of 0.25% for EGFP and 0.15% for cMyc. However, even with base-level interference, it can be observed that the frequency of EGFP gene point mutations in the dCas9-AIDX+sgRNA group is significantly higher than that in the AIDX group, which also proves that CRISPR-Cas9 improves the efficiency of AID point mutations. Moreover, these high-frequency mutation sites are mainly concentrated in the target site of sgRNA, while almost no point mutations occur in the cMyc gene. It is proved that after dCas9 is fused with AID, sgRNA targets dCas9-AID to the target site of sgRNA, so that AID will only act on the target site of sgRNA, resulting in point mutations, without causing significant changes to other gene sites; and it can greatly increase the frequency of point mutations.

[0171] Example 9: dCas9-AIDX randomly mutates C and G bases into three other bases

[0172] AIDX itself mutates C to T and G to A. After fusing dCas9 with AIDX, the mutation directions of C and G became more uniform compared to the AIDX group.

[0173] At the same time, the effect of AID itself is dependent on the WRCY of the hotspot motif (W represents A / T, R represents A / C, and Y represents C / T), among which the most preferred motif is AGCT. After fusing dCas9 with AIDX, the preference for this motif will obviously disappear. Therefore, the inventors put forward a hypothesis that under normal circumstances, AID will deaminize cytosine to form uracil, and through DNA replication and repair, this ug mismatch will be retained, and mutations from C to T and G to A will occur. In addition, the U base can be excised through base excision repair, and then the four bases can be inserted. Therefore, the fusion of dCas9 and AID is likely to inhibit the DNA replication pathway, promote base excision repair, and make the mutation direction more uniform ( Figure 10 , b).

[0174] In addition, statistical analysis of Miseq data showed that the types of point mutations in EGFP caused by AIDX and dCas9-AIDX+sgRNA groups were basically consistent with the reports, with C and G base mutations accounting for the majority, and A and T accounting for a smaller proportion. In addition, G mainly mutated to T, and C mutated to A. However, in the dCas9-AIDX group, the proportion of G mutations to T and C increased, and the proportion of C mutations to G or A increased. Therefore, dCas9-AIDX can produce more uniform mutation types ( Figure 10 , a).

[0175] Example 10: UGI increases the base substitution frequency of the dCas9-AIDX system, reveals the action trajectory of dCas9-AIDX on the gene, and makes the base mutation direction more unified.

[0176] UGI is an inhibitor of UNG, a phage protein that protects its own genome from the repair of host UNG when the phage invades E. coli. Figure 11 , a). Three plasmids were co-transfected into reporter cells to express dCas9-AIDX, a single sgRNA (target binding region is GCCTCGAACTTCACCTCGGCG, SEQ ID NO: 16), and UGI (protein sequence: UniProtKB-P14739) respectively, to improve the mutation efficiency of a single sgRNA in the entire system. The results showed that the highest point mutation efficiency was improved by 10 times ( Figure 11 , b).

[0177] In addition, after adding UGI, the mutation direction of the entire system became more uniform, from C to T and G to A. The action trajectory of dCas9-AIDX was also counted, and the mutation frequency caused by the entire system before and after the PAM sequence was also counted. Figure 11 (c) is a statistical analysis based on the data of four sgRNAs designed for the EGFP locus. All of them use the N in the NGG in the PAM sequence as the first base. Its upstream is - and its downstream is +. The statistical results of the two sets of data are consistent. Both cause mutations in the 20bp upstream of the PAM, that is, in the protospacer sequence region, and the highest mutation point is at the -12 / -13 position of the PAM. UGI can increase the overall mutation frequency of AID, but it will increase the proportion of base substitutions and reduce the conversion ratio ( Figure 11 , d).

[0178] Example 11: dCas9-AIDX can act not only on exogenous genes, but also on endogenous genes. The above experiments were all performed in reporter cells. In this example, the endogenous gene AAVS1 was selected as the target site, and three sgRNAs (SEQ ID NOs: 18-20) were designed. A vector expressing dCas9-AID and three sgRNAs targeting AAVS1 was co-transfected into 293T cells (as described in Example 7).

[0179] The results are as follows Figure 12 As shown in Figure 2, the dCas9-AID system can also produce base substitutions in the endogenous gene AAVS1, and this mutation is also concentrated in the sgRNA target site.

[0180] Example 12: Application of dCas9-AIDX in Gleevec Resistance Screening of K562 BCR-ABL Gene

[0181] K562 is a leukemia cell line derived from a patient with chronic myeloid leukemia. Within these cells, a chromosome called the ph chromosome exists. This chromosome is formed by transposition of the long arms of chromosomes 9 and 22. The ABL gene on chromosome 9 contains a tyrosine kinase active site that is normally inactive. However, upon transposition to the BCR locus, it becomes highly active. This triggers a series of signaling pathways, leading to cancer. Therefore, BCR-ABL is a proto-oncogene. A commonly used drug is Gleevec (the active ingredient is imatinib mesylate). Its primary mechanism of action is that Gleevec competes with ABL for ATP binding, thereby reducing ABL gene activity. However, point mutations, such as T315I, have been found in patient samples within the tyrosine kinase active domain, which can disrupt the domain's ability to bind Gleevec and confer Gleevec resistance. In addition, base substitutions at other sites can also contribute to Gleevec resistance. The dCas9-AIDX system can be used to screen Gleevec resistance sites and specific mutation types as the basis for designing the next generation of inhibitors.

[0182] First, to obtain K562 cells stably expressing dCas9-AIDX, we co-transfected 293T cells with the target plasmid MSCV-dCas9-AID-P182X-IRES-Thy1.1 and the viral packaging plasmid pcl-10A1. 1x10 6293T cells were cultured overnight in 2ml of antibiotic-free 10% FBS-containing DMEM. The next day, when the cells reached 80% confluency, 3µg of the target plasmid and 1µg of the viral packaging plasmid were transfected with 10µl of the transfection reagent LIPO2000. Twenty-four hours after transfection, the cells were cultured in 2ml of antibiotic-containing medium. Viruses were collected at 48 and 72 hours. The collected viruses were immediately centrifuged at 1000rpm for 5 minutes to remove cell debris. The supernatant was then added with 2µl of 10mg / ml Polybrene to infect 1x10 5 K562 cells were shaken at 37°C and 900g for 90 minutes. 4 hours after infection, the cells were centrifuged and the pellet was cultured with anti-infection medium. After two days of continuous infection, K562 cells were cultured for another two days and then flow cytometry was used to mark cells expressing Thy1.1 surface molecules as PE. + (antibody 1:200 dilution), and single cell sorting technology was used to obtain two 96-well plates of PE-Thy1.1 + After two weeks of culture, RNA from cell populations generated from each single-cell clone was collected and subjected to RT-qPCR. The cell line with the highest dCas9-AIDX expression was subsequently used to screen for Gleevec resistance sites and mutation types.

[0183] At the same time, to screen for sites of Gleevec resistance, we designed sgRNAs targeting the genomic region where the sixth exon of the ABL gene, Exon 6, is located. A total of 16 sgRNAs were designed (target region sequences are shown in SEQ ID NOs: 49-64), of which 6 targeted the intron region adjacent to exon Exon 6, and 10 directly targeted the Exon 6 region, covering 83% of the exon sequence. Because the T315I mutation has been recognized as one of the most important mutations causing Gleevec resistance, only one of our designed sgRNAs can cover the site of the T315I mutation (944C), which can serve as a positive control. At the same time, we designed three sgRNAs targeting the genomic sequence of the AAVS1 gene, which is not related to Gleevec resistance, as negative controls (target region sequences are shown in SEQ ID NOs: 18-20). These sgRNA sequences were chemically synthesized, double-digested with BamH1 and HindIII, and finally cloned into the pSUPER-sgRNA vector carrying the H1 promoter. We used the phenol chloroform-ethanol sedimentation method to precipitate an equal amount of 16 Exon6 sgRNA plasmids or 3 AAVS1 sgRNA plasmids, so that the final concentration of the mixed plasmid was above 1.5ug / ul. Subsequently, the K562 cell line stably expressing dCas9-AIDX was electroporated with the mixed sgRNA library of ABL-Exon6 and AAVS1, respectively. The instrument used was the Neo electroporator from Life Technology, USA. 12-24 hours before electroporation, K562 cells were cultured in IMDM medium without antibiotics and 10% FBS. On the day of electroporation, two 1.2x10 6 K562 cells were transfected with 8 μg of ABL-Exon6 or AAVS1 sgRNA. Since the pSUPER-sgRNA plasmid vector carries a puromycin resistance gene, 2 μg / ml puromycin was added 24 hours after transfection to select cells expressing sgRNA. Puromycin treatment was removed after 48 hours, and K562 cells were further cultured. 2 x 10 5 The cell DNA and RNA were subjected to high-throughput sequencing and used as input controls. The remaining cells were divided into two parts and treated with 10uM Gleevec or an equal volume of DMSO. Ficoll was applied every three days to remove dead cells until the cell count was less than 2x10 4Under Gleevec drug treatment, the control group cells transfected with AAVS1sgRNA basically died in about 7-10 days, while the experimental group cells transfected with ABL-Exon6sgRNA were able to continue to proliferate. Around 36-40 days after transfection, the experimental group cells proliferated to 10 7 Order of magnitude ( Figure 14 , b). DNA and RNA were collected from Gleevec-treated and DMSO-treated cells and analyzed by high-throughput sequencing. Sequencing results showed that 30% of the cells harbored the T315I mutation, a known drug-resistance mutation found in patients. In addition, several previously unreported point mutations were also found ( Figure 14 , c and d).

[0184] Example 13: Application of dCas9-AIDX to Improve Antibody Affinity and Specificity in Vitro

[0185] Antibodies can specifically recognize antigens and serve as therapeutic proteins for treating a variety of diseases. The affinity of an antibody is proportional to the somatic mutations it generates within the germinal centers of the body. Generally speaking, high-affinity antibodies harbor multiple high-frequency somatic mutations. Therefore, dCas9-AIDX can be used to mutate antibody genes and screen for antibodies with stronger affinity or other characteristics, such as improved specificity.

[0186] The usage scheme is as follows: the antibody molecules are stably expressed on the surface of 293T cells, and then sgRNA is designed for the antibody gene, and 293T cells are transfected simultaneously with dCas9-AIDX. The cell surface is then stained. The stronger the staining, the stronger the affinity of the mutated antibody molecule.

[0187] This example uses Invitrogen's Flp-In stably expressing a lacZ-ZeocinTM fusion locus. TM -293 cells. First, a low affinity mouse IgG1 antibody (K D =2.78E-09M) cDNA sequence and connected to the coding sequence of the H2Kk protein transmembrane region sequence to add the H2Kk protein transmembrane region sequence to the antibody end, and the resulting DNA sequence was cloned into the pcDNA5 / FRT / GOI vector (Life Science Technology, USA). TM -293 cells, using the Flp-In TM Flp-In contained in -293 cells TMThe system integrates the IgG1 coding sequence, containing the Flp recombination target site, into the lacZ-Zeocin™ fusion locus using Flp recombinase. Cells that fail to successfully integrate express the Zeocin-resistant protein. However, cells that successfully integrate fail to express the Zeocin-resistant protein due to the lack of the ATG start codon, but are able to express the hygromycin-resistant protein. Therefore, hygromycin is used to select 293 cells that have successfully integrated the IgG1. In these cells, only one copy of the anti-HEL-IgG1 gene is expressed per cell.

[0188] Next, 16 appropriate PAM sequences were selected for each of the three CDRs of the IgG1 heavy and light chains, and the following sgRNAs (SEQ ID NOs: 73-88) were designed, so that each CDR of the heavy or light chain was covered by at least two sgRNAs:

[0189] IgH

[0190] CDR1_1:TCCCTCACCTGTTCTGTCAC(SEQ ID NO:73);

[0191] CDR1_2:GCTCCAGTAATCACTGGTGA (SEQ ID NO:74);

[0192] CDR1_3:GATCCAGCTCCAGTAATCAC (SEQ ID NO:75);

[0193] CDR1_4: GTGATTACTGGAGCTGGATC (SEQ ID NO:76);

[0194] CDR2_1:ATGGGGTACGTAAGCTACAG (SEQ ID NO:77);

[0195] CDR2_2:GAGATTCGACTTTTGAGAGA(SEQ ID NO:78);

[0196] CDR3_1:TATTACTGTGCAAACTGGGA(SEQ ID NO:79);

[0197] CDR3_2:CAAACTGGGACGGTGATTAC(SEQ ID NO:80);

[0198] CDR3_3:GACGGTGATTACTGGGGCCA (SEQ ID NO:81);

[0199] IgL

[0200] CDR1_1:GTTGTTGCCAATACTTTGGC (SEQ ID NO:82);

[0201] CDR1_2:ATAGCGTCAGTCTTTCCTGC(SEQ ID NO:83);

[0202] CDR1_3:GTATTGGCAACAACCTACAC(SEQ ID NO:84);

[0203] CDR2_1:AGGGGATCCCAGAGATGGAC (SEQ ID NO:85);

[0204] CDR2_2:TATGCTTCCCAGTCCATCTC (SEQ ID NO:86);

[0205] CDR3_1:TCTGTCAACAGAGTAACAGC (SEQ ID NO:87);

[0206] CDR3_2:GTCCCCCCTCCGAACGTGTA (SEQ ID NO:88).

[0207] The sgRNA sequence was then cloned into the pSUPER-puro plasmid vector (Addgene). The MO91-dCas9 (3*flag, NLS)-AIDX plasmid constructed in Example 3 and the sgRNA library (i.e., 16 sgRNAs mixed in equal amounts) or the sgRNA of the control gene AAVS1 were co-transfected into the 293 cells expressing IgG1 obtained above. After being screened with puromycin and blasticidin antibiotics, PE anti-mouse IgG and Alex647-HEL surface staining were performed on the 7th day after transfection, and flow sorting was performed to sort out cells with unchanged IgG intensity but increased binding to the HEL antigen. After culture and proliferation, the mutations on the DNA were first analyzed by high-throughput sequencing, and the results were basically consistent with those of the mutations in the ABL gene or GFP gene in this article ( Figure 15 dCas9-AIDX induced base mutations in the variable region of anti-HEL IgG1 and reproducibly induced base mutations in IgG1 CDR ( Figure 16 ).

[0208] Then, using PE anti-mouse IgG1 and 647-HEL surface staining to detect the mutated cells on a flow cytometer, it was found that a small group of cells had unchanged IgG1 expression but increased binding to HEL. This group of cells was then flow sorted, amplified, and compared with the cells before mutation. It was found that the affinity of the mutated antibody for HEL increased by more than 10 times ( Figure 17 ).

[0209] Then, an appropriate amount of cells were collected to extract genomic DNA for sequencing. It was found that the main reason for the increase in affinity was the mutation of glycine at position 52 of the light chain to aspartic acid (the base was changed from GGT to GAT, Figure 15 ).

[0210] Example 14: Preparation of other fusion proteins

[0211] 1. Plasmid construction

[0212] (1) Synthesizing the XTEN linker sequence using gene synthesis;

[0213] (2) Using restriction endonucleases to digest the MO91-dCas9-AIDX plasmid constructed in Example 2, and recovering the vector, AIDX fragment, and dCas9 fragment;

[0214] (3) The AIDX fragment, dCas9 fragment, and XTEN linker sequence after enzyme digestion were ligated to the MO91 vector, and the ligation products were then transformed into Stbl3 competent cells;

[0215] (4) Select positive clones, extract plasmids, and send them for sequencing verification, thus completing the construction of MO91-dCas9-XTEN-AIDX plasmid;

[0216] Plasmids MO91-AIDX-XTEN-dCas9, MO91-dCas9-XTEN-AIDX (K10E T82I E156G), and MO91-nCas9-AIDX can be constructed according to the above steps and the methods of Examples 1 and 2.

[0217] When cloning in the 3*flag and / or NLS fragment is required, the 3*flag and / or NLS fragment can be cloned into the above plasmid according to the method of Example 3 to obtain plasmids expressing the fusion proteins represented by SEQ ID NOs: 66, 68, 70, and 72, respectively. The AIDX in these fusion proteins is an AID fragment truncated from amino acid residue 183 or a mutant thereof.

[0218] 2. Expression and purification of recombinant proteins

[0219] (1) Construct the plasmid pET-nCas9-AIDX-6His according to conventional methods, and then use the plasmid to transform Escherichia coli BL21STAR-competent cells;

[0220] (2) The resulting expression strain was grown overnight at 37°C in LB medium containing 100 μg / ml kanamycin. The cells were diluted 1:100 into 2xYT medium and grown at 37°C to an OD 600 of ~0.6. The culture was cooled to 4°C within 2 hours, and IPTG 0.5 mM was added to induce protein expression for ~16 hours.

[0221] (3) Cells were collected by centrifugation at 4000 g for 15 min and resuspended in lysis buffer;

[0222] (4) cells were lysed with a cell disruptor (Union) at 800 bar for 5 min, and the supernatant of the lysate was separated after centrifugation for 15 min;

[0223] (5) The lysate was incubated with Ni-NTA (1 ml slurry / L bacteria) (DP101, TransGen) for 1 hour at 4°C to capture the His-tagged fusion protein; the resin was transferred to a column and washed extensively with cold wash buffer (to the extent that no color change could be observed using Coomassie G250);

[0224] (6) The His-tagged fusion protein was eluted in elution buffer and concentrated to a total volume of 1 ml by ultrafiltration (Amicon-Millipore, 100 kDa molecular weight cutoff);

[0225] (7) The protein was diluted to 20 ml in buffer A, loaded onto a Hi-Trap SP column (29051324, GE Healthcare) and eluted with a 100 mM-1 M NaCl gradient;

[0226] (8) The eluted fraction containing nCas9-AIDX was concentrated to about 1 ml and purified by using a Superdex 20010 / 300GL column (17517501, GE Healthcare);

[0227] (9) The eluted protein was concentrated to approximately 3 mg / ml, quickly frozen in liquid nitrogen, and stored at −80°C.

[0228] The electrophoretic pattern of induced nCas9-AIDX expression in bacteria is shown in Figure 18 .

[0229] 3. Functional testing of different fusion proteins

[0230] The functions of the different fusion proteins in this example were tested using the same method as in Example 10. Figure 19 -21. Sequence Listing <110> Shanghai Institutes for Life Sciences, Chinese Academy of Sciences <120> Fusion protein capable of generating point mutation in cells, preparation and use thereof <130> 162593Z1 <160> 95 <170> PatentIn version 3.3 <210> 1 <211> 4989 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequence: dCas9-AID coding sequence <400> 1 atggactata aggaccacga cggagactac aaggatcatg atattgatta caaagacgat 60 gacgataaga tggccccaaa gaagaagcgg aaggtcggta tccacggagt cccagcagct 120 accatggaca agaagtattc tatcggactg gccatcggga ctaatagcgt cgggtgggcc 180 gtgatcactg acgagtacaa ggtgccctct aagaagttca aggtgctcgg gaacaccgac 240 cggcattcca tcaagaaaaa tctgatcgga gctctcctct ttgattcagg ggagaccgct 300 gaagcaaccc gcctcaagcg gactgctaga cggcggtaca ccaggaggaa gaaccggatt 360 tgttaccttc aagagatatt ctccaacgaa atggcaaagg tcgacgacag cttcttccat 420 aggctggaag aatcattcct cgtggaagag gataagaagc atgaacggca tcccatcttc 480 ggtaatatcg tcgacgaggt ggcctatcac gagaaatacc caaccatcta ccatcttcgc 540 aaaaagctgg tggactcaac cgacaaggca gacctccggc ttatctacct ggccctggcc 600 cacatgatca agttcagagg ccacttcctg atcgagggcg acctcaatcc tgacaatagc 660 gatgtggata aactgttcat ccagctggtg cagacttaca accagctctt tgaagagaac 720 cccatcaatg caagcggagt cgatgccaag gccattctgt cagcccggct gtcaaagagc 780 cgcagacttg agaatcttat cgctcagctg ccgggtgaaa agaaaaatgg actgttcggg 840 aacctgattg ctctttcact tgggctgact cccaatttca agtctaattt cgacctggca 900 gaggatgcca agctgcaact gtccaaggac acctatgatg acgatctcga caacctcctg 960 gcccagatcg gtgaccaata cgccgacctt ttccttgctg ctaagaatct ttctgacgcc 1020 atcctgctgt ctgacattct ccgcgtgaac actgaaatca ccaaggcccc tctttcagct 1080 tcaatgatta agcggtatga tgagcaccac caggacctga ccctgcttaa ggcactcgtc 1140 cggcagcagc ttccggagaa gtacaaggaa atcttctttg accagtcaaa gaatggatac 1200 gccggctaca tcgacggagg tgcctcccaa gaggaatttt ataagtttat caacctac 1260 cttgagaga tggacggcac cgaagagctc ctcgtgaac tgaatcggga ggatctgctg 1320 cggaagcagc gcacttcga caatgggagc attccccacc agatcatct tggggagctt 1380 cacgccatcc ttcggcgcca agaggacttc tacccctttc ttaggacaa cagggagaag 1440 attgagaaaa ttctcactt ccgcatcccc tactacgtgg gacccctcgc cagaggaaat 1500 agccggtttg cttggatgac cagaagtca gagaaacta tcactccctg gaacttcgaa 1560 gaggtggtgg acaagggagc cagcgctcag tcattcatcg aacggatgac taacttcgat 1620 aagaacctcc ccaatgagaa ggtcctgccg aacattccc tgcttacga gtactttacc 1680 gtgtacaacg agctgaccaa ggtgaaatat gtcaccgaag ggatgaggaa gcccgcattc 1740 ctgtcaggcg aaaaagaaaggcaattgtg gaccttctgt tcagaccaa tagaaggtg 1800 accgtgaagc agctgaagga ggactttc aagaaaattg aatgcttcga ctctgtggag 1860 attagcgggg tcgagatcg gttcaacgca agcctggta cctaccatga tctgcttaag 1920 atcatcagg acagatt tctgacat gaggagaacg aggacatcct tgaggacatt 1980 gtcctgactc tcactctt cgaggaccgg gaatgatcg aggagaggct tagacctac 2040 gcccatctgt tcgacgataa agtgatgaag caacttaac ggagaata taccggatgg 2100 ggacgcctta gccgcaact catcaacgga atccgggaca aacagcgg aagaccatt 2160 cttgatttcc ttagagcga cggattcgct aatcgcact tcatgcact tatccatgat 2220 gattccctga cctttaagga ggacatccag aaggcccaag tgtctggaca aggtgactca 2280 ctgcacgagc atatcgcaa tctggctggt tcaccgcta ttaagaaggg tattctccag 2340 accgtgaaag tcgtggacga gctggtcaag gtgatgggtc gccataacc agagaacatt 2400 gtcatcgaga tggccaggga aaaccagact acccagaagg gagagaga caggcagggag 2460 cggatgaaaa gattgagga agggattaag gagctcgggt cacagatccct taagagcac 2520 ccggtggaaa acacccagct tcagaatgag aagctctatc tgtacct tcaaatgga 2580 cgcgatatgt atgtggacca agagcttgat atcaacaggc tctcagacta cgacgtggac 2640 gccatcgtcc ctcagagctt cctcaaagac gactcaattg acaataaggt gctgactcgc 2700 tcagacaaga accggggaaa gtcagataac gtgccctcag aggaagtcgt gaaaaagatg 2760 aagaactatt ggcgccagct tctgaacgca aagctgatca ctcagcggaa gttcgacaat 2820 ctcactaagg ctgagagggg cggactgagc gaactggaca aagcaggatt cattaaacgg 2880 caacttgtgg agactcggca gattactaaa catgtcgccc aaatccttga ctcacgcatg 2940 aataccaagt acgacgaaaa cgacaaactt atccgcgagg tgaaggtgat taccctgaag 3000 tccaagctgg tcagcgattt cagaaaggac tttcaattct acaaagtgcg ggagatcaat 3060 aactatcatc atgctcatga cgcatatctg aatgccgtgg tgggaaccgc cctgatcaag 3120 aagtacccaa agctggaaag cgagttcgtg tacggagact acaaggtcta cgacgtgcgc 3180 aagatgattg ccaaatctga gcaggagatc ggaaaggcca ccgcaaagta cttcttctac 3240 agcaacatca tgaatttctt caagaccgaa atcacccttg caaacggtga gatccggaag 3300 aggccgctca tcgagactaa tggggagact ggcgaaatcg tgtgggacaa gggcagagat 3360 ttcgctaccg tgcgcaaagt gctttctatg cctcaagtga acatcgtgaa gaaaaccgag 3420 gtgcaaaccg gaggctttc taaggaatca atcctcccca agcgcaactc cgacaagctc 3480 attgcaagga agaaggattg ggaccctaag aagtacggcg gattcgattc accaactgtg 3540 gcttattctg tcctggtcgt ggctaaggtg gaaaaaggaa agtctaagaa gctcaagagc 3600 gtgaaggaac tgctgggtat caccattatg gagcgcagct ccttcgagaa gaacccaatt 3660 gactttctcg aagccaaagg ttacaaggaa gtcaagaagg accttatcat caagctccca 3720 aagtatagcc tgttcgaact ggagaatggg cggaagcgga tgctcgcctc cgctggcgaa 3780 cttcagaagg gtaatgagct ggctctcccc tccaagtacg tgaatttcct ctaccttgca 3840 agccattacg agaagctgaa ggggagcccc gaggacaacg agcaaaagca actgtttgtg 3900 gagcagcata agcattatct ggacgagatc attgagcaga tttccgagtt ttctaaacgc 3960 gtcattctcg ctgatgccaa cctcgataaa gtccttagcg catacaataa gcacagagac 4020 aaaccaattc gggagcaggc tgagaatatc atccacctgt tcaccctcac caatctttggt 4080 gcccctgccg cattcaagta cttcgacacc accatcgacc ggaaacgcta tacctccacc 4140 aaagaagtgc tggacgccac cctcatccac cagagcatca ccggacttta cgaaactcgg 4200 attgacctct cacagctcgg aggggatgag ggagctccca agaaaaagcg caaggtaggt 4260 agttccggat ctccgaaaaa gaaacgcaaa gttggtagtg atgctttaga cgattttgac 4320 ttagatatgc ttggttcaga cgcgttagac gacttcggtg gaggatccat ggacagcctc 4380 ttgatgaacc ggaggaagtt tctttaccaa ttcaaaaatg tccgctgggc taagggtcgg 4440 cgtgagacct acctgtgcta cgtagtgaag aggcgtgaca gtgctacatc cttttcactg 4500 gactttggtt atcttcgcaa taagaacggc tgccacgtgg aattgctctt cctccgctac 4560 atctcggact gggacctaga ccctggccgc tgctaccgcg tcacctggtt cacctcctgg 4620 agcccctgct acgactgtgc ccgacatgtg gccgactttc tgcgagggaa ccccaacctc 4680 agtctgagga tcttcaccgc gcgcctctac ttctgtgagg accgcaaggc tgagcccgag 4740 gggctgcggc ggctgcaccg cgccggggtg caaatagcca tcatgacctt caaagattat 4800 ttttactgct ggaatacttt tgtagaaaac catgaaagaa ctttcaaagc ctgggaaggg 4860 ctgcatgaaa attcagttcg tctctccaga cagcttcggc gcatcctttt gcccctgtat 4920 gaggttgatg acttacgaga cgcatttcgt acttggggac gtgattacaa agacgatgac 4980 gataagtga 4989 <210> 2 <211> 1662 <212> PRT <213> Artificial sequence <220> <223> Description of the artificial sequence: amino acid sequence of dCas9-AID <400> 2 Met Asp Tyr Lys Asp His Asp Gly Asp Tyr Lys Asp His Asp Ile Asp 1 5 10 15 Tyr Lys Asp Asp Asp Asp Lys Met Ala Pro Lys Lys Lys Arg Lys Val 20 25 30 Gly Ile His Gly Val Pro Ala Ala Thr Met Asp Lys Lys Tyr Ser Ile 35 40 45 Gly Leu Ala Ile Gly Thr Asn Ser Val Gly Trp Ala Val Ile Thr Asp 50 55 60 Glu Tyr Lys Val Pro Ser Lys Lys Phe Lys Val Leu Gly Asn Thr Asp 65 70 75 80 Arg His Ser Ile Lys Lys Asn Leu Ile Gly Ala Leu Leu Phe Asp Ser 85 90 95 Gly Glu Thr Ala Glu Ala Thr Arg Leu Lys Arg Thr Ala Arg Arg Arg 100 105 110 Tyr Thr Arg Arg Lys Asn Arg Ile Cys Tyr Leu Gln Glu Ile Phe Ser 115 120 125 Asn Glu Met Ala Lys Val Asp Asp Ser Phe Phe His Arg Leu Glu Glu 130 135 140 Ser Phe Leu Val Glu Glu Asp Lys Lys His Glu Arg His Pro Ile Phe 145 150 155 160 Gly Asn Ile Val Asp Glu Val Ala Tyr His Glu Lys Tyr Pro Thr Ile 165 170 175 Tyr His Leu Arg Lys Lys Leu Val Asp Ser Thr Asp Lys Ala Asp Leu 180 185 190 Arg Leu Ile Tyr Leu Ala Leu Ala His Met Ile Lys Phe Arg Gly His 195 200 205 Phe Leu Ile Glu Gly Asp Leu Asn Pro Asp Asn Ser Asp Val Asp Lys 210 215 220 Leu Phe Ile Gln Leu Val Gln Thr Tyr Asn Gln Leu Phe Glu Glu Asn 225 230 235 240 Pro Ile Asn Ala Ser Gly Val Asp Ala Lys Ala Ile Leu Ser Ala Arg 245 250 255 Leu Ser Lys Ser Arg Arg Leu Glu Asn Leu Ile Ala Gln Leu Pro Gly 260 265 270 Glu Lys Lys Asn Gly Leu Phe Gly Asn Leu Ile Ala Leu Ser Leu Gly 275 280 285 Leu Thr Pro Asn Phe Lys Ser Asn Phe Asp Leu Ala Glu Asp Ala Lys 290 295 300 Leu Gln Leu Ser Lys Asp Thr Tyr Asp Asp Asp Leu Asp Asn Leu Leu 305 310 315 320 Ala Gln Ile Gly Asp Gln Tyr Ala Asp Leu Phe Leu Ala Ala Lys Asn 325 330 335 Leu Ser Asp Ala Ile Leu Leu Ser Asp Ile Leu Arg Val Asn Thr Glu 340 345 350 Ile Thr Lys Ala Pro Leu Ser Ala Ser Met Ile Lys Arg Tyr Asp Glu 355 360 365 His His Gln Asp Leu Thr Leu Leu Lys Ala Leu Val Arg Gln Gln Leu 370 375 380 Pro Glu Lys Tyr Lys Glu Ile Phe Phe Asp Gln Ser Lys Asn Gly Tyr 385 390 395 400 Ala Gly Tyr Ile Asp Gly Gly Ala Ser Gln Glu Glu Phe Tyr Lys Phe 405 410 415 Ile Lys Pro Ile Leu Glu Lys Met Asp Gly Thr Glu Glu Leu Leu Val 420 425 430 Lys Leu Asn Arg Glu Asp Leu Leu Arg Lys Gln Arg Thr Phe Asp Asn 435 440 445 Gly Ser Ile Pro His Gln Ile His Leu Gly Glu Leu His Ala Ile Leu 450 455 460 Arg Arg Gln Glu Asp Phe Tyr Pro Phe Leu Lys Asp Asn Arg Glu Lys 465 470 475 480 Ile Glu Lys Ile Leu Thr Phe Arg Ile Pro Tyr Tyr Val Gly Pro Leu 485 490 495 Ala Arg Gly Asn Ser Arg Phe Ala Trp Met Thr Arg Lys Ser Glu Glu 500 505 510 Thr Ile Thr Pro Trp Asn Phe Glu Glu Val Val Asp Lys Gly Ala Ser 515 520 525 Ala Gln Ser Phe Ile Glu Arg Met Thr Asn Phe Asp Lys Asn Leu Pro 530 535 540 Asn Glu Lys Val Leu Pro Lys His Ser Leu Leu Tyr Glu Tyr Phe Thr 545 550 555 560 Val Tyr Asn Glu Leu Thr Lys Val Lys Tyr Val Thr Glu Gly Met Arg 565 570 575 Lys Pro Ala Phe Leu Ser Gly Glu Gln Lys Lys Ala Ile Val Asp Leu 580 585 590 Leu Phe Lys Thr Asn Arg Lys Val Thr Val Lys Gln Leu Lys Glu Asp 595 600 605 Tyr Phe Lys Lys Ile Glu Cys Phe Asp Ser Val Glu Ile Ser Gly Val 610 615 620 Glu Asp Arg Phe Asn Ala Ser Leu Gly Thr Tyr His Asp Leu Leu Lys 625 630 635 640 Ile Ile Lys Asp Lys Asp Phe Leu Asp Asn Glu Glu Asn Glu Asp Ile 645 650 655 Leu Glu Asp Ile Val Leu Thr Leu Thr Leu Phe Glu Asp Arg Glu Met 660 665 670 Ile Glu Glu Arg Leu Lys Thr Tyr Ala His Leu Phe Asp Asp Lys Val 675 680 685 Met Lys Gln Leu Lys Arg Arg Arg Tyr Thr Gly Trp Gly Arg Leu Ser 690 695 700 Arg Lys Leu Ile Asn Gly Ile Arg Asp Lys Gln Ser Gly Lys Thr Ile 705 710 715 720 Leu Asp Phe Leu Lys Ser Asp Gly Phe Ala Asn Arg Asn Phe Met Gln 725 730 735 Leu Ile His Asp Asp Ser Leu Thr Phe Lys Glu Asp Ile Gln Lys Ala 740 745 750 Gln Val Ser Gly Gln Gly Asp Ser Leu His Glu His Ile Ala Asn Leu 755 760 765 Ala Gly Ser Pro Ala Ile Lys Lys Gly Ile Leu Gln Thr Val Lys Val 770 775 780 Val Asp Glu Leu Val Lys Val Met Gly Arg His Lys Pro Glu Asn Ile 785 790 795 800 Val Ile Glu Met Ala Arg Glu Asn Gln Thr Thr Gln Lys Gly Gln Lys 805 810 815 Asn Ser Arg Glu Arg Met Lys Arg Ile Glu Glu Gly Ile Lys Glu Leu 820 825 830 Gly Ser Gln Ile Leu Lys Glu His Pro Val Glu Asn Thr Gln Leu Gln 835 840 845 Asn Glu Lys Leu Tyr Leu Tyr Tyr Leu Gln Asn Gly Arg Asp Met Tyr 850 855 860 Val Asp Gln Glu Leu Asp Ile Asn Arg Leu Ser Asp Tyr Asp Val Asp 865 870 875 880 Ala Ile Val Pro Gln Ser Phe Leu Lys Asp Asp Ser Ile Asp Asn Lys 885 890 895 Val Leu Thr Arg Ser Asp Lys Asn Arg Gly Lys Ser Asp Asn Val Pro 900 905 910 Ser Glu Glu Val Val Lys Lys Met Lys Asn Tyr Trp Arg Gln Leu Leu 915 920 925 Asn Ala Lys Leu Ile Thr Gln Arg Lys Phe Asp Asn Leu Thr Lys Ala 930 935 940 Glu Arg Gly Gly Leu Ser Glu Leu Asp Lys Ala Gly Phe Ile Lys Arg 945 950 955 960 Gln Leu Val Glu Thr Arg Gln Ile Thr Lys His Val Ala Gln Ile Leu 965 970 975 Asp Ser Arg Met Asn Thr Lys Tyr Asp Glu Asn Asp Lys Leu Ile Arg 980 985 990 Glu Val Lys Val Ile Thr Leu Lys Ser Lys Leu Val Ser Asp Phe Arg 995 1000 1005 Lys Asp Phe Gln Phe Tyr Lys Val Arg Glu Ile Asn Asn Tyr His 1010 1015 1020 His Ala His Asp Ala Tyr Leu Asn Ala Val Val Gly Thr Ala Leu 1025 1030 1035 Ile Lys Lys Tyr Pro Lys Leu Glu Ser Glu Phe Val Tyr Gly Asp 1040 1045 1050 Tyr Lys Val Tyr Asp Val Arg Lys Met Ile Ala Lys Ser Glu Gln 1055 1060 1065 Glu Ile Gly Lys Ala Thr Ala Lys Tyr Phe Phe Tyr Ser Asn Ile 1070 1075 1080 Met Asn Phe Phe Lys Thr Glu Ile Thr Leu Ala Asn Gly Glu Ile 1085 1090 1095 Arg Lys Arg Pro Leu Ile Glu Thr Asn Gly Glu Thr Gly Glu Ile 1100 1105 1110 Val Trp Asp Lys Gly Arg Asp Phe Ala Thr Val Arg Lys Val Leu 1115 1120 1125 Ser Met Pro Gln Val Asn Ile Val Lys Lys Thr Glu Val Gln Thr 1130 1135 1140 Gly Gly Phe Ser Lys Glu Ser Ile Leu Pro Lys Arg Asn Ser Asp 1145 1150 1155 Lys Leu Ile Ala Arg Lys Lys Asp Trp Asp Pro Lys Lys Tyr Gly 1160 1165 1170 Gly Phe Asp Ser Pro Thr Val Ala Tyr Ser Val Leu Val Val Ala 1175 1180 1185 Lys Val Glu Lys Gly Lys Ser Lys Lys Leu Lys Ser Val Lys Glu 1190 1195 1200 Leu Leu Gly Ile Thr Ile Met Glu Arg Ser Ser Phe Glu Lys Asn 1205 1210 1215 Pro Ile Asp Phe Leu Glu Ala Lys Gly Tyr Lys Glu Val Lys Lys 1220 1225 1230 Asp Leu Ile Ile Lys Leu Pro Lys Tyr Ser Leu Phe Glu Leu Glu 1235 1240 1245 Asn Gly Arg Lys Arg Met Leu Ala Ser Ala Gly Glu Leu Gln Lys 1250 1255 1260 Gly Asn Glu Leu Ala Leu Pro Ser Lys Tyr Val Asn Phe Leu Tyr 1265 1270 1275 Leu Ala Ser His Tyr Glu Lys Leu Lys Gly Ser Pro Glu Asp Asn 1280 1285 1290 Glu Gln Lys Gln Leu Phe Val Glu Gln His Lys His Tyr Leu Asp 1295 1300 1305 Glu Ile Ile Glu Gln Ile Ser Glu Phe Ser Lys Arg Val Ile Leu 1310 1315 1320 Ala Asp Ala Asn Leu Asp Lys Val Leu Ser Ala Tyr Asn Lys His 1325 1330 1335 Arg Asp Lys Pro Ile Arg Glu Gln Ala Glu Asn Ile Ile His Leu 1340 1345 1350 Phe Thr Leu Thr Asn Leu Gly Ala Pro Ala Ala Phe Lys Tyr Phe 1355 1360 1365 Asp Thr Thr Ile Asp Arg Lys Arg Tyr Thr Ser Thr Lys Glu Val 1370 1375 1380 Leu Asp Ala Thr Leu Ile His Gln Ser Ile Thr Gly Leu Tyr Glu 1385 1390 1395 Thr Arg Ile Asp Leu Ser Gln Leu Gly Gly Asp Glu Gly Ala Pro 1400 1405 1410 Lys Lys Lys Arg Lys Val Gly Ser Ser Gly Ser Pro Lys Lys Lys 1415 1420 1425 Arg Lys Val Gly Ser Asp Ala Leu Asp Asp Phe Asp Leu Asp Met 1430 1435 1440 Leu Gly Ser Asp Ala Leu Asp Asp Phe Gly Gly Gly Ser Met Asp 1445 1450 1455 Ser Leu Leu Met Asn Arg Arg Lys Phe Leu Tyr Gln Phe Lys Asn 1460 1465 1470 Val Arg Trp Ala Lys Gly Arg Arg Glu Thr Tyr Leu Cys Tyr Val 1475 1480 1485 Val Lys Arg Arg Asp Ser Ala Thr Ser Phe Ser Leu Asp Phe Gly 1490 1495 1500 Tyr Leu Arg Asn Lys Asn Gly Cys His Val Glu Leu Leu Phe Leu 1505 1510 1515 Arg Tyr Ile Ser Asp Trp Asp Leu Asp Pro Gly Arg Cys Tyr Arg 1520 1525 1530 Val Thr Trp Phe Thr Ser Trp Ser Pro Cys Tyr Asp Cys Ala Arg 1535 1540 1545 His Val Ala Asp Phe Leu Arg Gly Asn Pro Asn Leu Ser Leu Arg 1550 1555 1560 Ile Phe Thr Ala Arg Leu Tyr Phe Cys Glu Asp Arg Lys Ala Glu 1565 1570 1575 Pro Glu Gly Leu Arg Arg Leu His Arg Ala Gly Val Gln Ile Ala 1580 1585 1590 Ile Met Thr Phe Lys Asp Tyr Phe Tyr Cys Trp Asn Thr Phe Val 1595 1600 1605 Glu Asn His Glu Arg Thr Phe Lys Ala Trp Glu Gly Leu His Glu 1610 1615 1620 Asn Ser Val Arg Leu Ser Arg Gln Leu Arg Arg Ile Leu Leu Pro 1625 1630 1635 Leu Tyr Glu Val Asp Asp Leu Arg Asp Ala Phe Arg Thr Trp Gly 1640 1645 1650 Arg Asp Tyr Lys Asp Asp Asp Asp Lys 1655 1660 <210> 3 <211> 4941 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequence: Coding sequence of dCas9-AIDX <400> 3 atggactata aggaccacga cggagactac aaggatcatg atattgatta caaagacgat 60 gacgataaga tggccccaaa gaagaagcgg aaggtcggta tccacggagt cccagcagct 120 accatggaca agaagtattc tatcggactg gccatcggga ctaatagcgt cgggtgggcc 180 gtgatcactg acgagtacaa ggtgccctct aagaagttca aggtgctcgg gaacaccgac 240 cggcattcca tcaagaaaaa tctgatcgga gctctcctct ttgattcagg ggagaccgct 300 gaagcaaccc gcctcaagcg gactgctaga cggcggtaca ccaggaggaa gaaccggatt 360 tgttaccttc aagagatatt ctccaacgaa atggcaaagg tcgacgacag cttcttccat 420 aggctggaag aatcattcct cgtggaagag gataagaagc atgaacggca tcccatcttc 480 ggtaatatcg tcgacgaggt ggcctatcac gagaaatacc caaccatcta ccatcttcgc 540 aaaaagctgg tggactcaac cgacaaggca gacctccggc ttatctacct ggccctggcc 600 cacatgatca agttcagagg ccacttcctg atcgagggcg acctcaatcc tgacaatagc 660 gatgtggata aactgttcat ccagctggtg cagacttaca accagctctt tgaagagaac 720 cccatcaatg caagcggagt cgatgccaag gccattctgt cagcccggct gtcaaagagc 780 cgcagacttg agaatcttat cgctcagctg ccgggtgaaa agaaaaatgg actgttcggg 840 aacctgattg ctctttcact tgggctgact cccaatttca agtctaattt cgacctggca 900 gaggatgcca agctgcaact gtccaaggac acctatgatg acgatctcga caacctcctg 960 gcccagatcg gtgaccaata cgccgacctt ttccttgctg ctaagaatct ttctgacgcc 1020 atcctgctgt ctgacattct ccgcgtgaac actgaaatca ccaaggcccc tctttcagct 1080 tcaatgatta agcggtatga tgagcaccac caggacctga ccctgcttaa ggcactcgtc 1140 cggcagcagc ttccggagaa gtacaaggaa atctctttg accagtcaa gatggatac 1200 gccggctaca tcgacggagg tgcctcccaa gaggaatttt ataagtttat caacctac 1260 cttgagaga tggacggcac cgaagagctc ctcgtgaac tgaatcggga ggatctgctg 1320 cggaagcagc gcacttcga caatgggagc attccccacc agatcatct tggggagctt 1380 cacgccatcc ttcggcgcca agaggacttc tacccctttc ttaggacaa cagggagaag 1440 attgagaaaa ttctcactt ccgcatcccc tactacgtgg gacccctcgc cagaggaaat 1500 agccggtttg cttggatgac cagaagtca gagaaacta tcactccctg gaacttcgaa 1560 gaggtggtgg acaagggagc cagcgctcag tcattcatcg aacggatgac taacttcgat 1620 aagaacctcc ccaatgagaa ggtcctgccg aacattccc tgcttacga gtactttacc 1680 gtgtacaacg agctgaccaa ggtgaaatat gtcaccgaag ggatgaggaa gcccgcattc 1740 ctgtcaggcg aaaaagaaaggcaattgtg gaccttctgt tcagaccaa tagaaggtg 1800 accgtgaagc agctgaagga ggactttc aagaaaattg aatgcttcga ctctgtggag 1860 attagcgggg tcgagatcg gttcaacgca agcctggta cctaccatga tctgcttaag 1920 atcatcagg acagatt tctgacat gaggagaacg aggacatcct tgaggacatt 1980 gtcctgactc tcactctt cgaggaccgg gaatgatcg aggagaggct tagacctac 2040 gcccatctgt tcgacgataa agtgatgaag caacttaac ggagaata taccggatgg 2100 ggacgcctta gccgcaact catcaacgga atccgggaca aacagcgg aagaccatt 2160 cttgatttcc ttagagcga cggattcgct aatcgcact tcatgcact tatccatgat 2220 gattccctga cctttaagga ggacatccag aaggcccaag tgtctggaca aggtgactca 2280 ctgcacgagc atatcgcaa tctggctggt tcaccgcta ttaagaaggg tattctccag 2340 accgtgaaag tcgtggacga gctggtcaag gtgatgggtc gccataacc agagaacatt 2400 gtcatcgaga tggccaggga aaaccagact acccagaagg gagagaga caggcagggag 2460 cggatgaaaa gattgagga agggattaag gagctcgggt cacagatccct taagagcac 2520 ccggtggaaa acacccagct tcagaatgag aagctctatc tgtacct tcaaatgga 2580 cgcgatatgt atgtggacca agagcttgat atcaacaggc tctcagacta cgacgtggac 2640 gccatcgtcc ctcagagctt cctcaaagac gactcaattg acaataaggt gctgactcgc 2700 tcagacaaga accggggaaa gtcagataac gtgccctcag aggaagtcgt gaaaaagatg 2760 aagaactatt ggcgccagct tctgaacgca aagctgatca ctcagcggaa gttcgacaat 2820 ctcactaagg ctgagagggg cggactgagc gaactggaca aagcaggatt cattaaacgg 2880 caacttgtgg agactcggca gattactaaa catgtcgccc aaatccttga ctcacgcatg 2940 aataccaagt acgacgaaaa cgacaaactt atccgcgagg tgaaggtgat taccctgaag 3000 tccaagctgg tcagcgattt cagaaaggac tttcaattct acaaagtgcg ggagatcaat 3060 aactatcatc atgctcatga cgcatatctg aatgccgtgg tgggaaccgc cctgatcaag 3120 aagtacccaa agctggaaag cgagttcgtg tacggagact acaaggtcta cgacgtgcgc 3180 aagatgattg ccaaatctga gcaggagatc ggaaaggcca ccgcaaagta cttcttctac 3240 agcaacatca tgaatttctt caagaccgaa atcacccttg caaacggtga gatccggaag 3300 aggccgctca tcgagactaa tggggagact ggcgaaatcg tgtgggacaa gggcagagat 3360 ttcgctaccg tgcgcaaagt gctttctatg cctcaagtga acatcgtgaa gaaaaccgag 3420 gtgcaaaccg gaggctttc taaggaatca atcctcccca agcgcaactc cgacaagctc 3480 attgcaagga agaaggattg ggaccctaag aagtacggcg gattcgattc accaactgtg 3540 gcttattctg tcctggtcgt ggctaaggtg gaaaaaggaa agtctaagaa gctcaagagc 3600 gtgaaggaac tgctgggtat caccattatg gagcgcagct ccttcgagaa gaacccaatt 3660 gactttctcg aagccaaagg ttacaaggaa gtcaagaagg accttatcat caagctccca 3720 aagtatagcc tgttcgaact ggagaatggg cggaagcgga tgctcgcctc cgctggcgaa 3780 cttcagaagg gtaatgagct ggctctcccc tccaagtacg tgaatttcct ctaccttgca 3840 agccattacg agaagctgaa ggggagcccc gaggacaacg agcaaaagca actgtttgtg 3900 gagcagcata agcattatct ggacgagatc attgagcaga tttccgagtt ttctaaacgc 3960 gtcattctcg ctgatgccaa cctcgataaa gtccttagcg catacaataa gcacagagac 4020 aaaccaattc gggagcaggc tgagaatatc atccacctgt tcaccctcac caatcttggt 4080 gcccctgccg cattcaagta cttcgacacc accatcgacc ggaaacgcta tacctccacc 4140 aaagaagtgc tggacgccac cctcatccac cagagcatca ccggacttta cgaaactcgg 4200 attgacctct cacagctcgg aggggatgag ggagctccca agaaaaagcg caaggtaggt 4260 agttccggat ctccgaaaaa gaaacgcaaa gttggtagtg atgctttaga cgattttgac 4320 ttagatatgc ttggttcaga cgcgttagac gacttcggtg gaggatccat ggacagcctc 4380 ttgatgaacc ggaggaagtt tctttaccaa ttcaaaaatg tccgctgggc taagggtcgg 4440 cgtgagacct acctgtgcta cgtagtgaag aggcgtgaca gtgctacatc cttttcactg 4500 gactttggtt atcttcgcaa taagaacggc tgccacgtgg aattgctctt cctccgctac 4560 atctcggact gggacctaga ccctggccgc tgctaccgcg tcacctggtt cacctcctgg 4620 agcccctgct acgactgtgc ccgacatgtg gccgactttc tgcgagggaa ccccaacctc 4680 agtctgagga tcttcaccgc gcgcctctac ttctgtgagg accgcaaggc tgagcccgag 4740 gggctgcggc ggctgcaccg cgccggggtg caaatagcca tcatgacctt caaagattat 4800 ttttactgct ggaatacttt tgtagaaaac catgaaagaa ctttcaaagc ctgggaaggg 4860 ctgcatgaaa attcagttcg tctctccaga cagcttcggc gcatcctttt gcccgattac 4920 aaagacgatg acgataagtg a 4941 <210> 4 <211> 1646 <212> PRT <213> Artificial sequence <220> <223> Description of the artificial sequence: amino acid sequence of dCas9-AIDX <400> 4 Met Asp Tyr Lys Asp His Asp Gly Asp Tyr Lys Asp His Asp Ile Asp 1 5 10 15 Tyr Lys Asp Asp Asp Asp Lys Met Ala Pro Lys Lys Lys Arg Lys Val 20 25 30 Gly Ile His Gly Val Pro Ala Ala Thr Met Asp Lys Lys Tyr Ser Ile 35 40 45 Gly Leu Ala Ile Gly Thr Asn Ser Val Gly Trp Ala Val Ile Thr Asp 50 55 60 Glu Tyr Lys Val Pro Ser Lys Lys Phe Lys Val Leu Gly Asn Thr Asp 65 70 75 80 Arg His Ser Ile Lys Lys Asn Leu Ile Gly Ala Leu Leu Phe Asp Ser 85 90 95 Gly Glu Thr Ala Glu Ala Thr Arg Leu Lys Thr Arg Ala Arg Arg 100 105 110 Tyr Thyr Arg Lys Arg Asn with Cys Tyr Leu Gln Glu With Phe Ser 115 120 125 Asn Glu Met Ala Lys Val Asp Asp Ser Phe Phe His Arg Leu Glu Glu 130 135 140 Ser Phe Leu Val Glu Glu Asp Lys Lys His Glu Arg His Pro Ile Phe 145 150 155 160 Gly Asn Ile Val Asp Glu Val Ala Tyr His Glu Lys Tyr Pro Thr Ile 165 170 175 Tyr His Leu Arg Lys Lys Leu Val Asp Ser Thr Asp Lys Ala Asp Leu 180 185 190 Arg Leu Ile Tyr Leu Ala Leu Ala His Met Ile Lys Phe Arg Gly His 195 200 205 Phe Leu Ile Glu Gly Asp Leu Asn Pro Asp Asn Ser Asp Val Asp Lys 210 215 220 Leu Phe Ile Gln Leu Val Gln Thr Tyr Asn Gln Leu Phe Glu Glu Asn 225 230 235 240 Pro Ile Asn Ala Ser Gly Val Asp Ala Lys Ala Ile Leu Ser Ala Arg 245 250 255 Leu Ser Lys Ser Arg Arg Leu Glu Asn Leu Ile Ala Gln Leu Pro Gly 260 265 270 Glu Lys Lys Asn Gly Leu Phe Gly Asn Leu Ile Ala Leu Ser Leu Gly 275 280 285 Leu Thr Pro Asn Phe Lys Ser Asn Phe Asp Leu Ala Glu Asp Ala Lys 290 295 300 Leu Gln Leu Ser Lys Asp Thr Tyr Asp Asp Asp Leu Asp Asn Leu Leu 305 310 315 320 Ala Gln Ile Gly Asp Gln Tyr Ala Asp Leu Phe Leu Ala Ala Lys Asn 325 330 335 Leu Ser Asp Ala Ile Leu Leu Ser Asp Ile Leu Arg Val Asn Thr Glu 340 345 350 Ile Thr Lys Ala Pro Leu Ser Ala Ser Met Ile Lys Arg Tyr Asp Glu 355 360 365 His His Gln Asp Leu Thr Leu Leu Lys Ala Leu Val Arg Gln Gln Leu 370 375 380 Pro Glu Lys Tyr Lys Glu Ile Phe Phe Asp Gln Ser Lys Asn Gly Tyr 385 390 395 400 Ala Gly Tyr Ile Asp Gly Gly Ala Ser Gln Glu Glu Phe Tyr Lys Phe 405 410 415 Ile Lys Pro Ile Leu Glu Lys Met Asp Gly Thr Glu Glu Leu Leu Val 420 425 430 Lys Leu Asn Arg Glu Asp Leu Leu Arg Lys Gln Arg Thr Phe Asp Asn 435 440 445 Gly Ser Ile Pro His Gln Ile His Leu Gly Glu Leu His Ala Ile Leu 450 455 460 Arg Arg Gln Glu Asp Phe Tyr Pro Phe Leu Lys Asp Asn Arg Glu Lys 465 470 475 480 Ile Glu Lys Ile Leu Thr Phe Arg Ile Pro Tyr Tyr Val Gly Pro Leu 485 490 495 Ala Arg Gly Asn Ser Arg Phe Ala Trp Met Thr Arg Lys Ser Glu Glu 500 505 510 Thr Ile Thr Pro Trp Asn Phe Glu Glu Val Val Asp Lys Gly Ala Ser 515 520 525 Ala Gln Ser Phe Ile Glu Arg Met Thr Asn Phe Asp Lys Asn Leu Pro 530 535 540 Asn Glu Lys Val Leu Pro Lys His Ser Leu Leu Tyr Glu Tyr Phe Thr 545 550 555 560 Val Tyr Asn Glu Leu Thr Lys Val Lys Tyr Val Thr Glu Gly Met Arg 565 570 575 Lys Pro Ala Phe Leu Ser Gly Glu Gln Lys Lys Ala Ile Val Asp Leu 580 585 590 Leu Phe Lys Thr Asn Arg Lys Val Thr Val Lys Gln Leu Lys Glu Asp 595 600 605 Tyr Phe Lys Lys Ile Glu Cys Phe Asp Ser Val Glu Ile Ser Gly Val 610 615 620 Glu Asp Arg Phe Asn Ala Ser Leu Gly Thr Tyr His Asp Leu Leu Lys 625 630 635 640 Ile Ile Lys Asp Lys Asp Phe Leu Asp Asn Glu Glu Asn Glu Asp Ile 645 650 655 Leu Glu Asp Ile Val Leu Thr Leu Thr Leu Phe Glu Asp Arg Glu Met 660 665 670 Ile Glu Glu Arg Leu Lys Thr Tyr Ala His Leu Phe Asp Asp Lys Val 675 680 685 Met Lys Gln Leu Lys Arg Arg Arg Tyr Thr Gly Trp Gly Arg Leu Ser 690 695 700 Arg Lys Leu Ile Asn Gly Ile Arg Asp Lys Gln Ser Gly Lys Thr Ile 705 710 715 720 Leu Asp Phe Leu Lys Ser Asp Gly Phe Ala Asn Arg Asn Phe Met Gln 725 730 735 Leu Ile His Asp Asp Ser Leu Thr Phe Lys Glu Asp Ile Gln Lys Ala 740 745 750 Gln Val Ser Gly Gln Gly Asp Ser Leu His Glu His Ile Ala Asn Leu 755 760 765 Ala Gly Ser Pro Ala Ile Lys Lys Gly Ile Leu Gln Thr Val Lys Val 770 775 780 Val Asp Glu Leu Val Lys Val Met Gly Arg His Lys Pro Glu Asn Ile 785 790 795 800 Val Ile Glu Met Ala Arg Glu Asn Gln Thr Thr Gln Lys Gly Gln Lys 805 810 815 Asn Ser Arg Glu Arg Met Lys Arg Ile Glu Glu Gly Ile Lys Glu Leu 820 825 830 Gly Ser Gln Ile Leu Lys Glu His Pro Val Glu Asn Thr Gln Leu Gln 835 840 845 Asn Glu Lys Leu Tyr Leu Tyr Tyr Leu Gln Asn Gly Arg Asp Met Tyr 850 855 860 Val Asp Gln Glu Leu Asp Ile Asn Arg Leu Ser Asp Tyr Asp Val Asp 865 870 875 880 Ala Ile Val Pro Gln Ser Phe Leu Lys Asp Asp Ser Ile Asp Asn Lys 885 890 895 Val Leu Thr Arg Ser Asp Lys Asn Arg Gly Lys Ser Asp Asn Val Pro 900 905 910 Ser Glu Glu Val Val Lys Lys Met Lys Asn Tyr Trp Arg Gln Leu Leu 915 920 925 Asn Ala Lys Leu Ile Thr Gln Arg Lys Phe Asp Asn Leu Thr Lys Ala 930 935 940 Glu Arg Gly Gly Leu Ser Glu Leu Asp Lys Ala Gly Phe Ile Lys Arg 945 950 955 960 Gln Leu Val Glu Thr Arg Gln Ile Thr Lys His Val Ala Gln Ile Leu 965 970 975 Asp Ser Arg Met Asn Thr Lys Tyr Asp Glu Asn Asp Lys Leu Ile Arg 980 985 990 Glu Val Lys Val Ile Thr Leu Lys Ser Lys Leu Val Ser Asp Phe Arg 995 1000 1005 Lys Asp Phe Gln Phe Tyr Lys Val Arg Glu Ile Asn Asn Tyr His 1010 1015 1020 His Ala His Asp Ala Tyr Leu Asn Ala Val Val Gly Thr Ala Leu 1025 1030 1035 Ile Lys Lys Tyr Pro Lys Leu Glu Ser Glu Phe Val Tyr Gly Asp 1040 1045 1050 Tyr Lys Val Tyr Asp Val Arg Lys Met Ile Ala Lys Ser Glu Gln 1055 1060 1065 Glu Ile Gly Lys Ala Thr Ala Lys Tyr Phe Phe Tyr Ser Asn Ile 1070 1075 1080 Met Asn Phe Phe Lys Thr Glu Ile Thr Leu Ala Asn Gly Glu Ile 1085 1090 1095 Arg Lys Arg Pro Leu Ile Glu Thr Asn Gly Glu Thr Gly Glu Ile 1100 1105 1110 Val Trp Asp Lys Gly Arg Asp Phe Ala Thr Val Arg Lys Val Leu 1115 1120 1125 Ser Met Pro Gln Val Asn Ile Val Lys Lys Thr Glu Val Gln Thr 1130 1135 1140 Gly Gly Phe Ser Lys Glu Ser Ile Leu Pro Lys Arg Asn Ser Asp 1145 1150 1155 Lys Leu Ile Ala Arg Lys Lys Asp Trp Asp Pro Lys Lys Tyr Gly 1160 1165 1170 Gly Phe Asp Ser Pro Thr Val Ala Tyr Ser Val Leu Val Val Ala 1175 1180 1185 Lys Val Glu Lys Gly Lys Ser Lys Lys Leu Lys Ser Val Lys Glu 1190 1195 1200 Leu Leu Gly Ile Thr Ile Met Glu Arg Ser Ser Phe Glu Lys Asn 1205 1210 1215 Pro Ile Asp Phe Leu Glu Ala Lys Gly Tyr Lys Glu Val Lys Lys 1220 1225 1230 Asp Leu Ile Ile Lys Leu Pro Lys Tyr Ser Leu Phe Glu Leu Glu 1235 1240 1245 Asn Gly Arg Lys Arg Met Leu Ala Ser Ala Gly Glu Leu Gln Lys 1250 1255 1260 Gly Asn Glu Leu Ala Leu Pro Ser Lys Tyr Val Asn Phe Leu Tyr 1265 1270 1275 Leu Ala Ser His Tyr Glu Lys Leu Lys Gly Ser Pro Glu Asp Asn 1280 1285 1290 Glu Gln Lys Gln Leu Phe Val Glu Gln His Lys His Tyr Leu Asp 1295 1300 1305 Glu Ile Ile Glu Gln Ile Ser Glu Phe Ser Lys Arg Val Ile Leu 1310 1315 1320 Ala Asp Ala Asn Leu Asp Lys Val Leu Ser Ala Tyr Asn Lys His 1325 1330 1335 Arg Asp Lys Pro Ile Arg Glu Gln Ala Glu Asn Ile Ile His Leu 1340 1345 1350 Phe Thr Leu Thr Asn Leu Gly Ala Pro Ala Ala Phe Lys Tyr Phe 1355 1360 1365 Asp Thr Thr Ile Asp Arg Lys Arg Tyr Thr Ser Thr Lys Glu Val 1370 1375 1380 Leu Asp Ala Thr Leu Ile His Gln Ser Ile Thr Gly Leu Tyr Glu 1385 1390 1395 Thr Arg Ile Asp Leu Ser Gln Leu Gly Gly Asp Glu Gly Ala Pro 1400 1405 1410 Lys Lys Lys Arg Lys Val Gly Ser Ser Gly Ser Pro Lys Lys Lys 1415 1420 1425 Arg Lys Val Gly Ser Asp Ala Leu Asp Asp Phe Asp Leu Asp Met 1430 1435 1440 Leu Gly Ser Asp Ala Leu Asp Asp Phe Gly Gly Gly Ser Met Asp 1445 1450 1455 Ser Leu Leu Met Asn Arg Arg Lys Phe Leu Tyr Gln Phe Lys Asn 1460 1465 1470 Val Arg Trp Ala Lys Gly Arg Arg Glu Thr Tyr Leu Cys Tyr Val 1475 1480 1485 Val Lys Arg Arg Asp Ser Ala Thr Ser Phe Ser Leu Asp Phe Gly 1490 1495 1500 Tyr Leu Arg Asn Lys Asn Gly Cys His Val Glu Leu Leu Phe Leu 1505 1510 1515 Arg Tyr Ile Ser Asp Trp Asp Leu Asp Pro Gly Arg Cys Tyr Arg 1520 1525 1530 Val Thr Trp Phe Thr Ser Trp Ser Pro Cys Tyr Asp Cys Ala Arg 1535 1540 1545 His Val Ala Asp Phe Leu Arg Gly Asn Pro Asn Leu Ser Leu Arg 1550 1555 1560 Ile Phe Thr Ala Arg Leu Tyr Phe Cys Glu Asp Arg Lys Ala Glu 1565 1570 1575 Pro Glu Gly Leu Arg Arg Leu His Arg Ala Gly Val Gln Ile Ala 1580 1585 1590 Ile Met Thr Phe Lys Asp Tyr Phe Tyr Cys Trp Asn Thr Phe Val 1595 1600 1605 Glu Asn His Glu Arg Thr Phe Lys Ala Trp Glu Gly Leu His Glu 1610 1615 1620 Asn Ser Val Arg Leu Ser Arg Gln Leu Arg Arg Ile Leu Leu Pro 1625 1630 1635 Asp Tyr Lys Asp Asp Asp Asp Lys 1640 1645 <210> 5 <211> 28 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: primers <400> 5 gcggatccat ggacagcctc ttgatgaa 28 <210> 6 <211> 54 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: primers <400> 6 actcgagtca cttatcgtca tcgtctttgt aatcacgtcc ccaagtacga aatg 54 <210> 7 <211> 55 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: primers <400> 7 gactcgagtc acttatcgtc atcgtctttg taatcgggca aaaggatgcg ccgaa 55 <210> 8 <211> 34 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: primers <400> 8 gcagatctac catggacaag aagtattcta tcgg 34 <210> 9 <211> 35 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: primers <400> 9 gactcgagtc acttatcgtc atcgtctttg taatc 35 <210> 10 <211> 33 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: primers <400> 10 aaactcgagt gtacaaaaaa gcaggcttta aag 33 <210> 11 <211> 37 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: primers <220> <221> misc_feature <222> (2)..(20) <223> n is a, c, g or t <400> 11 gnnnnnnnnn nnnnnnnnnn ggtgtttcgt cctttcc 37 <210> 12 <211> 42 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: primers <220> <221> misc_feature <222> (2)..(20) <223> n is a, c, g or t <400> 12 gnnnnnnnnn nnnnnnnnnn gttttagagc tagaaatagc aa 42 <210> 13 <211> 36 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: primers <400> 13 aaagctagct aatgccaact ttgtacaaga aagctg 36 <210> 14 <211> twenty one <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: target binding region of sgRNA <400> 14 gcatgcccga aggctacgtc c 21 <210> 15 <211> twenty one <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: target binding region of sgRNA <400> 15 gcaactagta tacccgcgcc g 21 <210> 16 <211> twenty one <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: target binding region of sgRNA <400> 16 gcctcgaact tcacctcggc g 21 <210> 17 <211> twenty one <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: target binding region of sgRNA <400> 17 gtcagctcga tgcggttcac c 21 <210> 18 <211> twenty one <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: target binding region of sgRNA <400> 18 gattcccagg gccggttaat g 21 <210> 19 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: target binding region of sgRNA <400> 19 gtcccctcca ccccacagtg 20 <210> 20 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: target binding region of sgRNA <400> 20 ggggccacta gggacaggat 20 <210> twenty one <211> twenty one <212> PRT <213> Artificial sequence <220> <223> Description of artificial sequences: adapters <400> twenty one Gly Ser Gly Gly Gly Gly Gly Ser Gly Gly Gly Gly Ser Gly Gly Gly Leu 1 5 10 15 Gly Ser Thr Glu Phe 20 <210> twenty two <211> twenty one <212> PRT <213> Artificial sequence <220> <223> Description of artificial sequences: adapters <400> twenty two Arg Ser Thr Ser Gly Leu Gly Gly Gly Ser Gly Gly Gly Gly Ser Gly 1 5 10 15 Gly Gly Gly Ser Gly 20 <210> twenty three <211> twenty one <212> PRT <213> Artificial sequence <220> <223> Description of artificial sequences: adapters <400> twenty three Gln Leu Thr Ser Gly Leu Gly Gly Gly Ser Gly Gly Gly Gly Ser Gly 1 5 10 15 Gly Gly Gly Ser Gly 20 <210> twenty four <211> 4 <212> PRT <213> Artificial sequence <220> <223> Description of artificial sequences: adapters <400> twenty four Gly Gly Gly Ser 1 <210> 25 <211> 5 <212> PRT <213> Artificial sequence <220> <223> Description of artificial sequences: adapters <400> 25 Gly Gly Gly Gly Ser 1 5 <210> 26 <211> 5 <212> PRT <213> Artificial sequence <220> <223> Description of artificial sequences: adapters <400> 26 Ser Ser Ser Ser Gly 1 5 <210> 27 <211> 5 <212> PRT <213> Artificial sequence <220> <223> Description of artificial sequences: adapters <400> 27 Gly Ser Gly Ser Ala 1 5 <210> 28 <211> 20 <212> PRT <213> Artificial sequence <220> <223> Description of artificial sequences: adapters <400> 28 Gly Gly Ser Gly Gly Gly Gly Gly Gly Gly Ser Gly Gly Gly Gly Ser Gly 1 5 10 15 Gly Gly Gly Ser 20 <210> 29 <211> 15 <212> PRT <213> Artificial sequence <220> <223> Description of artificial sequences: adapters <400> 29 Ser Ser Ser Ser Gly Ser Ser Ser Ser Gly Ser Ser Ser Ser Gly 1 5 10 15 <210> 30 <211> 15 <212> PRT <213> Artificial sequence <220> <223> Description of artificial sequences: adapters <400> 30 Gly Ser Gly Ser Ala Gly Ser Gly Ser Ala Gly Ser Gly Ser Ala 1 5 10 15 <210> 31 <211> 15 <212> PRT <213> Artificial sequence <220> <223> Description of artificial sequences: adapters <400> 31 Gly Gly Ser Gly Gly Gly Gly Ser Gly Gly Gly Gly Ser Gly Gly 1 5 10 15 <210> 32 <211> 8 <212> PRT <213> Artificial sequence <220> <223> Description of artificial sequence: FLAG tag <400> 32 Asp Tyr Lys Asp Asp Asp Asp Lys 1 5 <210> 33 <211> 7 <212> PRT <213> Artificial sequence <220> <223> Description of artificial sequences: nuclear localization sequences <400> 33 Pro Lys Lys Lys Arg Lys Val 1 5 <210> 34 <211> 16 <212> PRT <213> Artificial sequence <220> <223> Description of artificial sequences: nuclear localization sequences <400> 34 Lys Arg Pro Ala Ala Thr Lys Lys Ala Gly Gln Ala Lys Lys Lys Lys 1 5 10 15 <210> 35 <211> 9 <212> PRT <213> Artificial sequence <220> <223> Description of artificial sequences: nuclear localization sequences <400> 35 Pro Ala Ala Lys Arg Val Lys Leu Asp 1 5 <210> 36 <211> 11 <212> PRT <213> Artificial sequence <220> <223> Description of artificial sequences: nuclear localization sequences <400> 36 Arg Gln Arg Arg Asn Glu Leu Lys Arg Ser Pro 1 5 10 <210> 37 <211> 38 <212> PRT <213> Artificial sequence <220> <223> Description of artificial sequences: nuclear localization sequences <400> 37 Asn Gln Ser Ser Asn Phe Gly Pro Met Lys Gly Gly Asn Phe Gly Gly 1 5 10 15 Arg Ser Ser Gly Pro Tyr Gly Gly Gly Gly Gln Tyr Phe Ala Lys Pro 20 25 30 Arg Asn Gln Gly Gly Tyr 35 <210> 38 <211> 42 <212> PRT <213> Artificial sequence <220> <223> Description of artificial sequences: nuclear localization sequences <400> 38 Arg Met Arg Ile Glx Phe Lys Asn Lys Gly Lys Asp Thr Ala Glu Leu 1 5 10 15 Arg Arg Arg Arg Val Glu Val Ser Val Glu Leu Arg Lys Ala Lys Lys 20 25 30 Asp Glu Gln Ile Leu Lys Arg Arg Asn Val 35 40 <210> 39 <211> 8 <212> PRT <213> Artificial sequence <220> <223> Description of artificial sequences: nuclear localization sequences <400> 39 Val Ser Arg Lys Arg Pro Arg Pro 1 5 <210> 40 <211> 8 <212> PRT <213> Artificial sequence <220> <223> Description of artificial sequences: nuclear localization sequences <400> 40 Pro Pro Lys Lys Ala Arg Glu Asp 1 5 <210> 41 <211> 12 <212> PRT <213> Artificial sequence <220> <223> Description of artificial sequences: nuclear localization sequences <400> 41 Ser Ala Leu Ile Lys Lys Lys Lys Lys Lys Met Ala Pro 1 5 10 <210> 42 <211> 5 <212> PRT <213> Artificial sequence <220> <223> Description of artificial sequences: nuclear localization sequences <400> 42 Asp Arg Leu Arg Arg 1 5 <210> 43 <211> 7 <212> PRT <213> Artificial sequence <220> <223> Description of artificial sequences: nuclear localization sequences <400> 43 Pro Lys Gln Lys Lys Arg Lys 1 5 <210> 44 <211> 10 <212> PRT <213> Artificial sequence <220> <223> Description of artificial sequences: nuclear localization sequences <400> 44 Arg Lys Leu Lys Lys Lys Ile Lys Lys Leu 1 5 10 <210> 45 <211> 10 <212> PRT <213> Artificial sequence <220> <223> Description of artificial sequences: nuclear localization sequences <400> 45 Arg Glu Lys Lys Lys Phe Leu Lys Arg Arg 1 5 10 <210> 46 <211> 20 <212> PRT <213> Artificial sequence <220> <223> Description of artificial sequences: nuclear localization sequences <400> 46 Lys Arg Lys Gly Asp Glu Val Asp Gly Val Asp Glu Val Ala Lys Lys 1 5 10 15 Lys Ser Lys Lys 20 <210> 47 <211> 17 <212> PRT <213> Artificial sequence <220> <223> Description of artificial sequences: nuclear localization sequences <400> 47 Arg Lys Cys Leu Gln Ala Gly Met Asn Leu Glu Ala Arg Lys Thr Lys 1 5 10 15 Lys <210> 48 <211> 644 <212> DNA <213> Homo sapiens <400> 48 acaagttcag cgtgtctggc gagggcgagg gcgatgccac ctacggcaag ctgaccctga 60 agttcatctg caccaccggc aagctgcccg tgccctggcc caccctcgtg accaccctga 120 cctacggcgt gcagtgcttc agccgctacc ccgaccacat gaagcagcac gacttcttca 180 agtccgccat gcccgaaggc tacgtccagg agcgcaccat cttcttcaag gacgacggca 240 actagtatac ccgcgccgag gtgaagttcg agggcgacac cctggtgaac cgcatcgagc 300 tgaagggcat cgacttcaag gaggacggca acatcctggg gcacaagctg gagtacaact 360 acaacagcca caacgtctat atcatggccg acaagcagaa gaacggcatc aaggcgaact 420 tcaagatccg ccacaacatc gaggacggca gcgtgcagct cgccgaccac taccagcaga 480 acacccccat cggcgacggc cccgtgctgc tgcccgacaa ccactacctg agcacccagt 540 ccgccctgag caaagacccc aacgagaagc gcgatcacat ggtcctgctg gagttcgtga 600 ccgccgccgg gatcactctc ggcatggacg agctgtacaa gtaa 644 <210> 49 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: target binding region of sgRNA <400> 49 tagacagttg tttgttcagt 20 <210> 50 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: target binding region of sgRNA <400> 50 gtcctcgttg tcttgttggc 20 <210> 51 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: target binding region of sgRNA <400> 51 gttggcaggg gtctgcaccc 20 <210> 52 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: target binding region of sgRNA <400> 52 tcactgagtt catgacctac 20 <210> 53 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: target binding region of sgRNA <400> 53 catgacctac gggaacctcc 20 <210> 54 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: target binding region of sgRNA <400> 54 cctgagggag tgcaaccggc 20 <210> 55 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: target binding region of sgRNA <400> 55 ccggcaggag gtgaacgccg 20 <210> 56 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: target binding region of sgRNA <400> 56 cgccgtggtg ctgctgtaca 20 <210> 57 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: target binding region of sgRNA <400> 57 ctcgtcagcc atggagtacc 20 <210> 58 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: target binding region of sgRNA <400> 58 aaaaacttca tccacaggta 20 <210> 59 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: target binding region of sgRNA <400> 59 agcctgcgcc atggagtcac 20 <210> 60 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: target binding region of sgRNA <400> 60 ggagtcacag ggcgtggagc 20 <210> 61 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: target binding region of sgRNA <400> 61 acaacgagga cttcaacacg 20 <210> 62 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: target binding region of sgRNA <400> 62 tcagtgatga tatagaacgg 20 <210> 63 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: target binding region of sgRNA <400> 63 tgcactccct caggtagtcc 20 <210> 64 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: target binding region of sgRNA <400> 64 gccctgtgac tccatggcgc 20 <210> 65 <211> 4731 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequence: Coding sequence of AIDX-XTEN-dCas9 <400> 65 atggacagcc tcttgatgaa ccggaggaag tttctttacc aattcaaaaa tgtccgctgg 60 gctaagggtc ggcgtgagac ctacctgtgc tacgtagtga agaggcgtga cagtgctaca 120 tccttttcac tggactttgg ttatcttcgc aataagaacg gctgccacgt ggaattgctc 180 ttcctccgct acatctcgga ctgggaccta gaccctggcc gctgctaccg cgtcacctgg 240 ttcacctcct ggagcccctg ctacgactgt gcccgacatg tggccgactt tctgcgaggg 300 aaccccaacc tcagtctgag gatcttcacc gcgcgcctct acttctgtga ggaccgcaag 360 gctgagcccg aggggctgcg gcggctgcac cgcgccgggg tgcaaatagc catcatgacc 420 ttcaaagatt atttttactg ctggaatact tttgtagaaa accatgaaag aactttcaaa 480 gcctgggaag ggctgcatga aaattcagtt cgtctctcca gacagcttcg gcgcatcctt 540 ttgcccagcg gcagcgagac tcccgggacc tcagagtccg ccacacccga aagtgataaa 600 aagtattcta ttggtttagc catcggcact aattccgttg gatgggctgt cataaccgat 660 gaatacaaag taccttcaaa gaaatttaag gtgttggga acacagaccg tcattcgatt 720 aaaaagaatc ttatcggtgc cctcctattc gatagtggcg aaacggcaga ggcgactcgc 780 ctgaaacgaa ccgctcggag aaggtataca cgtcgcaaga accgaatatg ttacttacaa 840 gaaatttta gcaatgagat ggccaaagtt gacgattctt tctttcaccg tttggaagag 900 960 gatgaggtgg catatcatga aaagtaccca acgatttatc acctcagaaa aaagctagtt 1020 gactcaactg ataaagcgga cctgaggtta atctacttgg ctcttgccca tatgataaag 1080 ttccgtgggc actttctcat tgagggtgat ctaaatccgg aaactcgga tgtcgacaaa 1140 ctgttcatcc agttagtaca aacctataat cagttgtttg aagagaaccc tataaatgca 1200 agtggcgtgg atgcgaaggc tattcttagc gcccgctctct ctaaatcccg acggctagaa 1260 aacctgatcg caaattacc cggagagaag aaaaatgggt tgttcggtaa ccttatagcg 1320 ctctcactag gcctgacacc aaattttaag tcgaacttcg acttagctga agatgccaaa 1380 ttgcagctta gtaggacac gtacgatgac gatctcgaca atctactggc acaaattgga 1440 gatcagtag cggacttatt ttggctgcc aaaaacctta gcgatgcaat cctcctatct 1500 gagatactga gagttaatac tgagattacc aaggcgccgt tatccgctc aatgatcaa 1560 aggtacgatg aacatcacca agacttgaca cttctcagg ccctagtccg tcagcaactg 1620 cctgagaaat ataghaat attctttgat cagtcgaaaa acgggtacgc aggttatatt 1680 gacggcggag cgagtcaga ggaattctac aagtttaca aacccatatt agagaagatg 1740 gatgggacgg aagagttgct tgtaaaactc atcgcgaag atctactgcg aaagcagcgg 1800 actttcgaca acggtagcat tccacatca atccacttag gcgaattgca tgctatactt 1860 agaaggcagg aggattta tccgttccctc aaagacaatc gtgaaaagat tgagaaaatc 1920 ctaacctttc gcatacctta ctagtggga cccctggccc gagggaactc tcggttcgca 1980 tggatgacaa gaaagtccga agaaacgatt actcatgga attttgagga agttgtcgat 2040 aaaggtgcgt cagctcaatc gttcatcgag aggatgacca actttgaca gatttaccg 2100 aacgaaaaag tattgcctaa gcacagttta ctttacgagt atttcacagt gtacaatgaa 2160 ctcacgaaag ttaagtatgt cactgagggc atgcgtaaac ccgcctttct aagcggagaa 2220 cagaaaaag caatagtaga tctgttattc aagaccaacc gcaaagtgac agttaagcaa 2280 ttgaaagagg actactttaa gaaattgaa tgcttcgatt ctgtcgagat ctccggggta 2340 gaagatcgat ttaatgcgtc acttggtacg tatcatgacc tcctaaagat aattaaagat 2400 areacttcc tggataacga area gatatcttag aagatatagt gttgactctt 2460 2520 gacgataagg ttatgaaaca gttaaagagg cgtcgctata cgggctgggg acgattgtcg 2580 cggaaactta tcaacgggat aagacaag caaagtggta aaactattct cgattttcta 2640 aagagcgacg gcttcgccaa taggaacttt atgcagctga tccatgatga ctctttaacc 2700 ttcaaagagg atatacaaaa ggcacaggtt tccggacaag gggactcatt gcacgaacat 2760 attgcgaatc ttgctggttc gccagccatc aaaaagggca tactccagac agtcaaagta 2820 gtggatgagc tagttaaggt catgggacgt cacaaaccgg aaaacattgt aatcgagatg 2880 gcacgcgaaa atcaaacgac tcagaagggg caaaaaaaca gtcgagagcg gatgaagaga 2940 atagaagagg gtattaaaga actgggcagc cagatcttaa aggagcatcc tgtggaaaat 3000 acccaattgc agaacgagaa actttacctc tattacctac aaaatggaag ggacatgtat 3060 gttgatcagg aactggacat aaaccgttta tctgattacg acgtcgatgc cattgtaccc 3120 caatcctttt tgaaggacga ttcaatcgac aataaagtgc ttacacgctc ggataagaac 3180 cgagggaaaa gtgacaatgt tccaagcgag gaagtcgtaa agaaaatgaa gaactattgg 3240 cggcagctcc taaatgcgaa actgataacg caaagaaagt tcgataactt aactaaagct 3300 gagaggggtg gcttgtctga acttgacaag gccggattta ttaaacgtca gctcgtggaa 3360 acccgccaaa tcacaaagca tgttgcacag atactagatt cccgaatgaa tacgaaatac 3420 gacgagaacg ataagctgat tcgggaagtc aaagtaatca ctttaaagtc aaaattggtg 3480 tcggacttca gaaaggattt tcaattctat aaagttaggg agataaataa ctaccaccat 3540 gcgcacgacg cttatcttaa tgccgtcgta gggaccgcac tcattaagaa atacccgaag ctagaaagtg agtttgtgta tggtgattac aaagtttatg acgtccgtaa gatgatcgcg aaaagcgaac aggagatagg caaggctaca gccaatact tcttttattc taacattatg aatttcttta agacggaat cactctggca aacggagaga tacgcaaacg acctttaatt gaaaccaatg gggagacagg tgaatcgta tgggataagg gccgggactt cgcgacggtg agaaaagttt tgtccatgcc ccaagtcaac atagtaaaga aaactgaggt gcagaccgga gggttttcaa aggaatcgat tcttccaaaa aggaatagtg ataagctcat cgctcgtaaa aaggactggg acccgaaaaa gtacggtggc ttcgatagcc ctacagttgc ctattctgtc 4020 ctagtagtgg caaaagttga gaagggaaaa tccaagaac tgaagtcagt caaagaatta ttggggataa cgattatgga gcgctcgtct tttgaaaaga accccatcga cttccttgag gcgaaaggtt acaaggaagt aaaaaaggt ctcataatta aactaccaaa gtatagtctg tttgagttag aaaatggccg aaaacggatg ttggctagcg ccggagagct tcaaaagggg 4260. aacgaactcg cactaccgtc taaatacgtg aatttcctgt atttagcgtc ccattacgag 4320 aagttgaaag gttcacctga agataacgaa cagaagcaac tttttgttga gcagcacaaa 4380 cattatctcg acgaaatcat agagcaaatt tcggaattca gtaagagagt catcctagct 4440 gatgccaatc tggacaaagt attaagcgca tacaacaagc acagggataa acccatacgt 4500 gagcaggcgg aaaatattat ccatttgttt actcttacca acctcggcgc tccagccgca 4560 ttcaagtatt ttgacacaac gatagatcgc aaacgataca cttctaccaa ggaggtgcta 4620 gacgcgacac tgattcacca atccatcacg ggattatatg aaactcggat agatttgtca 4680 cagcttgggg gtgactctgg tggttctccc aagaagaaga ggaaagtcta a 4731 <210> 66 <211> 1576 <212> PRT <213> Artificial Sequence <220> <223> Description of artificial sequence: Amino acid sequence of AIDX-XTEN-dCas9 <400> 66 Met Asp Ser Leu Leu Met Asn Arg Arg Lys Phe Leu Tyr Gln Phe Lys 1 5 10 15 Asn Val Arg Trp Ala Lys Gly Arg Arg Glu Thr Tyr Leu Cys Tyr Val 20 25 30 Val Lys Arg Arg Asp Ser Ala Thr Ser Phe Ser Leu Asp Phe Gly Tyr 35 40 45 Leu Arg Asn Lys Asn Gly Cys His Val Glu Leu Leu Phe Leu Arg Tyr 50 55 60 Ile Ser Asp Trp Asp Leu Asp Pro Gly Arg Cys Tyr Arg Val Thr Trp 65 70 75 80 Phe Thr Ser Trp Ser Pro Cys Tyr Asp Cys Ala Arg His Val Ala Asp 85 90 95 Phe Leu Arg Gly Asn Pro Asn Leu Ser Leu Arg Ile Phe Thr Ala Arg 100 105 110 Leu Tyr Phe Cys Glu Asp Arg Lys Ala Glu Pro Glu Gly Leu Arg Arg 115 120 125 Leu His Arg Ala Gly Val Gln Ile Ala Ile Met Thr Phe Lys Asp Tyr 130 135 140 Phe Tyr Cys Trp Asn Thr Phe Val Glu Asn His Glu Arg Thr Phe Lys 145 150 155 160 Ala Trp Glu Gly Leu His Glu Asn Ser Val Arg Leu Ser Arg Gln Leu 165 170 175 Arg Arg Ile Leu Leu Pro Ser Gly Ser Glu Thr Pro Gly Thr Ser Glu 180 185 190 Ser Ala Thr Pro Glu Ser Asp Lys Lys Tyr Ser Ile Gly Leu Ala Ile 195 200 205 Gly Thr Asn Ser Val Gly Trp Ala Val Ile Thr Asp Glu Tyr Lys Val 210 215 220 Pro Ser Lys Lys Phe Lys Val Leu Gly Asn Thr Asp Arg His Ser Ile 225 230 235 240 Lys Lys Asn Leu Ile Gly Ala Leu Leu Phe Asp Ser Gly Glu Thr Ala 245 250 255 Glu Ala Thr Arg Leu Lys Arg Thr Ala Arg Arg Arg Tyr Thr Arg Arg 260 265 270 Lys Asn Arg Ile Cys Tyr Leu Gln Glu Ile Phe Ser Asn Glu Met Ala 275 280 285 Lys Val Asp Asp Ser Phe Phe His Arg Leu Glu Glu Ser Phe Leu Val 290 295 300 Glu Glu Asp Lys Lys His Glu Arg His Pro Ile Phe Gly Asn Ile Val 305 310 315 320 Asp Glu Val Ala Tyr His Glu Lys Tyr Pro Thr Ile Tyr His Leu Arg 325 330 335 Lys Lys Leu Val Asp Ser Thr Asp Lys Ala Asp Leu Arg Leu Ile Tyr 340 345 350 Leu Ala Leu Ala His Met Ile Lys Phe Arg Gly His Phe Leu Ile Glu 355 360 365 Gly Asp Leu Asn Pro Asp Asn Ser Asp Val Asp Lys Leu Phe Ile Gln 370 375 380 Leu Val Gln Thr Tyr Asn Gln Leu Phe Glu Glu Asn Pro Ile Asn Ala 385 390 395 400 Ser Gly Val Asp Ala Lys Ala Ile Leu Ser Ala Arg Leu Ser Lys Ser 405 410 415 Arg Arg Leu Glu Asn Leu Ile Ala Gln Leu Pro Gly Glu Lys Lys Asn 420 425 430 Gly Leu Phe Gly Asn Leu Ile Ala Leu Ser Leu Gly Leu Thr Pro Asn 435 440 445 Phe Lys Ser Asn Phe Asp Leu Ala Glu Asp Ala Lys Leu Gln Leu Ser 450 455 460 Lys Asp Thr Tyr Asp Asp Asp Leu Asp Asn Leu Leu Ala Gln Ile Gly 465 470 475 480 Asp Gln Tyr Ala Asp Leu Phe Leu Ala Ala Lys Asn Leu Ser Asp Ala 485 490 495 Ile Leu Leu Ser Asp Ile Leu Arg Val Asn Thr Glu Ile Thr Lys Ala 500 505 510 Pro Leu Ser Ala Ser Met Ile Lys Arg Tyr Asp Glu His His Gln Asp 515 520 525 Leu Thr Leu Leu Lys Ala Leu Val Arg Gln Gln Leu Pro Glu Lys Tyr 530 535 540 Lys Glu Ile Phe Phe Asp Gln Ser Lys Asn Gly Tyr Ala Gly Tyr Ile 545 550 555 560 Asp Gly Gly Ala Ser Gln Glu Glu Phe Tyr Lys Phe Ile Lys Pro Ile 565 570 575 Leu Glu Lys Met Asp Gly Thr Glu Glu Leu Leu Val Lys Leu Asn Arg 580 585 590 Glu Asp Leu Leu Arg Lys Gln Arg Thr Phe Asp Asn Gly Ser Ile Pro 595 600 605 His Gln Ile His Leu Gly Glu Leu His Ala Ile Leu Arg Arg Gln Glu 610 615 620 Asp Phe Tyr Pro Phe Leu Lys Asp Asn Arg Glu Lys Ile Glu Lys Ile 625 630 635 640 Leu Thr Phe Arg Ile Pro Tyr Tyr Val Gly Pro Leu Ala Arg Gly Asn 645 650 655 Ser Arg Phe Ala Trp Met Thr Arg Lys Ser Glu Glu Thr Ile Thr Pro 660 665 670 Trp Asn Phe Glu Glu Val Val Asp Lys Gly Ala Ser Ala Gln Ser Phe 675 680 685 Ile Glu Arg Met Thr Asn Phe Asp Lys Asn Leu Pro Asn Glu Lys Val 690 695 700 Leu Pro Lys His Ser Leu Leu Tyr Glu Tyr Phe Thr Val Tyr Asn Glu 705 710 715 720 Leu Thr Lys Val Lys Tyr Val Thr Glu Gly Met Arg Lys Pro Ala Phe 725 730 735 Leu Ser Gly Glu Gln Lys Lys Ala Ile Val Asp Leu Leu Phe Lys Thr 740 745 750 Asn Arg Lys Val Thr Val Lys Gln Leu Lys Glu Asp Tyr Phe Lys Lys 755 760 765 Ile Glu Cys Phe Asp Ser Val Glu Ile Ser Gly Val Glu Asp Arg Phe 770 775 780 Asn Ala Ser Leu Gly Thr Tyr His Asp Leu Leu Lys Ile Ile Lys Asp 785 790 795 800 Lys Asp Phe Leu Asp Asn Glu Glu Asn Glu Asp Ile Leu Glu Asp Ile 805 810 815 Val Leu Thr Leu Thr Leu Phe Glu Asp Arg Glu Met Ile Glu Glu Arg 820 825 830 Leu Lys Thr Tyr Ala His Leu Phe Asp Asp Lys Val Met Lys Gln Leu 835 840 845 Lys Arg Arg Arg Tyr Thr Gly Trp Gly Arg Leu Ser Arg Lys Leu Ile 850 855 860 Asn Gly Ile Arg Asp Lys Gln Ser Gly Lys Thr Ile Leu Asp Phe Leu 865 870 875 880 Lys Ser Asp Gly Phe Ala Asn Arg Asn Phe Met Gln Leu Ile His Asp 885 890 895 Asp Ser Leu Thr Phe Lys Glu Asp Ile Gln Lys Ala Gln Val Ser Gly 900 905 910 Gln Gly Asp Ser Leu His Glu His Ile Ala Asn Leu Ala Gly Ser Pro 915 920 925 Ala Ile Lys Lys Gly Ile Leu Gln Thr Val Lys Val Val Asp Glu Leu 930 935 940 Val Lys Val Met Gly Arg His Lys Pro Glu Asn Ile Val Ile Glu Met 945 950 955 960 Ala Arg Glu Asn Gln Thr Thr Gln Lys Gly Gln Lys Asn Ser Arg Glu 965 970 975 Arg Met Lys Arg Ile Glu Glu Gly Ile Lys Glu Leu Gly Ser Gln Ile 980 985 990 Leu Lys Glu His Pro Val Glu Asn Thr Gln Leu Gln Asn Glu Lys Leu 995 1000 1005 Tyr Leu Tyr Tyr Leu Gln Asn Gly Arg Asp Met Tyr Val Asp Gln 1010 1015 1020 Glu Leu Asp Ile Asn Arg Leu Ser Asp Tyr Asp Val Asp Ala Ile 1025 1030 1035 Val Pro Gln Ser Phe Leu Lys Asp Asp Ser Ile Asp Asn Lys Val 1040 1045 1050 Leu Thr Arg Ser Asp Lys Asn Arg Gly Lys Ser Asp Asn Val Pro 1055 1060 1065 Ser Glu Glu Val Val Lys Lys Met Lys Asn Tyr Trp Arg Gln Leu 1070 1075 1080 Leu Asn Ala Lys Leu Ile Thr Gln Arg Lys Phe Asp Asn Leu Thr 1085 1090 1095 Lys Ala Glu Arg Gly Gly Leu Ser Glu Leu Asp Lys Ala Gly Phe 1100 1105 1110 Ile Lys Arg Gln Leu Val Glu Thr Arg Gln Ile Thr Lys His Val 1115 1120 1125 Only Gln With Arg Ser Asp With Thr Lys Tyr Asp Glu Asn 1130 1135 1140 Asp Lys With Arg Glu Val Lys Val With Thr Lys Ser Lys 1145 1150 1155 Leu Val Ser Asp Phe Arg Lys Asp Phe Gln Phe Tyr Lys Val Arg 1160 1165 1170 Glu Ile Asn Asn Tyr His Ala His Asp Ala Tyr Leu Asn Ala 1175 1180 1185 Val Val Gly Thr Ala Leu Ile Lys Tyr Pro Lys Leu Glu Ser 1190 1195 1200 Glu Phe Val Tyr Gly Asp Tyr Lys Val Tyr Asp Val Arg Lys Met 1205 1210 1215 I Ile Lys Sere Glu Gln Glu Ile Gly Lys Ike Thr Ike Lys Tyr 1220 1225 1230 Phe Phe Tyr Ser Asn Ile Met Asn Phe Phe Lys Thr Glu Ile Thr 1235 1240 1245 Only Asn Gly Glu With Arg Lys Arg Pro Only With Glu Thr Asn 1250 1255 1260 Gly Glu Thr Gly Glu Ile Val Trp Asp Lys Gly Arg Asp Phe Ala 1265 1270 1275 Thr Val Arg Lys Val Leu Ser Met Pro Gln Val Asn Ile Val Lys 1280 1285 1290 Lys Thr Glu Val Gln Thr Gly Gly Phe Ser Lys Glu Ser Ile Leu 1295 1300 1305 Pro Lys Arg Asn Ser Asp Lys Leu Ile Ala Arg Lys Lys Asp Trp 1310 1315 1320 Asp Pro Lys Lys Tyr Gly Gly Phe Asp Ser Pro Thr Val Ala Tyr 1325 1330 1335 Ser Val Leu Val Val Ala Lys Val Glu Lys Gly Lys Ser Lys Lys 1340 1345 1350 Leu Lys Ser Val Lys Glu Leu Leu Gly Ile Thr Ile Met Glu Arg 1355 1360 1365 Ser Ser Phe Glu Lys Asn Pro Ile Asp Phe Leu Glu Ala Lys Gly 1370 1375 1380 Tyr Lys Glu Val Lys Lys Asp Leu Ile Ile Lys Leu Pro Lys Tyr 1385 1390 1395 Ser Leu Phe Glu Leu Glu Asn Gly Arg Lys Arg Met Leu Ala Ser 1400 1405 1410 Ala Gly Glu Leu Gln Lys Gly Asn Glu Leu Ala Leu Pro Ser Lys 1415 1420 1425 Tyr Val Asn Phe Leu Tyr Leu Ala Ser His Tyr Glu Lys Leu Lys 1430 1435 1440 Gly Ser Pro Glu Asp Asn Glu Gln Lys Gln Leu Phe Val Glu Gln 1445 1450 1455 His Lys His Tyr Leu Asp Glu Ile Ile Glu Gln Ile Ser Glu Phe 1460 1465 1470 Ser Lys Arg Val Ile Leu Ala Asp Ala Asn Leu Asp Lys Val Leu 1475 1480 1485 Ser Ala Tyr Asn Lys His Arg Asp Lys Pro Ile Arg Glu Gln Ala 1490 1495 1500 Glu Asn Ile Ile His Leu Phe Thr Leu Thr Asn Leu Gly Ala Pro 1505 1510 1515 Ala Ala Phe Lys Tyr Phe Asp Thr Thr Ile Asp Arg Lys Arg Tyr 1520 1525 1530 Thr Ser Thr Lys Glu Val Leu Asp Ala Thr Leu Ile His Gln Ser 1535 1540 1545 Ile Thr Gly Leu Tyr Glu Thr Arg Ile Asp Leu Ser Gln Leu Gly 1550 1555 1560 Gly Asp Ser Gly Gly Ser Pro Lys Lys Lys Arg Lys Val 1565 1570 1575 <210> 67 <211> 4890 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Coding sequence of dCas9-XTEN-AIDX (K10E T82I E156G) <400> 67 atggactata aggaccacga cggagactac aaggatcatg atattgatta caaagacgat 60 gacgataaga tggccccaaa gaagaagcgg aaggtcggta tccacggagt cccagcagct 120 accatggaca agaagtattc tatcggactg gccatcggga ctaatagcgt cgggtgggcc 180 gtgatcactg acgagtacaa ggtgccctct aagaagttca aggtgctcgg gaacaccgac 240 cggcattcca tcaagaaaaa tctgatcgga gctctcctct ttgattcagg ggagaccgct 300 gaagcaaccc gcctcaagcg gactgctaga cggcggtaca ccaggaggaa gaaccggatt 360 tgttaccttc aagagatatt ctccaacgaa atggcaaagg tcgacgacag cttcttccat 420 aggctggaag aatcattcct cgtggaagag gataagaagc atgaacggca tcccatcttc 480 ggtaatatcg tcgacgaggt ggcctatcac gagaaatacc caaccatcta ccatcttcgc 540 aaaaagctgg tggactcaac cgacaaggca gacctccggc ttatctacct ggccctggcc 600 cacatgatca agttcagagg ccacttcctg atcgagggcg acctcaatcc tgacaatagc 660 gatgtggata aactgttcat ccagctggtg cagacttaca accagctctt tgaagagaac 720 cccatcaatg caagcggagt cgatgccaag gccattctgt cagcccggct gtcaaagagc 780 cgcagacttg agaatcttat cgctcagctg ccgggtgaaa agaaaaatgg actgttcggg 840 aacctgattg ctctttcact tgggctgact cccaatttca agtctaattt cgacctggca 900 gaggatgcca agctgcaact gtccaaggac acctatgatg acgatctcga caacctcctg 960 gcccagatcg gtgaccaata cgccgacctt ttccttgctg ctaagaatct ttctgacgcc 1020 atcctgctgt ctgacattct ccgcgtgaac actgaaatca ccaaggcccc tctttcagct 1080 tcaatgatta agcggtatga tgagcaccac caggacctga ccctgcttaa ggcactcgtc 1140 cggcagcagc ttccggagaa gtacaaggaa atcttctttg accagtcaaa gaatggatac 1200 gccggctaca tcgacggagg tgcctcccaa gaggaatttt ataagtttat caaacctatc 1260 cttgagaga tggacggcac cgaagagctc ctcgtgaac tgaatcggga ggatctgctg 1320 cggaagcagc gcacttcga caatgggagc attccccacc agatcatct tggggagctt 1380 cacgccatcc ttcggcgcca agaggacttc tacccctttc ttaggacaa cagggagaag 1440 attgagaaaa ttctcactt ccgcatcccc tactacgtgg gacccctcgc cagaggaaat 1500 agccggtttg cttggatgac cagaagtca gagaaacta tcactccctg gaacttcgaa 1560 gaggtggtgg acaagggagc cagcgctcag tcattcatcg aacggatgac taacttcgat 1620 aagaacctcc ccaatgagaa ggtcctgccg aacattccc tgcttacga gtactttacc 1680 gtgtacaacg agctgaccaa ggtgaaatat gtcaccgaag ggatgaggaa gcccgcattc 1740 ctgtcaggcg aaaaagaaaggcaattgtg gaccttctgt tcagaccaa tagaaggtg 1800 accgtgaagc agctgaagga ggactttc aagaaaattg aatgcttcga ctctgtggag 1860 attagcgggg tcgagatcg gttcaacgca agcctggta cctaccatga tctgcttaag 1920 atcatcagg acagatt tctgacat gaggagaag aggacatcct tgaggacatt 1980 gtcctgactc tcactctt cgaggaccgg gaatgatcg aggagaggct tagacctac 2040 gcccatctgt tcgacgataa agtgatgaag caacttaac ggagaata taccggatgg 2100 ggacgcctta gccgcaact catcaacgga atccgggaca aacagcgg aagaccatt 2160 cttgatttcc ttagagcga cggattcgct aatcgcact tcatgcact tatccatgat 2220 gattccctga cctttaagga ggacatccag aaggcccaag tgtctggaca aggtgactca 2280 ctgcacgagc atatcgcaa tctggctggt tcaccgcta ttaagaaggg tattctccag 2340 accgtgaaag tcgtggacga gctggtcaag gtgatgggtc gccataacc agagaacatt 2400 gtcatcgaga tggccaggga aaaccagact acccagaagg gagagaga caggcagggag 2460 cggatgaaaa gattgagga agggattaag gagctcgggt cacagatccct taagagcac 2520 ccggtggaaa acacccagct tcagaatgag aagctctatc tgtacct tcaaatgga 2580 cgcgatatgt atgtggacca agagcttgat atcaacaggc tctcagacta cgacgtggac 2640 gccatcgtcc ctcagagctt cctcaagac gactcattg acataggt gctgactcgc 2700 tcagacaaga accggggaaa gtcagataac gtgccctcag aggaagtcgt gaaaaagatg 2760 aagaactatt ggcgccagct tctgaacgca aagctgatca ctcagcggaa gttcgacaat 2820 ctcactaagg ctgagagggg cggactgagc gaactggaca aagcaggatt cattaaacgg 2880 caacttgtgg agactcggca gattactaaa catgtagccc aaatccttga ctcacgcatg 2940 aataccaagt acgacgaaaa cgacaaactt atccgcgagg tgaaggtgat taccctgaag 3000 tccaagctgg tcagcgattt cagaaaggac tttcaattct acaaagtgcg ggagatcaat 3060 aactatcatc atgctcatga cgcatatctg aatgccgtgg tgggaaccgc cctgatcaag 3120 aagtacccaa agctggaaag cgagttcgtg tacggagact acaaggtcta cgacgtgcgc 3180 aagatgattg ccaaatctga gcaggagatc ggaaaggcca ccgcaaagta cttcttctac 3240 agcaacatca tgaatttctt caagaccgaa atcacccttg caaacggtga gatccggaag 3300 aggccgctca tcgagactaa tggggagact ggcgaaatcg tgtgggacaa gggcagagat 3360 ttcgctaccg tgcgcaaagt gctttctatg cctcaagtga acatcgtgaa gaaaaccgag 3420 gtgcaaaccg gaggctttc taaggaatca atcctcccca agcgcaactc cgacaagctc 3480 attgcaagga agaaggattg ggaccctaag aagtacggcg gattcgattc accaactgtg 3540 gcttattctg tcctggtcgt ggctaaggtg gaaaaaggaa agtctaagaa gctcaagagc 3600 gtgaaggaac tgctgggtat caccattatg gagcgcagct ccttcgagaa gaacccaatt 3660 gactttctcg aagccaaagg ttacaaggaa gtcaagaagg accttatcat caagctccca 3720 aagtatagcc tgttcgaact ggagaatggg cggaagcgga tgctcgcctc cgctggcgaa 3780 cttcagaagg gtaatgagct ggctctcccc tccaagtacg tgaatttcct ctaccttgca 3840 agccattacg agaagctgaa ggggagcccc gaggacaacg agcaaaagca actgtttgtg 3900 gagcagcata agcattatct ggacgagatc attgagcaga tttccgagtt ttctaaacgc 3960 gtcattctcg ctgatgccaa cctcgataaa gtccttagcg catacaataa gcacagagac 4020 aaaccaattc gggagcaggc tgagaatatc atccacctgt tcaccctcac caatctttggt 4080 gccctgccg cattcaagta cttcgacacc accatcgacc ggaaacgcta tacctccacc 4140 aaagaagtgc tggacgccac cctcatccac cagagcatca ccggacttta cgaaactcgg 4200 attgacctct cacagctcgg aggggatgag ggagctccca agaaaaagcg caaggataggt 4260 agttccggat ctccgaaaaa gaaacgcaaa gttagcggca gcgagactcc cgggacctca 4320 gagtccgcca cacccgaaag tatggacagc ctcttgatga accggaggga gtttctttac 4380 caattcaaaa atgtccgctg ggctaagggt cggcgtgaga cctacctgtg ctacgtagtg 4440 aagaggcgtg acagtgctac atcctttca ctggactttg gttatcttcg caataagaac 4500 gggctcccg tggaattgct cttcctccgc tacatctcgg actgggacct agaccctggc 4560 cgctgctacc gcgtcacctg gttcatctcc tggagcccct gctacgactg tgcccgacat 4620 gtggccgact ttctgcgagg gaaccccaac ctcagtctga ggatcttcac cgcgcgcctc 4680 tacttctgtg aggaccgcaa ggctgagccc gaggggctgc ggcggctgca ccgcgccggg 4740 gtgcaaatag ccatcatgac cttcaaagat tatttttact gctggaatac ttttgtagaa 4800 aaccatggaa gaactttcaa agcctgggaa gggctgcatg aaaattcagt tcgtctctcc 4860 agacagcttc ggcgcatccttttgccctga 4890 <210> 68 <211> 1629 <212> PRT <213> Artificial sequence <220> <223> Description of artificial sequence: Amino acid sequence of dCas9-XTEN-AIDX (K10E T82I E156G) <400> 68 Met Asp Tyr Lys Asp His Asp Gly Asp Tyr Lys Asp His Asp Ile Asp 1 5 10 15 Tyr Lys Asp Asp Asp Asp Lys Met Ala Pro Lys Lys Lys Arg Lys Val 20 25 30 Gly Ile His Gly Val Pro Ala Ala Thr Met Asp Lys Lys Tyr Ser Ile 35 40 45 Gly Leu Ala Ile Gly Thr Asn Ser Val Gly Trp Ala Val Ile Thr Asp 50 55 60 Glu Tyr Lys Val Pro Ser Lys Lys Phe Lys Val Leu Gly Asn Thr Asp 65 70 75 80 Arg His Ser Ile Lys Lys Asn Leu Ile Gly Ala Leu Leu Phe Asp Ser 85 90 95 Gly Glu Thr Ala Glu Ala Thr Arg Leu Lys Arg Thr Ala Arg Arg Arg 100 105 110 Tyr Thr Arg Arg Lys Asn Arg Ile Cys Tyr Leu Gln Glu Ile Phe Ser 115 120 125 Asn Glu Met Ala Lys Val Asp Asp Ser Phe Phe His Arg Leu Glu Glu 130 135 140 Ser Phe Leu Val Glu Glu Asp Lys Lys His Glu Arg His Pro Ile Phe 145 150 155 160 Gly Asn Ile Val Asp Glu Val Ala Tyr His Glu Lys Tyr Pro Thr Ile 165 170 175 Tyr His Leu Arg Lys Lys Leu Val Asp Ser Thr Asp Lys Ala Asp Leu 180 185 190 Arg Leu Ile Tyr Leu Ala Leu Ala His Met Ile Lys Phe Arg Gly His 195 200 205 Phe Leu Ile Glu Gly Asp Leu Asn Pro Asp Asn Ser Asp Val Asp Lys 210 215 220 Leu Phe Ile Gln Leu Val Gln Thr Tyr Asn Gln Leu Phe Glu Glu Asn 225 230 235 240 Pro Ile Asn Ala Ser Gly Val Asp Ala Lys Ala Ile Leu Ser Ala Arg 245 250 255 Leu Ser Lys Ser Arg Arg Leu Glu Asn Leu Ile Ala Gln Leu Pro Gly 260 265 270 Glu Lys Lys Asn Gly Leu Phe Gly Asn Leu Ile Ala Leu Ser Leu Gly 275 280 285 Leu Thr Pro Asn Phe Lys Ser Asn Phe Asp Leu Ala Glu Asp Ala Lys 290 295 300 Leu Gln Leu Ser Lys Asp Thr Tyr Asp Asp Asp Leu Asp Asn Leu Leu 305 310 315 320 Ala Gln Ile Gly Asp Gln Tyr Ala Asp Leu Phe Leu Ala Ala Lys Asn 325 330 335 Leu Ser Asp Ala Ile Leu Leu Ser Asp Ile Leu Arg Val Asn Thr Glu 340 345 350 Ile Thr Lys Ala Pro Leu Ser Ala Ser Met Ile Lys Arg Tyr Asp Glu 355 360 365 His His Gln Asp Leu Thr Leu Leu Lys Ala Leu Val Arg Gln Gln Leu 370 375 380 Pro Glu Lys Tyr Lys Glu Ile Phe Phe Asp Gln Ser Lys Asn Gly Tyr 385 390 395 400 Ala Gly Tyr Ile Asp Gly Gly Ala Ser Gln Glu Glu Phe Tyr Lys Phe 405 410 415 Ile Lys Pro Ile Leu Glu Lys Met Asp Gly Thr Glu Glu Leu Leu Val 420 425 430 Lys Leu Asn Arg Glu Asp Leu Leu Arg Lys Gln Arg Thr Phe Asp Asn 435 440 445 Gly Ser Ile Pro His Gln Ile His Leu Gly Glu Leu His Ala Ile Leu 450 455 460 Arg Arg Gln Glu Asp Phe Tyr Pro Phe Leu Lys Asp Asn Arg Glu Lys 465 470 475 480 Ile Glu Lys Ile Leu Thr Phe Arg Ile Pro Tyr Tyr Val Gly Pro Leu 485 490 495 Ala Arg Gly Asn Ser Arg Phe Ala Trp Met Thr Arg Lys Ser Glu Glu 500 505 510 Thr Ile Thr Pro Trp Asn Phe Glu Glu Val Val Asp Lys Gly Ala Ser 515 520 525 Ala Gln Ser Phe Ile Glu Arg Met Thr Asn Phe Asp Lys Asn Leu Pro 530 535 540 Asn Glu Lys Val Leu Pro Lys His Ser Leu Leu Tyr Glu Tyr Phe Thr 545 550 555 560 Val Tyr Asn Glu Leu Thr Lys Val Lys Tyr Val Thr Glu Gly Met Arg 565 570 575 Lys Pro Ala Phe Leu Ser Gly Glu Gln Lys Lys Ala Ile Val Asp Leu 580 585 590 Leu Phe Lys Thr Asn Arg Lys Val Thr Val Lys Gln Leu Lys Glu Asp 595 600 605 Tyr Phe Lys Lys Ile Glu Cys Phe Asp Ser Val Glu Ile Ser Gly Val 610 615 620 Glu Asp Arg Phe Asn Ala Ser Leu Gly Thr Tyr His Asp Leu Leu Lys 625 630 635 640 Ile Ile Lys Asp Lys Asp Phe Leu Asp Asn Glu Glu Lys Glu Asp Ile 645 650 655 Leu Glu Asp Ile Val Leu Thr Leu Thr Leu Phe Glu Asp Arg Glu Met 660 665 670 Ile Glu Glu Arg Leu Lys Thr Tyr Ala His Leu Phe Asp Asp Lys Val 675 680 685 Met Lys Gln Leu Lys Arg Arg Arg Tyr Thr Gly Trp Gly Arg Leu Ser 690 695 700 Arg Lys Leu Ile Asn Gly Ile Arg Asp Lys Gln Ser Gly Lys Thr Ile 705 710 715 720 Leu Asp Phe Leu Lys Ser Asp Gly Phe Ala Asn Arg Asn Phe Met Gln 725 730 735 Leu Ile His Asp Asp Ser Leu Thr Phe Lys Glu Asp Ile Gln Lys Ala 740 745 750 Gln Val Ser Gly Gln Gly Asp Ser Leu His Glu His Ile Ala Asn Leu 755 760 765 Ala Gly Ser Pro Ala Ile Lys Lys Gly Ile Leu Gln Thr Val Lys Val 770 775 780 Val Asp Glu Leu Val Lys Val Met Gly Arg His Lys Pro Glu Asn Ile 785 790 795 800 Val Ile Glu Met Ala Arg Glu Asn Gln Thr Thr Gln Lys Gly Gln Lys 805 810 815 Asn Ser Arg Glu Arg Met Lys Arg Ile Glu Glu Gly Ile Lys Glu Leu 820 825 830 Gly Ser Gln Ile Leu Lys Glu His Pro Val Glu Asn Thr Gln Leu Gln 835 840 845 Asn Glu Lys Leu Tyr Leu Tyr Tyr Leu Gln Asn Gly Arg Asp Met Tyr 850 855 860 Val Asp Gln Glu Leu Asp Ile Asn Arg Leu Ser Asp Tyr Asp Val Asp 865 870 875 880 Ala Ile Val Pro Gln Ser Phe Leu Lys Asp Asp Ser Ile Asp Asn Lys 885 890 895 Val Leu Thr Arg Ser Asp Lys Asn Arg Gly Lys Ser Asp Asn Val Pro 900 905 910 Ser Glu Glu Val Val Lys Lys Met Lys Asn Tyr Trp Arg Gln Leu Leu 915 920 925 Asn Ala Lys Leu Ile Thr Gln Arg Lys Phe Asp Asn Leu Thr Lys Ala 930 935 940 Glu Arg Gly Gly Leu Ser Glu Leu Asp Lys Ala Gly Phe Ile Lys Arg 945 950 955 960 Gln Leu Val Glu Thr Arg Gln Ile Thr Lys His Val Ala Gln Ile Leu 965 970 975 Asp Ser Arg Met Asn Thr Lys Tyr Asp Glu Asn Asp Lys Leu Ile Arg 980 985 990 Glu Val Lys Val Ile Thr Leu Lys Ser Lys Leu Val Ser Asp Phe Arg 995 1000 1005 Lys Asp Phe Gln Phe Tyr Lys Val Arg Glu Ile Asn Asn Tyr His 1010 1015 1020 His Ala His Asp Ala Tyr Leu Asn Ala Val Val Gly Thr Ala Leu 1025 1030 1035 Ile Lys Lys Tyr Pro Lys Leu Glu Ser Glu Phe Val Tyr Gly Asp 1040 1045 1050 Tyr Lys Val Tyr Asp Val Arg Lys Met Ile Ala Lys Ser Glu Gln 1055 1060 1065 Glu Ile Gly Lys Ala Thr Ala Lys Tyr Phe Phe Tyr Ser Asn Ile 1070 1075 1080 Met Asn Phe Phe Lys Thr Glu Ile Thr Leu Ala Asn Gly Glu Ile 1085 1090 1095 Arg Lys Arg Pro Leu Ile Glu Thr Asn Gly Glu Thr Gly Glu Ile 1100 1105 1110 Val Trp Asp Lys Gly Arg Asp Phe Ala Thr Val Arg Lys Val Leu 1115 1120 1125 Ser Met Pro Gln Val Asn Ile Val Lys Lys Thr Glu Val Gln Thr 1130 1135 1140 Gly Gly Phe Ser Lys Glu Ser Ile Leu Pro Lys Arg Asn Ser Asp 1145 1150 1155 Lys Leu Ile Ala Arg Lys Lys Asp Trp Asp Pro Lys Lys Tyr Gly 1160 1165 1170 Gly Phe Asp Ser Pro Thr Val Ala Tyr Ser Val Leu Val Val Ala 1175 1180 1185 Lys Val Glu Lys Gly Lys Ser Lys Lys Leu Lys Ser Val Lys Glu 1190 1195 1200 Leu Leu Gly Ile Thr Ile Met Glu Arg Ser Ser Phe Glu Lys Asn 1205 1210 1215 Pro Ile Asp Phe Leu Glu Ala Lys Gly Tyr Lys Glu Val Lys Lys 1220 1225 1230 Asp Leu Ile Ile Lys Leu Pro Lys Tyr Ser Leu Phe Glu Leu Glu 1235 1240 1245 Asn Gly Arg Lys Arg Met Leu Ala Ser Ala Gly Glu Leu Gln Lys 1250 1255 1260 Gly Asn Glu Leu Ala Leu Pro Ser Lys Tyr Val Asn Phe Leu Tyr 1265 1270 1275 Leu Ala Ser His Tyr Glu Lys Leu Lys Gly Ser Pro Glu Asp Asn 1280 1285 1290 Glu Gln Lys Gln Leu Phe Val Glu Gln His Lys His Tyr Leu Asp 1295 1300 1305 Glu Ile Ile Glu Gln Ile Ser Glu Phe Ser Lys Arg Val Ile Leu 1310 1315 1320 Ala Asp Ala Asn Leu Asp Lys Val Leu Ser Ala Tyr Asn Lys His 1325 1330 1335 Arg Asp Lys Pro Ile Arg Glu Gln Ala Glu Asn Ile Ile His Leu 1340 1345 1350 Phe Thr Leu Thr Asn Leu Gly Ala Pro Ala Ala Phe Lys Tyr Phe 1355 1360 1365 Asp Thr Thr Ile Asp Arg Lys Arg Tyr Thr Ser Thr Lys Glu Val 1370 1375 1380 Leu Asp Ala Thr Leu Ile His Gln Ser Ile Thr Gly Leu Tyr Glu 1385 1390 1395 Thr Arg Ile Asp Leu Ser Gln Leu Gly Gly Asp Glu Gly Ala Pro 1400 1405 1410 Lys Lys Lys Arg Lys Val Gly Ser Ser Gly Ser Pro Lys Lys Lys 1415 1420 1425 Arg Lys Val Ser Gly Ser Glu Thr Pro Gly Thr Ser Glu Ser Ala 1430 1435 1440 Thr Pro Glu Ser Met Asp Ser Leu Leu Met Asn Arg Arg Glu Phe 1445 1450 1455 Leu Tyr Gln Phe Lys Asn Val Arg Trp Ala Lys Gly Arg Arg Glu 1460 1465 1470 Thr Tyr Leu Cys Tyr Val Val Lys Arg Arg Asp Ser Ala Thr Ser 1475 1480 1485 Phe Ser Leu Asp Phe Gly Tyr Leu Arg Asn Lys Asn Gly Cys His 1490 1495 1500 Val Glu Leu Leu Phe Leu Arg Tyr Ile Ser Asp Trp Asp Leu Asp 1505 1510 1515 Pro Gly Arg Cys Tyr Arg Val Thr Trp Phe Ile Ser Trp Ser Pro 1520 1525 1530 Cys Tyr Asp Cys Ala Arg His Val Ala Asp Phe Leu Arg Gly Asn 1535 1540 1545 Pro Asn Leu Ser Leu Arg Ile Phe Thr Ala Arg Leu Tyr Phe Cys 1550 1555 1560 Glu Asp Arg Lys Ala Glu Pro Glu Gly Leu Arg Arg Leu His Arg 1565 1570 1575 Ala Gly Val Gln Ile Ala Ile Met Thr Phe Lys Asp Tyr Phe Tyr 1580 1585 1590 Cys Trp Asn Thr Phe Val Glu Asn His Gly Arg Thr Phe Lys Ala 1595 1600 1605 Trp Glu Gly Leu His Glu Asn Ser Val Arg Leu Ser Arg Gln Leu 1610 1615 1620 Arg Arg Ile Leu Leu Pro 1625 <210> 69 <211> 4890 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequence: coding sequence of dCas9-XTEN-AIDX <400> 69 atggactata aggaccacga cggagactac aaggatcatg atattgatta caaagacgat 60 gacgataaga tggccccaaa gaagaagcgg aaggtcggta tccacggagt cccagcagct 120 accatggaca agaagtattc tatcggactg gccatcggga ctaatagcgt cgggtgggcc 180 gtgatcactg acgagtacaa ggtgccctct aagaagttca aggtgctcgg gaacaccgac 240 cggcattcca tcaagaaaaa tctgatcgga gctctcctct ttgattcagg ggagaccgct 300 gaagcaaccc gcctcaagcg gactgctaga cggcggtaca ccaggaggaa gaaccggatt 360 tgttaccttc aagagatatt ctccaacgaa atggcaaagg tcgacgacag cttcttccat 420 aggctggaag aatcattcct cgtggaagag gataagaagc atgaacggca tcccatcttc 480 ggtaatacg tcgacgaggt ggcctatcac gagaaatacc caaccatcta ccatcttcgc 540 aaaaagctgg tggactcaac cgacaaggca gacctccggc ttatctacct ggccctggcc 600 cacatgatca agttcagagg ccacttcctg atcgagggcg acctcaatcc tgacaatagc 660 gatgtggata aactgttcat ccagctggtg cagacttaca accagctctt tgaagagaac 720 cccatcaatg caagcggagt cgatgccaag gccattctgt cagcccggct gtcaaagagc 780 cgcagacttg agaatcttat cgctcagctg ccgggtgaaa agaaaaatgg actgttcggg 840 aacctgattg ctctttcact tgggctgact cccaatttca agtctaattt cgacctggca 900 gaggatgcca agctgcaact gtccaaggac acctatgatg acgatctcga caacctcctg 960 gcccagatcg gtgaccaata cgccgacctt ttccttgctg ctaagaatct ttctgacgcc 1020 atcctgctgt ctgacattct ccgcgtgaac actgaaatca ccaaggcccc tctttcagct 1080 tcaatgatta agcggtatga tgagcaccac caggacctga ccctgcttaa ggcactcgtc 1140 cggcagcagc ttccggagaa gtacaaggaa atcttctttg accagtcaaa gaatggatac 1200 gccggctaca tcgacggagg tgcctcccaa gaggaatttt ataagtttat caaacctatc 1260 cttgagaaga tggacggcac cgaagagctc ctcgtgaaac tgaatcggga ggatctgctg 1320 cggaagcagc gcactttcga caatgggagc attccccacc agatccatct tggggagctt 1380 cacgccatcc ttcggcgcca agaggacttc tacccctttc ttaaggacaa cagggagaag 1440 attgagaaaa ttctcacttt ccgcatcccc tactacgtgg gacccctcgc cagaggaaat agccggtttg cttggatgac cagaaagtca cagaaacta tcactccctg cagaaagtca gaggtggtgg acaagggagc cagcgctcag tcattcatcg aacggatgac taacttcgat aagaacctcc ccaatgagaa ggtcctgccg aaacattccc tgctctacga gtactttacc gtgtacaacg agctgaccaa ggtgaaatat gtcaccgaag ggatgagga gcccgcattc ctgtcaggcg grandfather ggcaattgtg gaccttctgt tcaagacca tagaaggtg accgtgaagc agctgaagga ggactatttc aagaaaattg aatgcttcga ctctgtggag attagcgggg tcgaagatcg gttcaacgca agcctgggta cctaccatga tctgcttaag atcatcaagg acaaggattt tctggacaat gaggagaag aggacatcct tgaggacatt gtcctgactc tcactctgtt cgaggaccgg gaaatgatcg aggagaggct tagcctac gcccatctgt tcgacgataa agtgatgaag caacttaaac ggagaagata taccggatgg ggacgcctta gccgcaaact catcaacgga atccggggaca aacagagcgg aaagaccatt cttgatttcc ttaagagcga cggattcgct aatcgcaact tcatgcaact tatccatgat 2220 gattccctga cctttaagga ggacatccag aaggcccaag tgtctggaca aggtgactca 2280 ctgcacgagc atatcgcaaa tctggctggt tcacccgcta ttaagaaggg tattctccag 2340 accgtgaaag tcgtggacga gctggtcaag gtgatgggtc gccataaacc agagaacatt 2400 gtcatcgaga tggccaggga aaaccagact acccagaagg gacagaagaa cagcaggagag 2460 cggatgaaaa gaattgagga agggattaag gagctcgggt cacagatcct taagagcac 2520 ccggtggaaa acacccagct tcagaatgag aagctctatc tgtactacct tcaaaatgga 2580 cgcgatatgt atgtggacca agagcttgat atcaacaggc tctcagacta cgacgtggac 2640 gccatcgtcc ctcagagctt cctcaaagac gactcaattg acaataaggt gctgactcgc 2700 tcagacaaga accggggaaa gtcagataac gtgccctcag aggaagtcgt gaaaaagatg 2760 aagaactatt ggcgccagct tctgaacgca aagctgatca ctcagcggaa gttcgacaat 2820 ctcactaagg ctgagagggg cggactgagc gaactggaca aagcaggatt cattaaacgg 2880 caacttgtgg agactcggca gattactaaa catgtagccc aaatccttga ctcacgcatg 2940 aataccaagt acgacgaaaa cgacaaactt atccgcgagg tgaaggtgat taccctgaag 3000 tccaagctgg tcagcgattt cagaaggac tttcaattct acaaagtgcg ggagatcaat 3060 aactatcatc atgctcatga cgcatatctg aatgccgtgg tgggaaccgc cctgatcaag 3120 aagtacccaa agctggaaag cgagttcgtg tacggagact acaaggtcta cgacgtgcgc 3180 aagatgattg ccaaatctga gcaggagatc ggaaaggcca ccgcaaagta cttcttctac 3240 agcaacatca tgaatttctt caagaccgaa atcacccttg caaacggtga gatccggaag 3300 aggccgctca tcgagactaa tggggagact ggcgaaatcg tgtgggacaa gggcagagat 3360 3420 gtgcaaaccg gaggcttttc taggaatca atcctcccca agcgcaactc cgacaagctc 3480 3540 gcttattctg tcctggtcgt ggctaaggtg gaaaaggaa agtctaagaa gctcaagagc 3600 gtgaaggaac tgctgggtat caccattatg gagcgcagct ccttcgagaa gaacccaatt 3660 gactttctcg aagccaaagg ttacaaggaa gtcaagaagg accttatcat caagctccca 3720 aagtatagcc tgttcgaact ggagaatggg cggaagcgga tgctcgcctc cgctggcgaa 3780 cttcagaagg gtaatgagct ggctctcccc tccaagtacg tgaatttcct ctaccttgca 3840 agccattacg agaagctgaa ggggagcccc gaggacaacg agcaaaagca actgtttgtg 3900 gagcagcata agcattatct ggacgagatc attgagcaga tttccgagtt ttctaaacgc 3960 gtcattctcg ctgatgccaa cctcgataaa gtccttagcg catacaataa gcacagagac 4020 aaaccaattc gggagcaggc tgagaatatc atccacctgt tcaccctcac caatcttggt 4080 gcccctgccg cattcaagta cttcgacacc accatcgacc ggaaacgcta tacctccacc 4140 aaagaagtgc tggacgccac cctcatccac cagagcatca ccggacttta cgaaactcgg 4200 attgacctct cacagctcgg aggggatgag ggagctccca agaaaaagcg caaggtaggt 4260 agttccggat ctccgaaaaa gaaacgcaaa gttagcggca gcgagactcc cgggacctca 4320 gagtccgcca cacccgaaag tatggacagc ctcttgatga accggaggaa gtttctttac 4380 caattcaaaa atgtccgctg ggctaagggt cggcgtgaga cctacctgtg ctacgtagtg 4440 aagaggcgtg acagtgctac atccttttca ctggactttg gttatcttcg caataagaac 4500 ggctgccacg tggaattgct cttcctccgc tacatctcgg actgggacct agaccctggc 4560 cgctgctacc gcgtcacctg gttcacctcc tggagcccct gctacgactg tgcccgacat 4620 gtggccgact ttctgcgagg gaaccccaac ctcagtctga ggatcttcac cgcgcgcctc 4680 tacttctgtg aggaccgcaa ggctgagccc gaggggctgc ggcggctgca ccgcgccggg 4740 gtgcaaatag ccatcatgac cttcaaagat tatttttact gctggaatac ttttgtagaa 4800 aaccatgaaa gaactttcaa agcctgggaa gggctgcatg aaaattcagt tcgtctctcc 4860 agacagcttc ggcgcatcct tttgccctga 4890 <210> 70 <211> 1629 <212> PRT <213> Artificial Sequence <220> <223> Description of artificial sequence: Amino acid sequence of dCas9-XTEN-AIDX <400> 70 Met Asp Tyr Lys Asp His Asp Gly Asp Tyr Lys Asp His Asp Ile Asp 1 5 10 15 Tyr Lys Asp Asp Asp Asp Lys Met Ala Pro Lys Lys Lys Arg Lys Val 20 25 30 Gly Ile His Gly Val Pro Ala Ala Thr Met Asp Lys Lys Tyr Ser Ile 35 40 45 Gly Leu Ala Ile Gly Thr Asn Ser Val Gly Trp Ala Val Ile Thr Asp 50 55 60 Glu Tyr Lys Val Pro Ser Lys Lys Phe Lys Val Leu Gly Asn Thr Asp 65 70 75 80 Arg His Ser Ile Lys Lys Asn Leu Ile Gly Ala Leu Leu Phe Asp Ser 85 90 95 Gly Glu Thr Ala Glu Ala Thr Arg Leu Lys Arg Thr Ala Arg Arg Arg 100 105 110 Tyr Thr Arg Arg Lys Asn Arg Ile Cys Tyr Leu Gln Glu Ile Phe Ser 115 120 125 Asn Glu Met Ala Lys Val Asp Asp Ser Phe Phe His Arg Leu Glu Glu 130 135 140 Ser Phe Leu Val Glu Glu Asp Lys Lys His Glu Arg His Pro Ile Phe 145 150 155 160 Gly Asn Ile Val Asp Glu Val Ala Tyr His Glu Lys Tyr Pro Thr Ile 165 170 175 Tyr His Leu Arg Lys Lys Leu Val Asp Ser Thr Asp Lys Ala Asp Leu 180 185 190 Arg Leu Ile Tyr Leu Ala Leu Ala His Met Ile Lys Phe Arg Gly His 195 200 205 Phe Leu Ile Glu Gly Asp Leu Asn Pro Asp Asn Ser Asp Val Asp Lys 210 215 220 Leu Phe Ile Gln Leu Val Gln Thr Tyr Asn Gln Leu Phe Glu Glu Asn 225 230 235 240 Pro Ile Asn Ala Ser Gly Val Asp Ala Lys Ala Ile Leu Ser Ala Arg 245 250 255 Leu Ser Lys Ser Arg Arg Leu Glu Asn Leu Ile Ala Gln Leu Pro Gly 260 265 270 Glu Lys Lys Asn Gly Leu Phe Gly Asn Leu Ile Ala Leu Ser Leu Gly 275 280 285 Leu Thr Pro Asn Phe Lys Ser Asn Phe Asp Leu Ala Glu Asp Ala Lys 290 295 300 Leu Gln Leu Ser Lys Asp Thr Tyr Asp Asp Asp Leu Asp Asn Leu Leu 305 310 315 320 Ala Gln Ile Gly Asp Gln Tyr Ala Asp Leu Phe Leu Ala Ala Lys Asn 325 330 335 Leu Ser Asp Ala Ile Leu Leu Ser Asp Ile Leu Arg Val Asn Thr Glu 340 345 350 Ile Thr Lys Ala Pro Leu Ser Ala Ser Met Ile Lys Arg Tyr Asp Glu 355 360 365 His His Gln Asp Leu Thr Leu Leu Lys Ala Leu Val Arg Gln Gln Leu 370 375 380 Pro Glu Lys Tyr Lys Glu Ile Phe Phe Asp Gln Ser Lys Asn Gly Tyr 385 390 395 400 Ala Gly Tyr Ile Asp Gly Gly Ala Ser Gln Glu Glu Phe Tyr Lys Phe 405 410 415 Ile Lys Pro Ile Leu Glu Lys Met Asp Gly Thr Glu Glu Leu Leu Val 420 425 430 Lys Leu Asn Arg Glu Asp Leu Leu Arg Lys Gln Arg Thr Phe Asp Asn 435 440 445 Gly Ser Ile Pro His Gln Ile His Leu Gly Glu Leu His Ala Ile Leu 450 455 460 Arg Arg Gln Glu Asp Phe Tyr Pro Phe Leu Lys Asp Asn Arg Glu Lys 465 470 475 480 Ile Glu Lys Ile Leu Thr Phe Arg Ile Pro Tyr Tyr Val Gly Pro Leu 485 490 495 Ala Arg Gly Asn Ser Arg Phe Ala Trp Met Thr Arg Lys Ser Glu Glu 500 505 510 Thr Ile Thr Pro Trp Asn Phe Glu Glu Val Val Asp Lys Gly Ala Ser 515 520 525 Ala Gln Ser Phe Ile Glu Arg Met Thr Asn Phe Asp Lys Asn Leu Pro 530 535 540 Asn Glu Lys Val Leu Pro Lys His Ser Leu Leu Tyr Glu Tyr Phe Thr 545 550 555 560 Val Tyr Asn Glu Leu Thr Lys Val Lys Tyr Val Thr Glu Gly Met Arg 565 570 575 Lys Pro Ala Phe Leu Ser Gly Glu Gln Lys Lys Ala Ile Val Asp Leu 580 585 590 Leu Phe Lys Thr Asn Arg Lys Val Thr Val Lys Gln Leu Lys Glu Asp 595 600 605 Tyr Phe Lys Lys Ile Glu Cys Phe Asp Ser Val Glu Ile Ser Gly Val 610 615 620 Glu Asp Arg Phe Asn Ala Ser Leu Gly Thr Tyr His Asp Leu Leu Lys 625 630 635 640 Ile Ile Lys Asp Lys Asp Phe Leu Asp Asn Glu Glu Lys Glu Asp Ile 645 650 655 Leu Glu Asp Ile Val Leu Thr Leu Thr Leu Phe Glu Asp Arg Glu Met 660 665 670 Ile Glu Glu Arg Leu Lys Thr Tyr Ala His Leu Phe Asp Asp Lys Val 675 680 685 Met Lys Gln Leu Lys Arg Arg Arg Tyr Thr Gly Trp Gly Arg Leu Ser 690 695 700 Arg Lys Leu Ile Asn Gly Ile Arg Asp Lys Gln Ser Gly Lys Thr Ile 705 710 715 720 Leu Asp Phe Leu Lys Ser Asp Gly Phe Ala Asn Arg Asn Phe Met Gln 725 730 735 Leu Ile His Asp Asp Ser Leu Thr Phe Lys Glu Asp Ile Gln Lys Ala 740 745 750 Gln Val Ser Gly Gln Gly Asp Ser Leu His Glu His Ile Ala Asn Leu 755 760 765 Ala Gly Ser Pro Ala Ile Lys Lys Gly Ile Leu Gln Thr Val Lys Val 770 775 780 Val Asp Glu Leu Val Lys Val Met Gly Arg His Lys Pro Glu Asn Ile 785 790 795 800 Val Ile Glu Met Ala Arg Glu Asn Gln Thr Thr Gln Lys Gly Gln Lys 805 810 815 Asn Ser Arg Glu Arg Met Lys Arg Ile Glu Glu Gly Ile Lys Glu Leu 820 825 830 Gly Ser Gln Ile Leu Lys Glu His Pro Val Glu Asn Thr Gln Leu Gln 835 840 845 Asn Glu Lys Leu Tyr Leu Tyr Tyr Leu Gln Asn Gly Arg Asp Met Tyr 850 855 860 Val Asp Gln Glu Leu Asp Ile Asn Arg Leu Ser Asp Tyr Asp Val Asp 865 870 875 880 Ala Ile Val Pro Gln Ser Phe Leu Lys Asp Asp Ser Ile Asp Asn Lys 885 890 895 Val Leu Thr Arg Ser Asp Lys Asn Arg Gly Lys Ser Asp Asn Val Pro 900 905 910 Ser Glu Glu Val Val Lys Lys Met Lys Asn Tyr Trp Arg Gln Leu Leu 915 920 925 Asn Ala Lys Leu Ile Thr Gln Arg Lys Phe Asp Asn Leu Thr Lys Ala 930 935 940 Glu Arg Gly Gly Leu Ser Glu Leu Asp Lys Ala Gly Phe Ile Lys Arg 945 950 955 960 Gln Leu Will Glu Thr Arg Gln Ile Thr Lys His Will Ala Gln Ile Leu 965,970,975 Asp Ser Arg With Thr Lys Tyr Asp Glu Asp Lys With Arg 980,985,990 Glu Val Lys Val Ile Thr Leu Lys Ser Lys Leu Val Ser Asp Phe Arg 995 1000 1005 Lys Asp Phe Gln Phe Tyr Lys Val Arg Glu Ile Asn Asn Tyr His 1010 1015 1020 His Ala His Asp Ala Tyr Leu Asn Ala Val Val Gly Thr Ala Leu 1025 1030 1035 Ile Lys Lys Pro Lys Leu Glu Ser Glu Phe Val Tyr Gly Asp 1040 1045 1050 Tyr Lys Val Tyr Asp Val Arg Lys Met Ile Ala Lys Ser Glu Gln 1055 1060 1065 Glu Ile Gly Lys Ala Thr Ala Lys Tyr Phe Phe Tyr Ser Asn Ile 1070 1075 1080 Met Asn Phe Phe Lys Thr Glu Ile Thr Leu Ala Asn Gly Glu Ile 1085 1090 1095 Arg Lys Arg Pro Leu Ile Glu Thr Asn Gly Glu Thr Gly Glu Ile 1100 1105 1110 Val Trp Asp Lys Gly Arg Asp Phe Ala Thr Val Arg Lys Val Leu 1115 1120 1125 Ser Met Pro Gln Val Asn Ile Val Lys Lys Thr Glu Val Gln Thr 1130 1135 1140 Gly Gly Phe Ser Lys Glu Ser Ile Leu Pro Lys Arg Asn Ser Asp 1145 1150 1155 Lys Leu Ile Ala Arg Lys Lys Asp Trp Asp Pro Lys Lys Tyr Gly 1160 1165 1170 Gly Phe Asp Ser Pro Thr Val Ala Tyr Ser Val Leu Val Val Ala 1175 1180 1185 Lys Val Glu Lys Gly Lys Ser Lys Lys Leu Lys Ser Val Lys Glu 1190 1195 1200 Leu Leu Gly Ile Thr Ile Met Glu Arg Ser Ser Phe Glu Lys Asn 1205 1210 1215 Pro Ile Asp Phe Leu Glu Ala Lys Gly Tyr Lys Glu Val Lys Lys 1220 1225 1230 Asp Leu Ile Ile Lys Leu Pro Lys Tyr Ser Leu Phe Glu Leu Glu 1235 1240 1245 Asn Gly Arg Lys Arg Met Leu Ala Ser Ala Gly Glu Leu Gln Lys 1250 1255 1260 Gly Asn Glu Leu Ala Leu Pro Ser Lys Tyr Val Asn Phe Leu Tyr 1265 1270 1275 Leu Ala Ser His Tyr Glu Lys Leu Lys Gly Ser Pro Glu Asp Asn 1280 1285 1290 Glu Gln Lys Gln Leu Phe Val Glu Gln His Lys His Tyr Leu Asp 1295 1300 1305 Glu Ile Ile Glu Gln Ile Ser Glu Phe Ser Lys Arg Val Ile Leu 1310 1315 1320 Ala Asp Ala Asn Leu Asp Lys Val Leu Ser Ala Tyr Asn Lys His 1325 1330 1335 Arg Asp Lys Pro Ile Arg Glu Gln Ala Glu Asn Ile Ile His Leu 1340 1345 1350 Phe Thr Leu Thr Asn Leu Gly Ala Pro Ala Ala Phe Lys Tyr Phe 1355 1360 1365 Asp Thr Thr Ile Asp Arg Lys Arg Tyr Thr Ser Thr Lys Glu Val 1370 1375 1380 Leu Asp Ala Thr Leu Ile His Gln Ser Ile Thr Gly Leu Tyr Glu 1385 1390 1395 Thr Arg Ile Asp Leu Ser Gln Leu Gly Gly Asp Glu Gly Ala Pro 1400 1405 1410 Lys Lys Lys Arg Lys Val Gly Ser Ser Gly Ser Pro Lys Lys Lys 1415 1420 1425 Arg Lys Val Ser Gly Ser Glu Thr Pro Gly Thr Ser Glu Ser Ala 1430 1435 1440 Thr Pro Glu Ser Met Asp Ser Leu Leu Met Asn Arg Arg Lys Phe 1445 1450 1455 Leu Tyr Gln Phe Lys Asn Val Arg Trp Ala Lys Gly Arg Arg Glu 1460 1465 1470 Thr Tyr Leu Cys Tyr Val Val Lys Arg Arg Asp Ser Ala Thr Ser 1475 1480 1485 Phe Ser Leu Asp Phe Gly Tyr Leu Arg Asn Lys Asn Gly Cys His 1490 1495 1500 Val Glu Leu Leu Phe Leu Arg Tyr Ile Ser Asp Trp Asp Leu Asp 1505 1510 1515 Pro Gly Arg Cys Tyr Arg Val Thr Trp Phe Thr Ser Trp Ser Pro 1520 1525 1530 Cys Tyr Asp Cys Ala Arg His Val Ala Asp Phe Leu Arg Gly Asn 1535 1540 1545 Pro Asn Leu Ser Leu Arg Ile Phe Thr Ala Arg Leu Tyr Phe Cys 1550 1555 1560 Glu Asp Arg Lys Ala Glu Pro Glu Gly Leu Arg Arg Leu His Arg 1565 1570 1575 Ala Gly Val Gln Ile Ala Ile Met Thr Phe Lys Asp Tyr Phe Tyr 1580 1585 1590 Cys Trp Asn Thr Phe Val Glu Asn His Glu Arg Thr Phe Lys Ala 1595 1600 1605 Trp Glu Gly Leu His Glu Asn Ser Val Arg Leu Ser Arg Gln Leu 1610 1615 1620 Arg Arg Ile Leu Leu Pro 1625 <210> 71 <211> 4917 <212> DNA <213> Artificial sequence <220> <223> Description of the artificial sequence: Coding sequence of nCas9-AIDX <400> 71 atggactata aggaccacga cggagactac aaggatcatg atattgatta caaagacgat 60 gacgataaga tggccccaaa gaagaagcgg aaggtcggta tccacggagt cccagcagct 120 accatggaca agaagtattc tatcggactg gccatcggga ctaatagcgt cgggtgggcc 180 gtgatcactg acgagtacaa ggtgccctct aagaagttca aggtgctcgg gaacaccgac 240 cggcattcca tcaagaaaaa tctgatcgga gctctcctct ttgattcagg ggagaccgct 300 gaagcaaccc gcctcaagcg gactgctaga cggcggtaca ccaggaggaa gaaccggatt 360 tgttaccttc aagagatatt ctccaacgaa atggcaaagg tcgacgacag cttcttccat 420 aggctggaag aatcattcct cgtggaagag gataagaagc atgaacggca tcccatcttc 480 ggtaatatcg tcgacgaggt ggcctatcac gagaaatacc caaccatcta ccatcttcgc 540 aaaaagctgg tggactcaac cgacaaggca gacctccggc ttatctacct ggccctggcc 600 cacatgatca agttcagagg ccacttcctg atcgagggcg acctcaatcc tgacaatagc 660 gatgtggata aactgttcat ccagctggtg cagacttaca accagctctt tgaagagaac 720 cccatcaatg caagcggagt cgatgccaag gccattctgt cagcccggct gtcaaagagc 780 cgcagacttg agaatcttat cgctcagctg ccgggtgaaa agaaaaatgg actgttcggg 840 aacctgattg ctctttcact tgggctgact cccaatttca agtctaattt cgacctggca 900 gaggatgcca agctgcaact gtccaaggac acctatgatg acgatctcga caacctcctg 960 gcccagatcg gtgaccaata cgccgacctt ttccttgctg ctaagaatct ttctgacgcc 1020 atcctgctgt ctgacattct ccgcgtgaac actgaatca ccaggcccc tcttcagct 1080 tcaatgatta agcggtatga tgagcaccac caggacctga ccctgcttaa ggcactcgtc 1140 cggcagcagc ttccggagaa gtacaaggaa atctctttg accagtcaa gatggatac 1200 gccggctaca tcgacggagg tgcctcccaa gaggaatttt ataagtttat caacctac 1260 cttgagaga tggacggcac cgaagagctc ctcgtgaac tgaatcggga ggatctgctg 1320 cggaagcagc gcacttcga caatgggagc attccccacc agatcatct tggggagctt 1380 cacgccatcc ttcggcgcca agaggacttc tacccctttc ttaggacaa cagggagaag 1440 attgagaaaa ttctcactt ccgcatcccc tactacgtgg gacccctcgc cagaggaaat 1500 agccggtttg cttggatgac cagaagtca gagaaacta tcactccctg gaacttcgaa 1560 gaggtggtgg acaagggagc cagcgctcag tcattcatcg aacggatgac taacttcgat 1620 aagaacctcc ccaatgagaa ggtcctgccg aaacattccc tgctctacga gtactttacc gtgtacaacg agctgaccaa ggtgaaatat gtcaccgaag ggatgagga gcccgcattc ctgtcaggcg grandfather ggcaattgtg gaccttctgt tcaagacca tagaaggtg accgtgaagc agctgaagga ggactatttc aagaaaattg aatgcttcga ctctgtggag attagcgggg tcgaagatcg gttcaacgca agcctgggta cctaccatga tctgcttaag atcatcaagg acaaggattt tctggacaat gaggagaag aggacatcct tgaggacatt gtcctgactc tcactctgtt cgaggaccgg gaaatgatcg aggagaggct tagcctac gcccatctgt tcgacgataa agtgatgaag caacttaaac ggagaagata taccggatgg ggacgcctta gccgcaaact catcaacgga atccggggaca aacagagcgg aaagaccatt cttgatttcc ttaagagcga cggattcgct aatcgcaact tcatgcaact tatccatgat gattccctga cctttaagga ggacatccag aaggcccag tgtctggaca aggtgactca ctgcacgagc atatcgcaaa tctggctggt tcacccgcta ttaagaaggg tattctccag accgtgaaag tcgtggacga gctggtcaag gtgatgggtc gccataaacc agagaacatt 2400 gtcatcgaga tggccaggga aaaccagact acccagaagg gacagaagaa cagcaggagag 2460 cggatgaaaa gaattgagga agggattaag gagctcgggt cacagatcct taagagcac 2520 ccggtggaaa acacccagct tcagaatgag aagctctatc tgtactacct tcaaaatgga 2580 cgcgatatgt atgtggacca agagcttgat atcaacaggc tctcagacta cgacgtggac 2640 catatcgtcc ctcagagctt cctcaaagac gactcaattg acaataaggt gctgactcgc 2700 tcagacaaga accggggaaa gtcagataac gtgccctcag aggaagtcgt gaaaaagatg 2760 aagaactatt ggcgccagct tctgaacgca aagctgatca ctcagcggaa gttcgacaat 2820 ctcactaagg ctgagagggg cggactgagc gaactggaca aagcaggatt cattaaacgg 2880 caacttgtgg agactcggca gattactaaa catgtagccc aaatccttga ctcacgcatg 2940 aataccaagt acgacgaaaa cgacaaactt atccgcgagg tgaaggtgat taccctgaag 3000 tccaagctgg tcagcgattt cagaaggac tttcaattct acaaagtgcg ggagatcaat 3060 aactatcatc atgctcatga cgcatatctg aatgccgtgg tgggaaccgc cctgatcaag 3120 aagtacccaa agctggaaag cgagttcgtg tacggagact acaaggtcta cgacgtgcgc 3180 aagatgattg ccaaatctga gcaggagatc ggaaaggcca ccgcaaagta cttcttctac 3240 agcaacatca tgaatttctt caagaccgaa atcacccttg caaacggtga gatccggaag 3300 aggccgctca tcgagactaa tggggagact ggcgaaatcg tgtgggacaa gggcagagat 3360 3420 gtgcaaaccg gaggcttttc taggaatca atcctcccca agcgcaactc cgacaagctc 3480 3540 gcttattctg tcctggtcgt ggctaaggtg gaaaaggaa agtctaagaa gctcaagagc 3600 gtgaaggaac tgctgggtat caccattatg gagcgcagct ccttgagagaa gaacccaatt 3660 gactttctcg aagccaaagg ttacaaggaa gtcaagaagg accttatcat caagctccca 3720 aagtatagcc tgttcgaact ggagaatggg cggaagcgga tgctcgcctc cgctggcgaa 3780 cttcagaagg gtaatgagct ggctctcccc tccaagtacg tgaatttcct ctaccttgca 3840 agccattacg agaagctgaa ggggagcccc gaggacaacg agcaaaagca actgtttgtg 3900 gagcagcata agcattatct ggacgagatc attgagcaga tttccgagtt ttctaaacgc 3960 gtcattctcg ctgatgccaa cctcgataaa gtccttagcg catacaataa gcacagagac 4020 aaaccaattc gggagcaggc tgagaatatc atccacctgt tcaccctcac caatctttggt 4080 gccctgccg cattcaagta cttcgacacc accatcgacc ggaaacgcta tacctccacc 4140 aaagaagtgc tggacgccac cctcatccac cagagcatca ccggacttta cgaaactcgg 4200 attgacctct cacagctcgg aggggatgag ggagctccca agaaaaagcg caaggtaggt 4260 agttccggat ctccgaaaaa gaaacgcaaa gttggtagtg atgctttaga cgattttgac 4320 ttagatatgc ttggttcaga cgcgttagac gacttcggtg gaggatccat ggacagcctc 4380 ttgatgaacc ggaggaagtt tctttaccaa ttcaaaaatg tccgctgggc taagggtcgg 4440 cgtgagacct acctgtgcta cgtagtgaag aggcgtgaca gtgctacatc cttttcactg 4500 gactttggtt atcttcgcaa taagaacggc tgccacgtgg aattgctctt cctccgctac 4560 atctcggact gggacctaga ccctggccgc tgctaccgcg tcacctggtt cacctcctgg 4620 agcccctgct acgactgtgc ccgacatgtg gccgactttc tgcgagggaa ccccaacctc 4680 agtctgagga tcttcaccgc gcgcctctac ttctgtgagg accgcaaggc tgagcccgag 4740 gggctgcggc ggctgcaccg cgccggggtg caaatagcca tcatgacctt caaagattat 4800 ttttactgct ggaatacttt tgtagaaaac catgaaagaa ctttcaaagc ctgggaaggg 4860 ctgcatgaaa attcagttcg tctctccaga cagcttcggc gcatcctttt gccctga 4917 <210> 72 <21​​​​​​​​​​​​​​​​​​​Gly Ile His Gly Val Pro Ala Ala Thr Met Asp Lys Lys Tyr Ser Ile 35 40 45 Gly Leu Ala Ile Gly Thr Asn Ser Val Gly Trp Ala Val Ile Thr Asp 50 55 60 Glu Tyr Lys Val Pro Ser Lys Lys Phe Lys Val Leu Gly Asn Thr Asp 65 70 75 80 Arg His Ser Ile Lys Lys Asn Leu Ile Gly Ala Leu Leu Phe Asp Ser 85 90 95 Gly Glu Thr Ala Glu Ala Thr Arg Leu Lys Arg Thr Ala Arg Arg Arg 100 105 110 Tyr Thr Arg Arg Lys Asn Arg Ile Cys Tyr Leu Gln Glu Ile Phe Ser 115 120 125 Asn Glu Met Ala Lys Val Asp Asp Ser Phe Phe His Arg Leu Glu Glu 130 135 140 Ser Phe Leu Val Glu Glu Asp Lys Lys His Glu Arg His Pro Ile Phe 145 150 155 160 Gly Asn Ile Val Asp Glu Val Ala Tyr His Glu Lys Tyr Pro Thr Ile 165 170 175 Tyr His Leu Arg Lys Lys Leu Val Asp Ser Thr Asp Lys Ala Asp Leu 180 185 190 Arg Leu Ile Tyr Leu Ala Leu Ala His Met Ile Lys Phe Arg Gly His 195 200 205 Phe Leu Ile Glu Gly Asp Leu Asn Pro Asp Asn Ser Asp Val Asp Lys 210 215 220 Leu Phe Ile Gln Leu Val Gln Thr Tyr Asn Gln Leu Phe Glu Glu Asn 225 230 235 240 Pro Ile Asn Ala Ser Gly Val Asp Ala Lys Ala Ile Leu Ser Ala Arg 245 250 255 Leu Ser Lys Ser Arg Arg Leu Glu Asn Leu Ile Ala Gln Leu Pro Gly 260 265 270 Glu Lys Lys Asn Gly Leu Phe Gly Asn Leu Ile Ala Leu Ser Leu Gly 275 280 285 Leu Thr Pro Asn Phe Lys Ser Asn Phe Asp Leu Ala Glu Asp Ala Lys 290 295 300 Leu Gln Leu Ser Lys Asp Thr Tyr Asp Asp Asp Leu Asp Asn Leu Leu 305 310 315 320 Ala Gln Ile Gly Asp Gln Tyr Ala Asp Leu Phe Leu Ala Ala Lys Asn 325 330 335 Leu Ser Asp Ala Ile Leu Leu Ser Asp Ile Leu Arg Val Asn Thr Glu 340 345 350 Ile Thr Lys Ala Pro Leu Ser Ala Ser Met Ile Lys Arg Tyr Asp Glu 355 360 365 His His Gln Asp Leu Thr Leu Leu Lys Ala Leu Val Arg Gln Gln Leu 370 375 380 Pro Glu Lys Tyr Lys Glu Ile Phe Phe Asp Gln Ser Lys Asn Gly Tyr 385 390 395 400 Ala Gly Tyr Ile Asp Gly Gly Ala Ser Gln Glu Glu Phe Tyr Lys Phe 405 410 415 Ile Lys Pro Ile Leu Glu Lys Met Asp Gly Thr Glu Glu Leu Leu Val 420 425 430 Lys Leu Asn Arg Glu Asp Leu Leu Arg Lys Gln Arg Thr Phe Asp Asn 435 440 445 Gly Ser Ile Pro His Gln Ile His Leu Gly Glu Leu His Ala Ile Leu 450 455 460 Arg Arg Gln Glu Asp Phe Tyr Pro Phe Leu Lys Asp Asn Arg Glu Lys 465 470 475 480 Ile Glu Lys Ile Leu Thr Phe Arg Ile Pro Tyr Tyr Val Gly Pro Leu 485 490 495 Ala Arg Gly Asn Ser Arg Phe Ala Trp Met Thr Arg Lys Ser Glu Glu 500 505 510 Thr Ile Thr Pro Trp Asn Phe Glu Glu Val Val Asp Lys Gly Ala Ser 515 520 525 Ala Gln Ser Phe Ile Glu Arg Met Thr Asn Phe Asp Lys Asn Leu Pro 530 535 540 Asn Glu Lys Val Leu Pro Lys His Ser Leu Leu Tyr Glu Tyr Phe Thr 545 550 555 560 Val Tyr Asn Glu Leu Thr Lys Val Lys Tyr Val Thr Glu Gly Met Arg 565 570 575 Lys Pro Ala Phe Leu Ser Gly Glu Gln Lys Lys Ala Ile Val Asp Leu 580 585 590 Leu Phe Lys Thr Asn Arg Lys Val Thr Val Lys Gln Leu Lys Glu Asp 595 600 605 Tyr Phe Lys Lys Ile Glu Cys Phe Asp Ser Val Glu Ile Ser Gly Val 610 615 620 Glu Asp Arg Phe Asn Ala Ser Leu Gly Thr Tyr His Asp Leu Leu Lys 625 630 635 640 Ile Ile Lys Asp Lys Asp Phe Leu Asp Asn Glu Glu Lys Glu Asp Ile 645 650 655 Leu Glu Asp Ile Val Leu Thr Leu Thr Leu Phe Glu Asp Arg Glu Met 660 665 670 Ile Glu Glu Arg Leu Lys Thr Tyr Ala His Leu Phe Asp Asp Lys Val 675 680 685 Met Lys Gln Leu Lys Arg Arg Arg Tyr Thr Gly Trp Gly Arg Leu Ser 690 695 700 Arg Lys Leu Ile Asn Gly Ile Arg Asp Lys Gln Ser Gly Lys Thr Ile 705 710 715 720 Leu Asp Phe Leu Lys Ser Asp Gly Phe Ala Asn Arg Asn Phe Met Gln 725 730 735 Leu Ile His Asp Asp Ser Leu Thr Phe Lys Glu Asp Ile Gln Lys Ala 740 745 750 Gln Val Ser Gly Gln Gly Asp Ser Leu His Glu His Ile Ala Asn Leu 755 760 765 Ala Gly Ser Pro Ala Ile Lys Lys Gly Ile Leu Gln Thr Val Lys Val 770 775 780 Val Asp Glu Leu Val Lys Val Met Gly Arg His Lys Pro Glu Asn Ile 785 790 795 800 Val Ile Glu Met Ala Arg Glu Asn Gln Thr Thr Gln Lys Gly Gln Lys 805 810 815 Asn Ser Arg Glu Arg Met Lys Arg Ile Glu Glu Gly Ile Lys Glu Leu 820 825 830 Gly Ser Gln Ile Leu Lys Glu His Pro Val Glu Asn Thr Gln Leu Gln 835 840 845 Asn Glu Lys Leu Tyr Leu Tyr Tyr Leu Gln Asn Gly Arg Asp Met Tyr 850 855 860 Val Asp Gln Glu Leu Asp Ile Asn Arg Leu Ser Asp Tyr Asp Val Asp 865 870 875 880 His Ile Val Pro Gln Ser Phe Leu Lys Asp Asp Ser Ile Asp Asn Lys 885 890 895 Val Leu Thr Arg Ser Asp Lys Asn Arg Gly Lys Ser Asp Asn Val Pro 900 905 910 Ser Glu Glu Val Val Lys Lys Met Lys Asn Tyr Trp Arg Gln Leu Leu 915 920 925 Asn Ala Lys Leu Ile Thr Gln Arg Lys Phe Asp Asn Leu Thr Lys Ala 930 935 940 Glu Arg Gly Gly Leu Ser Glu Leu Asp Lys Ala Gly Phe Ile Lys Arg 945 950 955 960 Gln Leu Val Glu Thr Arg Gln Ile Thr Lys His Val Ala Gln Ile Leu 965 970 975 Asp Ser Arg With Thr Lys Tyr Asp Glu Asp Lys With Arg 980,985,990 Glu Val Lys Val Ile Thr Leu Lys Ser Lys Leu Val Ser Asp Phe Arg 995 1000 1005 Lys Asp Phe Gln Phe Tyr Lys Val Arg Glu Ile Asn Asn Tyr His 1010 1015 1020 His Ala His Asp Ala Tyr Leu Asn Ala Val Val Gly Thr Ala Leu 1025 1030 1035 Ile Lys Lys Pro Lys Leu Glu Ser Glu Phe Val Tyr Gly Asp 1040 1045 1050 Tyr Lys Val Tyr Asp Val Arg Lys Met Ile Ala Lys Ser Glu Gln 1055 1060 1065 Glu Ile Gly Lys Ala Thr Ala Lys Tyr Phe Phe Tyr Ser Asn Ile 1070 1075 1080 Met Asn Phe Phe Lys Thr Glu Ile Thr Leu Ala Asn Gly Glu Ile 1085 1090 1095 Arg Lys Arg Pro Leu Ile Glu Thr Asn Gly Glu Thr Gly Glu Ile 1100 1105 1110 Val Trp Asp Lys Gly Arg Asp Phe Ala Thr Val Arg Lys Val Leu 1115 1120 1125 Ser Met Pro Gln Val Asn Ile Val Lys Lys Thr Glu Val Gln Thr 1130 1135 1140 Gly Gly Phe Ser Lys Glu Ser Ile Leu Pro Lys Arg Asn Ser Asp 1145 1150 1155 Lys Leu Ile Ala Arg Lys Lys Asp Trp Asp Pro Lys Lys Tyr Gly 1160 1165 1170 Gly Phe Asp Ser Pro Thr Val Ala Tyr Ser Val Leu Val Val Ala 1175 1180 1185 Lys Val Glu Lys Gly Lys Ser Lys Lys Leu Lys Ser Val Lys Glu 1190 1195 1200 Leu Leu Gly Ile Thr Ile Met Glu Arg Ser Ser Phe Glu Lys Asn 1205 1210 1215 Pro Ile Asp Phe Leu Glu Ala Lys Gly Tyr Lys Glu Val Lys Lys 1220 1225 1230 Asp Leu Ile Ile Lys Leu Pro Lys Tyr Ser Leu Phe Glu Leu Glu 1235 1240 1245 Asn Gly Arg Lys Arg Met Leu Ala Ser Ala Gly Glu Leu Gln Lys 1250 1255 1260 Gly Asn Glu Leu Ala Leu Pro Ser Lys Tyr Val Asn Phe Leu Tyr 1265 1270 1275 Leu Ala Ser His Tyr Glu Lys Leu Lys Gly Ser Pro Glu Asp Asn 1280 1285 1290 Glu Gln Lys Gln Leu Phe Val Glu Gln His Lys His Tyr Leu Asp 1295 1300 1305 Glu Ile Ile Glu Gln Ile Ser Glu Phe Ser Lys Arg Val Ile Leu 1310 1315 1320 Ala Asp Ala Asn Leu Asp Lys Val Leu Ser Ala Tyr Asn Lys His 1325 1330 1335 Arg Asp Lys Pro Ile Arg Glu Gln Ala Glu Asn Ile Ile His Leu 1340 1345 1350 Phe Thr Leu Thr Asn Leu Gly Ala Pro Ala Ala Phe Lys Tyr Phe 1355 1360 1365 Asp Thr Thr Ile Asp Arg Lys Arg Tyr Thr Ser Thr Lys Glu Val 1370 1375 1380 Leu Asp Ala Thr Leu Ile His Gln Ser Ile Thr Gly Leu Tyr Glu 1385 1390 1395 Thr Arg Ile Asp Leu Ser Gln Leu Gly Gly Asp Glu Gly Ala Pro 1400 1405 1410 Lys Lys Lys Arg Lys Val Gly Ser Ser Gly Ser Pro Lys Lys Lys 1415 1420 1425 Arg Lys Val Gly Ser Asp Ala Leu Asp Asp Phe Asp Leu Asp Met 1430 1435 1440 Leu Gly Ser Asp Ala Leu Asp Asp Phe Gly Gly Gly Ser Met Asp 1445 1450 1455 Ser Leu Leu Met Asn Arg Arg Lys Phe Leu Tyr Gln Phe Lys Asn 1460 1465 1470 Val Arg Trp Ala Lys Gly Arg Arg Glu Thr Tyr Leu Cys Tyr Val 1475 1480 1485 Val Lys Arg Arg Asp Ser Ala Thr Ser Phe Ser Leu Asp Phe Gly 1490 1495 1500 Tyr Leu Arg Asn Lys Asn Gly Cys His Val Glu Leu Leu Phe Leu 1505 1510 1515 Arg Tyr Ile Ser Asp Trp Asp Leu Asp Pro Gly Arg Cys Tyr Arg 1520 1525 1530 Val Thr Trp Phe Thr Ser Trp Ser Pro Cys Tyr Asp Cys Ala Arg 1535 1540 1545 His Val Ala Asp Phe Leu Arg Gly Asn Pro Asn Leu Ser Leu Arg 1550 1555 1560 Ile Phe Thr Ala Arg Leu Tyr Phe Cys Glu Asp Arg Lys Ala Glu 1565 1570 1575 Pro Glu Gly Leu Arg Arg Leu His Arg Ala Gly Val Gln Ile Ala 1580 1585 1590 Ile Met Thr Phe Lys Asp Tyr Phe Tyr Cys Trp Asn Thr Phe Val 1595 1600 1605 Glu Asn His Glu Arg Thr Phe Lys Ala Trp Glu Gly Leu His Glu 1610 1615 1620 Asn Ser Val Arg Leu Ser Arg Gln Leu Arg Arg Ile Leu Leu Pro 1625 1630 1635 <210> 73 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: target binding region of sgRNA <400> 73 tccctcacct gttctgtcac 20 <210> 74 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: target binding region of sgRNA <400> 74 gctccagtaa tcactggtga 20 <210> 75 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: target binding region of sgRNA <400> 75 gatccagctc cagtaatcac 20 <210> 76 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: target binding region of sgRNA <400> 76 gtgattactg gagctggatc 20 <210> 77 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: target binding region of sgRNA <400> 77 atggggtacg taagctacag 20 <210> 78 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: target binding region of sgRNA <400> 78 gagattcgac ttttgagaga 20 <210> 79 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: target binding region of sgRNA <400> 79 tattactgtg caaactggga 20 <210> 80 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: target binding region of sgRNA <400> 80 caaactggga cggtgattac 20 <210> 81 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: target binding region of sgRNA <400> 81 gacggtgatt actggggcca 20 <210> 82 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: target binding region of sgRNA <400> 82 gttgttgcca atactttggc 20 <210> 83 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: target binding region of sgRNA <400> 83 atagcgtcag tctttcctgc 20 <210> 84 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: target binding region of sgRNA <400> 84 gtattggcaa caacctacac 20 <210> 85 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: target binding region of sgRNA <400> 85 aggggatccc agagatggac 20 <210> 86 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: target binding region of sgRNA <400> 86 tatgcttccc agtccatctc 20 <210> 87 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: target binding region of sgRNA <400> 87 tctgtcaaca gagtaacagc 20 <210> 88 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Description of artificial sequences: target binding region of sgRNA <400> 88 gtcccccctc cgaacgtgta 20 <210> 89 <211> 4 <212> PRT <213> Artificial sequence <220> <223> Description of artificial sequences: repeating motifs of linkers <400> 89 Ser Gly Gly Ser 1 <210> 90 <211> 5 <212> PRT <213> Artificial sequence <220> <223> Description of artificial sequences: repeating motifs of linkers <400> 90 Gly Ser Ser Gly Ser 1 5 <210> 91 <211> 4 <212> PRT <213> Artificial sequence <220> <223> Description of artificial sequences: repeating motifs of linkers <400> 91 Gly Gly Gly Ser 1 <210> 92 <211> 5 <212> PRT <213> Artificial sequence <220> <223> Description of artificial sequences: repeating motifs of linkers <400> 92 Gly Gly Gly Gly Ser 1 5 <210> 93 <211> 5 <212> PRT <213> Artificial sequence <220> <223> Description of artificial sequences: repeating motifs of linkers <400> 93 Ser Ser Ser Ser Gly 1 5 <210> 94 <211> 5 <212> PRT <213> Artificial sequence <220> <223> Description of artificial sequences: repeating motifs of linkers <400> 94 Gly Ser Gly Ser Ala 1 5 <210> 95 <211> 5 <212> PRT <213> Artificial sequence <220> <223> Description of artificial sequences: repeating motifs of linkers <400> 95 Gly Gly Ser Gly Gly 1 5

Claims

1. A fusion protein, characterized in that The amino acid sequence of the fusion protein is shown in SEQ ID NO:

4.

2. A nucleic acid molecule, characterized in that The nucleic acid molecule encodes the fusion protein of claim 1.

3. The nucleic acid molecule according to claim 2, wherein The polynucleotide sequence of the nucleic acid molecule is shown in SEQ ID NO:

3.

4. A nucleic acid construct comprising the nucleic acid molecule according to claim 2 or 3.

5. The nucleic acid construct according to claim 4, wherein The nucleic acid construct is an expression vector for expressing the fusion protein of claim 1 in a host cell.

6. A host cell containing or expressing the fusion protein according to claim 1, or containing the nucleic acid molecule according to claim 2 or 3, or the nucleic acid construct according to claim 4 or 5.

7. A method for generating point mutations in cells, characterized in that: The method comprises: a step of expressing the fusion protein of claim 1 and sgRNA in the cell, wherein the sgRNA comprises a target binding region and a Cas protein recognition region, the target binding region can specifically bind to the nucleic acid sequence to be mutated, and the Cas protein recognition region can be recognized and bound by the Cas enzyme in the fusion protein.

8. The method according to claim 7, wherein The method includes the steps of transferring the fusion protein or its expression vector and the sgRNA or its expression vector into the cell, and then screening to obtain the desired mutant nucleic acid sequence.

9. The method according to claim 8, wherein The target binding region of the sgRNA specifically binds to the template chain of the nucleic acid sequence to be mutated, and the opposite region of the sgRNA binding region on the template chain is adjacent to the protospacer sequence motif recognized by the Cas protein, or is separated by no more than 10 bases.

10. The method according to claim 8, wherein The nucleic acid sequence to be mutated encodes a functional protein.

11. The method according to claim 10, wherein The functional protein is selected from the group consisting of antibodies, enzymes, lipoproteins, hormone proteins, transport and storage proteins, motor proteins, receptor proteins, and membrane proteins.

12. A kit, characterized in that The kit contains the fusion protein according to claim 1, the nucleic acid molecule according to claim 2 or 3, or the nucleic acid construct according to claim 4 or 5.

Citation Information

Patent Citations

  • Delivery, engineering and optimization of systems, methods and compositions for sequence manipulation and therapeutic applications

    CN105164264A

  • Inducible DNA binding proteins and genome perturbation tools and applications thereof

    CN105188767A

  • Nuclease-independent targeted gene editing platform and uses thereof

    CN108291218A

  • CAS variants for gene editing

    WO2015089406A1

Cited By

  • Method for regulating RNA splicing by inducing splice site base mutation or polypyrimidine region base substitution

    CN109295053A

  • Method for modulating RNA splicing by inducing base mutation at splice site or base substitution in polypyrimidine region

    WO2019020007A1