Method for modifying genomic sequences by specifically converting nucleic acid bases of a targeted DNA sequence, and molecular complex used therefor

By combining deaminase with the CRISPR-Cas system for base transformation without cleaving double-stranded DNA, the side effects caused by DNA double-stranded cleavage in existing genome editing technologies are solved, and efficient and safe genome editing is achieved.

CN111500570BActive Publication Date: 2025-06-20KOBE UNIV
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202010159184.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2014-09-30
Filing Date
2015-03-04
Publication Date
2025-06-20
Estimated Expiration
2035-03-04

AI Technical Summary

Technical Problem

Existing genome editing technologies usually rely on DNA double-strand cleavage, resulting in side effects such as cytotoxicity and chromosomal rearrangement, and it is difficult to genetically modify in certain cell types such as primate egg cells and single-cell microorganisms.

Method used

Using a method without cleavage of double-stranded DNA, the deaminase that catalyzes the deaminase reaction and combines with the CRISPR-Cas system to achieve base transformation of specific DNA sequences and modify genomic sequences.

Benefits of technology

This method avoids the side effects of DNA double-strand cleavage, improves the safety and efficiency of genetic modification, and can effectively perform genome editing in various cell types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111500570B_ABST
    Figure CN111500570B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for modifying a genomic sequence by specifically converting nucleic acid bases of a target DNA sequence, and a molecular complex used therefor. The present invention provides a method for modifying a target site of double-stranded DNA, which comprises: contacting a complex formed by linking a nucleic acid base conversion enzyme and a nucleic acid sequence recognition module capable of specifically binding to a target nucleotide sequence in a selected double-stranded DNA with the double-stranded DNA, without cleaving at least one strand of the double-stranded DNA at the target site, and causing one or more nucleotides at the target site to be deleted or converted into one or more other nucleotides, or inserting one or more nucleotides into the target site.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the invention with the application number CN201580023875.6, the filing date of March 4, 2015, and the invention title of "Method for Modifying Genomic Sequences of Nucleic Acid Bases Targeting DNA Sequences Specifically and Molecular Complexes Used Therein". Technical Field

[0002] The present invention relates to a method for modifying genomic sequences and a complex of a nucleic acid sequence recognition module and a nucleic acid base conversion enzyme used therein. The method for modifying genomic sequences can modify nucleic acid bases in a specific region of the genome without accompanying cleavage of double-stranded DNA (no cleavage or cleavage of one strand) and without inserting foreign DNA fragments. Background Art

[0003] In recent years, genome editing has attracted attention as a technique for modifying genes and genomic regions of interest in various biological species. Conventionally, as a method for genome editing, a method using an artificial nuclease that combines a molecule having sequence-independent DNA cleavage ability and a molecule having sequence recognition ability has been proposed (Non-Patent Document 1).

[0004] For example, there have been reports of a method of recombining a target locus in DNA in a plant cell or an insect cell as a host using a zinc finger nuclease (ZFN) formed by linking a zinc finger DNA binding domain and a non-specific DNA cleavage domain (Patent Document 1); a method of cutting and modifying a target gene within or adjacent to a specific nucleotide sequence using a TALEN formed by linking a transcriptional activator-like (TAL) effector, which is a DNA binding module possessed by the plant pathogen Xanthomonas, and a DNA endonuclease (Patent Document 2); or a method using a CRISPR-Cas9 system formed by combining a DNA sequence CRISPR (Clustered Regularly interspaced short palindromic repeats) that functions in the acquired immune systems of eubacteria and archaea and a nuclease Cas (CRISPR-associated) protein family that plays an important role together with CRISPR (Patent Document 3). Further, a method of cutting a target gene near the specific sequence using an artificial nuclease formed by linking a PPR protein and a nuclease has also been reported (Patent Document 4), wherein the PPR protein is constructed to recognize a specific nucleotide sequence using tandem PPR motifs each including 35 amino acids and recognizing one nucleic acid base.

[0005] Prior Art Documents

[0006] Patent Documents

[0007] Patent Document 1: Japanese Patent No. 4968498 Gazette

[0008] Patent Document 2: Japanese Patent Application Laid-Open No. 2013-513389 Gazette

[0009] Patent Document 3: Japanese Patent Application Laid-Open No. 2010-519929 Gazette

[0010] Patent Document 4: Japanese Unexamined Patent Application Publication No. 2013-128413 Gazette

[0011] Non-Patent Document

[0012] Non-Patent Document 1: Kelvin M Esvelt, Harris H Wang (2013) Genome-scale engineering for systems and synthetic biology, Molecular Systems Biology 9:641 Summary of the Invention

[0013] Problems to be Solved by the Invention

[0014] Genome editing techniques proposed so far are basically premised on double-DNA breaks (DSB). However, due to unexpected genome modifications, there are side effects such as strong cytotoxicity and chromosomal rearrangements, and there are common problems such as: the reliability in gene therapy is impaired, the number of surviving cells with nucleotide modifications is extremely small, and gene modification itself is difficult in primate oocytes and single-celled microorganisms.

[0015] Therefore, an object of the present invention is to provide a novel genome editing method that does not involve DSB and does not involve the insertion of exogenous DNA fragments, that is, a method for modifying the nucleic acid bases of a specific sequence of a gene by not cleaving double-stranded DNA or cleaving one strand, and a complex of a nucleic acid sequence recognition module and a nucleic acid base conversion enzyme for the same.

[0016] Means for Solving the Problems

[0017] The inventors of the present invention repeatedly conducted in-depth research to solve the above problems, and as a result, conceived a base conversion that does not involve DSB and uses a base conversion reaction based on DNA bases. Although the base conversion reaction itself based on the deamination reaction of DNA bases is known, it has not been achieved to identify a specific sequence of DNA with it, target any site, and specifically modify the targeted DNA by base conversion of DNA bases.

[0018] Therefore, by using a deaminase that catalyzes a deamination reaction as an enzyme that performs such a nucleic acid base conversion and linking it to a molecule having a DNA sequence recognition ability, a genome sequence modification based on a nucleic acid base conversion is performed in a region containing a specific DNA sequence.

[0019] Specifically, the CRISPR-Cas system (CRISPR-mutant Cas) is used. That is, a DNA for encoding an RNA (trans-activating crRNA: tracrRNA) for recruiting Cas proteins is connected to a sequence complementary to the target sequence of the gene to be modified, a genome-specific CRISPR-RNA: crRNA (gRNA) RNA molecule. On the other hand, a DNA (dCas) encoding a mutant Cas protein whose cutting ability of any one or both chains of double-stranded DNA has been inactivated and connected to a deaminase gene is made. These DNAs are introduced into a host yeast cell containing a gene to be modified. As a result, mutations are successfully introduced randomly within a few hundred nucleotides of the gene of interest including the target sequence. Compared with the case of using a double mutant Cas protein that does not cut the DNA chain of the double-stranded DNA, the efficiency of mutation introduction increases when using a mutant Cas protein that cuts any chain. In addition, it is clear that the area of ​​the mutation region and the diversity of mutations also change depending on which of the DNA double strands is cut. In addition, by targeting multiple regions in the gene of interest, mutations can be effectively introduced. That is, it was confirmed that the host cells into which the DNA had been introduced were inoculated into a non-selective medium, and the sequence of the gene of interest was checked for randomly selected colonies, and the mutation was introduced into almost all colonies. In addition, it was also confirmed that genome editing of multiple sites can be performed simultaneously by targeting the regions in more than two genes of interest. It was further confirmed that this method can simultaneously introduce mutations into alleles of a diploid or polyploid genome; it can introduce mutations not only into eukaryotic cells, but also into prokaryotic cells such as Escherichia coli; it can be widely used regardless of the species of the organism. In addition, by transiently performing a nucleic acid base conversion reaction at a desired period, the editing of essential genes, which has been less efficient so far, can be effectively performed.

[0020] The present inventors have further conducted studies based on these findings and have completed the present invention.

[0021] That is, the present invention is as follows.

[0022] [1] A method for modifying a target site of double-stranded DNA, comprising: contacting a complex formed by ligating a nucleic acid base-converting enzyme and a nucleic acid sequence recognition module that specifically binds to a target nucleotide sequence in a selected double-stranded DNA with the double-stranded DNA, without cleaving at least one strand of the double-stranded DNA at the target site, and causing one or more nucleotides at the target site to be deleted or converted into one or more other nucleotides, or inserting one or more nucleotides into the target site.

[0023] [2] The method according to [1], wherein the nucleic acid sequence recognition module is selected from: at least one CRISPR-Cas system in which the DNA cleavage ability of Cas has been inactivated, zinc finger motifs, TAL effectors, and PPR motifs.

[0024] [3] The method according to [1], wherein the nucleic acid sequence recognition module is at least one CRISPR-Cas system in which the DNA cleavage ability of Cas has been inactivated.

[0025] [4] The method according to any one of [1] to [3], wherein two or more nucleic acid sequence recognition modules that specifically bind to different target nucleotide sequences are used.

[0026] [5] The method according to [4], wherein the different target nucleotide sequences are present in different genes.

[0027] [6] The method according to any one of [1] to [5], wherein the nucleic acid base-converting enzyme is a deaminase.

[0028] [7] The method according to [6], wherein the deaminase is AID (AICDA).

[0029] [8] The method according to any one of [1] to [7], wherein the contact between the double-stranded DNA and the complex is carried out by introducing a nucleic acid encoding the complex into a cell having the double-stranded DNA.

[0030] [9] The method according to [8], wherein the cell is a prokaryotic cell.

[0031]

[10] The method according to [8], wherein the cell is a eukaryotic cell.

[0032]

[11] The method according to [8], wherein the cell is a microbial cell.

[0033]

[12] The method according to [8], wherein the cell is a plant cell.

[0034]

[13] The method according to [8], wherein the cell is an insect cell.

[0035]

[14] According to the method described in [8], wherein the cell is an animal cell.

[0036]

[15] According to the method described in [8], wherein the cell is a vertebrate cell.

[0037]

[16] According to the method described in [8], wherein the cell is a mammalian cell.

[0038]

[17] According to the method described in any one of [9] to

[16] , wherein the cell is a polyploid cell, and sites within all targeted alleles on homologous chromosomes are modified.

[0039]

[18] According to the method described in any one of [8] to

[17] , it includes: a step of introducing an expression vector containing a nucleic acid encoding a complex in a form capable of controlling the expression period into the cell, and a step of inducing the expression of the nucleic acid at a time when it is necessary to stabilize the modification of the target site of double-stranded DNA.

[0040]

[19] According to the method described in

[18] , wherein the target nucleotide sequence in the double-stranded DNA is present within a gene essential for the cell.

[0041]

[20] A nucleic acid modification enzyme complex, which is composed of a nucleic acid sequence recognition module that specifically binds to a target nucleotide sequence in double-stranded DNA and a nucleic acid base conversion enzyme, and in the targeted site, does not cleave at least one strand of the double-stranded DNA, causes deletion or conversion of one or more nucleotides at the targeted site into one or more other nucleotides, or inserts one or more nucleotides into the targeted site.

[0042]

[21] A nucleic acid encoding the nucleic acid modification enzyme complex described in

[20] .

[0043] Advantages of the Invention

[0044] According to the genome editing technology of the present invention, since it is not associated with the insertion of foreign DNA or double-stranded DNA cleavage, it is excellent in terms of safety. Even in cases where there are biological or legal controversies in the form of gene recombination in conventional methods, the technology may provide solutions. In addition, theoretically, the range in which mutations can be introduced can be widely set from a single-base pinpoint to a range of several hundred bases, and it can also be applied to local evolution induction by randomly introducing mutations into specific limited regions where it has been almost impossible to do so hitherto. Brief Description of the Drawings

[0045] Figure 1 shows a schematic illustration of the mechanism of the gene modification method of the present invention using the CRISPR-Cas system.​

[0046] Figure 2 shows the results of verifying the effect of the gene modification method of the present invention, which comprises a combination of a CRISPR-Cas system and the PmCDA1 deaminase from Petromyzon marinus, using budding yeast.

[0047] Figure 3 shows a graph of the change in the number of viable cells after expression induction in the case of combining the CRISPR-Cas9 system using the D10A mutant of Cas9 with nickase activity, deaminase, and PmCDA1 (nCas9D10A-PmCDA1), and in the case of using a conventional Cas9 with DNA double-strand cleavage ability.

[0048] Figure 4 shows a graph of the results of introducing an expression construct together with two gRNAs (targeting the sequences of target4 and target5) into budding yeast, where the expression construct was constructed by connecting the human-derived AID deaminase in a manner of being linked to dCas9 via an SH3 domain and its binding ligand.

[0049] Figure 5 shows a graph in which the mutagenesis efficiency is increased by using Cas9 that cleaves either DNA single strand.

[0050] Figure 6 shows a graph in which, without cleaving double-stranded DNA, the size and frequency of the mutagenesis region change depending on which single strand is cleaved.

[0051] Figure 7 shows a graph in which extremely high mutagenesis efficiency can be achieved by targeting two adjacent regions.

[0052] Figure 8 shows a graph that the gene modification method of the present invention does not require selection based on a marker. All colonies with determined sequences were introduced with mutations.

[0053] Figure 9 ​​​​​​​​Figure showing a gene modification method according to the present invention, which can simultaneously edit multiple sites in the genome. The upper figure shows the nucleotide sequences and amino acid sequences of the targeting sites of each gene. The arrows above the nucleotide sequences indicate the target nucleotide sequences. The numbers at the arrow tails or arrowheads represent the positions of the ends of the target nucleotide sequences on the ORF. The lower figure shows the sequencing results of the targeting sites in each of the 5 clones of red (R) and white (W) colonies. It shows that base conversions occurred in the nucleotides represented by letters in outline font in the displayed sequences. It should be noted that for the reactivity to canavanine (Can R ), R shows resistance and S shows sensitivity.

[0054] Figure 10 Figure showing that according to the gene modification method of the present invention, mutations can be simultaneously introduced into two alleles on homologous chromosomes of a diploid genome. Figure 10 A shows the respective homologous mutation introduction efficiencies for the ade1 gene (upper figure) and the can1 gene, Figure 10 B shows that homologous mutations were actually introduced in red colonies (lower figure). In addition, it shows that heterologous mutations also occurred in white colonies (upper figure).

[0055] Figure 11 Figure showing that according to the gene modification method of the present invention, genome editing of Escherichia coli, a prokaryotic cell, can be performed. Figure 11 A is a schematic illustration of the plasmid used. Figure 11 B shows that the region within the galK gene can be targeted to effectively introduce a mutation (CAA→TAA). Figure 11 C shows the results of sequencing analysis of 2 clones each for colonies on media without selection (none), media containing 25 μg / ml rifampicin (Rif 25), or media containing 50 μg / ml rifampicin (Rif 50). It was confirmed that mutations conferring rifampicin resistance were introduced (upper figure). The appearance frequency of rifampicin-resistant strains was estimated to be about 10% (lower figure).

[0056] Figure 12 Figure showing the regulation of the edited base sites based on the length of the guide RNA. Figure 12 A shows a conceptual diagram of the edited base sites when the target nucleotide sequence is 20 bases long and 24 bases long. Figure 12 ​​​Panel B shows the results of editing by targeting the gsiA gene and changing the length of the target nucleotide sequence. The mutated sites are in bold. "T" and "A" indicate that the mutation was fully introduced into the clone (C→T or G→A); "t" indicates that the mutation (C→T) was introduced into the clone at a ratio of more than 50% (incomplete clone); "c" indicates that the introduction efficiency of the mutation (C→T) in the clone was less than 50%.

[0057] Figure 13 Schematic illustration of the temperature-sensitive plasmid for mutation introduction used in Example 11.

[0058] Figure 14 Figure showing the protocol for mutation introduction in Example 11.

[0059] Figure 15 Figure showing the results of introducing a mutation into the rpoB gene in Example 11.

[0060] Figure 16 Figure showing the results of introducing a mutation into the galK gene in Example 11. Detailed implementation mode

[0061] The present invention provides a method for modifying a target site of a double-stranded DNA by converting the target nucleotide sequence and the nucleotides in its vicinity in the double-stranded DNA into other nucleotides without cleaving at least one strand of the double-stranded DNA to be modified. The method includes the step of converting the target site, that is, the target nucleotide sequence and the nucleotides in its vicinity, into other nucleotides by bringing a complex formed by a nucleic acid base-converting enzyme and a nucleic acid sequence recognition module capable of specifically binding to the target nucleotide sequence in the double-stranded DNA into contact with the double-stranded DNA.

[0062] In the present invention, "modifying" a double-stranded DNA means deleting a certain nucleotide (for example, dC) on the DNA strand or converting it into other nucleotides (for example, dT, dA or dG), or inserting a nucleotide or a nucleotide sequence between the nucleotides on the DNA strand. Here, there is no particular limitation on the double-stranded DNA to be modified, and genomic DNA is preferably used. In addition, the "target site" of the double-stranded DNA refers to all or part of the "target nucleotide sequence" that can be specifically recognized and bound by the nucleic acid sequence recognition module, or refers to it and the vicinity (any one or both of the 5' upstream and 3' downstream) of the target nucleotide sequence. Depending on the purpose, the range can be appropriately adjusted between 1 base and several hundred bases in length.

[0063] ​​​​In the present invention, the "nucleic acid sequence recognition module" refers to a molecule or molecular complex that has the ability to specifically recognize and bind to a specific nucleotide sequence (i.e., the target nucleotide sequence) on a DNA strand. The nucleic acid sequence recognition module can bind to the target nucleotide sequence, enabling the nucleic acid base conversion enzyme linked to this module to exert a specific function at the targeted site of double-stranded DNA.

[0064] In the present invention, the "nucleic acid base conversion enzyme" refers to an enzyme that can convert a target nucleotide into another nucleotide without cleaving the DNA strand by catalyzing a reaction in which a substituent on the purine or pyrimidine ring of a DNA base is converted into another group or atom.

[0065] In the present invention, the "nucleic acid modification enzyme complex" refers to a complex formed by linking the above-mentioned nucleic acid sequence recognition module and nucleic acid base conversion enzyme, wherein the complex has nucleic acid base conversion enzyme activity and is endowed with the ability to recognize a specific nucleotide sequence. The "complex" used in this application includes not only forms composed of multiple molecules, but also single molecules such as fusion proteins that have a nucleic acid sequence recognition module and a nucleic acid base conversion enzyme.

[0066] The nucleic acid base conversion enzyme used in the present invention is not particularly limited as long as it can catalyze the above reaction. Examples include deaminases belonging to the nucleic acid / nucleotide deaminase superfamily that catalyze deamination reactions in which an amino group is converted into a carbonyl group. Preferred nucleic acid base conversion enzymes include: cytidine deaminases that can convert cytosine or 5-methylcytosine into uracil or thymine respectively, adenosine deaminases that can convert adenine into hypoxanthine, guanosine deaminases that can convert guanine into xanthine, etc. As cytidine deaminases, more preferably activation-induced cytidine deaminase (hereinafter also referred to as AID), which is an enzyme that introduces mutations into immunoglobulin genes in the acquired immunity of vertebrates, etc.

[0067] There is no particular limitation on the source of the nucleic acid base conversion enzyme, and for example, PmCDA1 (Petromyzon marinus cytosine deaminase 1) from lamprey, AID (Activation-induced cytidine deaminase; AICDA) from mammals (such as humans, pigs, cows, horses, monkeys, etc.) can be used. The base sequence and amino acid sequence of the CDS of PmCDA1 are shown in SEQ ID NOs: 1 and 2 respectively, and the base sequence and amino acid sequence of the CDS of human AID are shown in SEQ ID NOs: 3 and 4 respectively.

[0068] For the target nucleotide sequence in double-stranded DNA that is recognized by the nucleic acid sequence recognition module of the nucleic acid modification enzyme complex of the present invention, there is no particular limitation as long as it can specifically bind to this module, and it can be any sequence in double-stranded DNA. Regarding the length of the target nucleotide sequence, as long as it is sufficient to specifically bind to the nucleic acid sequence recognition module, for example: when introducing a mutation into a specific site in the genomic DNA of a mammal, according to its genomic size, it is 12 nucleotides or more, preferably 15 nucleotides or more, more preferably 17 nucleotides or more. There is no particular limitation on the upper limit of the length, preferably 25 nucleotides or less, more preferably 22 nucleotides or less.

[0069] As the nucleic acid sequence recognition module of the nucleic acid modification enzyme complex of the present invention, for example, a CRISPR-Cas system (CRISPR-mutant Cas) in which at least one DNA cleavage ability of Cas has been inactivated, a zinc finger motif, a TAL effector, a PPR motif, etc. can be used. In addition, fragments containing DNA binding domains of proteins such as restriction enzymes, transcription factors, and RNA polymerases that can specifically bind to DNA and do not have DNA double-strand cleavage ability, etc., are not limited to these. The module preferably includes: CRISPR-mutant Cas, zinc finger motifs, TAL effectors, PPR motifs, etc.

[0070] The zinc finger motif is formed by connecting 3 to 6 different types of Cys2His2 zinc finger units (one finger recognizes approximately 3 bases), and can recognize a target nucleotide sequence of 9 to 18 bases. The zinc finger motif can be prepared by known methods such as the Modular assembly method (Nat Biotechnol (2002) 20: 135-141), the OPEN method (Mol Cell (2008) 31: 294-301), the CoDA method (Nat Methods (2011) 8: 67-69), the Escherichia coli one-hybrid method (Nat Biotechnol (2008) 26: 695-701), etc. For the details of preparing the zinc finger motif, reference can be made to the above-mentioned Patent Document 1.

[0071] For TAL effectors, they have a repeat structure with a module of approximately 34 amino acids as a unit, and the binding stability and base specificity are determined by the 12th and 13th amino acid residues (referred to as RVD) of one module. Due to the relatively high independence of each module, TAL effectors specific to a target nucleotide sequence can be produced by simply connecting the modules. TAL effectors can be constructed using open resource production methods (such as the REAL method (Curr Protoc Mol Biol (2012) Chapter 12: Unit 12.15), the FLASH method (Nat Biotechnol (2012) 30: 460 - 465), the Golden Gate method (Nucleic Acids Res (2011) 39: e82), etc.), and it is relatively easy to design TAL effectors relative to the target nucleotide sequence. For the details of producing TAL effectors, reference can be made to the above-mentioned Patent Document 2.

[0072] For PPR motifs, they are constructed such that one nucleic acid base - recognizing PPR motif containing 35 amino acids in series recognizes a specific nucleotide sequence, and only the 1st, 4th, and ii (the second - last) amino acids of each motif recognize the targeted base. Since there is no dependence on the motif configuration and it is not interfered by the motifs on both sides, similar to TAL effectors, PPR proteins specific to a target nucleotide sequence can be produced by simply connecting the PPR motifs. For the details of producing PPR motifs, reference can be made to the above-mentioned Patent Document 4.

[0073] In addition, in the case of using fragments such as restriction enzymes, transcription factors, RNA polymerases, etc., since the DNA - binding domains of their proteins are well - known, fragments containing the said domains and without DNA double - strand cleavage ability can be easily designed and constructed.

[0074] Any of the above nucleic acid sequence - recognizing modules can also be provided in the form of a fusion protein with the above nucleic acid base - converting enzyme. Alternatively, protein - binding domains such as the SH3 domain, PDZ domain, GK domain, GB domain, etc. and their binding partners can be fused with the nucleic acid sequence - recognizing module and the nucleic acid base - converting enzyme respectively, and provided in the form of a protein complex through the interaction of the domain and their binding ligands. Or, the nucleic acid sequence - recognizing module can be fused with the nucleic acid base - converting enzyme and an intein respectively, and the two can be connected by the conjugation after the synthesis of each protein.

[0075] Regarding the contact between the nucleic acid modification enzyme complex of the present invention, which is a complex (including a fusion protein) formed by linking a nucleic acid sequence recognition module and a nucleic acid base-converting enzyme, and double-stranded DNA, although it can be carried out in the form of an enzyme reaction in a cell-free system, in view of the main object of the present invention, it is necessary to carry out the contact by introducing a nucleic acid encoding the complex into a cell having the double-stranded DNA of interest (for example, genomic DNA).

[0076] Therefore, for the nucleic acid sequence recognition module and the nucleic acid base-converting enzyme, it is preferably prepared in the form of a nucleic acid encoding their fusion protein, or in the form of nucleic acids encoding them separately so that after translation into proteins using a binding domain, an intein, etc., a complex can be formed in the host cell. Here, the nucleic acid can be either DNA or RNA. When it is DNA, it is preferably double-stranded DNA and is provided in the form of an expression vector configured under the control of a functional promoter in the host cell. When it is RNA, it is preferably single-stranded RNA.

[0077] Since the complex of the present invention formed by linking a nucleic acid sequence recognition module and a nucleic acid base-converting enzyme is not associated with DNA double-strand cleavage (DSB), a genome editing with lower toxicity can be carried out, and the gene modification method of the present invention can be applied to a wide range of biological materials. Therefore, for the cells into which the nucleic acid encoding the nucleic acid sequence recognition module and / or the nucleic acid base-converting enzyme is introduced, it can include cells of all biological species, from bacteria such as Escherichia coli as prokaryotes, cells of microorganisms such as yeast as lower eukaryotes, to cells of higher eukaryotes such as insects and plants, and cells of vertebrates such as mammals including humans.

[0078] For the DNA encoding nucleic acid sequence recognition modules such as zinc finger motifs, TAL effectors, and PPR motifs, each module can be obtained by any of the above methods. For the DNA encoding sequence recognition modules such as restriction enzymes, transcription factors, and RNA polymerases, cloning can be carried out as follows: for example, based on their cDNA sequence information, oligonucleotide DNA primers covering the region encoding the desired part (the part containing the DNA binding domain) of the protein are synthesized, and using the total RNA or mRNA fraction prepared from the cells producing the protein as a template, amplification is carried out by RT-PCR method.

[0079] DNA encoding a nucleic acid base-converting enzyme can also be cloned in the same way: oligonucleotide DNA primers are synthesized based on the cDNA sequence information of the enzyme to be used, and the total RNA or mRNA fraction prepared from the cells producing the enzyme is used as a template, and amplified by RT-PCR. For example, for the DNA encoding PmCDA1 of lamprey, appropriate primers can be designed for the upstream and downstream of the CDS based on the cDNA sequence (accession No. EF094822) registered in the NCBI database, and cloned from the mRNA of lamprey by RT-PCR. In addition, for the DNA encoding human AID, appropriate primers can be designed for the upstream and downstream of the CDS based on the cDNA sequence (accession No. AB040431) registered in the NCBI database, and cloned from the mRNA of human lymph nodes by, for example, RT-PCR.

[0080] The cloned DNA can be digested directly or as needed with restriction enzymes, or after adding appropriate linkers and / or nuclear localization signals (in the case where the double-stranded DNA of interest is mitochondrial or chloroplast DNA, the transfer signals for each organelle), ligated to the DNA encoding the nucleic acid sequence recognition module to prepare the DNA encoding the fusion protein. Alternatively, the DNA encoding the nucleic acid sequence recognition module and the DNA encoding the nucleic acid base-converting enzyme can be fused separately with the DNA encoding the binding domain or its binding partner, or the two DNAs can also be fused with the DNA encoding the split intein, so that the nucleic acid sequence recognition conversion module and the nucleic acid base-converting enzyme are translated in the host cell and then form a complex. In these cases, linkers and / or nuclear localization signals can also be ligated at appropriate positions of one or both DNAs as needed.

[0081] For the DNA encoding the nucleic acid sequence recognition module and the DNA encoding the nucleic acid base-converting enzyme, the DNA strand can be chemically synthesized, or the partially overlapping synthetic oligonucleotide short chains can be ligated by PCR or Gibson Assembly to construct the DNA encoding its full length. The advantage of constructing the full-length DNA by combinatorial chemical synthesis or PCR or Gibson Assembly is that codons compatible with the host into which the DNA is to be introduced can be designed throughout the entire CDS full length. By converting the DNA sequence to codons with a high usage frequency in the host organism when expressing heterologous DNA, an increase in protein expression level can be expected. For data on the codon usage frequency in the host used, the genetic code usage frequency database publicly available on the homepage of the (public interest incorporated foundation) Kazusa DNA Research Institute can be used.

[0082] (http: / / www.kazusa.or.jp / codon / index.html), and it is also possible to refer to the literature recording the codon usage frequencies in each host. As long as the obtained data and the DNA sequence to be introduced are referred to, the codons with low usage frequencies in the host among the codons used in the DNA sequence can be changed to the codons with high usage frequencies that encode the same amino acid.

[0083] A DNA expression vector containing a nucleic acid sequence recognition module and / or a nucleic acid base conversion enzyme can be produced, for example, by ligating the DNA downstream of the promoter of an appropriate expression vector.

[0084] As the expression vector, plasmids from Escherichia coli (e.g., pBR322, pBR325, pUC12, pUC13); plasmids from Bacillus subtilis (e.g., pUB110, pTP5, pC194); plasmids from yeast (e.g., pSH19, pSH15); insect cell expression plasmids (e.g.: pFast-Bac); animal cell expression plasmids (e.g.: pA1-11, pXT1, pRc / CMV, pRc / RSV, pcDNAI / Neo); phages such as λ phage; insect virus vectors such as baculovirus (e.g.: BmNPV, AcNPV); animal virus vectors such as retrovirus, vaccinia virus, adenovirus, etc. can be used.

[0085] As the promoter, any appropriate promoter corresponding to the host for gene expression can be used. Due to the toxicity of the conventional method involving DSB, the survival rate of host cells sometimes decreases significantly. Therefore, it is desirable to increase the number of cells until the induction starts by using an inducible promoter. However, since sufficient cell proliferation can also be achieved by expressing the nucleic acid modification enzyme complex of the present invention, the constitutive promoter can be used without limitation.

[0086] For example, when the host is an animal cell, the SRα promoter, SV40 promoter, LTR promoter, CMV (cytomegalovirus) promoter, RSV (Rous sarcoma virus) promoter, MoMuLV (Moloney murine leukemia virus) LTR, HSV-TK (herpes simplex virus thymidine kinase) promoter, etc. can be used. Among them, the CMV promoter, SRα promoter, etc. are preferred.

[0087] When the host is Escherichia coli, the trp promoter, lac promoter, recA promoter, λPL promoter, lpp promoter, T7 promoter, etc. are preferred.

[0088] When the host is a Bacillus species, the SPO1 promoter, SPO2 promoter, penP promoter, etc. are preferred.

[0089] When the host is yeast, the Gal1 / 10 promoter, PHO5 promoter, PGK promoter, GAP promoter, ADH promoter, etc. are preferred.

[0090] When the host is an insect cell, the polyhedrin promoter, P10 promoter, etc. are preferred.

[0091] When the host is a plant cell, the CaMV35S promoter, CaMV19S promoter, NOS promoter, etc. are preferred.

[0092] As an expression vector, in addition to the above, a vector containing an enhancer, splicing signal, terminator, polyA addition signal, selection markers such as drug resistance genes, auxotrophic complementation genes, etc., origin of replication, etc. can be used as needed.

[0093] RNA encoding a nucleic acid sequence recognition module and / or a nucleic acid base conversion enzyme can be prepared, for example, by using a DNA encoding vector encoding the above nucleic acid sequence recognition module and / or nucleic acid base conversion enzyme as a template and transcribing it into mRNA through an in vitro transcription system known per se.

[0094] A complex of a nucleic acid sequence recognition module and a nucleic acid base conversion enzyme can be expressed intracellularly by introducing a DNA expression vector containing the nucleic acid sequence recognition module and / or nucleic acid base conversion enzyme into a host cell and culturing the host cell.

[0095] As a host, for example, Escherichia, Bacillus, yeast, insect cells, insects, animal cells, etc. can be used.

[0096] As Escherichia, for example, Escherichia coli K12?DH1 [Proc. Natl. Acad. Sci. USA, 60, 160 (1968)], Escherichia coli JM103 [Nucleic Acids Research, 9, 309 (1981)], Escherichia coli JA221 [Journal of Molecular Biology, 120, 517 (1978)], Escherichia coli HB101 [Journal of Molecular Biology, 41, 459 (1969)], Escherichia coli C600 [Genetics, 39, 440 (1954)], etc. can be used.

[0097] As the genus Bacillus, for example, Bacillus subtilis MI114 [Gene, 24, 255 (1983)], Bacillus subtilis 207-21 [Journal of Biochemistry, 95, 87 (1984)], etc. can be used.

[0098] As yeast, for example, Saccharomyces cerevisiae AH22, AH22R-, NA87-11A, DKD-5D, 20B-12, Schizosaccharomyces pombe NCYC1913, NCYC2036, Pichia pastoris KM71, etc. can be used.

[0099] As insect cells, for example, in the case where the virus is AcNPV, cell lines derived from larvae of Spodoptera frugiperda (Spodoptera frugiperda cell; Sf cell), midgut MG1 cells from Trichoplusia ni, High FiveTM cells from eggs of Trichoplusia ni, cells from Mamestra brassicae, cells from Estigmena acrea, etc. can be used. In the case where the virus is BmNPV, as insect cells, cell lines derived from silkworms (Bombyx mori N cell; BmN cell), etc. can be used. As the Sf cell, for example, Sf9 cell (ATCC CRL1711), Sf21 cell [above, In Vivo, 13, 213-217 (1977)], etc. can be used.

[0100] As insects, for example, larvae of silkworms, Drosophila, crickets, etc. [Nature, 315, 592 (1985)] can be used.

[0101] As animal cells, for example, cell lines such as monkey COS-7 cells, monkey Vero cells, Chinese hamster ovary (CHO) cells, dhfr gene-deficient CHO cells, mouse L cells, mouse AtT-20 cells, mouse myeloma cells, rat GH3 cells, human FL cells, etc., pluripotent stem cells such as iPS cells and ES cells of humans and other mammals, and primary cultured cells prepared from various tissues can be used. Further, zebrafish embryos, Xenopus laevis oocytes, etc. can also be used.

[0102] As plant cells, suspension-cultured cells, calli, protoplasts, leaf slices, root slices, etc. prepared from various plants (e.g., cereals such as rice, wheat, corn, etc., commercial crops such as tomato, cucumber, eggplant, etc., horticultural plants such as carnation, lisianthus, etc., experimental plants such as tobacco, Arabidopsis, etc.) can be used.

[0103] Any of the above host cells can be haploid (monoploid), polyploid (e.g., diploid, triploid, tetraploid, etc.). In conventional mutagenesis methods, as a principle, mutations are introduced only into one of the homologous chromosomes to form a heterozygous genotype. Therefore, if it is not a dominant mutation, the desired trait will not be expressed, and it is time-consuming and laborious to make it homozygous, which is inconvenient in most cases. In contrast, according to the present invention, since mutations can be introduced into all alleles on homologous chromosomes in the genome, even if it is a recessive mutation, the desired trait can be expressed in this representative ( Figure 10 ), and it is extremely useful in terms of overcoming the crux of the conventional method.

[0104] For the introduction of the expression vector, it can be carried out according to known methods (e.g., lysozyme method, competent method, PEG method, CaCl2 co-precipitation method, electroporation method, microinjection method, particle gun method, lipid transfection method, Agrobacterium method, etc.) according to the type of host.

[0105] For Escherichia coli, transformation can be carried out according to the methods described, for example, in Proc. Natl. Acad. Sci. USA, 69, 2110 (1972), Gene, 17, 107 (1982), etc.

[0106] The vector can be introduced into the genus Bacillus according to the methods described, for example, in Molecular & General Genetics, 168, 111 (1979), etc.

[0107] The vector can be introduced into yeast according to the methods described, for example, in Methods in Enzymology, 194, 182 - 187 (1991), Proc. Natl. Acad. Sci. USA, 75, 1929 (1978), etc.

[0108] The vector can be introduced into insect cells and insects according to the methods described, for example, in Bio / Technology, 6, 47 - 55 (1988), etc.

[0109] The vector can be introduced into animal cells according to the methods described, for example, in Cell Engineering Supplement 8 New Cell Engineering Experimental Protocols, 263 - 267 (1995) (published by Xiurun Company), Virology, 52, 456 (1973).

[0110] Regarding the cultivation of cells into which a vector has been introduced, it can be carried out according to known methods depending on the type of host.

[0111] For example, in the case of cultivating Escherichia coli or Bacillus, a liquid medium is preferably used as the medium for cultivation. In addition, the medium preferably contains a carbon source, a nitrogen source, inorganic substances, etc. necessary for the growth of the transformant. Here, examples of the carbon source include glucose, dextrin, soluble starch, sucrose, etc.; examples of the nitrogen source include ammonium salts, nitrates, corn steep liquor, peptone, casein, meat extract, soybean cake, potato extract, etc., which are inorganic or organic substances; examples of the inorganic substances include calcium chloride, sodium dihydrogen phosphate, magnesium chloride, etc. In addition, yeast extract, vitamins, growth promoting factors, etc. can also be added to the medium. The pH of the medium is preferably about 5 to about 8.

[0112] As the medium for cultivating Escherichia coli, for example, the M9 medium [Journal of Experiments in Molecular Genetics, 431 - 433, Cold Spring Harbor Laboratory, New York 1972] containing glucose and casein amino acids is preferred. If necessary, in order to make the promoter work effectively, a reagent such as 3β - indolylacrylic acid can also be added to the medium, for example. The cultivation of Escherichia coli is usually carried out at about 15 to about 43 °C. Aeration and stirring can be carried out as needed.

[0113] The cultivation of Bacillus is usually carried out at about 30 to about 40 °C. Aeration and stirring can be carried out as needed.

[0114] As the medium for cultivating yeast, for example, Burkholder minimal medium [Proc. Natl. Acad. Sci. USA, 77, 4505 (1980)], 0.5% SD medium containing casein amino acids [Proc. Natl. Acad. Sci. USA, 81, 5330 (1984)], etc. can be cited. The pH of the medium is preferably about 5 to about 8. The cultivation is usually carried out at about 20 °C to about 35 °C. Aeration and stirring can be carried out as needed.

[0115] As the medium for cultivating insect cells or insects, for example, a medium obtained by appropriately adding additives such as inactivated 10% bovine serum to Grace's Insect Medium [Nature, 195, 788 (1962)] can be used. The pH of the medium is preferably about 6.2 to about 6.4. The cultivation is usually carried out at about 27 °C. Aeration and stirring can be carried out as needed.

[0116] As a culture medium for culturing animal cells, for example, Minimal Essential Medium (MEM) [Science, 122, 501 (1952)] containing about 5 to about 20% fetal bovine serum, Dulbecco's Modified Eagle Medium (DMEM) [Virology, 8, 396 (1959)], RPMI 1640 medium [The Journal of the American Medical Association, 199, 519 (1967)], Medium 199 [Proceeding of the Society for the Biological Medicine, 73, 1 (1950)], etc. can be used. The pH of the culture medium is preferably about 6 to about 8. The culture is usually carried out at about 30°C to about 40°C. Aeration and stirring can be carried out as needed.

[0117] As a culture medium for culturing plant cells, MS medium, LS medium, B5 medium, etc. can be used. The pH of the culture medium is preferably about 5 to about 8. The culture is usually carried out at about 20°C to about 30°C. Aeration and stirring can be carried out as needed.

[0118] As described above, a complex of a nucleic acid sequence recognition module and a nucleic acid base conversion enzyme, that is, a nucleic acid modification enzyme complex, can be expressed in cells.

[0119] The introduction of the RNA encoding the nucleic acid sequence recognition module and / or the nucleic acid base conversion enzyme into the host cell can be carried out by microinjection, lipid transfection, etc. The introduction of the RNA can be carried out once or multiple times at appropriate intervals (for example: 2 to 5 times).

[0120] When expressing a complex of a nucleic acid sequence recognition module and a nucleic acid base conversion enzyme from an expression vector or an RNA molecule introduced into a cell, the nucleic acid sequence recognition module specifically recognizes and binds to a target nucleotide sequence in the target double-stranded DNA (for example, genomic DNA). Using the action of the nucleic acid base conversion enzyme linked to the nucleic acid sequence recognition module, a base conversion occurs in the sense strand or the antisense strand of the target site (which can be appropriately adjusted within a range of several hundred bases including all or part of the target nucleotide sequence or in the vicinity thereof), generating a mismatch in the double-stranded DNA (for example: when using a cytidine deaminase such as PmCDA1 or AID as the nucleic acid base conversion enzyme, cytosine on the sense strand or the antisense strand of the target site is converted to uracil, generating a U:G or G:U mismatch). By introducing various mutations as follows: the mismatch is not correctly repaired, or the repair causes the base of the opposite strand to pair with the base of the converted strand (in the above example, T-A or A-T), and further substitution with other nucleotides during repair (for example: U→A, G), or a deletion or insertion of 1 to dozens of bases occurs.

[0121] For zinc finger motifs, since the production efficiency of zinc fingers that specifically bind to the target nucleotide sequence is not high, and in addition, it is not easy to screen zinc fingers with high binding specificity, it is not easy to produce multiple zinc finger motifs that actually function. For TAL effectors and PPR motifs, they have a higher degree of freedom in recognizing target nucleic acid sequences than zinc finger motifs, but they require designing and constructing huge proteins each time according to the target nucleotide sequence, so there are problems in terms of efficiency.

[0122] In contrast, since the CRISPR-Cas system recognizes the sequence of double-stranded DNA of interest through a guide RNA complementary to the target nucleotide sequence, it can target any sequence by simply synthesizing oligonucleotides that can form a specific hybrid with the target nucleotide sequence.

[0123] Therefore, in a more preferred embodiment of the present invention, a CRISPR-Cas system (CRISPR-mutant Cas) in which at least one DNA cleavage ability of Cas has been inactivated is used as a nucleic acid sequence recognition module.

[0124] Figure 1 Schematic illustration of the method for modifying double-stranded DNA of the present invention showing the use of CRISPR-mutant Cas as a nucleic acid sequence recognition module.

[0125] The nucleic acid sequence recognition module of the present invention using CRISPR-mutant Cas is provided in the form of a complex of an RNA molecule and a mutant Cas protein, wherein the RNA molecule consists of a guide RNA complementary to the target nucleotide sequence and a tracrRNA necessary for the recruitment of the mutant Cas protein.

[0126] The Cas protein used in the present invention only needs to belong to the CRISPR system and is not particularly limited, and Cas9 is preferably used. Examples of Cas9 include, but are not limited to, Cas9 from Streptococcus pyogenes (SpCas9), Cas9 from Streptococcus thermophilus (StCas9), etc. SpCas9 is preferred. As the mutant Cas used in the present invention, a Cas having the ability to inactivate both strands of double-stranded DNA or a Cas having nickase activity with the ability to inactivate only one strand can be used. For example, in the case of SpCas9, a D10A mutant in which the 10th Asp residue is changed to an Ala residue and lacks the ability to cleave the opposite strand of the strand complementary to the guide RNA, or an H840A mutant in which the 840th His residue is changed to an Ala residue and lacks the ability to cleave the strand complementary to the guide RNA can be used, and a double mutant thereof can further be used, but other mutant Cas can also be used.

[0127] For the nucleic acid base conversion enzyme, it is provided in the form of a complex with mutant Cas by the same method as the connection method with the above zinc finger, etc. Alternatively, the nucleic acid base conversion enzyme and mutant Cas can also be linked using RNA scaffolds composed of RNA aptamers such as MS2F6 and PP7 and their binding proteins. The guide RNA forms a complementary strand with the target nucleotide sequence, and then tracrRNA recruits mutant Cas to recognize the DNA cleavage site recognition sequence PAM (protospacer adjacent motif) (when SpCas9 is used, PAM is 3 bases of NGG (N is any base), and theoretically, any site on the genome can be targeted), without cleaving one or both DNA strands, and using the action of the nucleic acid base conversion enzyme linked to the mutant Cas, a base conversion is generated at the targeted site (which can be appropriately adjusted within a range of several hundred bases including all or part of the target nucleotide sequence), resulting in a mismatch in the double-stranded DNA. By introducing various mutations as follows: the mismatch is not correctly repaired, the repair causes the base of the opposite strand to pair with the base of the converted strand, and is further converted to another nucleotide during repair, or a deletion or insertion of 1 to several tens of bases occurs (for example: refer to Figure 2 ).

[0128] When using CRISPR-mutant Cas as a nucleic acid sequence recognition module, similar to the case of using zinc fingers, etc. as nucleic acid sequence recognition modules, it is preferable to introduce the nucleic acid sequence recognition module and the nucleic acid base conversion enzyme into the cell having the double-stranded DNA of interest in the form of nucleic acids encoding them.

[0129] The DNA encoding Cas can be cloned from the cells producing the enzyme by the same method as the above method for the DNA encoding the nucleic acid base converting enzyme. In addition, mutant Cas can be obtained by the following method: in the cloned Cas-encoding DNA, a site-specific mutagenesis method known per se is used to introduce a mutation in such a way that the amino acid residues at the sites important for DNA cleavage activity (for example, in the case of Cas9, the 10th Asp residue and the 840th His residue can be mentioned, but not limited to these) are converted into other amino acids.

[0130] Alternatively, for the DNA encoding mutant Cas, for the DNA encoding the nucleic acid sequence recognition module and the DNA encoding the nucleic acid base converting enzyme, a DNA form having a codon usage suitable for expression in the host cell to be used can be constructed by using the same method as above in combination with chemical synthesis, PCR method or Gibson Assembly method. For example, the optimized CDS sequence and amino acid sequence for expressing SpCas9 in eukaryotic cells are shown in SEQ ID NOs: 5 and 6. If the base at position 29 of the sequence shown in SEQ ID NO: 5 is changed from "A" to "C", DNA encoding the D10A mutant can be obtained, and if the bases at positions 2518-2519 are changed from "CA" to "GC", DNA encoding the H840A mutant can be obtained.

[0131] For the DNA encoding mutant Cas and the DNA encoding the nucleic acid base converting enzyme, they can be ligated so as to be expressed as a fusion protein, or can be designed so as to be expressed separately using a binding domain, an intein, etc., and a complex is formed in the host cell by protein-protein interaction and protein ligation.

[0132] For the obtained DNA encoding mutant Cas and / or nucleic acid base converting enzyme, depending on the host, it can be inserted downstream of the promoter of the same expression vector as above.

[0133] On the other hand, the DNA encoding the guide RNA and tracrRNA can be designed as an oligonucleotide sequence in which a guide RNA sequence complementary to the target nucleotide sequence is ligated to a known tracrRNA sequence (for example, gttttagagctagaaatagcaagttaaaataaggctagtccgttatcaacttgaaaaagtggcaccgagtcggtggtg ctttt; SEQ ID NO: 7), and chemically synthesized using a DNA / RNA synthesizer.

[0134] Regarding the length of the guide RNA sequence, there is no particular limitation as long as it can specifically bind to the target nucleotide sequence, for example, it is 15 to 30 nucleotides, preferably 18 to 24 nucleotides.

[0135] The DNA encoding the guide RNA and tracrRNA can also be inserted into the same expression vector as described above according to the host. As a promoter, a class III pol promoter (for example, SNR6, SNR52, SCR1, RPR1, U6, H1 promoter, etc.) and a terminator (for example, T6 sequence) are preferably used.

[0136] Regarding the RNA encoding the mutant Cas and / or the nucleic acid base converter, it can be prepared, for example, by transcribing into mRNA using an in vitro transcription system known per se with a vector encoding the DNA encoding the mutant Cas and / or the nucleic acid base converter as a template.

[0137] The guide RNA-tracrRNA can be designed as an oligoribonucleotide sequence formed by linking a sequence complementary to the target nucleotide sequence and a known tracrRNA sequence, and chemically synthesized using a DNA / RNA synthesizer.

[0138] Regarding the DNA or RNA encoding the mutant Cas and / or the nucleic acid base converter, the guide RNA-tracrRNA or the DNA encoding it, it can be introduced into the host cell by the same method as described above according to the host.

[0139] In conventional artificial nucleases, since DNA double-strand cleavage (DSB) is accompanied, when targeting a sequence in the genome, proliferation disorders and cell death caused by random cleavage (off-target cleavage) of chromosomes occur. This effect is particularly fatal to many microorganisms and prokaryotes, hindering their applicability. In the present invention, since mutagenesis is achieved by a reaction for converting substituents on DNA bases (especially deamination reaction) without cleaving DNA, a significant reduction in toxicity can be achieved. In fact, as shown in the following examples, in a comparative experiment using budding yeast as a host, it was confirmed that in the case of using Cas9 having conventional DSB activity, the number of viable cells decreased due to expression induction. In contrast, in the technique of the present invention combining the mutant Cas and the nucleic acid base converter, the cells continued to proliferate and the number of viable cells increased ( Figure 3 ).

[0140] It should be noted that the modification of the double-stranded DNA of the present invention does not prevent the cleavage of the double-stranded DNA from occurring outside the target site (which can be appropriately adjusted within the range of several hundred bases including all or part of the target nucleotide sequence). However, one of the greatest advantages of the present invention is the avoidance of toxicity caused by off-target cleavage. Considering that it can be applied to any biological species in principle, in a preferred embodiment, the modification of the double-stranded DNA of the present invention is not only not associated with the cleavage of the DNA strand at the target site of the selected double-stranded DNA, but also at other sites.

[0141] In addition, as shown in the examples below, compared with the case of using a mutant Cas that cannot cleave both strands, when a Cas having nickase activity that can cleave only one of the double-stranded DNAs is used as the mutant Cas ( Figure 5 ), the mutation introduction efficiency is increased. Therefore, for example, by further linking a protein having nickase activity in addition to the nucleic acid sequence recognition module and the nucleic acid base conversion enzyme, only one strand of DNA can be cleaved near the target nucleotide sequence, while avoiding the strong toxicity based on DSB and increasing the mutation introduction efficiency.

[0142] Furthermore, comparing the effects of mutant Cas having two nickase activities that cleave different strands, the results show that using one mutant Cas causes the mutation sites to accumulate near the center of the target nucleotide sequence, and using the other mutant Cas causes various mutations to be randomly introduced into the region reaching several hundred bases from the target nucleotide sequence ( Figure 6 ). Therefore, by selecting the strand to be cleaved by the nickase, point mutations can be introduced into specific nucleotides or nucleotide regions, or various mutations can be randomly introduced into a relatively wide range, and it can be appropriately used according to different purposes. For example, if the former technique is used for gene disease iPS cells, after repairing the mutations of the pathogenic genes in the iPS cells made from the patient's own cells, by differentiating them into the somatic cells of interest, a cell transplantation therapy agent with a lower rejection risk can be produced.

[0143] In Example 7 and subsequent examples below, it is shown that mutations can be introduced almost at the pinpoint in specific nucleotides. As such, in order to introduce mutations pinpointedly into a desired nucleotide, it is only necessary to set the target nucleotide sequence so that there is a regularity in the positional relationship between the nucleotide into which the mutation is desired to be introduced and the target nucleotide sequence. Using the CRIPR-Cas system as a nucleic acid sequence recognition module and AID as a nucleic acid base-converting enzyme, the target nucleotide sequence can be designed so that the C (or G on the opposite strand) into which the mutation is desired to be introduced is at a position 2 to 5 nucleotides from the 5'-end of the target nucleotide sequence. As described above, the length of the guide RNA sequence can be appropriately set between 15 and 30 nucleotides, preferably between 18 and 24 nucleotides. Since the guide RNA sequence is a sequence complementary to the target nucleotide sequence, changing the length of the guide RNA sequence causes a change in the length of the target nucleotide sequence as well, but regardless of the nucleotide length, the regularity of the possible introduction of mutations into C or G existing at a position 2 to 5 nucleotides from the 5'-end is maintained( Figure 12 ). Therefore, by appropriately selecting the length of the target nucleotide sequence (guide RNA as its complementary strand), the site where the mutant base can be introduced can be moved. Thus, the restriction caused by the DNA cleavage site recognition sequence PAM (NGG) can also be lifted, further increasing the freedom of mutant introduction.

[0144] As shown in the examples below, as compared with when a single nucleotide sequence is used as the target, by preparing a sequence recognition module with respect to multiple adjacent target nucleotide sequences and using them simultaneously, the mutant introduction efficiency is significantly increased( Figure 7 ). The effect is that mutation induction is equally achieved from the case where a part of two target nucleotide sequences is repeated to the case where they are separated by about 600 bp. In addition, it occurs in both cases where the target nucleotide sequences are in the same direction (the target nucleotide sequences are on the same strand)( Figure 7 ), and in the opposite direction (the target nucleotide sequences are on respective strands of double-stranded DNA)( Figure 4 ).

[0145] As shown in the examples below, with respect to the method for modifying the genomic sequence of the present invention, if an appropriate target nucleotide sequence is selected, mutations can be introduced into almost all cells expressing the nucleic acid modifying enzyme complex of the present invention( Figure 8 ). Therefore, the insertion and selection of a selectable marker gene, which are necessary in conventional genome editing, are not required. This significantly simplifies gene manipulation, and since there is no organism with recombinant foreign DNA, the applicability to crop breeding and the like is significantly broadened.

[0146] In addition, since the mutagenesis introduction efficiency is extremely high and selection using a marker is not required, in the method for modifying a genomic sequence of the present invention, multiple DNA regions at completely different positions can be targeted for modification( Figure 9 ). Therefore, in a preferred embodiment of the present invention, two or more nucleic acid sequence recognition modules that specifically bind to different target nucleotide sequences (which may be within one gene of interest or within two or more different genes of interest. These target genes may be located on the same chromosome or on different chromosomes) can be used. In this case, one of each of these nucleic acid sequence recognition modules is formed into a nucleic acid modification enzyme complex with a nucleic acid base conversion enzyme. Here, a common enzyme can be used as the nucleic acid base editing enzyme. For example, when using the CRISPR-Cas system as the nucleic acid sequence recognition module, a common substance can be used for the complex of the Cas protein and the nucleic acid base conversion enzyme (including the fusion protein), and two or more chimeric RNAs are produced and used as guide RNA-tracrRNA, where the chimeric RNAs are each formed by each of two or more guide RNAs that form complementary strands with different target nucleotide sequences respectively and tracrRNA. On the other hand, when using a zinc finger motif, a TAL effector, etc. as the nucleic acid sequence recognition module, for example, each nucleic acid sequence recognition module that specifically binds to a different target nucleotide can be fused with the nucleic acid base conversion enzyme.

[0147] In order to express the nucleic acid modification enzyme complex of the present invention in a host cell, as described above, an expression vector containing DNA encoding the nucleic acid modification enzyme complex and RNA encoding the nucleic acid modification enzyme complex are introduced into the host cell. However, in order to effectively introduce mutations, it is preferable to maintain the expression of the nucleic acid modification enzyme complex at a prescribed level or above for a prescribed period or longer. From this perspective, an expression vector (such as a plasmid) capable of autonomous replication is introduced into the host cell. However, since such a plasmid or the like is foreign DNA, it is preferably removed promptly after the introduction of mutations is successfully achieved. Therefore, although it varies depending on the type of host cell, etc., it is preferable, for example: from 6 hours to 2 days after the introduction of the expression vector, various plasmid removal methods well-known in the art are used to remove the introduced plasmid from the host cell.

[0148] Alternatively, as long as the expression of the nucleic acid modification enzyme complex sufficient to introduce mutations can be achieved, it is also preferable to use an expression vector (for example: a vector lacking an origin of replication that functions in the host cell and / or a gene encoding a protein essential for replication, etc.) or RNA that does not have the ability of autonomous replication in the host cell, and introduce mutations into the double-stranded DNA of interest by transient expression.

[0149] Since the expression of the target gene is suppressed during the period when the nucleic acid-modifying enzyme complex of the present invention is expressed in a host cell and a nucleic acid base conversion reaction is carried out, it has been difficult to directly edit a gene essential for the survival of the host cell as a target gene (side effects such as growth impairment of the host, unstable mutation introduction efficiency, and mutations at sites different from the target) to date. The present invention transiently expresses the nucleic acid-modifying enzyme complex of the present invention in a host cell at a period when a nucleic acid base conversion reaction occurs and at a period required for stabilizing the modification of the target site, thereby successfully and effectively achieving direct editing of an essential gene. The period required for the nucleic acid base conversion reaction to occur and for stabilizing the modification of the target site varies depending on the type of host cell, culture conditions, etc., but is generally considered to be 2 to 20 generations. For example, in the case where the host cell is yeast or bacteria (e.g., Escherichia coli), it is necessary to induce the expression of the nucleic acid-modifying enzyme complex between 5 and 10 generations. Those skilled in the art can appropriately determine an appropriate expression induction period based on the doubling time of the host cell under the culture conditions used. For example, when budding yeast is liquid-cultured in a 0.02% galactose induction medium, the expression induction period can be, for example, 20 to 40 hours. Regarding the expression induction period of the nucleic acid encoding the nucleic acid-modifying enzyme complex of the present invention, it can also be extended beyond the above-mentioned "period required for stabilizing the modification of the target site" within a range that does not cause side effects to the host cell.

[0150] As a method for transiently expressing the nucleic acid-modifying enzyme complex of the present invention at a desired period, a method including preparing a construct (expression vector) that contains a nucleic acid encoding the nucleic acid-modifying enzyme complex (in the CRISPR-Cas system, DNA encoding a guide RNA-tracrRNA and DNA encoding a mutant Cas and a nucleic acid base substitution enzyme) in a form capable of controlling the expression period and introducing the construct into a host cell can be used. The "form capable of controlling the expression period" specifically includes a form in which the nucleic acid encoding the nucleic acid-modifying enzyme complex of the present invention is placed under the control of an inducible regulatory region. There is no particular limitation on the "inducible regulatory region". For example, in microbial cells such as bacteria (e.g., Escherichia coli) and yeast, an operon of a temperature-sensitive (ts) mutant repressor and an operator controlling the same can be exemplified. As the ts mutant repressor, for example, a ts mutant of the cI repressor derived from bacteriophage λ can be exemplified, but it is not limited to these. In the case of the λ bacteriophage cI repressor (ts), it is linked to the operator at 30°C or lower (e.g., 28°C) to inhibit the gene expression downstream, but dissociates from the operator at a high temperature of 37°C or higher (e.g., 42°C), and thus gene expression can be induced ( Figure 13 and 14)。Therefore, by culturing a host cell into which a nucleic acid encoding a nucleic acid-modifying enzyme complex has been introduced, generally at a temperature below 30°C, raising the temperature to 37°C or higher at an appropriate time and culturing for a certain period to cause a nucleic acid base conversion reaction, and quickly lowering the temperature back below 30°C after introducing a mutation into the target gene, it is possible to minimize the period of suppressing the expression of the target gene. Even when targeting a gene essential for the host cell, side effects can be suppressed and editing can be effectively carried out ( Figure 15 )。

[0151] In the case of using a temperature-sensitive mutation, for example, a temperature-sensitive mutant of a protein essential for the autonomous replication of a vector can be included in a DNA vector containing the DNA encoding the nucleic acid-modifying enzyme complex of the present invention. After the expression of the nucleic acid-modifying enzyme complex, the vector cannot replicate autonomously quickly and is naturally shed during cell division. Examples of such temperature-sensitive mutant proteins include, but are not limited to, the temperature-sensitive mutant of Rep101 ori, which is essential for the replication of pSC101 ori. For Rep101 ori(ts), although it can act on pSC101 ori and carry out autonomous replication of the plasmid at a temperature below 30°C (e.g., 28°C), it loses its function at a temperature above 37°C (e.g., 42°C), and the plasmid cannot replicate autonomously. Therefore, by using it in combination with the above-mentioned cI repressor (ts) of λ phage, it is possible to simultaneously achieve transient expression of the nucleic acid-modifying enzyme complex of the present invention and plasmid removal.

[0152] On the other hand, when using higher eukaryotic cells such as animal cells, insect cells, and plant cells as host cells, the DNA encoding the nucleic acid-modifying enzyme complex of the present invention is introduced into the host cell under the control of an inducible promoter (e.g., a metallothionein promoter (induced by heavy metal ions), a heat shock protein promoter (induced by heat shock), a Tet-ON / Tet-OFF type promoter (induced by adding or removing tetracycline or its derivatives), a steroid-responsive promoter (induced by a steroid hormone or its derivatives), etc.). At an appropriate time, an inducer is added to the medium (or removed from the medium) to induce the expression of the nucleic acid-modifying enzyme complex, cultured for a certain period, and a nucleic acid base conversion reaction is carried out, so as to achieve transient expression of the nucleic acid-modifying enzyme complex after introducing a mutation into the target gene.

[0153] It should be noted that inducible promoters can also be used in prokaryotic cells such as Escherichia coli. Examples of such inducible promoters include, but are not limited to, the lac promoter (induced by IPTG), the cspA promoter (induced by cold shock), the araBAD promoter (induced by arabinose), etc.

[0154] Alternatively, the above-inducible promoter can also be used as a method for removing vectors when using higher eukaryotic cells such as animal cells, insect cells, and plant cells as host cells. That is, by loading it with an origin of replication that functions in the host cell and a nucleic acid encoding a protein essential for its replication (for example, if it is an animal cell, SV40 ori and large T antigen, oriP and EBNA-1, etc.), the expression of the nucleic acid encoding the protein is controlled by the above-inducible promoter. Although the vector can replicate autonomously in the presence of the inducing substance, it cannot replicate autonomously if the inducing substance is removed. During cell division, the vector will naturally fall off (in the case of Tet-OFF type vectors, conversely, the addition of tetracycline or doxycycline results in the inability to replicate autonomously).

[0155] Hereinafter, the present invention will be described based on examples. However, the present invention is not limited to these examples.

[0156] Examples

[0157] In Examples 1 to 6 below, experiments were conducted as described below.

[0158] <Cell line, culture, transformation, expression induction>

[0159] The budding yeast Saccharomyces cerevisiae BY4741 strain (requiring leucine and uracil) was used and cultured using a standard YPDA medium or an SD medium composed of dropout components that meet the nutritional requirements. The culture was carried out between 25°C and 30°C, either by static culture using an agar plate or by shaking culture using a liquid medium. For transformation, the lithium acetate method was used, and selection was carried out using an SD medium that meets the appropriate nutritional requirements. For galactose-based expression induction, it was pre-cultured overnight using an appropriate SD medium, then inoculated into an SR medium in which the carbon source was replaced from 2% glucose to 2% raffinose and cultured overnight, and further inoculated into an SGal medium in which the carbon source was replaced with 0.2 - 2% galactose and cultured for about 3 hours to two nights for expression induction.

[0160] For the determination of the number of viable cells and the Can1 mutation rate, the cell suspension was appropriately diluted and spread on SD plate medium and SD-Arg + 60 mg / l canavanine plate medium or SD + 300 mg / l canavanine plate medium. The number of colonies that appeared after 3 days was counted as the number of viable cells. The number of surviving colonies in the SD plate was used as the total number of cells, and the number of surviving colonies in the canavanine plate was used as the number of resistant mutant strains, and the mutation rate was calculated and evaluated. For the mutation introduction site, the DNA fragment containing the target gene region of each strain was amplified by colony PCR method, and then DNA sequencing was performed. Alignment analysis was carried out based on the sequence of the yeast genome database (http: / / www.yeastgenome.org / ), and identification was performed.

[0161] <Nucleic acid manipulation>

[0162] For DNA, it was processed and constructed by any one of PCR method, restriction enzyme treatment, ligation, Gibson Assembly method / artificial chemical synthesis. For plasmids, as yeast-E. coli shuttle vectors, pRS315 for leucine selection and pRS426 for uracil selection were used as the backbones. For plasmids, they were amplified using E. coli strain XL-10gold or DH5α and introduced into yeast by the lithium acetate method.

[0163] <Construct>

[0164] For inducible expression, the Saccharomyces cerevisiae pGal1 / 10 (SEQ ID NO: 8), which is galactose-inducible and a bidirectional promoter, was used. A nuclear localization signal (ccc aag aag aag agg aag gtg; SEQ ID NO: 9 (PKKKRV; coding SEQ ID NO: 10)) was added to the ORF of the Cas9 gene (SEQ ID NO: 5) from Streptococcus pyogenes that had been optimized for eukaryotic expression codons. The resulting product was ligated downstream of the promoter and linked to the ORF (SEQ ID NO: 1 or 3) of a deaminase gene (PmCDA1 from Petromyzon marinus or hAID from humans) via a linker sequence for expression in the form of a fusion protein. As the linker sequence, the GS linker (ggt gga gga ggt tct; SEQ ID NO: 11 (encoding GGGGS; SEQ ID NO: 12) repeats), Flag tag (gac tat aag gac cac gac gga gac tac aag gat cat gat att gat tac aaa gac gat gac gat aag; SEQ ID NO: 13 (encoding DYKDHDGDYKDHDIDYKDDDDK; SEQ ID NO: 14)), Strep-tag (tgg agc cac ccg cag ttc gaa aaa; SEQ ID NO: 15 (encoding WSHPQFEK; SEQ ID NO: 16)), and other domains were selected and combined for use. Here, in particular, a combination of 2xGS, SH3 domain (SEQ ID NOs: 17 and 18), and Flag tag was used. As the terminator, it was linked to the ADH1 terminator (SEQ ID NO: 19) and Top2 terminator (SEQ ID NO: 20) from Saccharomyces cerevisiae. Additionally, in the domain binding mode, the Cas9 gene ORF was linked to the SH3 domain via a 2xGS linker as one protein, and the SH3 ligand sequences (SEQ ID NOs: 21 and 22) were added to the deaminase as another protein and connected to both directions of the Gal1 / 10 promoter. And they were expressed simultaneously. They were assembled into the pRS315 plasmid.

[0165] To remove the cleavage ability of each single-stranded DNA, mutations that convert aspartic acid at position 10 to alanine (D10A, corresponding DNA sequence mutation a29c), and histidine at position 840 to alanine (H840A, corresponding DNA sequence mutation ca2518gc) were introduced into Cas9.

[0166] For the gRNA, it was configured in the form of a chimeric structure with tracrRNA (from Streptococcus pyogenes; SEQ ID NO: 7) between the SNR52 promoter (SEQ ID NO: 23) and the Sup4 terminator (SEQ ID NO: 24), and assembled into the pRS426 plasmid. As the gRNA target base sequence, the 187 - 206 of the CAN1 gene ORF (gatacgttctctatggagga; SEQ ID NO: 25) (target1), 786 - 805 (ttggagaaacccaggtgcct; SEQ ID NO: 26) (target3), 793 - 812 (aacccaggtgcctggggtcc; SEQ ID NO: 27) (target4), 563 - 582 (ttggccaagtcattcaattt; SEQ ID NO: 28) (target2), and the complementary strand sequence of 767 - 786 (ataacggaatccaactgggc; SEQ ID NO: 29) (target5r) were used. In the case of co-expressing multiple targets, the sequence from the promoter to the terminator was used as one group, and multiple groups were assembled into the same plasmid. It was introduced into cells together with the plasmid for Cas9-deaminase expression and expressed intracellularly to form a complex of gRNA-tracrRNA and Cas9-deaminase.

[0167] Example 1 Modifying gene group sequences

[0168] To test the effect of the genomic sequence modification technique of the present invention that utilizes the nucleic acid sequence recognition ability of deaminase and CRISPR-Cas, an attempt was made to introduce mutations into the CAN1 gene encoding the canavanine transporter, and the gene defect results in canavanine resistance. The sequence complementary to the 187 - 206 (target1) of the CAN1 gene ORF was used as the gRNA, an expression vector of a chimeric RNA formed by ligating this gRNA with tracrRNA from Streptococcus pyogenes was constructed, and a vector expressing a protein formed by fusing dCas9, which was mutated (D10A and H840A) in Cas9 from Streptococcus pyogenes to lose nuclease activity, with PmCDA1 from lamprey as the deaminase was constructed, and they were introduced into budding yeast by the lithium acetate method and co-expressed. The results are shown in by linking the DNA sequence recognition ability of CRISPR-Cas with the deaminase PmCDA1When cultured on SD plates containing canavanine, only the cells that had introduced and expressed gRNA-tracrRNA and dCas9-PmCDA1 formed colonies. Resistant colonies were selected and the sequences of the CAN1 gene region were sequenced, and as a result, it was confirmed that mutations were introduced within and near the target nucleotide sequence (target1).

[0169] Example 2 Figure 2

[0170] In conventional Cas9 and other artificial nucleases (ZFN, TALEN), when targeting sequences within the genome, proliferation disorders and cell death caused by random cleavage of chromosomes are generated. This effect is particularly lethal in many microorganisms and prokaryotes, hindering applicability.

[0171] Therefore, in order to verify the safety and cytotoxicity of the genomic sequence modification technology of the present invention, a comparative experiment of conventional CRISPR-Cas9 was conducted. By using the sequences within the CAN1 gene (target 3, 4) as gRNA targets, the colony-forming ability of viable cells on SD plates from immediately after the start of galactose expression induction to 6 hours after induction was measured. The results are shown in Greatly reducing side effects and toxicity 。In conventional Cas9, cell death is induced due to hindered proliferation, so the number of surviving cells decreases. In contrast, in this technology (nCas9 D10A-PmCDA1), cells can proliferate continuously and the number of surviving cells increases significantly.

[0172] Example 3 Figure 3

[0173] It was examined whether mutations could also be introduced into the target gene if Cas9 and the deaminase were not in the form of a fusion protein but formed a nucleic acid-modifying enzyme complex through a binding domain and its ligand. dCas9 used in Example 1 was used as Cas9, and human AID was used as the deaminase instead of PmCDA1. The former was fused with the SH3 domain and the latter with its binding ligand to produce Use of different ligation types the various constructs shown. In addition, the sequences within the CAN1 gene (target 4, 5r) were used as gRNA targets. Their constructs were introduced into budding yeast. As a result, even when dCas9 and the deaminase were linked through a binding domain, mutations were effectively introduced into the target sites of the CAN1 gene ( Figure 4 ). By introducing multiple binding domains into dCas9, the mutation introduction efficiency was significantly improved. The main mutation introduction site was the 782nd position (g782c) of the ORF.

[0174] Example 4 Figure 4

[0175] Instead of dCas9, the D10A mutant nCas9 (D10A) that only cleaves the strand complementary to the gRNA or the H840A mutant nCas9 (H840A) that only cleaves the opposite strand of the strand complementary to the gRNA was used. Otherwise, the procedure was the same as in Example 1. Mutations were introduced into the CAN1 gene, and the sequence of the CAN1 gene region of the colonies generated on the SD plate containing canavanine was examined. As a result, in the former (nCas9 (D10A)), the efficiency was higher than that of dCas9 ( Enhancement of efficiency and change in mutation pattern caused by using Nickase ), and the mutations were concentrated in the center of the target sequence ( Figure 5 ). Therefore, site-directed mutagenesis can be performed according to this method. On the other hand, it was found that in the latter (nCas9 (H840A)), the efficiency was higher than that of dCas9 ( Figure 6 ), and multiple random mutations were introduced into the region of several hundred bases from the targeted nucleotide ( Figure 5 ).

[0176] Even when the target nucleotide sequence was changed, similarly significant mutagenesis could be confirmed. In this genome editing system using the CRISPR-Cas9 system and cytidine deaminase, as shown in Table 1, cytosine present in the range of about 2 to 5 bp from the 5' side of the target nucleotide sequence (20 bp) was preferentially deaminated. Therefore, by setting the target nucleotide sequence based on this regularity and further combining it with nCas9 (D10A), precise genome editing at the single nucleotide unit level can be performed. On the other hand, if nCas9 (H840A) is used, multiple mutations can be inserted simultaneously in the range of about several hundred bp near the target nucleotide sequence. Further, there is a possibility of further changing the site specificity by changing the connection type of the deaminase.

[0177] These results indicate that the type of Cas9 protein can be appropriately used according to the purpose.

[0178] [Table 1]

[0179]

[0180] Example 5 Figure 6

[0181] Compared with a single target, by simultaneously using multiple adjacent targets, the efficiency was significantly increased ( By targeting multiple adjacent DNA sequences, the efficiency is synergistically increased)。In fact, 10-20% of the cells have canavanine resistance mutations (target3, 4). gRNA1 and gRNA2 in the figure target target3 and target4 respectively. PmCDA1 is used as a deaminase. It was confirmed that this effect is produced not only when a part of the sequence is repeated (target3, 4), but also when 600 bp is isolated (target1, 3). In addition, this effect is produced in both cases where the DNA sequences are in the same direction (target3, 4) and in opposite directions (target4, 5)( Figure 7 ).

[0182] Example 6 Figure 4

[0183] For the cells (Target3, 4) targeting target3 and target4 in Example 5, 10 colonies randomly selected from the colonies grown on a non-selective (canavanine-free) plate (SD plate without Leu and Ura) were sequenced for the sequence of the CAN1 gene region. As a result, mutations were introduced at the targeting sites of the CAN1 gene in all the colonies examined( Gene modification without the need for selection markers ). That is, according to the present invention, if an appropriate targeting sequence is selected, it can be expected that editing will occur in substantially all of the expressed cells. Therefore, the insertion and selection of marker genes necessary in conventional genetic manipulations are not required. This not only significantly simplifies genetic manipulation, but also greatly broadens the applicability to crop breeding and the like because it is not a foreign DNA recombinant organism.

[0184] In the following examples, the same experimental techniques as in Examples 1-6 were carried out in the same manner as above.

[0185] Example 7 Figure 8

[0186] In the usual genetic manipulation methods, due to various limitations, usually only one site can be mutated in one operation. Therefore, it was tested whether multiple site mutation operations can be carried out simultaneously using the method of the present invention.

[0187] The 3-22 positions of the ORF of the Ade1 gene of the budding yeast strain BY4741 were used as the first target nucleotide sequence (Ade1target5: GTCAATTACGAAGACTGAAC; SEQ ID NO: 30), and the 767-786 positions (complementary strand) of the ORF of the Can1 gene were used as the second target nucleotide sequence (Can1target8(786-767; ATAACGGAATCCAACTGGGC; SEQ ID NO: 29). DNA of two chimeric RNAs containing gRNAs encoding nucleotide sequences complementary to them and tracrRNA (SEQ ID NO: 7) were both carried on the same plasmid (pRS426), and together with the plasmid nCas9 D10A-PmCDA1 containing the nucleic acid encoding the fusion protein of mutant Cas9 and PmCDA1, were introduced into the BY4741 strain and expressed to verify the introduction of mutations into the two genes. Based on the SD dropout component medium (Dropout, uracil and leucine deficient; SD-UL) for maintaining the plasmid, the cells were cultured. The cells were appropriately diluted and spread on SD-UL and the medium supplemented with canavanine to form colonies. After culturing at 28 °C for 2 days, the colonies were observed, and the incidence of red colonies caused by ade1 mutations and the survival rate in the canavanine medium were counted separately. The results are shown in Table 2.

[0188] [Table 2]

[0189]

[0190] As a phenotype, the proportion of mutants introduced into both the Ade1 gene and the Can1 gene was high, at about 31%.

[0191] Next, the colonies on the SD-UL medium were amplified by PCR and subjected to sequencing analysis. The ORF regions containing Ade1 and Can1 were broadened respectively to obtain the sequencing information of sequences about 500b around the target sequences. Specifically, 5 red colonies and 5 white colonies were analyzed. As a result, the C at the 5th position of the ORF of the Ade1 gene in all red colonies was changed to G, and the C at the 5th position in all white colonies was changed to T ( Simultaneous editing of multiple sites (different genes) ). Although the mutation rate of the target was 100%, since the mutation rate aimed at gene disruption required the C at the 5th position to be changed to G to become a stop codon, the expected mutation rate was considered to be 50%. Similarly, for the Can1 gene, although it was confirmed that the G at the 782nd position of the ORF occurred mutations in all clones ( Figure 9 ), since only the mutation to C could provide canavanine resistance, the expected mutation rate was 70%. Within the scope of the examination, the proportion of clones that simultaneously obtained the expected mutations for the two genes was 40% (4 out of 10 clones), and a relatively high efficiency was actually obtained.

[0192] Example 8 Figure 9

[0193] Although many organisms have diploid or polyploid genomes, in conventional mutagenesis methods, as a rule, mutations are introduced only into one of the homologous chromosomes to form a heterozygous genotype. Therefore, if the mutation is not a dominant mutation, the desired trait cannot be obtained, and making it homozygous is time-consuming and laborious. Therefore, it was tested whether mutations could be introduced into all target alleles on homologous chromosomes within the genome according to the technology of the present invention.

[0194] That is, in the diploid budding yeast strain YPH501, simultaneous editing of the Ade1 and Can1 genes was performed. Since the phenotypes (red colonies and canavanine resistance) of these gene mutations are disadvantageous phenotypes, these phenotypes will not be manifested unless mutations in both genes (homologous mutations) are introduced.

[0195] The 1173-1154 positions (complementary strand) (Ade1 target1: GTCAATAGGATCCCCTTTT; SEQ ID NO: 31) or 3-22 positions (Ade1 target5: GTCAATTACGAAGACTGAAC; SEQ ID NO: 30) of the Ade1 gene ORF were used as the first target nucleotide sequence, and the 767-786 positions (complementary strand) of the Can1 gene ORF were used as the second target nucleotide sequence (Can1 target8: ATAACGGAATCCAACTGGGC; SEQ ID NO: 29). DNAs of two chimeric RNAs of gRNA and tracrRNA (SEQ ID NO: 7) respectively containing nucleotide sequences complementary to them were loaded onto the same plasmid (pRS426), and together with the plasmid nCas9 D10A-PmCDA1 containing the nucleic acid encoding the fusion protein of mutant Cas9 and PmCDA1, they were introduced into the BY4741 strain and expressed, and the introduction of mutations into each gene was verified.

[0196] The results of colony counting showed that the characteristics of each phenotype could be obtained with a high probability (40% - 70%) ( Editing of polyploid genomes A).

[0197] Furthermore, in order to confirm the mutations, sequencing of the Ade1 target region of white colonies and red colonies was performed. As a result, overlapping of sequencing signals showing heterozygous mutations was confirmed at the targeted sites in white colonies ( Figure 10 Upper figure of B, overlapping of signals of G and T at ↓). It was confirmed that no phenotype appeared in colonies with heterozygous mutations. On the other hand, in red colonies, no overlapping signals were confirmed, indicating homologous mutations ( Figure 10 Lower figure of B, signal of T at ↓).

[0198] Example 9 Figure 10

[0199] In this example, it was verified that the present technology functions effectively in Escherichia coli, a representative model organism of bacteria. In particular, since conventional nuclease-type genome editing technologies are lethal and difficult to apply in bacteria, the superiority of the present technology was emphasized. In addition, together with yeast, a model cell of eukaryotes, it was shown that the technology can be widely applied in both prokaryotic and eukaryotic species.

[0200] An amino acid mutation (dCas9) of D10A and H840A was introduced into the Streptococcus pyogenes Cas9 gene containing a bidirectional promoter region, and a construct was constructed to express it in the form of a fusion protein with PmCDA1 through a linker sequence. Furthermore, a plasmid containing a chimeric gRNA encoding a sequence complementary to each target nucleotide sequence was prepared (the full-length nucleotide sequence is shown in SEQ ID NO: 32. Sequences complementary to each target sequence were introduced into the n 20 portion of the sequence) ( Genome editing in Escherichia coli A).

[0201] First, a plasmid with the target nucleotide sequence introduced at positions 426 - 445 (T CAA TGG GCT AAC TAC GTT C; SEQ ID NO: 33) of the ORF of the Escherichia coli galK gene was transformed into various Escherichia coli strains (XL10-gold, DH5a, MG1655, BW25113) using the calcium method or electroporation method. After adding SOC medium and culturing overnight for recovery, cells carrying the plasmid were selected by LB medium containing ampicillin to form colonies. The introduction of mutations was verified by direct sequencing of colony PCR. The results are shown in Figure 11 B.

[0202] Randomly selected independent colonies (1 - 3) were subjected to sequencing analysis. As a result, it was confirmed that there was a probability of more than 60% that the C at position 427 of the ORF was changed to T (clones 2 and 3), resulting in the disruption of the gene that generates a stop codon (TAA).

[0203] Next, using the complementary sequence (5’-GGTCCATAAACTGAGACAGC-3’; SEQ ID NO: 34) of the 1530-1549 base region of the rpoB gene ORF, which is an essential gene, as a target, specific point mutations were introduced by the same method as above, and an attempt was made to confer rifampicin resistance function on Escherichia coli. Sequencing analysis of colonies selected using media containing non-selective medium (none), 25 μg / ml rifampicin (Rif25), and 50 μg / ml rifampicin (Rif50) confirmed that by changing G at position 1546 of the ORF to A, an amino acid mutation that changed Asp (GAC) to Asn (AAC) was introduced, conferring rifampicin resistance ( Figure 11 C, upper panel). A 10-fold dilution series of the transformed cell suspension was spotted onto media containing non-selective medium (none), 25 μg / ml rifampicin (Rif25), and 50 μg / ml rifampicin (Rif50) and cultured, and as a result, rifampicin-resistant strains were obtained at a frequency of approximately 10% ( Figure 11 C, lower panel).

[0204] In this way, according to the present technology, not only can genes be disrupted, but new functions can also be added by specific point mutations. In addition, the superiority of the present technology was also demonstrated in directly editing essential genes.

[0205] Example 10 Figure 11

[0206] Previously, the length of the gRNA for the target nucleotide sequence was based on 20b, and cytosine (or guanine on the opposite strand) present at the site from its 5’ end to 2-5b (15-19b upstream of the PAM sequence) was used as the mutation target. It was examined whether the site of the base corresponding to the nucleic acid whose gRNA length was changed was shifted if the nucleic acid was expressed ( Regulating the editing base site using the gRNA length A).

[0207] An experimental example using Escherichia coli is illustrated in Figure 12 B. Sites with multiple cytosine arrangements on the Escherichia coli genome were searched, and experiments were conducted using the gsiA gene, which is a putative ABC-transporter. The substituted cytosines were examined when the target length was changed to 24bp, 22bp, 20bp, and 18bp. As a result, in the case of 20bp (standard length), the cytosines at positions 898 and 899 were substituted with thymine. It was found that when the targeting site was longer than 20bp, the cytosines at positions 896 and 897 were also substituted, and when the targeting site was shorter, the cytosines at positions 900 and 901 could also be substituted. In fact, the targeting site could be shifted by changing the length of the gRNA.

[0208] Example 11 Figure 12

[0209] The nucleic acid modifying enzyme complex of the present invention was designed as a plasmid for inducing expression under high temperature conditions. The aim is to optimize the efficiency by controlling the expression state and to reduce side effects (host growth disorders, unstable mutagenesis efficiency, mutations at sites different from the target). At the same time, by combining the mechanism of stopping plasmid replication using high temperature, it is intended to easily remove the plasmid simultaneously after editing. The details of the experiment are shown below.

[0210] The temperature-sensitive plasmid pSC101-Rep101 system (the sequence of pSC101 ori is shown in SEQ ID NO: 35, and the sequence of temperature-sensitive Rep101 is shown in SEQ ID NO: 36) was used as the backbone, and expression was induced using the temperature-sensitive λ repressor (cI857) system. A RecA-resistant G113E mutation was introduced into the λ repressor for genome editing so that it functions normally under SOS response (SEQ ID NO: 37). dCas9-PmCDA1 (SEQ ID NO: 38) was ligated downstream of the Right Operator (SEQ ID NO: 39), and gRNA (SEQ ID NO: 40) was ligated downstream of the Left Operator (SEQ ID NO: 41) to control expression (all nucleotide sequences of the constructed expression vector are shown in SEQ ID NO: 42). During the cultivation period at 30 °C or lower, since the transcription of gRNA and the expression of dCas9-PmCDA1 were respectively inhibited, the cells could proliferate normally. When cultured at 37 °C or higher, the transcription of gRNA and the expression of dCas9-PmCDA1 were induced, and at the same time, plasmid replication was inhibited. Therefore, the nucleic acid modifying enzyme complex necessary for genome editing was transiently supplied, and the plasmid could be easily removed after editing ( Development of a temperature-dependent genome editing plasmid )

[0211] The specific protocol for base substitution is shown in Figure 13 .

[0212] The culture temperature during plasmid construction was set to about 28 °C, and first, Escherichia coli colonies retaining the required plasmid were established. Next, the colonies were used directly, or when the strain was changed, the plasmid was extracted, and transformation into the target strain was performed again, and the obtained colonies were used. Liquid culture was carried out overnight at 28 °C. Subsequently, it was diluted with the medium, and induction culture was carried out at 42 °C for about 1 hour to overnight. The cell suspension was appropriately diluted and spread or spotted on the plate to obtain single colonies.

[0213] As a verification experiment, point mutations were introduced into rpoB, which is an essential gene. If rpoB, which is one of the components of RNA polymerase, is deleted or disabled, Escherichia coli will not survive. On the other hand, it is known that resistance to rifampicin (Rif), an antibiotic, can be obtained by introducing point mutations into specific sites. Therefore, target sites were selected to introduce such point mutations and measurements were carried out.

[0214] The results are shown in Figure 14 . On the left of the upper left figure are LB (containing chloramphenicol) plates, and on the right are LB (containing chloramphenicol) plates supplemented with rifampicin. Using these plates, samples with or without chloramphenicol were prepared for culturing at 28 °C to 42 °C. Although the proportion of Rif resistance was low when cultured at 28 °C, rifampicin resistance was obtained with extremely high efficiency when cultured at 42 °C. Sequencing of 8 colonies of the actual colonies (non-selected) obtained on LB showed that in the strains cultured at 42 °C, more than 60% of the guanine (g) at position 1546 was replaced by adenine (A) (lower left and upper right figures). It can be seen that in the actual sequencing spectrum (lower right figure), the base was also completely replaced.

[0215] Similarly, base substitutions were carried out on galK, which is one of the factors involved in galactose metabolism. Since for Escherichia coli, galK is lethal by metabolizing 2-deoxygalactose (2DOG), which is an analogue of galactose, it was used as a selection method. Target sites were set respectively such that a missense mutation was caused in target 8 and target 12 became a stop codon ( Figure 15 lower right).

[0216] The results are shown in Figure 16 Figure 16 . In the upper left and lower left figures, on the left are LB (containing chloramphenicol) plates, and on the right are LB (containing chloramphenicol) plates supplemented with 2-DOG. Using these plates, samples with or without chloramphenicol were prepared for culturing at 28 °C and 42 °C. In target 8, only a few colonies were slightly produced on the plate supplemented with 2-DOG (upper left figure), but sequencing of 3 colonies on LB (red frame) determined that the cytosine (C) at position 61 in all colonies was replaced by thymine (T) (upper right). It is speculated that this mutation is not sufficient to cause the disability of galK. On the other hand, in target 12, colonies were obtained on the 2-DOG-added plates in any case of culturing at 28 °C and 42 °C (lower left figure). Sequencing of 3 colonies on LB determined that the cytosine at position 271 in all colonies was replaced by thymine (lower right). It shows that mutations can be introduced more stably and efficiently even in such different targets.

[0217] The entire contents of the patents and patent application specifications mentioned herein are hereby incorporated by reference into this specification to the same extent as if fully set forth herein.

[0218] This application is based on Japanese Patent Application No. 2014-43348 filed in Japan on March 5, 2014, and Japanese Patent Application No. 2014-201859 filed on September 30, 2014 of the same year, the entire contents of which are hereby incorporated by reference into this specification.

[0219] Industrial Applicability

[0220] According to the present invention, site-specific mutations can be safely introduced into any biological species without the insertion of foreign DNA and without the cleavage of DNA double strands. In addition, the range in which mutations can be introduced can be widely set from a pinpoint of one base to several hundred bases, and it can also be applied to the induction of local evolution by randomly introducing mutations into a specific limited region where it has been almost impossible until now, which is extremely useful. SEQUENCE LISTING <110> NATIONAL UNIVERSITY CORPORATION KOBE UNIVERSITY <120> Method of modifying genomic sequence comprising specifically converting nucleobase in targeted DNA sequence and molecular complex used therefor converting nucleobase in targeted DNA sequence and molecular complex used therefor <130> 092301 <150> JP 2014-043348 <151> 2014-03-05 <150> JP 2014-201859 <151> 2014-09-30 <160> 42 <170> PatentIn version 3.5 <210> 1 <211> 624 <212> DNA <213> Petromyzon marinus <220> <221> CDS <222> (1)..(624) <400> 1 atg acc gac gct gag tac gtg aga atc cat gag aag ttg gac atc tac 48 Met Thr Asp Ala Glu Tyr Val Arg Ile His Glu Lys Leu Asp Ile Tyr 1 5 10 15 acg ttt aag aaa cag ttt ttc aac aac aaa aaa tcc gtg tcg cat aga 96 Thr Phe Lys Lys Gln Phe Phe Asn Asn Lys Lys Ser Val Ser His Arg 20 25 30 tgc tac gtt ctc ttt gaa tta aaa cga cgg ggt gaa cgt aga gcg tgt 144 Cys Tyr Val Leu Phe Glu Leu Lys Arg Arg Gly Glu Arg Arg Ala Cys 35 40 45 ttt tgg ggc tat gct gtg aat aaa cca cag agc ggg aca gaa cgt ggc 192 Phe Trp Gly Tyr Ala Val Asn Lys Pro Gln Ser Gly Thr Glu Arg Gly 50 55 60 att cac gcc gaa atc ttt agc att aga aaa gtc gaa gaa tac ctg cgc 240 Ile His Ala Glu Ile Phe Ser Ile Arg Lys Val Glu Glu Tyr Leu Arg 65 70 75 80 gac aac ccc gga caa ttc acg ata aat tgg tac tca tcc tgg agt cct 288 Asp Asn Pro Gly Gln Phe Thr Ile Asn Trp Tyr Ser Ser Trp Ser Pro 85 90 95 tgt gca gat tgc gct gaa aag atc tta gaa tgg tat aac cag gag ctg 336 Cys Ala Asp Cys Ala Glu Lys Ile Leu Glu Trp Tyr Asn Gln Glu Leu 100 105 110 cgg ggg aac ggc cac act ttg aaa atc tgg gct tgc aaa ctc tat tac 384 Arg Gly Asn Gly His Thr Leu Lys Ile Trp Ala Cys Lys Leu Tyr Tyr 115 120 125 gag aaa aat gcg agg aat caa att ggg ctg tgg aac ctc aga gat aac 432 Glu Lys Asn Ala Arg Asn Gln Ile Gly Leu Trp Asn Leu Arg Asp Asn 130 135 140 ggg gtt ggg ttg aat gta atg gta agt gaa cac tac caa tgt tgc agg 480 Gly Val Gly Leu Asn Val Met Val Ser Glu His Tyr Gln Cys Cys Arg 145 150 155 160 aaa ata ttc atc caa tcg tcg cac aat caa ttg aat gag aat aga tgg 528 Lys Ile Phe Ile Gln Ser Ser His Asn Gln Leu Asn Glu Asn Arg Trp 165 170 175 ctt gag aag act ttg aag cga gct gaa aaa cga cgg agc gag ttg tcc 576 Leu Glu Lys Thr Leu Lys Arg Ala Glu Lys Arg Arg Ser Glu Leu Ser 180 185 190 att atg att cag gta aaa ata ctc cac acc act aag agt cct gct gtt 624 Ile Met Ile Gln Val Lys Ile Leu His Thr Thr Lys Ser Pro Ala Val 195 200 205 <210> 2 <211> 208 <212> PRT <213> Lamprey <400> 2 Met Thr Asp Ala Glu Tyr Val Arg Ile His Glu Lys Leu Asp Ile Tyr 1 5 10 15 Thr Phe Lys Lys Gln Phe Phe Asn Asn Lys Lys Ser Val Ser His Arg 20 25 30 Cys Tyr Val Leu Phe Glu Leu Lys Arg Arg Gly Glu Arg Arg Ala Cys 35 40 45 Phe Trp Gly Tyr Ala Val Asn Lys Pro Gln Ser Gly Thr Glu Arg Gly 50 55 60 Ile His Ala Glu Ile Phe Ser Ile Arg Lys Val Glu Glu Tyr Leu Arg 65 70 75 80 Asp Asn Pro Gly Gln Phe Thr Ile Asn Trp Tyr Ser Ser Trp Ser Pro 85 90 95 Cys Ala Asp Cys Ala Glu Lys Ile Leu Glu Trp Tyr Asn Gln Glu Leu 100 105 110 Arg Gly Asn Gly His Thr Leu Lys Ile Trp Ala Cys Lys Leu Tyr Tyr 115 120 125 Glu Lys Asn Ala Arg Asn Gln Ile Gly Leu Trp Asn Leu Arg Asp Asn 130 135 140 Gly Val Gly Leu Asn Val Met Val Ser Glu His Tyr Gln Cys Cys Arg 145 150 155 160 Lys Ile Phe Ile Gln Ser Ser His Asn Gln Leu Asn Glu Asn Arg Trp 165 170 175 Leu Glu Lys Thr Leu Lys Arg Ala Glu Lys Arg Arg Ser Glu Leu Ser 180 185 190 Ile Met Ile Gln Val Lys Ile Leu His Thr Thr Lys Ser Pro Ala Val 195 200 205 <210> 3 <211> 600 <212> DNA <213> Homo sapiens <220> <221> CDS <222> (1)..(600) <400> 3 atg gac agc ctc ttg atg aac cgg agg aag ttt ctt tac caa ttc aaa 48 Met Asp Ser Leu Leu Met Asn Arg Arg Lys Phe Leu Tyr Gln Phe Lys 1 5 10 15 aat gtc cgc tgg gct aag ggt cgg cgt gag acc tac ctg tgc tac gta 96 Asn Val Arg Trp Ala Lys Gly Arg Arg Glu Thr Tyr Leu Cys Tyr Val 20 25 30 gtg aag agg cgt gac agt gct aca tcc ttt tca ctg gac ttt ggt tat 144 Val Lys Arg Arg Asp Ser Ala Thr Ser Phe Ser Leu Asp Phe Gly Tyr 35 40 45 ctt cgc aat aag aac ggc tgc cac gtg gaa ttg ctc ttc ctc cgc tac 192 Leu Arg Asn Lys Asn Gly Cys His Val Glu Leu Leu Phe Leu Arg Tyr 50 55 60 atc tcg gac tgg gac cta gac cct ggc cgc tgc tac cgc gtc acc tgg 240 Ile Ser Asp Trp Asp Leu Asp Pro Gly Arg Cys Tyr Arg Val Thr Trp 65 70 75 80 ttc acc tcc tgg agc ccc tgc tac gac tgt gcc cga cat gtg gcc gac 288 Phe Thr Ser Trp Ser Pro Cys Tyr Asp Cys Ala Arg His Val Ala Asp 85 90 95 ttt ctg cga ggg aac ccc tac ctc agt ctg agg atc ttc acc gcg cgc 336 Phe Leu Arg Gly Asn Pro Tyr Leu Ser Leu Arg Ile Phe Thr Ala Arg 100 105 110 ctc tac ttc tgt gag gac cgc aag gct gag ccc gag ggg ctg cgg cgg 384 Leu Tyr Phe Cys Glu Asp Arg Lys Ala Glu Pro Glu Gly Leu Arg Arg 115 120 125 ctg cac cgc gcc ggg gtg caa ata gcc atc atg acc ttc aaa gat tat 432 Leu His Arg Ala Gly Val Gln Ile Ala Ile Met Thr Phe Lys Asp Tyr 130 135 140 ttt tac tgc tgg aat act ttt gta gaa aac cat gaa aga act ttc aaa 480 Phe Tyr Cys Trp Asn Thr Phe Val Glu Asn His Glu Arg Thr Phe Lys 145 150 155 160 gcc tgg gaa ggg ctg cat gaa aat tca gtt cgt ctc tcc aga cag ctt 528 Ala Trp Glu Gly Leu His Glu Asn Ser Val Arg Leu Ser Arg Gln Leu 165 170 175 cgg cgc atc ctt ttg ccc ctg tat gag gtt gat gac tta cga gac gca 576 Arg Arg Ile Leu Leu Pro Leu Tyr Glu Val Asp Asp Leu Arg Asp Ala 180 185 190 ttt cgt act ttg gga ctt ctc gac 600 Phe Arg Thr Leu Gly Leu Leu Asp 195 200 <210> 4 <211> 200 <212> PRT <213> Human <400> 4 Met Asp Ser Leu Leu Met Asn Arg Arg Lys Phe Leu Tyr Gln Phe Lys 1 5 10 15 Asn Val Arg Trp Ala Lys Gly Arg Arg Glu Thr Tyr Leu Cys Tyr Val 20 25 30 Val Lys Arg Arg Asp Ser Ala Thr Ser Phe Ser Leu Asp Phe Gly Tyr 35 40 45 Leu Arg Asn Lys Asn Gly Cys His Val Glu Leu Leu Phe Leu Arg Tyr 50 55 60 Ile Ser Asp Trp Asp Leu Asp Pro Gly Arg Cys Tyr Arg Val Thr Trp 65 70 75 80 Phe Thr Ser Trp Ser Pro Cys Tyr Asp Cys Ala Arg His Val Ala Asp 85 90 95 Phe Leu Arg Gly Asn Pro Tyr Leu Ser Leu Arg Ile Phe Thr Ala Arg 100 105 110 Leu Tyr Phe Cys Glu Asp Arg Lys Ala Glu Pro Glu Gly Leu Arg Arg 115 120 125 Leu His Arg Ala Gly Val Gln Ile Ala Ile Met Thr Phe Lys Asp Tyr 130 135 140 Phe Tyr Cys Trp Asn Thr Phe Val Glu Asn His Glu Arg Thr Phe Lys 145 150 155 160 Ala Trp Glu Gly Leu His Glu Asn Ser Val Arg Leu Ser Arg Gln Leu 165 170 175 Arg Arg Ile Leu Leu Pro Leu Tyr Glu Val Asp Asp Leu Arg Asp Ala 180 185 190 Phe Arg Thr Leu Gly Leu Leu Asp 195 200 <210> 5 <211> 4116 <212> DNA <213> Artificial Sequence <220> <223> Cas9 CDS from Streptococcus pyogenes, optimized for eukaryotic expression <220> <221> CDS <222> (1)..(4116) <400> 5 atg gac aag aag tac tcc att ggg ctc gat atc ggc aca aac agc gtc 48 Met Asp Lys Lys Tyr Ser Ile Gly Leu Asp Ile Gly Thr Asn Ser Val 1 5 10 15 ggt tgg gcc gtc att acg gac gag tac aag gtg ccg agc aaa aaa ttc 96 Gly Trp Ala Val Ile Thr Asp Glu Tyr Lys Val Pro Ser Lys Lys Phe 20 25 30 aaa gtt ctg ggc aat acc gat cgc cac agc ata aag aag aac ctc att 144 Lys Val Leu Gly Asn Thr Asp Arg His Ser Ile Lys Lys Asn Leu Ile 35 40 45 ggc gcc ctc ctg ttc gac tcc ggg gag acg gcc gaa gcc acg cgg ctc 192 Gly Ala Leu Leu Phe Asp Ser Gly Glu Thr Ala Glu Ala Thr Arg Leu 50 55 60 aaa aga aca gca cgg cgc aga tat acc cgc aga aag aat cgg atc tgc 240 Lys Arg Thr Ala Arg Arg Arg Tyr Thr Arg Arg Lys Asn Arg Ile Cys 65 70 75 80 tac ctg cag gag atc ttt agt aat gag atg gct aag gtg gat gac tct 288 Tyr Leu Gln Glu Ile Phe Ser Asn Glu Met Ala Lys Val Asp Asp Ser 85 90 95 ttc ttc cat agg ctg gag gag tcc ttt ttg gtg gag gag gat aaa aag 336 Phe Phe His Arg Leu Glu Glu Ser Phe Leu Val Glu Glu Asp Lys Lys 100 105 110 cac gag cgc cac cca atc ttt ggc aat atc gtg gac gag gtg gcg tac 384 His Glu Arg His Pro Ile Phe Gly Asn Ile Val Asp Glu Val Ala Tyr 115 120 125 cat gaa aag tac cca acc ata tat cat ctg agg aag aag ctt gta gac 432 His Glu Lys Tyr Pro Thr Ile Tyr His Leu Arg Lys Lys Leu Val Asp 130 135 140 agt act gat aag gct gac ttg cgg ttg atc tat ctc gcg ctg gcg cat 480 Ser Thr Asp Lys Ala Asp Leu Arg Leu Ile Tyr Leu Ala Leu Ala His 145 150 155 160 atg atc aaa ttt cgg gga cac ttc ctc atc gag ggg gac ctg aac cca 528 Met Ile Lys Phe Arg Gly His Phe Leu Ile Glu Gly Asp Leu Asn Pro 165 170 175 gac aac agc gat gtc gac aaa ctc ttt atc caa ctg gtt cag act tac 576 Asp Asn Ser Asp Val Asp Lys Leu Phe Ile Gln Leu Val Gln Thr Tyr 180 185 190 aat cag ctt ttc gaa gag aac ccg atc aac gca tcc gga gtt gac gcc 624 Asn Gln Leu Phe Glu Glu Asn Pro Ile Asn Ala Ser Gly Val Asp Ala 195 200 205 aaa gca atc ctg agc gct agg ctg tcc aaa tcc cgg cgg ctc gaa aac 672 Lys Ala Ile Leu Ser Ala Arg Leu Ser Lys Ser Arg Arg Leu Glu Asn 210 215 220 ctc atc gca cag ctc cct ggg gag aag aag aac ggc ctg ttt ggt aat 720 Leu Ile Ala Gln Leu Pro Gly Glu Lys Lys Asn Gly Leu Phe Gly Asn 225 230 235 240 ctt atc gcc ctg tca ctc ggg ctg acc ccc aac ttt aaa tct aac ttc 768 Leu Ile Ala Leu Ser Leu Gly Leu Thr Pro Asn Phe Lys Ser Asn Phe 245 250 255 gac ctg gcc gaa gat gcc aag ctt caa ctg agc aaa gac acc tac gat 816 Asp Leu Ala Glu Asp Ala Lys Leu Gln Leu Ser Lys Asp Thr Tyr Asp 260 265 270 gat gat ctc gac aat ctg ctg gcc cag atc ggc gac cag tac gca gac 864 Asp Asp Leu Asp Asn Leu Leu Ala Gln Ile Gly Asp Gln Tyr Ala Asp 275 280 285 ctt ttt ttg gcg gca aag aac ctg tca gac gcc att ctg ctg agt gat 912 Leu Phe Leu Ala Ala Lys Asn Leu Ser Asp Ala Ile Leu Leu Ser Asp 290 295 300 att ctg cga gtg aac acg gag atc acc aaa gct ccg ctg agc gct agt 960 Ile Leu Arg Val Asn Thr Glu Ile Thr Lys Ala Pro Leu Ser Ala Ser 305 310 315 320 atg atc aag cgc tat gat gag cac cac caa gac ttg act ttg ctg aag 1008 Met Ile Lys Arg Tyr Asp Glu His His Gln Asp Leu Thr Leu Leu Lys 325 330 335 gcc ctt gtc aga cag caa ctg cct gag aag tac aag gaa att ttc ttc 1056 Ala Leu Val Arg Gln Gln Leu Pro Glu Lys Tyr Lys Glu Ile Phe Phe 340 345 350 gat cag tct aaa aat ggc tac gcc gga tac att gac ggc gga gca agc 1104 Asp Gln Ser Lys Asn Gly Tyr Ala Gly Tyr Ile Asp Gly Gly Ala Ser 355 360 365 cag gag gaa ttt tac aaa ttt att aag ccc atc ttg gaa aaa atg gac 1152 Gln Glu Glu Phe Tyr Lys Phe Ile Lys Pro Ile Leu Glu Lys Met Asp 370 375 380 ggc acc gag gag ctg ctg gta aag ctt aac aga gaa gat ctg ttg cgc 1200 Gly Thr Glu Glu Leu Leu Val Lys Leu Asn Arg Glu Asp Leu Leu Arg 385 390 395 400 aaa cag cgc act ttc gac aat gga agc atc ccc cac cag att cac ctg 1248 Lys Gln Arg Thr Phe Asp Asn Gly Ser Ile Pro His Gln Ile His Leu 405 410 415 ggc gaa ctg cac gct atc ctc agg cgg caa gag gat ttc tac ccc ttt 1296 Gly Glu Leu His Ala Ile Leu Arg Arg Gln Glu Asp Phe Tyr Pro Phe 420 425 430 ttg aaa gat aac agg gaa aag att gag aaa atc ctc aca ttt cgg ata 1344 Leu Lys Asp Asn Arg Glu Lys Ile Glu Lys Ile Leu Thr Phe Arg Ile 435 440 445 ccc tac tat gta ggc ccc ctc gcc cgg gga aat tcc aga ttc gcg tgg 1392 Pro Tyr Tyr Val Gly Pro Leu Ala Arg Gly Asn Ser Arg Phe Ala Trp 450 455 460 atg act cgc aaa tca gaa gag acc atc act ccc tgg aac ttc gag gaa 1440 Met Thr Arg Lys Ser Glu Glu Thr Ile Thr Pro Trp Asn Phe Glu Glu 465 470 475 480 gtc gtg gat aag ggg gcc tct gcc cag tcc ttc atc gaa agg atg act 1488 Val Val Asp Lys Gly Ala Ser Ala Gln Ser Phe Ile Glu Arg Met Thr 485 490 495 aac ttt gat aaa aat ctg cct aac gaa aag gtg ctt cct aaa cac tct 1536 Asn Phe Asp Lys Asn Leu Pro Asn Glu Lys Val Leu Pro Lys His Ser 500 505 510 ctg ctg tac gag tac ttc aca gtt tat aac gag ctc acc aag gtc aaa 1584 Leu Leu Tyr Glu Tyr Phe Thr Val Tyr Asn Glu Leu Thr Lys Val Lys 515 520 525 tac gtc aca gaa ggg atg aga aag cca gca ttc ctg tct gga gag cag 1632 Tyr Val Thr Glu Gly Met Arg Lys Pro Ala Phe Leu Ser Gly Glu Gln 530 535 540 aag aaa gct atc gtg gac ctc ctc ttc aag acg aac cgg aaa gtt acc 1680 Lys Lys Ala Ile Val Asp Leu Leu Phe Lys Thr Asn Arg Lys Val Thr 545 550 555 560 gtg aaa cag ctc aaa gaa gac tat ttc aaa aag att gaa tgt ttc gac 1728 Val Lys Gln Leu Lys Glu Asp Tyr Phe Lys Lys Ile Glu Cys Phe Asp 565 570 575 tct gtt gaa atc agc gga gtg gag gat cgc ttc aac gca tcc ctg gga 1776 Ser Val Glu Ile Ser Gly Val Glu Asp Arg Phe Asn Ala Ser Leu Gly 580 585 590 acg tat cac gat ctc ctg aaa atc att aaa gac aag gac ttc ctg gac 1824 Thr Tyr His Asp Leu Leu Lys Ile Ile Lys Asp Lys Asp Phe Leu Asp 595 600 605 aat gag gag aac gag gac att ctt gag gac att gtc ctc acc ctt acg 1872 Asn Glu Glu Asn Glu Asp Ile Leu Glu Asp Ile Val Leu Thr Leu Thr 610 615 620 ttg ttt gaa gat agg gag atg att gaa gaa cgc ttg aaa act tac gct 1920 Leu Phe Glu Asp Arg Glu Met Ile Glu Glu Arg Leu Lys Thr Tyr Ala 625 630 635 640 cat ctc ttc gac gac aaa gtc atg aaa cag ctc aag agg cgc cga tat 1968 His Leu Phe Asp Asp Lys Val Met Lys Gln Leu Lys Arg Arg Arg Tyr 645 650 655 aca gga tgg ggg cgg ctg tca aga aaa ctg atc aat ggg atc cga gac 2016 Thr Gly Trp Gly Arg Leu Ser Arg Lys Leu Ile Asn Gly Ile Arg Asp 660 665 670 aag cag agt gga aag aca atc ctg gat ttt ctt aag tcc gat gga ttt 2064 Lys Gln Ser Gly Lys Thr Ile Leu Asp Phe Leu Lys Ser Asp Gly Phe 675 680 685 gcc aac cgg aac ttc atg cag ttg atc cat gat gac tct ctc acc ttt 2112 Ala Asn Arg Asn Phe Met Gln Leu Ile His Asp Asp Ser Leu Thr Phe 690 695 700 aag gag gac atc cag aaa gca caa gtt tct ggc cag ggg gac agt ctt 2160 Lys Glu Asp Ile Gln Lys Ala Gln Val Ser Gly Gln Gly Asp Ser Leu 705 710 715 720 cac gag cac atc gct aat ctt gca ggt agc cca gct atc aaa aag gga 2208 His Glu His Ile Ala Asn Leu Ala Gly Ser Pro Ala Ile Lys Lys Gly 725 730 735 ata ctg cag acc gtt aag gtc gtg gat gaa ctc gtc aaa gta atg gga 2256 Ile Leu Gln Thr Val Lys Val Val Asp Glu Leu Val Lys Val Met Gly 740 745 750 agg cat aag ccc gag aat atc gtt atc gag atg gcc cga gag aac caa 2304 Arg His Lys Pro Glu Asn Ile Val Ile Glu Met Ala Arg Glu Asn Gln 755 760 765 act acc cag aag gga cag aag aac agt agg gaa agg atg aag agg att 2352 Thr Thr Gln Lys Gly Gln Lys Asn Ser Arg Glu Arg Met Lys Arg Ile 770 775 780 gaa gag ggt ata aaa gaa ctg ggg tcc caa atc ctt aag gaa cac cca 2400 Glu Glu Gly Ile Lys Glu Leu Gly Ser Gln Ile Leu Lys Glu His Pro 785 790 795 800 gtt gaa aac acc cag ctt cag aat gag aag ctc tac ctg tac tac ctg 2448 Val Glu Asn Thr Gln Leu Gln Asn Glu Lys Leu Tyr Leu Tyr Tyr Leu 805 810 815 cag aac ggc agg gac atg tac gtg gat cag gaa ctg gac atc aat cgg 2496 Gln Asn Gly Arg Asp Met Tyr Val Asp Gln Glu Leu Asp Ile Asn Arg 820 825 830 ctc tcc gac tac gac gtg gat cat atc gtg ccc cag tct ttt ctc aaa 2544 Leu Ser Asp Tyr Asp Val Asp His Ile Val Pro Gln Ser Phe Leu Lys 835 840 845 gat gat tct att gat aat aaa gtg ttg aca aga tcc gat aaa aat aga 2592 Asp Asp Ser Ile Asp Asn Lys Val Leu Thr Arg Ser Asp Lys Asn Arg 850 855 860 ggg aag agt gat aac gtc ccc tca gaa gaa gtt gtc aag aaa atg aaa 2640 Gly Lys Ser Asp Asn Val Pro Ser Glu Glu Val Val Lys Lys Met Lys 865 870 875 880 aat tat tgg cgg cag ctg ctg aac gcc aaa ctg atc aca caa cgg aag 2688 Asn Tyr Trp Arg Gln Leu Leu Asn Ala Lys Leu Ile Thr Gln Arg Lys 885 890 895 ttc gat aat ctg act aag gct gaa cga ggt ggc ctg tct gag ttg gat 2736 Phe Asp Asn Leu Thr Lys Ala Glu Arg Gly Gly Leu Ser Glu Leu Asp 900 905 910 aaa gcc ggc ttc atc aaa agg cag ctt gtt gag aca cgc cag atc acc 2784 Lys Ala Gly Phe Ile Lys Arg Gln Leu Val Glu Thr Arg Gln Ile Thr 915 920 925 aag cac gtg gcc caa att ctc gat tca cgc atg aac acc aag tac gat 2832 Lys His Val Ala Gln Ile Leu Asp Ser Arg Met Asn Thr Lys Tyr Asp 930 935 940 gaa aat gac aaa ctg att cga gag gtg aaa gtt att act ctg aag tct 2880 Glu Asn Asp Lys Leu Ile Arg Glu Val Lys Val Ile Thr Leu Lys Ser 945 950 955 960 aag ctg gtc tca gat ttc aga aag gac ttt cag ttt tat aag gtg aga 2928 Lys Leu Val Ser Asp Phe Arg Lys Asp Phe Gln Phe Tyr Lys Val Arg 965 970 975 gag atc aac aat tac cac cat gcg cat gat gcc tac ctg aat gca gtg 2976 Glu Ile Asn Asn Tyr His His Ala His Asp Ala Tyr Leu Asn Ala Val 980 985 990 gta ggc act gca ctt atc aaa aaa tat ccc aag ctt gaa tct gaa ttt 3024 Val Gly Thr Ala Leu Ile Lys Lys Tyr Pro Lys Leu Glu Ser Glu Phe 995 1000 1005 gtt tac gga gac tat aaa gtg tac gat gtt agg aaa atg atc gca 3069 Val Tyr Gly Asp Tyr Lys Val Tyr Asp Val Arg Lys Met Ile Ala 1010 1015 1020 aag tct gag cag gaa ata ggc aag gcc acc gct aag tac ttc ttt 3114 Lys Ser Glu Gln Glu Ile Gly Lys Ala Thr Ala Lys Tyr Phe Phe 1025 1030 1035 tac agc aat att atg aat ttt ttc aag acc gag att aca ctg gcc 3159 Tyr Ser Asn Ile Met Asn Phe Phe Lys Thr Glu Ile Thr Leu Ala 1040 1045 1050 aat gga gag att cgg aag cga cca ctt atc gaa aca aac gga gaa 3204 Asn Gly Glu Ile Arg Lys Arg Pro Leu Ile Glu Thr Asn Gly Glu 1055 1060 1065 aca gga gaa atc gtg tgg gac aag ggt agg gat ttc gcg aca gtc 3249 Thr Gly Glu Ile Val Trp Asp Lys Gly Arg Asp Phe Ala Thr Val 1070 1075 1080 cgg aag gtc ctg tcc atg ccg cag gtg aac atc gtt aaa aag acc 3294 Arg Lys Val Leu Ser Met Pro Gln Val Asn Ile Val Lys Lys Thr 1085 1090 1095 gaa gta cag acc gga ggc ttc tcc aag gaa agt atc ctc ccg aaa 3339 Glu Val Gln Thr Gly Gly Phe Ser Lys Glu Ser Ile Leu Pro Lys 1100 1105 1110 agg aac agc gac aag ctg atc gca cgc aaa aaa gat tgg gac ccc 3384 Arg Asn Ser Asp Lys Leu Ile Ala Arg Lys Lys Asp Trp Asp Pro 1115 1120 1125 aag aaa tac ggc gga ttc gat tct cct aca gtc gct tac agt gta 3429 Lys Lys Tyr Gly Gly Phe Asp Ser Pro Thr Val Ala Tyr Ser Val 1130 1135 1140 ctg gtt gtg gcc aaa gtg gag aaa ggg aag tct aaa aaa ctc aaa 3474 Leu Val Val Ala Lys Val Glu Lys Gly Lys Ser Lys Lys Leu Lys 1145 1150 1155 agc gtc aag gaa ctg ctg ggc atc aca atc atg gag cga tca agc 3519 Ser Val Lys Glu Leu Leu Gly Ile Thr Ile Met Glu Arg Ser Ser 1160 1165 1170 ttc gaa aaa aac ccc atc gac ttt ctc gag gcg aaa gga tat aaa 3564 Phe Glu Lys Asn Pro Ile Asp Phe Leu Glu Ala Lys Gly Tyr Lys 1175 1180 1185 gag gtc aaa aaa gac ctc atc att aag ctt ccc aag tac tct ctc 3609 Glu Val Lys Lys Asp Leu Ile Ile Lys Leu Pro Lys Tyr Ser Leu 1190 1195 1200 ttt gag ctt gaa aac ggc cgg aaa cga atg ctc gct agt gcg ggc 3654 Phe Glu Leu Glu Asn Gly Arg Lys Arg Met Leu Ala Ser Ala Gly 1205 1210 1215 gag ctg cag aaa ggt aac gag ctg gca ctg ccc tct aaa tac gtt 3699 Glu Leu Gln Lys Gly Asn Glu Leu Ala Leu Pro Ser Lys Tyr Val 1220 1225 1230 aat ttc ttg tat ctg gcc agc cac tat gaa aag ctc aaa ggg tct 3744 Asn Phe Leu Tyr Leu Ala Ser His Tyr Glu Lys Leu Lys Gly Ser 1235 1240 1245 ccc gaa gat aat gag cag aag cag ctg ttc gtg gaa caa cac aaa 3789 Pro Glu Asp Asn Glu Gln Lys Gln Leu Phe Val Glu Gln His Lys 1250 1255 1260 cac tac ctt gat gag atc atc gag caa ata agc gaa ttc tcc aaa 3834 His Tyr Leu Asp Glu Ile Ile Glu Gln Ile Ser Glu Phe Ser Lys 1265 1270 1275 aga gtg atc ctc gcc gac gct aac ctc gat aag gtg ctt tct gct 3879 Arg Val Ile Leu Ala Asp Ala Asn Leu Asp Lys Val Leu Ser Ala 1280 1285 1290 tac aat aag cac agg gat aag ccc atc agg gag cag gca gaa aac 3924 Tyr Asn Lys His Arg Asp Lys Pro Ile Arg Glu Gln Ala Glu Asn 1295 1300 1305 att atc cac ttg ttt act ctg acc aac ttg ggc gcg cct gca gcc 3969 Ile Ile His Leu Phe Thr Leu Thr Asn Leu Gly Ala Pro Ala Ala 1310 1315 1320 ttc aag tac ttc gac acc acc ata gac aga aag cgg tac acc tct 4014 Phe Lys Tyr Phe Asp Thr Thr Ile Asp Arg Lys Arg Tyr Thr Ser 1325 1330 1335 aca aag gag gtc ctg gac gcc aca ctg att cat cag tca att acg 4059 Thr Lys Glu Val Leu Asp Ala Thr Leu Ile His Gln Ser Ile Thr 1340 1345 1350 ggg ctc tat gaa aca aga atc gac ctc tct cag ctc ggt gga gac 4104 Gly Leu Tyr Glu Thr Arg Ile Asp Leu Ser Gln Leu Gly Gly Asp 1355 1360 1365 agc agg gct gac 4116 Ser Arg Ala Asp 1370 <210> 6 <211> 1372 <212> PRT <213> Artificial Sequence <220> <223> Synthetic Construct <400> 6 Met Asp Lys Lys Tyr Ser Ile Gly Leu Asp Ile Gly Thr Asn Ser Val 1 5 10 15 Gly Trp Ala Val Ile Thr Asp Glu Tyr Lys Val Pro Ser Lys Lys Phe 20 25 30 Lys Val Leu Gly Asn Thr Asp Arg His Ser Ile Lys Lys Asn Leu Ile 35 40 45 Gly Ala Leu Leu Phe Asp Ser Gly Glu Thr Ala Glu Ala Thr Arg Leu 50 55 60 Lys Arg Thr Ala Arg Arg Arg Tyr Thr Arg Arg Lys Asn Arg Ile Cys 65 70 75 80 Tyr Leu Gln Glu Ile Phe Ser Asn Glu Met Ala Lys Val Asp Asp Ser 85 90 95 Phe Phe His Arg Leu Glu Glu Ser Phe Leu Val Glu Glu Asp Lys Lys 100 105 110 His Glu Arg His Pro Ile Phe Gly Asn Ile Val Asp Glu Val Ala Tyr 115 120 125 His Glu Lys Tyr Pro Thr Ile Tyr His Leu Arg Lys Lys Leu Val Asp 130 135 140 Ser Thr Asp Lys Ala Asp Leu Arg Leu Ile Tyr Leu Ala Leu Ala His 145 150 155 160 Met Ile Lys Phe Arg Gly His Phe Leu Ile Glu Gly Asp Leu Asn Pro 165 170 175 Asp Asn Ser Asp Val Asp Lys Leu Phe Ile Gln Leu Val Gln Thr Tyr 180 185 190 Asn Gln Leu Phe Glu Glu Asn Pro Ile Asn Ala Ser Gly Val Asp Ala 195 200 205 Lys Ala Ile Leu Ser Ala Arg Leu Ser Lys Ser Arg Arg Leu Glu Asn 210 215 220 Leu Ile Ala Gln Leu Pro Gly Glu Lys Lys Asn Gly Leu Phe Gly Asn 225 230 235 240 Leu Ile Ala Leu Ser Leu Gly Leu Thr Pro Asn Phe Lys Ser Asn Phe 245 250 255 Asp Leu Ala Glu Asp Ala Lys Leu Gln Leu Ser Lys Asp Thr Tyr Asp 260 265 270 Asp Asp Leu Asp Asn Leu Leu Ala Gln Ile Gly Asp Gln Tyr Ala Asp 275 280 285 Leu Phe Leu Ala Ala Lys Asn Leu Ser Asp Ala Ile Leu Leu Ser Asp 290 295 300 Ile Leu Arg Val Asn Thr Glu Ile Thr Lys Ala Pro Leu Ser Ala Ser 305 310 315 320 Met Ile Lys Arg Tyr Asp Glu His His Gln Asp Leu Thr Leu Leu Lys 325 330 335 Ala Leu Val Arg Gln Gln Leu Pro Glu Lys Tyr Lys Glu Ile Phe Phe 340 345 350 Asp Gln Ser Lys Asn Gly Tyr Ala Gly Tyr Ile Asp Gly Gly Ala Ser 355 360 365 Gln Glu Glu Phe Tyr Lys Phe Ile Lys Pro Ile Leu Glu Lys Met Asp 370 375 380 Gly Thr Glu Glu Leu Leu Val Lys Leu Asn Arg Glu Asp Leu Leu Arg 385 390 395 400 Lys Gln Arg Thr Phe Asp Asn Gly Ser Ile Pro His Gln Ile His Leu 405 410 415 Gly Glu Leu His Ala Ile Leu Arg Arg Gln Glu Asp Phe Tyr Pro Phe 420 425 430 Leu Lys Asp Asn Arg Glu Lys Ile Glu Lys Ile Leu Thr Phe Arg Ile 435 440 445 Pro Tyr Tyr Val Gly Pro Leu Ala Arg Gly Asn Ser Arg Phe Ala Trp 450 455 460 Methionine Threonine Arginine Lysine Serine Glutamic acid Glutamic acid Threonine Isoleucine Threonine Proline Tryptophan Asparagine Phenylalanine Glutamic acid Glutamic acid 465 470 475 480 Valine Valine Aspartic acid Lysine Glycine Alanine Serine Alanine Glutamine Serine Phenylalanine Isoleucine Glutamic acid Arginine Methionine Threonine 485 490 495 Asparagine Phenylalanine Aspartic acid Lysine Asparagine Leucine Proline Asparagine Glutamic acid Lysine Valine Leucine Proline Lysine Histidine Serine 500 505 510 Leucine Leucine Tyrosine Glutamic acid Tyrosine Phenylalanine Threonine Valine Tyrosine Asparagine Glutamic acid Leucine Threonine Lysine Valine Lysine 515 520 525 Tyrosine Valine Threonine Glutamic acid Glycine Methionine Arginine Lysine Proline Alanine Phenylalanine Leucine Serine Glycine Glutamic acid Glutamine 530 535 540 Lysine Lysine Alanine Isoleucine Valine Aspartic acid Leucine Leucine Phenylalanine Lysine Threonine Asparagine Arginine Lysine Valine Threonine 545 550 555 560 Valine Lysine Glutamine Leucine Lysine Glutamic acid Aspartic acid Tyrosine Phenylalanine Lysine Lysine Isoleucine Glutamic acid Cysteine Phenylalanine Aspartic acid 565 570 575 Serine Valine Glutamic acid Isoleucine Serine Glycine Valine Glutamic acid Aspartic acid Arginine Phenylalanine Asparagine Alanine Serine Leucine Glycine 580 585 590 Threonine Tyrosine Histidine Aspartic acid Leucine Leucine Lysine Isoleucine Isoleucine Lysine Aspartic acid Lysine Aspartic acid Phenylalanine Leucine Aspartic acid 595 600 605 Asparagine Glutamic acid Glutamic acid Asparagine Glutamic acid Aspartic acid Isoleucine Leucine Glutamic acid Aspartic acid Isoleucine Valine Leucine Threonine Leucine Threonine 610 615 620 Leu Phe Glu Asp Arg Glu Met Ile Glu Glu Arg Leu Lys Thr Tyr Ala 625 630 635 640 His Leu Phe Asp Asp Lys Val Met Lys Gln Leu Lys Arg Arg Arg Tyr 645 650 655 Thr Gly Trp Gly Arg Leu Ser Arg Lys Leu Ile Asn Gly Ile Arg Asp 660 665 670 Lys Gln Ser Gly Lys Thr Ile Leu Asp Phe Leu Lys Ser Asp Gly Phe 675 680 685 Ala Asn Arg Asn Phe Met Gln Leu Ile His Asp Asp Ser Leu Thr Phe 690 695 700 Lys Glu Asp Ile Gln Lys Ala Gln Val Ser Gly Gln Gly Asp Ser Leu 705 710 715 720 His Glu His Ile Ala Asn Leu Ala Gly Ser Pro Ala Ile Lys Lys Gly 725 730 735 Ile Leu Gln Thr Val Lys Val Val Asp Glu Leu Val Lys Val Met Gly 740 745 750 Arg His Lys Pro Glu Asn Ile Val Ile Glu Met Ala Arg Glu Asn Gln 755 760 765 Thr Thr Gln Lys Gly Gln Lys Asn Ser Arg Glu Arg Met Lys Arg Ile 770 775 780 Glu Glu Gly Ile Lys Glu Leu Gly Ser Gln Ile Leu Lys Glu His Pro 785 790 795 800 Val Glu Asn Thr Gln Leu Gln Asn Glu Lys Leu Tyr Leu Tyr Tyr Leu 805 810 815 Gln Asn Gly Arg Asp Met Tyr Val Asp Gln Glu Leu Asp Ile Asn Arg 820 825 830 Leu Ser Asp Tyr Asp Val Asp His Ile Val Pro Gln Ser Phe Leu Lys 835 840 845 Asp Asp Ser Ile Asp Asn Lys Val Leu Thr Arg Ser Asp Lys Asn Arg 850 855 860 Gly Lys Ser Asp Asn Val Pro Ser Glu Glu Val Val Lys Lys Met Lys 865 870 875 880 Asn Tyr Trp Arg Gln Leu Leu Asn Ala Lys Leu Ile Thr Gln Arg Lys 885 890 895 Phe Asp Asn Leu Thr Lys Ala Glu Arg Gly Gly Leu Ser Glu Leu Asp 900 905 910 Lys Ala Gly Phe Ile Lys Arg Gln Leu Val Glu Thr Arg Gln Ile Thr 915 920 925 Lys His Val Ala Gln Ile Leu Asp Ser Arg Met Asn Thr Lys Tyr Asp 930 935 940 Glu Asn Asp Lys Leu Ile Arg Glu Val Lys Val Ile Thr Leu Lys Ser 945 950 955 960 Lys Leu Val Ser Asp Phe Arg Lys Asp Phe Gln Phe Tyr Lys Val Arg 965 970 975 Glu Ile Asn Asn Tyr His His Ala His Asp Ala Tyr Leu Asn Ala Val 980 985 990 Val Gly Thr Ala Leu Ile Lys Lys Tyr Pro Lys Leu Glu Ser Glu Phe 995 1000 1005 Val Tyr Gly Asp Tyr Lys Val Tyr Asp Val Arg Lys Met Ile Ala 1010 1015 1020 Lys Ser Glu Gln Glu Ile Gly Lys Ala Thr Ala Lys Tyr Phe Phe 1025 1030 1035 Tyr Ser Asn Ile Met Asn Phe Phe Lys Thr Glu Ile Thr Leu Ala 1040 1045 1050 Asn Gly Glu Ile Arg Lys Arg Pro Leu Ile Glu Thr Asn Gly Glu 1055 1060 1065 Thr Gly Glu Ile Val Trp Asp Lys Gly Arg Asp Phe Ala Thr Val 1070 1075 1080 Arg Lys Val Leu Ser Met Pro Gln Val Asn Ile Val Lys Lys Thr 1085 1090 1095 Glu Val Gln Thr Gly Gly Phe Ser Lys Glu Ser Ile Leu Pro Lys 1100 1105 1110 Arg Asn Ser Asp Lys Leu Ile Ala Arg Lys Lys Asp Trp Asp Pro 1115 1120 1125 Lys Lys Tyr Gly Gly Phe Asp Ser Pro Thr Val Ala Tyr Ser Val 1130 1135 1140 Leu Val Val Ala Lys Val Glu Lys Gly Lys Ser Lys Lys Leu Lys 1145 1150 1155 Ser Val Lys Glu Leu Leu Gly Ile Thr Ile Met Glu Arg Ser Ser 1160 1165 1170 Phe Glu Lys Asn Pro Ile Asp Phe Leu Glu Ala Lys Gly Tyr Lys 1175 1180 1185 Glu Val Lys Lys Asp Leu Ile Ile Lys Leu Pro Lys Tyr Ser Leu 1190 1195 1200 Phe Glu Leu Glu Asn Gly Arg Lys Arg Met Leu Ala Ser Ala Gly 1205 1210 1215 Glu Leu Gln Lys Gly Asn Glu Leu Ala Leu Pro Ser Lys Tyr Val 1220 1225 1230 Asn Phe Leu Tyr Leu Ala Ser His Tyr Glu Lys Leu Lys Gly Ser 1235 1240 1245 Pro Glu Asp Asn Glu Gln Lys Gln Leu Phe Val Glu Gln His Lys 1250 1255 1260 His Tyr Leu Asp Glu Ile Ile Glu Gln Ile Ser Glu Phe Ser Lys 1265 1270 1275 Arg Val Ile Leu Ala Asp Ala Asn Leu Asp Lys Val Leu Ser Ala 1280 1285 1290 Tyr Asn Lys His Arg Asp Lys Pro Ile Arg Glu Gln Ala Glu Asn 1295 1300 1305 Ile Ile His Leu Phe Thr Leu Thr Asn Leu Gly Ala Pro Ala Ala 1310 1315 1320 Phe Lys Tyr Phe Asp Thr Thr Ile Asp Arg Lys Arg Tyr Thr Ser 1325 1330 1335 Thr Lys Glu Val Leu Asp Ala Thr Leu Ile His Gln Ser Ile Thr 1340 1345 1350 Gly Leu Tyr Glu Thr Arg Ile Asp Leu Ser Gln Leu Gly Gly Asp 1355 1360 1365 Ser Arg Ala Asp 1370 <210> 7 <211> 83 <212> DNA <213> Streptococcus pyogenes <400> 7 gttttagagc tagaaatagc aagttaaaat aaggctagtc cgttatcaac ttgaaaaagt 60 ggcaccgagt cggtggtgct ttt 83 <210> 8 <211> 665 <212> DNA <213> Saccharomyces cerevisiae <400> 8 tttcaaaaat tcttactttt tttttggatg gacgcaaaga agtttaataa tcatattaca 60 tggcattacc accatataca tatccatata catatccata tctaatctta cttatatgtt 120 gtggaaatgt aaagagcccc attatcttag cctaaaaaaa ccttctcttt ggaactttca 180 gtaatacgct taactgctca ttgctatatt gaagtacgga ttagaagccg ccgagcgggt 240 gacagccctc cgaaggaaga ctctcctccg tgcgtcctcg tcttcaccgg tcgcgttcct 300 gaaacgcaga tgtgcctcgc gccgcactgc tccgaacaat aaagattcta caatactagc 360 ttttatggtt atgaagagga aaaattggca gtaacctggc cccacaaacc ttcaaatgaa 420 cgaatcaaat taacaaccat aggatgataa tgcgattagt tttttagcct tatttctggg 480 gtaattaatc agcgaagcga tgatttttga tctattaaca gatatataaa tgcaaaaact 540 gcataaccac tttaactaat actttcaaca ttttcggttt gtattacttc ttattcaaat 600 gtaataaaag tatcaacaaa aaattgttaa tatacctcta tactttaacg tcaaggagaa 660 aaaac 665 <210> 9 <211> 21 <212> DNA <213> Artificial sequence <220> <223> Nuclear transition signal. <220> <221> CDS <222> (1)..(21) <400> 9 ccc aag aag aag agg aag gtg 21 Pro Lys Lys Lys Arg Lys Val 1 5 <210> 10 <211> 7 <212> PRT <213> Artificial sequence <220> <223> Synthetic construct <400> 10 Pro Lys Lys Lys Arg Lys Val 1 5 <210> 11 <211> 15 <212> DNA <213> Artificial sequence <220> <223> GS linker <220> <221> CDS <222> (1)..(15) <400> 11 ggt gga gga ggt tct 15 Gly Gly Gly Gly Ser 1 5 <210> 12 <211> 5 <212> PRT <213> Artificial sequence <220> <223> Synthetic construct <400> 12 Gly Gly Gly Gly Ser 1 5 <210> 13 <211> 66 <212> DNA <213> Artificial sequence <220> <223> Flag tag <220> <221> CDS <222> (1)..(66) <400> 13 gac tat aag gac cac gac gga gac tac aag gat cat gat att gat tac 48 Asp Tyr Lys Asp His Asp Gly Asp Tyr Lys Asp His Asp Ile Asp Tyr 1 5 10 15 aaa gac gat gac gat aag 66 Lys Asp Asp Asp Asp Lys 20 <210> 14 <211> 22 <212> PRT <213> Artificial sequence <220> <223> Synthetic construct <400> 14 Asp Tyr Lys Asp His Asp Gly Asp Tyr Lys Asp His Asp Ile Asp Tyr 1 5 10 15 Lys Asp Asp Asp Asp Lys 20 <210> 15 <211> 24 <212> DNA <213> Artificial sequence <220> <223> Strep-tag <220> <221> CDS <222> (1)..(24) <400> 15 tgg agc cac ccg cag ttc gaa aaa 24 Trp Ser His Pro Gln Phe Glu Lys 1 5 <210> 16 <211> 8 <212> PRT <213> Artificial sequence <220> <223> Synthetic construct <400> 16 Trp Ser His Pro Gln Phe Glu Lys 1 5 <210> 17 <211> 171 <212> DNA <213> Artificial sequence <220> <223> SH3 domain <220> <221> CDS <222> (1)..(171) <400> 17 gca gag tat gtg cgg gcc ctc ttt gac ttt aat ggg aat gat gaa gaa 48 Ala Glu Tyr Val Arg Ala Leu Phe Asp Phe Asn Gly Asn Asp Glu Glu 1 5 10 15 gat ctt ccc ttt aag aaa gga gac atc ctg aga atc cgg gat aag cct 96 Asp Leu Pro Phe Lys Lys Gly Asp Ile Leu Arg Ile Arg Asp Lys Pro 20 25 30 gaa gag cag tgg tgg aat gca gag gac agc gaa gga aag agg ggg atg 144 Glu Glu Gln Trp Trp Asn Ala Glu Asp Ser Glu Gly Lys Arg Gly Met 35 40 45 att cct gtc cct tac gtg gag aag tat 171 Ile Pro Val Pro Tyr Val Glu Lys Tyr 50 55 <210> 18 <211> 57 <212> PRT <213> Artificial Sequence <220> <223> Synthetic Construct <400> 18 Ala Glu Tyr Val Arg Ala Leu Phe Asp Phe Asn Gly Asn Asp Glu Glu 1 5 10 15 Asp Leu Pro Phe Lys Lys Gly Asp Ile Leu Arg Ile Arg Asp Lys Pro 20 25 30 Glu Glu Gln Trp Trp Asn Ala Glu Asp Ser Glu Gly Lys Arg Gly Met 35 40 45 Ile Pro Val Pro Tyr Val Glu Lys Tyr 50 55 <210> 19 <211> 188 <212> DNA <213> Saccharomyces cerevisiae <400> 19 gcgaatttct tatgatttat gatttttatt attaaataag ttataaaaaa aataagtgta 60 tacaaatttt aaagtgactc ttaggtttta aaacgaaaat tcttattctt gagtaactct 120 ttcctgtagg tcaggttgct ttctcaggta tagcatgagg tcgctcttat tgaccacacc 180 tctaccgg 188 <210> 20 <211> 417 <212> DNA <213> Saccharomyces cerevisiae <400> 20 ataccaggca tggagcttat ctggtccgtt cgagttttcg acgagtttgg agacattctt 60 tatagatgtc cttttttttt aatgatattc gttaaagaac aaaaagtcaa agcagtttaa 120 cctaacacct gttgttgatg ctacttgaaa caaggcttct aggcgaatac ttaaaaaggt 180 aatttcaata gcggtttata tatctgtttg cttttcaaga tattatgtaa acgcacgatg 240 tttttcgccc aggctttatt ttttttgttg ttgttgtctt ctcgaagaat tttctcgggc 300 agatctttgt cggaatgtaa aaaagcgcgt aattaaactt tctattatgc tgactaaaat 360 ggaagtgatc accaaaggct atttctgatt atataatcta gtcattactc gctcgag 417 <210> 21 <211> 33 <212> DNA <213> Artificial sequence <220> <223> SH3 - binding ligand <220> <221> CDS <222> (1)..(33) <400> 21 cct cca cct gct ctg cca cct aag aga agg aga 33 Pro Pro Pro Ala Leu Pro Pro Lys Arg Arg Arg 1 5 10 <210> 22 <211> 11 <212> PRT <213> Artificial sequence <220> <223> Synthetic construct <400> 22 Pro Pro Pro Ala Leu Pro Pro Lys Arg Arg Arg 1 5 10 <210> 23 <211> 269 <212> DNA <213> Saccharomyces cerevisiae <400> 23 tctttgaaaa gataatgtat gattatgctt tcactcatat ttatacagaa acttgatgtt 60 ttctttcgag tatatacaag gtgattacat gtacgtttga agtacaactc tagattttgt 120 agtgccctct tgggctagcg gtaaaggtgc gcattttttc acaccctaca atgttctgtt 180 caaaagattt tggtcaaacg ctgtagaagt gaaagttggt gcgcatgttt cggcgttcga 240 aacttctccg cagtgaaaga taaatgatc 269 <210> 24 <211> 14 <212> DNA <213> Saccharomyces cerevisiae <400> 24 tgttttttat gtct 14 <210> 25 <211> 20 <212> DNA <213> Saccharomyces cerevisiae <400> 25 gatacgttct ctatggagga 20 <210> 26 <211> 20 <212> DNA <213> Saccharomyces cerevisiae <400> 26 ttggagaaac ccaggtgcct 20 <210> 27 <211> 20 <212> DNA <213> Saccharomyces cerevisiae <400> 27 aacccaggtg cctggggtcc 20 <210> 28 <211> 20 <212> DNA <213> Saccharomyces cerevisiae <400> 28 ttggccaagt cattcaattt 20 <210> 29 <211> 20 <212> DNA <213> Saccharomyces cerevisiae <400> 29 ataacggaat ccaactgggc 20 <210> 30 <211> 20 <212> DNA <213> Saccharomyces cerevisiae <400> 30 gtcaattacg aagactgaac 20 <210> 31 <211> 19 <212> DNA <213> Saccharomyces cerevisiae <400> 31 gtcaatagga tcccctttt 19 <210> 32 <211> 10126 <212> DNA <213> Artificial Sequence <220> <223> Plasmid carrying dCas9-PmCDA1 fusion protein and chimeric RNA Targeting the galK gene of E.coli <220> <221> misc_feature <222> (5561)..(5580) <223> n is a, c, g, or t <400> 32 atcgccattc gccattcagg ctgcgcaact gttgggaagg gcgatcggtg cgggcctctt 60 cgctattacg ccagctggcg aaagggggat gtgctgcaag gcgattaagt tgggtaacgc 120 cagggttttc ccagtcacga cgttgtaaaa cgacggccag tgaattcgag ctcggtaccc 180 ggccgcaaac aacagataaa acgaaaggcc cagtctttcg actgagcctt tcgttttatt 240 tgatgcctgt caagtaacag caggactctt agtggtgtgg agtattttta cctgaatcat 300 aatggacaac tcgctccgtc gtttttcagc tcgcttcaaa gtcttctcaa gccatctatt 360 ctcattcaat tgattgtgcg acgattggat gaatattttc ctgcaacatt ggtagtgttc 420 acttaccatt acattcaacc caaccccgtt atctctgagg ttccacagcc caatttgatt 480 cctcgcattt ttctcgtaat agagtttgca agcccagatt ttcaaagtgt ggccgttccc 540 ccgcagctcc tggttatacc attctaagat cttttcagcg caatctgcac aaggactcca 600 ggatgagtac caatttatcg tgaattgtcc ggggttgtcg cgcaggtatt cttcgacttt 660 tctaatgcta aagatttcgg cgtgaatgcc acgttctgtc ccgctctgtg gtttattcac 720 agcatagccc caaaaacacg ctctacgttc accccgtcgt tttaattcaa agagaacgta 780 gcatctatgc gacacggatt ttttgttgtt gaaaaactgt ttcttaaacg tgtagatgtc 840 caacttctca tggattctca cgtactcagc gtcggtcatc ctagacttat cgtcatcgtc 900 tttgtaatca atatcatgat ccttgtagtc tccgtcgtgg tccttatagt ctccggactc 960 gagcctagac ttatcgtcat cgtctttgta atcaatatca tgatccttgt agtctccgtc 1020 gtggtcctta tagtctccgg aatacttctc cacgtaaggg acaggaatca tccccctctt 1080 tccttcgctg tcctctgcat tccaccactg ctcctcaggc ttatcccgga ttctcaggat 1140 gtctcctttc ttaaagggaa gatcctcttc atcattccca ttaaagtcaa agagggctcg 1200 cacatactca gcagaacctc cacctccaga acctcctcca ccgtcacctc ctagctgact 1260 caaatcaatg cgtgtttcat aaagaccagt gatggattga tggataagag tggcatctaa 1320 aacttctttt gtagacgtat atcgtttacg atcaattgtt gtatcaaaat atttaaaagc 1380 agcgggagct ccaagattcg tcaacgtaaa taaatgaata atattttctg cttgttcacg 1440 tattggtttg tctctatgtt tgttatatgc actaagaact ttatctaaat tggcatctgc 1500 taaaataaca cgcttagaaa attcactgat ttgctcaata atctcatcta aataatgctt 1560 atgctgctcc acaaacaatt gtttttgttc gttatcttct ggactaccct tcaacttttc 1620 ataatgacta gctaaatata aaaaattcac atatttgctt ggcagagcca gctcatttcc 1680 tttttgtaat tctccggcac tagccagcat ccgtttacga ccgttttcta actcaaaaag 1740 actatattta ggtagtttaa tgattaagtc ttttttaact tccttatatc ctttagcttc 1800 taaaaagtca atcggatttt tttcaaagga acttctttcc ataattgtga tccctagtaa 1860 ctctttaacg gattttaact tcttcgattt ccctttttcc accttagcaa ccactaggac 1920 tgaataagct accgttggac tatcaaaacc accatatttt tttggatccc agtctttttt 1980 acgagcaata agcttgtccg aatttctttt tggtaaaatt gactccttgg agaatccgcc 2040 tgtctgtact tctgttttct tgacaatatt gacttggggc atggacaata ctttgcgcac 2100 tgtggcaaaa tctcgccctt tatcccagac aatttctcca gtttccccat tagtttcgat 2160 tagagggcgt ttgcgaatct ctccatttgc aagtgtaatt tctgttttga agaagttcat 2220 gatattagag taaaagaaat attttgcggt tgctttgcct atttcttgct cagacttagc 2280 aatcatttta cgaacatcat aaactttata atcaccatag acaaactccg attcaagttt 2340 tggatatttc ttaatcaaag cagttccaac gacggcattt agatacgcat catgggcatg 2400 atggtaattg ttaatctcac gtactttata gaattggaaa tcttttcgga agtcagaaac 2460 taatttagat tttaaggtaa tcactttaac ctctcgaata agtttatcat tttcatcgta 2520 tttagtattc atgcgactat ccaaaatttg tgccacatgc ttagtgattt ggcgagtttc 2580 aaccaattgg cgtttgataa aaccagcttt atcaagttca ctcaaacctc cacgttcagc 2640 tttcgttaaa ttatcaaact tacgttgagt gattaacttg gcgtttagaa gttgtctcca 2700 atagtttttc atctttttga ctacttcttc acttggaacg ttatccgatt taccacgatt 2760 tttatcagaa cgcgttaaga ccttattgtc tattgaatcg tctttaagga aactttgtgg 2820 aacaatggca tcgacatcat aatcacttaa acgattaata tctaattctt ggtccacata 2880 catgtctctt ccattttgga gataatagag atagagcttt tcattttgca attgagtatt 2940 ttcaacagga tgctctttaa gaatctgact tcctaattct ttgatacctt cttcgattcg 3000 tttcatacgc tctcgcgaat ttttctggcc cttttgagtt gtctgatttt cacgtgccat 3060 ttcaataacg atattttctg gcttatgccg ccccattact ttgaccaatt catcaacaac 3120 ttttacagtc tgtaaaatac cttttttaat agcagggcta ccagctaaat ttgcaatatg 3180 ttcatgtaaa ctatcgcctt gtccagacac ttgtgctttt tgaatgtctt ctttaaatgt 3240 caaactatca tcatggatca gctgcataaa attgcgattg gcaaaaccat ctgatttcaa 3300 aaaatctaat attgttttgc cagattgctt atccctaata ccattaatca attttcgaga 3360 caaacgtccc caaccagtat aacggcgacg tttaagctgt ttcatcacct tatcatcaaa 3420 gaggtgagca tatgttttaa gtctttcctc aatcatctcc ctatcttcaa ataaggtcaa 3480 tgttaaaaca atatcctcta agatatcttc attttcttca ttatccaaaa aatctttatc 3540 tttaataatt tttagcaaat catggtaggt acctaatgaa gcattaaatc tatcttcaac 3600 tcctgaaatt tcaacactat caaaacattc tatttttttg aaataatctt cttttaattg 3660 cttaacggtt acttttcgat ttgttttgaa gagtaaatca acaatggctt tcttctgttc 3720 acctgaaaga aatgctggtt ttcgcattcc ttcagtaaca tatttgacct ttgtcaattc 3780 gttataaacc gtaaaatact cataaagcaa actatgtttt ggtagtactt tttcatttgg 3840 aagattttta tcaaagtttg tcatgcgttc aataaatgat tgagctgaag cacctttatc 3900 gacaacttct tcaaaattcc atggggtaat tgtttcttca gacttccgag tcatccatgc 3960 aaaacgacta ttgccacgcg ccaatggacc aacataataa ggaattcgaa aagtcaagat 4020 tttttcaatc ttctcacgat tgtcttttaa aaatggataa aagtcttctt gtcttctcaa 4080 aatagcatgc agctcaccca agtgaatttg atggggaata gagccgttgt caaaggtccg 4140 ttgcttgcgc agcaaatctt cacgatttag tttcaccaat aattcctcag taccatccat 4200 tttttctaaa attggtttga taaatttata aaattcttct tggctagctc ccccatcaat 4260 ataacctgca tatccgtttt ttgattgatc aaaaaagatt tctttatact tttctggaag 4320 ttgttgtcga actaaagctt ttaaaagagt caagtcttga tgatgttcat cgtagcgttt 4380 aatcattgaa gctgataggg gagccttagt tatttcagta tttactctta ggatatctga 4440 aagtaaaata gcatctgata aattcttagc tgccaaaaac aaatcagcat attgatctcc 4500 aatttgcgcc aataaattat ctaaatcatc atcgtaagta tcttttgaaa gctgtaattt 4560 agcatcttct gccaaatcaa aatttgattt aaaattaggg gtcaaaccca atgacaaagc 4620 aatgagattc ccaaataagc catttttctt ctcaccgggg agctgagcaa tgagattttc 4680 taatcgtctt gatttactca atcgtgcaga aagaatcgct ttagcatcta ctccacttgc 4740 gttaataggg ttttcttcaa ataattgatt gtaggtttgt accaactgga taaatagttt 4800 gtccacatca ctattatcag gatttaaatc tccctcaatc aaaaaatgac cacgaaactt 4860 aatcatatgc gctaaggcca aatagattaa gcgcaaatcc gctttatcag tagaatctac 4920 caattttttt cgcagatgat agatagttgg atatttctca tgataagcaa cttcatctac 4980 tatatttcca aaaataggat gacgttcatg cttcttgtct tcttccacca aaaaagactc 5040 ttcaagtcga tgaaagaaac tatcatctac tttcgccatc tcatttgaaa aaatctcctg 5100 tagataacaa atacgattct tccgacgtgt ataccttcta cgagctgtcc gtttgagacg 5160 agtcgcttcc gctgtctctc cactgtcaaa taaaagagcc cctataagat tttttttgat 5220 actgtggcgg tctgtatttc ccagaacctt gaacttttta gacggaacct tatattcatc 5280 agtgatcacc gcccatccga cgctatttgt gccgatagct aagcctattg agtatttctt 5340 atccattttt gcctcctaaa atgggccctt taaattaaat ccataatgag tttgatgatt 5400 tcaataatag ttttaatgac ctccgaaatt agtttaatat gctttaattt ttctttttca 5460 aaatatctct tcaaaaaata ttacccaata cttaataata aatagattat aacacaaaat 5520 tcttttgaca agtagtttat tttgttataa ttctatagta nnnnnnnnnn nnnnnnnnnn 5580 gttttagagc tagaaatagc aagttaaaat aaggctagtc cgttatcaac ttgaaaaagt 5640 ggcaccgagt cggtgctttt tttgatactt ctattctact ctgactgcaa accaaaaaaa 5700 caagcgcttt caaaacgctt gttttatcat ttttagggaa attaatctct taatcctttt 5760 atcattctac atttaggcgc tgccatcttg ctaaacctac taagctccac aggatgattt 5820 cgtaatcccg caagaggccc ggcagtaccg gcataaccaa gcctatgcct acagcatcca 5880 gggtgacggt gccgaggatg acgatgagcg cattgttaga tttcatacac ggtgcctgac 5940 tgcgttagca atttaactgt gataaactac cgcattaaag cttatcgatg ataagctgtc 6000 aaacatgaga attacaactt atatcgtatg gggctgactt caggtgctac atttgaagag 6060 ataaattgca ctgaaatcta gtcggatcct cgctcactga ctcgctgcgc tcggtcgttc 6120 ggctgcggcg agcggtatca gctcactcaa aggcggtaat acggttatcc acagaatcag 6180 gggataacgc aggaaagaac atgtgagcaa aaggccagca aaaggccagg aaccgtaaaa 6240 aggccgcgtt gctggcgttt ttccataggc tccgcccccc tgacgagcat cacaaaaatc 6300 gacgctcaag tcagaggtgg cgaaacccga caggactata aagataccag gcgtttcccc 6360 ctggaagctc cctcgtgcgc tctcctgttc cgaccctgcc gcttaccgga tacctgtccg 6420 cctttctccc ttcgggaagc gtggcgcttt ctcatagctc acgctgtagg tatctcagtt 6480 cggtgtaggt cgttcgctcc aagctgggct gtgtgcacga accccccgtt cagcccgacc 6540 gctgcgcctt atccggtaac tatcgtcttg agtccaaccc ggtaagacac gacttatcgc 6600 cactggcagc agccactggt aacaggatta gcagagcgag gtatgtaggc ggtgctacag 6660 agttcttgaa gtggtggcct aactacggct acactagaag gacagtattt ggtatctgcg 6720 ctctgctgaa gccagttacc ttcggaaaaa gagttggtag ctcttgatcc ggcaaacaaa 6780 ccaccgctgg tagcggtggt ttttttgttt gcaagcagca gattacgcgc agaaaaaaag 6840 gatctcaaga agatcctttg atcttttcta cggggtctga cgctcagtgg aacgaaaact 6900 cacgttaagg gattttggtc atgagattat caaaaaggat cttcacctag atccttttaa 6960 attaaaaatg aagttttaaa tcaatctaaa gtatatatga gtaaacttgg tctgacagtt 7020 accaatgctt aatcagtgag gcacctatct cagcgatctg tctatttcgt tcatccatag 7080 ttgcctgact ccccgtcgtg tagataacta cgatacggga gggcttacca tctggcccca 7140 gtgctgcaat gataccgcga gaaccacgct caccggctcc agatttatca gcaataaacc 7200 agccagccgg aagggccgag cgcagaagtg gtcctgcaac tttatccgcc tccatccagt 7260 ctattaattg ttgccgggaa gctagagtaa gtagttcgcc agttaatagt ttgcgcaacg 7320 ttgttgccat tgctgcaggc atcgtggtgt cacgctcgtc gtttggtatg gcttcattca 7380 gctccggttc ccaacgatca aggcgagtta catgatcccc catgttgtgc aaaaaagcgg 7440 ttagctcctt cggtcctccg atcgttgtca gaagtaagtt ggccgcagtg ttatcactca 7500 tggttatggc agcactgcat aattctctta ctgtcatgcc atccgtaaga tgcttttctg 7560 tgactggtga gtactcaacc aagtcattct gagaatagtg tatgcggcga ccgagttgct 7620 cttgcccggc gtcaacacgg gataataccg cgccacatag cagaacttta aaagtgctca 7680 tcattggaaa acgttcttcg gggcgaaaac tctcaaggat cttaccgctg ttgagatcca 7740 gttcgatgta acccactcgt gcacccaact gatcttcagc atcttttact ttcaccagcg 7800 tttctgggtg agcaaaaaca ggaaggcaaa atgccgcaaa aaagggaata agggcgacac 7860 ggaaatgttg aatactcata ctcttccttt ttcaatatta ttgaagcatt tatcagggtt 7920 attgtctcat gagcggatac atatttgaat gtatttagaa aaataaacaa ataggggttc 7980 cgcgcacatt tccccgaaaa gtgccacctg acgtcaatgc cgagcgaaag cgagccgaag 8040 ggtagcattt acgttagata accccctgat atgctccgac gctttatata gaaaagaaga 8100 ttcaactagg taaaatctta atataggttg agatgataag gtttataagg aatttgtttg 8160 ttctaatttt tcactcattt tgttctaatt tcttttaaca aatgttcttt tttttttaga 8220 acagttatga tatagttaga atagtttaaa ataaggagtg agaaaaagat gaaagaaaga 8280 tatggaacag tctataaagg ctctcagagg ctcatagacg aagaaagtgg agaagtcata 8340 gaggtagaca agttataccg taaacaaacg tctggtaact tcgtaaaggc atatatagtg 8400 caattaataa gtatgttaga tatgattggc ggaaaaaaac ttaaaatcgt taactatatc 8460 ctagataatg tccacttaag taacaataca atgatagcta caacaagaga aatagcaaaa 8520 gctacaggaa caagtctaca aacagtaata acaacactta aaatcttaga agaaggaaat 8580 attataaaaa gaaaaactgg agtattaatg ttaaaccctg aactactaat gagaggcgac 8640 gaccaaaaac aaaaatacct cttactcgaa tttgggaact ttgagcaaga ggcaaatgaa 8700 atagattgac ctcccaataa caccacgtag ttattgggag gtcaatctat gaaatgcgat 8760 taagcttttt ctaattcaca taagcgtgca ggtttaaagt acataaaaaa tataatgaaa 8820 aaaagcatca ttatactaac gttataccaa cattatactc tcattatact aattgcttat 8880 tccaatttcc tattggttgg aaccaacagg cgttagtgtg ttgttgagtt ggtactttca 8940 tgggattaat cccatgaaac ccccaaccaa ctcgccaaag ctttggctaa cacacacgcc 9000 attccaacca atagttttct cggcataaag ccatgctctg acgcttaaat gcactaatgc 9060 cttaaaaaaa cattaaagtc taacacacta gacttattta cttcgtaatt aagtcgttaa 9120 accgtgtgct ctacgaccaa aagtataaaa cctttaagaa ctttcttttt tcttgtaaaa 9180 aaagaaacta gataaatctc tcatatcttt tattcaataa tcgcatcaga ttgcagtata 9240 aatttaacga tcactcatca tgttcatatt tatcagagct cgtgctataa ttatactaat 9300 tttataagga ggaaaaaata aagagggtta taatgaacga gaaaaatata aaacacagtc 9360 aaaactttat tacttcaaaa cataatatag ataaaataat gacaaatata agattaaatg 9420 aacatgataa tatctttgaa atcggctcag gaaaagggca ttttaccctt gaattagtac 9480 agaggtgtaa tttcgtaact gccattgaaa tagaccataa attatgcaaa actacagaaa 9540 ataaacttgt tgatcacgat aatttccaag ttttaaacaa ggatatattg cagtttaaat 9600 ttcctaaaaa ccaatcctat aaaatatttg gtaatatacc ttataacata agtacggata 9660 taatacgcaa aattgttttt gatagtatag ctgatgagat ttatttaatc gtggaatacg 9720 ggtttgctaa aagattatta aatacaaaac gctcattggc attattttta atggcagaag 9780 ttgatatttc tatattaagt atggttccaa gagaatattt tcatcctaaa cctaaagtga 9840 atagctcact tatcagatta aatagaaaaa aatcaagaat atcacacaaa gataaacaga 9900 agtataatta tttcgttatg aaatgggtta acaaagaata caagaaaata tttacaaaaa 9960 atcaatttaa caattcctta aaacatgcag gaattgacga tttaaacaat attagctttg 10020 aacaattctt atctcttttc aatagctata aattatttaa taagtaagtt aagggatgca 10080 taaactgcat cccttaactt gtttttcgtg tacctatttt ttgtga 10126 <210> 33 <211> 20 <212> DNA <213> Escherichia coli <400> 33 tcaatgggct aactacgttc 20 <210> 34 <211> 20 <212> DNA <213> Escherichia coli <400> 34 ggtccataaa ctgagacagc 20 <210> 35 <211> 223 <212> DNA <213> Escherichia coli <400> 35 gagttataca cagggctggg atctattctt tttatctttt tttattcttt ctttattcta 60 taaattataa ccacttgaat ataaacaaaa aaaacacaca aaggtctagc ggaatttaca 120 gagggtctag cagaatttac aagttttcca gcaaaggtct agcagaattt acagataccc 180 acaactcaaa ggaaaaggac tagtaattat cattgactag ccc 223 <210> 36 <211> 951 <212> DNA <213> Escherichia coli <400> 36 atgtctgaat tagttgtttt caaagcaaat gaactagcga ttagtcgcta tgacttaacg 60 gagcatgaaa ccaagctaat tttatgctgt gtggcactac tcaaccccac gattgaaaac 120 cctacaatga aagaacggac ggtatcgttc acttataacc aatacgttca gatgatgaac 180 atcagtaggg aaaatgctta tggtgtatta gctaaagcaa ccagagagct gatgacgaga 240 actgtggaaa tcaggaatcc tttggttaaa ggctttgaga ttttccagtg gacaaactat 300 gccaagttct caagcgaaaa attagaatta gtttttagtg aagagatatt gccttatctt 360 ttccagttaa aaaaattcat aaaatataat ctggaacatg ttaagtcttt tgaaaacaaa 420 tactctatga ggatttatga gtggttatta aaagaactaa cacaaaagaa aactcacaag 480 gcaaatatag agattagcct tgatgaattt aagttcatgt taatgcttga aaataactac 540 catgagttta aaaggcttaa ccaatgggtt ttgaaaccaa taagtaaaga tttaaacact 600 tacagcaata tgaaattggt ggttgataag cgaggccgcc cgactgatac gttgattttc 660 caagttgaac tagatagaca aatggatctc gtaaccgaac ttgagaacaa ccagataaaa 720 atgaatggtg acaaaatacc aacaaccatt acatcagatt cctacctaca taacggacta 780 agaaaaacac tacacgatgc tttaactgca aaaattcagc tcaccagttt tgaggcaaaa 840 tttttgagtg acatgcaaag taagcatgat ctcaatggtt cgttctcatg gctcacgcaa 900 aaacaacgaa ccacactaga gaacatactg gctaaatacg gaaggatctg a 951 <210> 37 <211> 714 <212> DNA <213> Bacteriophage lambda <400> 37 tcagccaaac gtctcttcag gccactgact agcgataact ttccccacaa cggaacaact 60 ctcattgcat gggatcattg ggtactgtgg gtttagtggt tgtaaaaaca cctgaccgct 120 atccctgatc agtttcttga aggtaaactc atcaccccca agtctggcta tgcagaaatc 180 acctggctca acagcctgct cagggtcaac gagaattaac attccgtcag gaaagcttgg 240 cttggagcct gttggtgcgg tcatggaatt accttcaacc tcaagccaga atgcagaatc 300 actggctttt ttggttgtgc ttacccatct ctccgcatca cctttggtaa aggttctaag 360 cttaggtgag aacatccctg cctgaacatg agaaaaaaca gggtactcat actcacttct 420 aagtgacggc tgcatactaa ccgcttcata catctcgtag atttctctgg cgattgaagg 480 gctaaattct tcaacgctaa ctttgagaat ttttgtaagc aatgcggcgt tataagcatt 540 taatgcattg atgccattaa ataaagcacc aacgcctgac tgccccatcc ccatcttgtc 600 tgcgacagat tcctgggata agccaagttc atttttcttt ttttcataaa ttgctttaag 660 gcgacgtgcg tcctcaagct gctcttgtgt taatggtttc ttttttgtgc tcat 714 <210> 38 <211> 5097 <212> DNA <213> Artificial Sequence <220> <223> dCas9-PmCDA1 <400> 38 atggataaga aatactcaat aggcttagct atcggcacaa atagcgtcgg atgggcggtg 60 atcactgatg aatataaggt tccgtctaaa aagttcaagg ttctgggaaa tacagaccgc 120 cacagtatca aaaaaaatct tataggggct cttttatttg acagtggaga gacagcggaa 180 gcgactcgtc tcaaacggac agctcgtaga aggtatacac gtcggaagaa tcgtatttgt 240 tatctacagg agattttttc aaatgagatg gcgaaagtag atgatagttt ctttcatcga 300 cttgaagagt cttttttggt ggaagaagac aagaagcatg aacgtcatcc tatttttgga 360 aatatagtag atgaagttgc ttatcatgag aaatatccaa ctatctatca tctgcgaaaa 420 aaattggtag attctactga taaagcggat ttgcgcttaa tctatttggc cttagcgcat 480 atgattaagt ttcgtggtca ttttttgatt gagggagatt taaatcctga taatagtgat 540 gtggacaaac tatttatcca gttggtacaa acctacaatc aattatttga agaaaaccct 600 attaacgcaa gtggagtaga tgctaaagcg attctttctg cacgattgag taaatcaaga 660 cgattagaaa atctcattgc tcagctcccc ggtgagaaga aaaatggctt atttgggaat 720 ctcattgctt tgtcattggg tttgacccct aattttaaat caaattttga tttggcagaa 780 gatgctaaat tacagctttc aaaagatact tacgatgatg atttagataa tttattggcg 840 caaattggag atcaatatgc tgatttgttt ttggcagcta agaatttatc agatgctatt 900 ttactttcag atatcctaag agtaaatact gaaataacta aggctcccct atcagcttca 960 atgattaaac gctacgatga acatcatcaa gacttgactc ttttaaaagc tttagttcga 1020 caacaacttc cagaaaagta taaagaaatc ttttttgatc aatcaaaaaa cggatatgca 1080 ggttatattg atgggggagc tagccaagaa gaattttata aatttatcaa accaatttta 1140 gaaaaaatgg atggtactga ggaattattg gtgaaactaa atcgtgaaga tttgctgcgc 1200 aagcaacgga cctttgacaa cggctctatt ccccatcaaa ttcacttggg tgagctgcat 1260 gctattttga gaagacaaga agacttttat ccatttttaa aagacaatcg tgagaagatt 1320 gaaaaaatct tgacttttcg aattccttat tatgttggtc cattggcgcg tggcaatagt 1380 cgttttgcat ggatgactcg gaagtctgaa gaaacaatta ccccatggaa ttttgaagaa 1440 gttgtcgata aaggtgcttc agctcaatca tttattgaac gcatgacaaa ctttgataaa 1500 aatcttccaa atgaaaaagt actaccaaaa catagtttgc tttatgagta ttttacggtt 1560 tataacgaat tgacaaaggt caaatatgtt actgaaggaa tgcgaaaacc agcatttctt 1620 tcaggtgaac agaagaaagc cattgttgat ttactcttca aaacaaatcg aaaagtaacc 1680 gttaagcaat taaaagaaga ttatttcaaa aaaatagaat gttttgatag tgttgaaatt 1740 tcaggagttg aagatagatt taatgcttca ttaggtacct accatgattt gctaaaaatt 1800 attaaagata aagatttttt ggataatgaa gaaaatgaag atatcttaga ggatattgtt 1860 ttaacattga ccttatttga agatagggag atgattgagg aaagacttaa aacatatgct 1920 cacctctttg atgataaggt gatgaaacag cttaaacgtc gccgttatac tggttgggga 1980 cgtttgtctc gaaaattgat taatggtatt agggataagc aatctggcaa aacaatatta 2040 gattttttga aatcagatgg ttttgccaat cgcaatttta tgcagctgat ccatgatgat 2100 agtttgacat ttaaagaaga cattcaaaaa gcacaagtgt ctggacaagg cgatagttta 2160 catgaacata ttgcaaattt agctggtagc cctgctatta aaaaaggtat tttacagact 2220 gtaaaagttg ttgatgaatt ggtcaaagta atggggcggc ataagccaga aaatatcgtt 2280 attgaaatgg cacgtgaaaa tcagacaact caaaagggcc agaaaaattc gcgagagcgt 2340 atgaaacgaa tcgaagaagg tatcaaagaa ttaggaagtc agattcttaa agagcatcct 2400 gttgaaaata ctcaattgca aaatgaaaag ctctatctct attatctcca aaatggaaga 2460 gacatgtatg tggaccaaga attagatatt aatcgtttaa gtgattatga tgtcgatgcc 2520 attgttccac aaagtttcct taaagacgat tcaatagaca ataaggtctt aacgcgttct 2580 gataaaaatc gtggtaaatc ggataacgtt ccaagtgaag aagtagtcaa aaagatgaaa 2640 aactattgga gacaacttct aaacgccaag ttaatcactc aacgtaagtt tgataattta 2700 acgaaagctg aacgtggagg tttgagtgaa cttgataaag ctggttttat caaacgccaa 2760 ttggttgaaa ctcgccaaat cactaagcat gtggcacaaa ttttggatag tcgcatgaat 2820 actaaatacg atgaaaatga taaacttatt cgagaggtta aagtgattac cttaaaatct 2880 aaattagttt ctgacttccg aaaagatttc caattctata aagtacgtga gattaacaat 2940 taccatcatg cccatgatgc gtatctaaat gccgtcgttg gaactgcttt gattaagaaa 3000 tatccaaaac ttgaatcgga gtttgtctat ggtgattata aagtttatga tgttcgtaaa 3060 atgattgcta agtctgagca agaaataggc aaagcaaccg caaaatattt cttttactct 3120 aatatcatga acttcttcaa aacagaaatt acacttgcaa atggagagat tcgcaaacgc 3180 cctctaatcg aaactaatgg ggaaactgga gaaattgtct gggataaagg gcgagatttt 3240 gccacagtgc gcaaagtatt gtccatgccc caagtcaata ttgtcaagaa aacagaagta 3300 cagacaggcg gattctccaa ggagtcaatt ttaccaaaaa gaaattcgga caagcttatt 3360 gctcgtaaaa aagactggga tccaaaaaaa tatggtggtt ttgatagtcc aacggtagct 3420 tattcagtcc tagtggttgc taaggtggaa aaagggaaat cgaagaagtt aaaatccgtt 3480 aaagagttac tagggatcac aattatggaa agaagttcct ttgaaaaaaa tccgattgac 3540 tttttagaag ctaaaggata taaggaagtt aaaaaagact taatcattaa actacctaaa 3600 tatagtcttt ttgagttaga aaacggtcgt aaacggatgc tggctagtgc cggagaatta 3660 caaaaaggaa atgagctggc tctgccaagc aaatatgtga attttttata tttagctagt 3720 cattatgaaa agttgaaggg tagtccagaa gataacgaac aaaaacaatt gtttgtggag 3780 cagcataagc attatttaga tgagattatt gagcaaatca gtgaattttc taagcgtgtt 3840 attttagcag atgccaattt agataaagtt cttagtgcat ataacaaaca tagagacaaa 3900 ccaatacgtg aacaagcaga aaatattatt catttattta cgttgacgaa tcttggagct 3960 cccgctgctt ttaaatattt tgatacaaca attgatcgta aacgatatac gtctacaaaa 4020 gaagttttag atgccactct tatccatcaa tccatcactg gtctttatga aacacgcatt 4080 gatttgagtc agctaggagg tgacggtgga ggaggttctg gaggtggagg ttctgctgag 4140 tatgtgcgag ccctctttga ctttaatggg aatgatgaag aggatcttcc ctttaagaaa 4200 ggagacatcc tgagaatccg ggataagcct gaggagcagt ggtggaatgc agaggacagc 4260 gaaggaaaga gggggatgat tcctgtccct tacgtggaga agtattccgg agactataag 4320 gaccacgacg gagactacaa ggatcatgat attgattaca aagacgatga cgataagtct 4380 aggctcgagt ccggagacta taaggaccac gacggagact acaaggatca tgatattgat 4440 tacaaagacg atgacgataa gtctaggatg accgacgctg agtacgtgag aatccatgag 4500 aagttggaca tctacacgtt taagaaacag tttttcaaca acaaaaaatc cgtgtcgcat 4560 agatgctacg ttctctttga attaaaacga cggggtgaac gtagagcgtg tttttggggc 4620 tatgctgtga ataaaccaca gagcgggaca gaacgtggca ttcacgccga aatctttagc 4680 attagaaaag tcgaagaata cctgcgcgac aaccccggac aattcacgat aaattggtac 4740 tcatcctgga gtccttgtgc agattgcgct gaaaagatct tagaatggta taaccaggag 4800 ctgcggggga acggccacac tttgaaaatc tgggcttgca aactctatta cgagaaaaat 4860 gcgaggaatc aaattgggct gtggaacctc agagataacg gggttgggtt gaatgtaatg 4920 gtaagtgaac actaccaatg ttgcaggaaa atattcatcc aatcgtcgca caatcaattg 4980 aatgagaata gatggcttga gaagactttg aagcgagctg aaaaacgacg gagcgagttg 5040 tccattatga ttcaggtaaa aatactccac accactaaga gtcctgctgt tacttga 5097 <210> 39 <211> 105 <212> DNA <213> Escherichia coli <400> 39 acgttaaatc tatcaccgca agggataaat atctaacacc gtgcgtgttg actattttac 60 ctctggcggt gataatggtt gcagggccca ttttaggagg caaaa 105 <210> 40 <211> 247 <212> DNA <213> Artificial Sequence <220> <223> gRNA <400> 40 ggtttagcaa gatggcagcg cctaaatgta gaatgataaa aggattaaga gattaatttc 60 cctaaaaatg ataaaacaag cgttttgaaa gcgcttgttt ttttggtttg cagtcagagt 120 agaatagaag tatcaaaaaa agcaccgact cggtgccact ttttcaagtt gataacggac 180 tagccttatt ttaacttgct atttctagct ctaaaactga gaccatcccg ggtctctact 240 gcagaat 247 <210> 41 <211> 64 <212> DNA <213> Escherichia coli <400> 41 tatcaccgcc agtggtattt atgtcaacac cgccagagat aatttatcac cgcagatggt 60 tatc 64 <210> 42 <211> 10867 <212> DNA <213> Artificial Sequence <220> <223> Plasmid <400> 42 gtcggaactg actaaagtag tgagttatac acagggctgg gatctattct ttttatcttt 60 ttttattctt tctttattct ataaattata accacttgaa tataaacaaa aaaaacacac 120 aaaggtctag cggaatttac agagggtcta gcagaattta caagttttcc agcaaaggtc 180 tagcagaatt tacagatacc cacaactcaa aggaaaagga ctagtaatta tcattgacta 240 gcccatctca attggtatag tgattaaaat cacctagacc aattgagatg tatgtctgaa 300 ttagttgttt tcaaagcaaa tgaactagcg attagtcgct atgacttaac ggagcatgaa 360 accaagctaa ttttatgctg tgtggcacta ctcaacccca cgattgaaaa ccctacaatg 420 aaagaacgga cggtatcgtt cacttataac caatacgttc agatgatgaa catcagtagg 480 gaaaatgctt atggtgtatt agctaaagca accagagagc tgatgacgag aactgtggaa 540 atcaggaatc ctttggttaa aggctttgag attttccagt ggacaaacta tgccaagttc 600 tcaagcgaaa aattagaatt agtttttagt gaagagatat tgccttatct tttccagtta 660 aaaaaattca taaaatataa tctggaacat gttaagtctt ttgaaaacaa atactctatg 720 aggatttatg agtggttatt aaaagaacta acacaaaaga aaactcacaa ggcaaatata 780 gagattagcc ttgatgaatt taagttcatg ttaatgcttg aaaataacta ccatgagttt 840 aaaaggctta accaatgggt tttgaaacca ataagtaaag atttaaacac ttacagcaat 900 atgaaattgg tggttgataa gcgaggccgc ccgactgata cgttgatttt ccaagttgaa 960 ctagatagac aaatggatct cgtaaccgaa cttgagaaca accagataaa aatgaatggt 1020 gacaaaatac caacaaccat tacatcagat tcctacctac ataacggact aagaaaaaca 1080 ctacacgatg ctttaactgc aaaaattcag ctcaccagtt ttgaggcaaa atttttgagt 1140 gacatgcaaa gtaagcatga tctcaatggt tcgttctcat ggctcacgca aaaacaacga 1200 accacactag agaacatact ggctaaatac ggaaggatct gaggttctta tggctcttgt 1260 atctatcagt gaagcatcaa gactaacaaa caaaagtaga acaactgttc accgttacat 1320 atcaaaggga aaactgtcca tatgcacaga gataatctca tgaccaaaac cggtagctag 1380 aggggccgca ttaggcaccc caggctttac actttatgct tccggctcgt ataatgtgtg 1440 gattttgagt taggatccgg cgagattttc aggagctaag gaagctaaaa tggagaaaaa 1500 aatcactgga tataccaccg ttgatatatc ccaatggcat cgtaaagaac attttgaggc 1560 atttcagtca gttgctcaat gtacctataa ccagaccgtt cagctggata ttacggcctt 1620 tttaaagacc gtaaagaaaa ataagcacaa gttttatccg gcctttattc acattcttgc 1680 ccgcctgatg aatgctcatc cggaattccg tatggcaatg aaagacggtg agctggtgat 1740 atgggatagt gttcaccctt gttacaccgt tttccatgag caaactgaaa cgttttcatc 1800 gctctggagt gaataccacg acgatttccg gcagtttcta cacatatatt cgcaagatgt 1860 ggcgtgttac ggtgaaaacc tggcctattt ccctaaaggg tttattgaga atatgttttt 1920 cgtctcagcc aatccctggg tgagtttcac cagttttgat ttaaacgtgg ccaatatgga 1980 caacttcttc gcccccgttt tcaccatggg caaatattat acgcaaggcg acaaggtgct 2040 gatgccgctg gcgattcagg ttcatcatgc cgtctgtgat ggcttccatg tcggcagaat 2100 gcttaatgaa ttacaacagt actgcgatga gtggcagggc ggggcgtaaa cgcgtggatc 2160 cggcttacta aaagccagat aacagtatgc gtatttgcgc gctgattttt gcggtctaga 2220 ggtttagcaa gatggcagcg cctaaatgta gaatgataaa aggattaaga gattaatttc 2280 cctaaaaatg ataaaacaag cgttttgaaa gcgcttgttt ttttggtttg cagtcagagt 2340 agaatagaag tatcaaaaaa agcaccgact cggtgccact ttttcaagtt gataacggac 2400 tagccttatt ttaacttgct atttctagct ctaaaactga gaccatcccg ggtctctact 2460 gcagaattat caccgccagt ggtatttatg tcaacaccgc cagagataat ttatcaccgc 2520 agatggttat cgatgaagat tcttgctcaa ttgttatcag ctatgcgccg accagaacac 2580 cttgccgatc agccaaacgt ctcttcaggc cactgactag cgataacttt ccccacaacg 2640 gaacaactct cattgcatgg gatcattggg tactgtgggt ttagtggttg taaaaacacc 2700 tgaccgctat ccctgatcag tttcttgaag gtaaactcat cacccccaag tctggctatg 2760 cagaaatcac ctggctcaac agcctgctca gggtcaacga gaattaacat tccgtcagga 2820 aagcttggct tggagcctgt tggtgcggtc atggaattac cttcaacctc aagccagaat 2880 gcagaatcac tggctttttt ggttgtgctt acccatctct ccgcatcacc tttggtaaag 2940 gttctaagct taggtgagaa catccctgcc tgaacatgag aaaaaacagg gtactcatac 3000 tcacttctaa gtgacggctg catactaacc gcttcataca tctcgtagat ttctctggcg 3060 attgaagggc taaattcttc aacgctaact ttgagaattt ttgtaagcaa tgcggcgtta 3120 taagcattta atgcattgat gccattaaat aaagcaccaa cgcctgactg ccccatcccc 3180 atcttgtctg cgacagattc ctgggataag ccaagttcat ttttcttttt ttcataaatt 3240 gctttaaggc gacgtgcgtc ctcaagctgc tcttgtgtta atggtttctt ttttgtgctc 3300 atacgttaaa tctatcaccg caagggataa atatctaaca ccgtgcgtgt tgactatttt 3360 acctctggcg gtgataatgg ttgcagggcc cattttagga ggcaaaaatg gataagaaat 3420 actcaatagg cttagctatc ggcacaaata gcgtcggatg ggcggtgatc actgatgaat 3480 ataaggttcc gtctaaaaag ttcaaggttc tgggaaatac agaccgccac agtatcaaaa 3540 aaaatcttat aggggctctt ttatttgaca gtggagagac agcggaagcg actcgtctca 3600 aacggacagc tcgtagaagg tatacacgtc ggaagaatcg tatttgttat ctacaggaga 3660 ttttttcaaa tgagatggcg aaagtagatg atagtttctt tcatcgactt gaagagtctt 3720 ttttggtgga agaagacaag aagcatgaac gtcatcctat ttttggaaat atagtagatg 3780 aagttgctta tcatgagaaa tatccaacta tctatcatct gcgaaaaaaa ttggtagatt 3840 ctactgataa agcggatttg cgcttaatct atttggcctt agcgcatatg attaagtttc 3900 gtggtcattt tttgattgag ggagatttaa atcctgataa tagtgatgtg gacaaactat 3960 ttatccagtt ggtacaaacc tacaatcaat tatttgaaga aaaccctatt aacgcaagtg 4020 gagtagatgc taaagcgatt ctttctgcac gattgagtaa atcaagacga ttagaaaatc 4080 tcattgctca gctccccggt gagaagaaaa atggcttatt tgggaatctc attgctttgt 4140 cattgggttt gacccctaat tttaaatcaa attttgattt ggcagaagat gctaaattac 4200 agctttcaaa agatacttac gatgatgatt tagataattt attggcgcaa attggagatc 4260 aatatgctga tttgtttttg gcagctaaga atttatcaga tgctatttta ctttcagata 4320 tcctaagagt aaatactgaa ataactaagg ctcccctatc agcttcaatg attaaacgct 4380 acgatgaaca tcatcaagac ttgactcttt taaaagcttt agttcgacaa caacttccag 4440 aaaagtataa agaaatcttt tttgatcaat caaaaaacgg atatgcaggt tatattgatg 4500 ggggagctag ccaagaagaa ttttataaat ttatcaaacc aattttagaa aaaatggatg 4560 gtactgagga attattggtg aaactaaatc gtgaagattt gctgcgcaag caacggacct 4620 ttgacaacgg ctctattccc catcaaattc acttgggtga gctgcatgct attttgagaa 4680 gacaagaaga cttttatcca tttttaaaag acaatcgtga gaagattgaa aaaatcttga 4740 cttttcgaat tccttattat gttggtccat tggcgcgtgg caatagtcgt tttgcatgga 4800 tgactcggaa gtctgaagaa acaattaccc catggaattt tgaagaagtt gtcgataaag 4860 gtgcttcagc tcaatcattt attgaacgca tgacaaactt tgataaaaat cttccaaatg 4920 aaaaagtact accaaaacat agtttgcttt atgagtattt tacggtttat aacgaattga 4980 caaaggtcaa atatgttact gaaggaatgc gaaaaccagc atttctttca ggtgaacaga 5040 agaaagccat tgttgattta ctcttcaaaa caaatcgaaa agtaaccgtt aagcaattaa 5100 aagaagatta tttcaaaaaa atagaatgtt ttgatagtgt tgaaatttca ggagttgaag 5160 atagatttaa tgcttcatta ggtacctacc atgatttgct aaaaattatt aaagataaag 5220 attttttgga taatgaagaa aatgaagata tcttagagga tattgtttta acattgacct 5280 tatttgaaga tagggagatg attgaggaaa gacttaaaac atatgctcac ctctttgatg 5340 ataaggtgat gaaacagctt aaacgtcgcc gttatactgg ttggggacgt ttgtctcgaa 5400 aattgattaa tggtattagg gataagcaat ctggcaaaac aatattagat tttttgaaat 5460 cagatggttt tgccaatcgc aattttatgc agctgatcca tgatgatagt ttgacattta 5520 aagaagacat tcaaaaagca caagtgtctg gacaaggcga tagtttacat gaacatattg 5580 caaatttagc tggtagccct gctattaaaa aaggtatttt acagactgta aaagttgttg 5640 atgaattggt caaagtaatg gggcggcata agccagaaaa tatcgttatt gaaatggcac 5700 gtgaaaatca gacaactcaa aagggccaga aaaattcgcg agagcgtatg aaacgaatcg 5760 aagaaggtat caaagaatta ggaagtcaga ttcttaaaga gcatcctgtt gaaaatactc 5820 aattgcaaaa tgaaaagctc tatctctatt atctccaaaa tggaagagac atgtatgtgg 5880 accaagaatt agatattaat cgtttaagtg attatgatgt cgatgccatt gttccacaaa 5940 gtttccttaa agacgattca atagacaata aggtcttaac gcgttctgat aaaaatcgtg 6000 gtaaatcgga taacgttcca agtgaagaag tagtcaaaaa gatgaaaaac tattggagac 6060 aacttctaaa cgccaagtta atcactcaac gtaagtttga taatttaacg aaagctgaac 6120 gtggaggttt gagtgaactt gataaagctg gttttatcaa acgccaattg gttgaaactc 6180 gccaaatcac taagcatgtg gcacaaattt tggatagtcg catgaatact aaatacgatg 6240 aaaatgataa acttattcga gaggttaaag tgattacctt aaaatctaaa ttagtttctg 6300 acttccgaaa agatttccaa ttctataaag tacgtgagat taacaattac catcatgccc 6360 atgatgcgta tctaaatgcc gtcgttggaa ctgctttgat taagaaatat ccaaaacttg 6420 aatcggagtt tgtctatggt gattataaag tttatgatgt tcgtaaaatg attgctaagt 6480 ctgagcaaga aataggcaaa gcaaccgcaa aatatttctt ttactctaat atcatgaact 6540 tcttcaaaac agaaattaca cttgcaaatg gagagattcg caaacgccct ctaatcgaaa 6600 ctaatgggga aactggagaa attgtctggg ataaagggcg agattttgcc acagtgcgca 6660 aagtattgtc catgccccaa gtcaatattg tcaagaaaac agaagtacag acaggcggat 6720 tctccaagga gtcaatttta ccaaaaagaa attcggacaa gcttattgct cgtaaaaaag 6780 actgggatcc aaaaaaatat ggtggttttg atagtccaac ggtagcttat tcagtcctag 6840 tggttgctaa ggtggaaaaa gggaaatcga agaagttaaa atccgttaaa gagttactag 6900 ggatcacaat tatggaaaga agttcctttg aaaaaaatcc gattgacttt ttagaagcta 6960 aaggatataa ggaagttaaa aaagacttaa tcattaaact acctaaatat agtctttttg 7020 agttagaaaa cggtcgtaaa cggatgctgg ctagtgccgg agaattacaa aaaggaaatg 7080 agctggctct gccaagcaaa tatgtgaatt ttttatattt agctagtcat tatgaaaagt 7140 tgaagggtag tccagaagat aacgaacaaa aacaattgtt tgtggagcag cataagcatt 7200 atttagatga gattattgag caaatcagtg aattttctaa gcgtgttatt ttagcagatg 7260 ccaatttaga taaagttctt agtgcatata acaaacatag agacaaacca atacgtgaac 7320 aagcagaaaa tattattcat ttatttacgt tgacgaatct tggagctccc gctgctttta 7380 aatattttga tacaacaatt gatcgtaaac gatatacgtc tacaaaagaa gttttagatg 7440 ccactcttat ccatcaatcc atcactggtc tttatgaaac acgcattgat ttgagtcagc 7500 taggaggtga cggtggagga ggttctggag gtggaggttc tgctgagtat gtgcgagccc 7560 tctttgactt taatgggaat gatgaagagg atcttccctt taagaaagga gacatcctga 7620 gaatccggga taagcctgag gagcagtggt ggaatgcaga ggacagcgaa ggaaagaggg 7680 ggatgattcc tgtcccttac gtggagaagt attccggaga ctataaggac cacgacggag 7740 actacaagga tcatgatatt gattacaaag acgatgacga taagtctagg ctcgagtccg 7800 gagactataa ggaccacgac ggagactaca aggatcatga tattgattac aaagacgatg 7860 acgataagtc taggatgacc gacgctgagt acgtgagaat ccatgagaag ttggacatct 7920 acacgtttaa gaaacagttt ttcaacaaca aaaaatccgt gtcgcataga tgctacgttc 7980 tctttgaatt aaaacgacgg ggtgaacgta gagcgtgttt ttggggctat gctgtgaata 8040 aaccacagag cgggacagaa cgtggcattc acgccgaaat ctttagcatt agaaaagtcg 8100 aagaatacct gcgcgacaac cccggacaat tcacgataaa ttggtactca tcctggagtc 8160 cttgtgcaga ttgcgctgaa aagatcttag aatggtataa ccaggagctg cgggggaacg 8220 gccacacttt gaaaatctgg gcttgcaaac tctattacga gaaaaatgcg aggaatcaaa 8280 ttgggctgtg gaacctcaga gataacgggg ttgggttgaa tgtaatggta agtgaacact 8340 accaatgttg caggaaaata ttcatccaat cgtcgcacaa tcaattgaat gagaatagat 8400 ggcttgagaa gactttgaag cgagctgaaa aacgacggag cgagttgtcc attatgattc 8460 aggtaaaaat actccacacc actaagagtc ctgctgttac ttgacaggca tcaaataaaa 8520 cgaaaggctc agtcgaaaga ctgggccttt cgttttatct gttgtttgcg gccgggtacc 8580 gagctcgaat tcactggccg tcgttttaca acgtcgtgac tgggaaaacc ctggcgttac 8640 ccaacttaat cgccttgcag cacatccccc tttcgccagc tggcgtaata gcgaagaggc 8700 ccgcaccgat cgcccttccc aacagttgcg cagcctgaat ggcgaatggc gattcacaaa 8760 aaataggtac acgaaaaaca agttaaggga tgcagtttat gcatccctta acttacttat 8820 taaataattt atagctattg aaaagagata agaattgttc aaagctaata ttgtttaaat 8880 cgtcaattcc tgcatgtttt aaggaattgt taaattgatt ttttgtaaat attttcttgt 8940 attctttgtt aacccatttc ataacgaaat aattatactt ctgtttatct ttgtgtgata 9000 ttcttgattt ttttctattt aatctgataa gtgagctatt cactttaggt ttaggatgaa 9060 aatattctct tggaaccata cttaatatag aaatatcaac ttctgccatt aaaaataatg 9120 ccaatgagcg ttttgtattt aataatcttt tagcaaaccc gtattccacg attaaataaa 9180 tctcatcagc tatactatca aaaacaattt tgcgtattat atccgtactt atgttataag 9240 gtatattacc aaatatttta taggattggt ttttaggaaa tttaaactgc aatatatcct 9300 tgtttaaaac ttggaaatta tcgtgatcaa caagtttatt ttctgtagtt ttgcataatt 9360 tatggtctat ttcaatggca gttacgaaat tacacctctg tactaattca agggtaaaat 9420 gcccttttcc tgagccgatt tcaaagatat tatcatgttc atttaatctt atatttgtca 9480 ttattttatc tatattatgt tttgaagtaa taaagttttg actgtgtttt atatttttct 9540 cgttcattat aaccctcttt attttttcct ccttataaaa ttagtataat tatagcacga 9600 gctctgataa atatgaacat gatgagtgat cgttaaattt atactgcaat ctgatgcgat 9660 tattgaataa aagatatgag agatttatct agtttctttt tttacaagaa aaaagaaagt 9720 tcttaaaggt tttatacttt tggtcgtaga gcacacggtt taacgactta attacgaagt 9780 aaataagtct agtgtgttag actttaatgt ttttttaagg cattagtgca tttaagcgtc 9840 agagcatggc tttatgccga gaaaactatt ggttggaatg gcgtgtgtgt tagccaaagc 9900 tttggcgagt tggttggggg tttcatggga ttaatcccat gaaagtacca actcaacaac 9960 acactaacgc ctgttggttc caaccaatag gaaattggaa taagcaatta gtataatgag 10020 agtataatgt tggtataacg ttagtataat gatgcttttt ttcattatat tttttatgta 10080 ctttaaacct gcacgcttat gtgaattaga aaaagcttaa tcgcatttca tagattgacc 10140 tcccaataac tacgtggtgt tattgggagg tcaatctatt tcatttgcct cttgctcaaa 10200 gttcccaaat tcgagtaaga ggtatttttg tttttggtcg tcgcctctca ttagtagttc 10260 agggtttaac attaatactc cagtttttct ttttataata tttccttctt ctaagatttt 10320 aagtgttgtt attactgttt gtagacttgt tcctgtagct tttgctattt ctcttgttgt 10380 agctatcatt gtattgttac ttaagtggac attatctagg atatagttaa cgattttaag 10440 tttttttccg ccaatcatat ctaacatact tattaattgc actatatatg cctttacgaa 10500 gttaccagac gtttgtttac ggtataactt gtctacctct atgacttctc cactttcttc 10560 gtctatgagc ctctgagagc ctttatagac tgttccatat ctttctttca tctttttctc 10620 actccttatt ttaaactatt ctaactatat cataactgtt ctaaaaaaaa aagaacattt 10680 gttaaaagaa attagaacaa aatgagtgaa aaattagaac aaacaaattc cttataaacc 10740 ttatcatctc aacctatatt aagattttac ctagttgaat cttcttttct atataaagcg 10800 tcggagcata tcagggggtt atctaacgta aatgctaccc ttcggctcgc tttcgctcgg 10860 cattgac 10867

Claims

1. A method for modifying a target site of double-stranded DNA for non-disease treatment purposes, which includes: A step of contacting a complex formed by linking a nucleic acid base converting enzyme and two or more nucleic acid sequence recognition modules that specifically bind to different target nucleotide sequences in a selected double-stranded DNA, with the double-stranded DNA, without cleaving at least one strand of the double-stranded DNA at the target site, but causing deletion of one or more nucleotides at the target site or conversion of one or more nucleotides to other one or more nucleotides, or inserting one or more nucleotides into the target site; wherein, the nucleic acid sequence recognition module is at least one CRISPR-Cas system with inactivated DNA cleavage ability of Cas; wherein, the nucleic acid base converting enzyme is a deaminase, and wherein, the Cas is Cas9 with the Asp residue at position 10 converted to an Ala residue and the His residue at position 840 converted to an Ala residue.

2. A method for modifying a target site of double-stranded DNA for non-disease treatment purposes, which includes: A step of contacting a complex formed by linking a nucleic acid base converting enzyme and a nucleic acid sequence recognition module that specifically binds to a target nucleotide sequence in a selected double-stranded DNA, with the double-stranded DNA, without cleaving at least one strand of the double-stranded DNA at the target site, but causing deletion of one or more nucleotides at the target site or conversion of one or more nucleotides to other one or more nucleotides, or inserting one or more nucleotides into the target site; wherein, the contact between the double-stranded DNA and the complex is carried out by introducing a nucleic acid encoding the complex into a cell having the double-stranded DNA; wherein, the nucleic acid sequence recognition module is at least one CRISPR-Cas system with inactivated DNA cleavage ability of Cas; wherein, the nucleic acid base converting enzyme is a deaminase, and wherein, the Cas is Cas9 with the Asp residue at position 10 converted to an Ala residue and / or the His residue at position 840 converted to an Ala residue.

3. The method according to claim 2, wherein, Use two or more nucleic acid sequence recognition modules that specifically bind to different target nucleotide sequences respectively.

4. The method according to claim 1 or 3, wherein, The different target nucleotide sequences are present in different genes.

5. The method according to claim 1 or 2, wherein, The deaminase is AID (AICDA).

6. The method according to claim 1, wherein, The contact between the double-stranded DNA and the complex is carried out by introducing a nucleic acid encoding the complex into a cell having the double-stranded DNA.

7. The method according to claim 2 or 6, wherein, The cell is a prokaryotic cell.

8. The method according to claim 2 or 6, wherein, The cell is a eukaryotic cell.

9. The method according to claim 2 or 6, wherein, The cell is a microbial cell.

10. The method according to claim 2 or 6, wherein, The cell is a plant cell.

11. The method according to claim 2 or 6, wherein, The cell is an insect cell.

12. The method according to claim 2 or 6, wherein, The cell is an animal cell.

13. The method according to claim 2 or 6, wherein, The cell is a vertebrate cell.

14. The method according to claim 2 or 6, wherein, The cell is a mammalian cell.

15. The method according to claim 7, wherein, The cell is a polyploid cell, and modifies sites within all targeted alleles on homologous chromosomes.

16. The method according to claim 2 or 6, which includes: A step of introducing an expression vector containing a nucleic acid encoding the complex in a form capable of controlling the expression period into the cell, and a step of inducing the expression of the nucleic acid at a time when it is necessary to stabilize the modification of the target site of the double-stranded DNA.

17. The method according to claim 2 or 6, wherein, The target nucleotide sequence in the double-stranded DNA is present in a gene essential for the cell.

18. A nucleic acid modifying enzyme complex, which is composed of a nucleic acid sequence recognition module that specifically binds to a target nucleotide sequence in double-stranded DNA and a nucleic acid base conversion enzyme. In the target site, at least one strand of the double-stranded DNA is not cleaved, and one or more nucleotides at the target site are deleted or converted into one or more other nucleotides, or one or more nucleotides are inserted into the target site, wherein, The nucleic acid sequence recognition module is at least one CRISPR-Cas system with inactivated DNA cleavage ability of Cas; wherein, the nucleic acid base converting enzyme is a deaminase, and Among them, the Cas is Cas9 in which the Asp residue at the 10th position is changed to an Ala residue and / or the His residue at the 840th position is changed to an Ala residue.

19. A nucleic acid encoding the nucleic acid modifying enzyme complex according to claim 18.

20. Use of a nucleic acid sequence recognition module in the preparation of a reagent for modifying a target site of double-stranded DNA, wherein the target site of double-stranded DNA is modified by a method comprising the following steps: contacting a complex formed by linking a nucleic acid base conversion enzyme and two or more nucleic acid sequence recognition modules that specifically bind to different target nucleotide sequences in a selected double-stranded DNA with the double-stranded DNA, without cleaving at least one strand of the double-stranded DNA at the target site, and deleting one or more nucleotides at the target site or converting them into one or more other nucleotides, or inserting one or more nucleotides into the target site; wherein, The nucleic acid sequence recognition module is a CRISPR-Cas system in which at least one DNA cleavage ability of Cas is inactivated; Among them, the nucleic acid base conversion enzyme is a deaminase, and Among them, the Cas is Cas9 in which the Asp residue at the 10th position is changed to an Ala residue and the His residue at the 840th position is changed to an Ala residue.

21. Use of a nucleic acid sequence recognition module in the preparation of a reagent for modifying a target site of double-stranded DNA, wherein the target site of double-stranded DNA is modified by a method comprising the following steps: contacting a complex formed by linking a nucleic acid base conversion enzyme and a nucleic acid sequence recognition module that specifically binds to a target nucleotide sequence in a selected double-stranded DNA with the double-stranded DNA, without cleaving at least one strand of the double-stranded DNA at the target site, and deleting one or more nucleotides at the target site or converting them into one or more other nucleotides, or inserting one or more nucleotides into the target site; wherein, The contact between the double-stranded DNA and the complex is carried out by introducing a nucleic acid encoding the complex into a cell having the double-stranded DNA; The nucleic acid sequence recognition module is a CRISPR-Cas system in which at least one DNA cleavage ability of Cas is inactivated; Among them, the nucleic acid base conversion enzyme is a deaminase, and Among them, the Cas is Cas9 in which the Asp residue at the 10th position is changed to an Ala residue and / or the His residue at the 840th position is changed to an Ala residue.

22. The use according to claim 21, wherein, Use more than two nucleic acid sequence recognition modules that specifically bind to different target nucleotide sequences respectively.

23. The use according to claim 20 or 22, wherein, The different target nucleotide sequences are present in different genes.

24. The use according to claim 20 or 21, wherein The deaminase is AID (AICDA).

25. The use according to claim 20, wherein The contact between the double-stranded DNA and the complex is carried out by introducing a nucleic acid encoding the complex into a cell having the double-stranded DNA.

26. The use according to claim 21 or 25, wherein The cell is a prokaryotic cell.

27. The use according to claim 21 or 25, wherein The cell is a eukaryotic cell.

28. The use according to claim 21 or 25, wherein The cell is a microbial cell.

29. The use according to claim 21 or 25, wherein The cell is a plant cell.

30. The use according to claim 21 or 25, wherein The cell is an insect cell.

31. The use according to claim 21 or 25, wherein The cell is an animal cell.

32. The use according to claim 21 or 25, wherein The cell is a vertebrate cell.

33. The use according to claim 21 or 25, wherein The cell is a mammalian cell.

34. The use according to claim 26, wherein The cell is a polyploid cell, and modifies sites within all targeted alleles on homologous chromosomes.

35. The use according to claim 21 or 25, comprising: The step of introducing an expression vector containing a nucleic acid encoding the complex in a form capable of controlling the expression period into the cell, and the step of inducing the expression of the nucleic acid at a time when it is necessary to stabilize the modification of the targeted site of the double-stranded DNA.

36. The use according to claim 21 or 25, wherein The target nucleotide sequence in the double-stranded DNA is present in a gene essential for the cell.

Citation Information

Patent Citations

  • JP1974068498A

  • Hydraulic system

    JP1997002301A

  • Cultures with improved phage resistance

    JP2010519929A

  • Method for modifying RNA-binding protein using PPR motif

    JP2013128413A

  • DNA modification mediated by TAL effectors

    JP2013513389A