Methods for modifying genomic sequences that specifically target nucleic acid bases in DNA sequences, and the molecular complexes used therein.
By employing a base conversion method that does not cut double-stranded DNA, and using the CRISPR-Cas system and deaminase for genome editing, the cytotoxicity problem caused by DNA double-strand cutting in existing technologies has been solved, achieving safe and efficient gene modification applicable to a wide range of biological species.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2015-03-04
- Publication Date
- 2026-03-13
AI Technical Summary
Existing genome editing technologies typically rely on DNA double-strand cutting, which leads to side effects such as strong cytotoxicity and chromosomal rearrangement, making it difficult to achieve reliable gene modification in gene therapy and single-celled microorganisms.
A base conversion method that does not cut double-stranded DNA is employed, in which a deaminase catalyzing the deamination reaction is linked to a molecule with DNA sequence recognition capabilities, and the genome sequence is modified using the CRISPR-Cas system. Specifically, this involves using an RNA molecule linked with an inactivated Cas protein and a deaminase gene, which is then introduced into the host cell for nucleic acid base conversion.
It achieves safe and efficient genome editing, avoiding the side effects of foreign DNA insertion and double-strand DNA cutting, and can introduce mutations into specific regions in a wide range of biological species, improving the efficiency and diversity of mutation introduction.
Smart Images

Figure CN111500569B_ABST
Abstract
Description
[0001] This application is a divisional application of the invention with application number CN201580023875.6, application date March 4, 2015, entitled "Method for modifying a genomic sequence for specifically changing nucleic acid bases of a target DNA sequence, and a molecular complex thereof". Technical Field
[0002] This invention relates to a method for modifying genome sequences and a complex of a nucleic acid sequence recognition module and a nucleic acid base-converting enzyme used therein. The method for modifying genome sequences can modify nucleic acid bases in specific regions of the genome without cleaving the double strand of DNA (no cleavage or cleavage of one strand) or inserting exogenous DNA fragments. Background Technology
[0003] In recent years, genome editing has attracted attention as a technology for modifying genes and genomic regions of interest in various biological species. Previously, as a method for genome editing, a method using artificial nucleases that combine molecules with sequence-independent DNA cutting capabilities with molecules with sequence recognition capabilities has been proposed (Non-Patent Literature 1).
[0004] Existing reports include, for example, a method for recombination of target loci in DNA in plant or insect cells as hosts using zinc finger nucleases (ZFNs) consisting of zinc finger DNA-binding domains linked to non-specific DNA-cutting domains (Patent Document 1); a method for cutting and modifying target genes within or adjacent to specific nucleotide sequences using TALENs consisting of transcription activator-like (TAL) effectors (a DNA-binding module of plant pathogens such as Xanthomonas) linked to DNA endonucleases (Patent Document 2); or a method using the CRISPR-Cas9 system, which consists of CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats), a DNA sequence that functions in the acquired immune system of eubacteria and archaea, and the CRISPR-associated nuclease protein family, which plays an important role in conjunction with CRISPR (Patent Document 3). Furthermore, a method for cleaving target genes near a specific sequence using an artificial nuclease formed by linking a PPR protein with a nuclease has also been reported (Patent Document 4), wherein the PPR protein is constructed such that a specific nucleotide sequence can be recognized by using a tandem PPR motif comprising 35 amino acids that recognizes one nucleic acid base.
[0005] Existing technical documents
[0006] Patent documents
[0007] Patent Document 1: Japanese Patent No. 4968498
[0008] Patent Document 2: Japanese Patent Publication No. 2013-513389
[0009] Patent Document 3: Japanese Patent Publication No. 2010-519929
[0010] Patent Document 4: Japanese Patent Application Publication No. 2013-128413
[0011] Non-patent literature
[0012] Non-patent literature 1: Kelvin M Esvelt, Harris H Wang (2013) Genome-scale engineering for systems and synthetic biology, Molecular Systems Biology 9:641 Summary of the Invention
[0013] The problem the invention aims to solve
[0014] To date, genome editing technologies have been based primarily on double-DNA breaks (DSB). However, due to the unintended genome modifications involved, they have side effects such as strong cytotoxicity and chromosomal rearrangements, and share common problems such as: impaired reliability in gene therapy, very low survival rate of cells modified with nucleotides, and difficulty in gene modification itself in primate oocytes and single-celled microorganisms.
[0015] Therefore, the purpose of this invention is to provide a novel genome editing method that modifies a specific sequence of nucleic acid bases of a gene without DSB or the insertion of exogenous DNA fragments, i.e. without cutting double-stranded DNA or cutting one strand, as well as a nucleic acid sequence recognition module and a complex of nucleic acid base conversion enzymes for use therein.
[0016] Problem Solving Methods
[0017] To address the aforementioned issues, the inventors conducted extensive research and conceived a base transition method that does not involve DSB and employs a DNA base transition reaction. Although the base transition reaction itself, based on the deamination of DNA bases, is known, it has not yet been realized to use it to identify specific sequences in DNA and target arbitrary sites, and to specifically modify the targeted DNA through base transitions.
[0018] Therefore, by using deaminases that catalyze deamination reactions as enzymes to carry out such nucleic acid base transitions and linking them to molecules with DNA sequence recognition capabilities, genomic sequence modifications based on nucleic acid base transitions can be performed in regions containing specific DNA sequences.
[0019] Specifically, the CRISPR-Cas system (CRISPR-mutant Cas) was used. That is, DNA was created to encode RNA molecules containing a genome-specific CRISPR-RNA:crRNA (gRNA) that encodes RNA (trans-activating crRNA:tracrRNA) used to recruit Cas proteins, linked to a sequence complementary to the target sequence of the gene to be modified. On the other hand, DNA was created by linking DNA (dCas) encoding a mutant Cas protein whose cleavage ability of one or both strands of the double-stranded DNA was inactivated to a deaminase gene. These DNAs were introduced into host yeast cells containing the gene to be modified. As a result, mutations were successfully introduced randomly within a few hundred nucleotides of the gene of interest, including the target sequence. The mutation introduction efficiency was increased when using a mutant Cas protein that cleaves either strand of the double-stranded DNA compared to using a double-mutant Cas protein that does not cleave either strand of the double-stranded DNA. Furthermore, it was determined that the area and diversity of the mutated region changed depending on which strand of the DNA double strand was cleaved. Moreover, mutations could be effectively introduced by targeting multiple regions of the gene of interest. In other words, it was confirmed that inoculating host cells with the introduced DNA into a non-selective culture medium and examining the sequence of the gene of interest in randomly selected colonies resulted in mutations being introduced into almost all colonies. Furthermore, it was confirmed that by targeting regions present in two or more genes of interest, genome editing at multiple sites can be performed simultaneously. This method further demonstrates that it can simultaneously introduce mutations into alleles of diploid or polyploid genomes; it can introduce mutations not only into eukaryotic cells but also into prokaryotic cells such as *E. coli*; and it has broad applicability regardless of the biological species. Additionally, by transiently performing a nucleic acid base conversion reaction at the desired time, the editing of essential genes, which has been relatively inefficient until now, can be effectively performed.
[0020] Based on these insights, the inventors conducted further and repeated research, thereby completing this invention.
[0021] That is, the present invention is as follows.
[0022] [1] A method for modifying a target site of double-stranded DNA, comprising: contacting a complex consisting of a nucleic acid base-converting enzyme and a nucleic acid sequence recognition module that specifically binds to a target nucleotide sequence in a selected double-stranded DNA with the double-stranded DNA, wherein the double-stranded DNA is not cleaved at least one strand of the double-stranded DNA at the target site, but one or more nucleotides at the target site are deleted or transformed into one or more other nucleotides, or one or more nucleotides are inserted into the target site.
[0023] [2] According to the method described in [1], the nucleic acid sequence recognition module is selected from: at least one CRISPR-Cas system with inactivated DNA cutting ability of Cas, zinc finger motif, TAL effector and PPR motif.
[0024] [3] According to the method described in [1], wherein the nucleic acid sequence recognition module is a CRISPR-Cas system in which at least one DNA cutting ability of Cas has been inactivated.
[0025] [4] The method according to any one of [1] to [3], wherein two or more nucleic acid sequence recognition modules that specifically bind to different target nucleotide sequences are used.
[0026] [5] According to the method described in [4], different target nucleotide sequences are present in different genes.
[0027] [6] The method according to any one of [1] to [5], wherein the nucleic acid base conversion enzyme is a deaminase.
[0028] [7] According to the method described in [6], wherein the deaminase is AID (AICDA).
[0029] [8] The method according to any one of [1] to [7], wherein the contact between the double-stranded DNA and the complex is carried out by introducing a nucleic acid encoding the complex into a cell having the double-stranded DNA.
[0030] [9] According to the method described in [8], the cell is a prokaryotic cell.
[0031]
[10] The method described in [8] is a eukaryotic cell.
[0032]
[11] According to the method described in [8], the cell is a microbial cell.
[0033]
[12] The method described in [8] is wherein the cell is a plant cell.
[0034]
[13] The method described in [8] is wherein the cell is an insect cell.
[0035]
[14] The method according to [8], wherein the cell is an animal cell.
[0036]
[15] The method according to [8], wherein the cell is a vertebrate cell.
[0037]
[16] The method according to [8], wherein the cell is a mammalian cell.
[0038]
[17] The method according to any one of [9] to
[16] , wherein the cell is a polyploid cell and the sites within all targeted alleles on the homologous chromosome are modified.
[0039]
[18] The method according to any one of [8] to
[17] includes the steps of introducing an expression vector containing a nucleic acid encoding a complex in a form capable of controlling the expression period into a cell, and inducing the expression of the nucleic acid at a period in which the modification of the target site of the double-stranded DNA needs to be stabilized.
[0040]
[19] According to the method of
[18] , wherein the target nucleotide sequence in the double-stranded DNA is present in a gene essential to the cell.
[0041]
[20] A nucleic acid modifying enzyme complex, which is composed of a nucleic acid sequence recognition module that specifically binds to a target nucleotide sequence in double-stranded DNA and a nucleic acid base conversion enzyme, which, at the target site, does not cleave at least one strand of the double-stranded DNA, causes the deletion or conversion of one or more other nucleotides at the target site, or inserts one or more nucleotides into the target site.
[0042]
[21] The nucleic acid encoding the nucleic acid-modifying enzyme complex described in
[20] .
[0043] The effects of the invention
[0044] The genome editing technology of the present invention offers excellent safety because it is independent of the insertion of foreign DNA or the cutting of double-stranded DNA. Even in cases where conventional methods involve gene recombination that are biologically or legally controversial, the technology may provide a solution. Furthermore, the theoretically broad range of mutation introduction, from a single base pin point to hundreds of bases, can be applied to localized evolutionary induction by randomly introducing mutations into specific, defined regions that have been virtually impossible to perform until now. Attached Figure Description
[0045] [ Figure 1 This illustration shows a schematic depiction of the mechanism of the gene modification method of the present invention using a CRISPR-Cas system.
[0046] [ Figure 2 The results show the effectiveness of the gene modification method of the present invention, which uses budding yeast to verify the combination of a CRISPR-Cas system and a PmCDA1 deaminase from the petrmyzon marinus.
[0047] [ Figure 3 The graph shows the changes in cell viability after expression induction using a CRISPR-Cas9 system with a D10A mutant of Cas9 containing nickase activity, in combination with deaminase and PmCDA1 (nCas9D10A-PmCDA1), and in the case of using a conventional Cas9 system with DNA double-strand cleavage capability.
[0048] [ Figure 4 The figure shows the results of constructing a multiplex expression construct using human-derived AID deaminase, which is linked to dCas9 via the SH3 domain and its binding ligand. The expression construct was then introduced into budding yeast along with two gRNAs (targeting sequences target4 and target5).
[0049] [ Figure 5 The figure shows the improved mutation introduction efficiency by using Cas9 to cut any single strand of DNA.
[0050] [ Figure 6 The diagram shows how the size and frequency of the mutation-introduced region change depending on which single strand is cut, without cutting the double-stranded DNA.
[0051] [ Figure 7 The figure shows that extremely high mutation introduction efficiency can be achieved by targeting two neighboring regions.
[0052] [ Figure 8 The diagram illustrates that the gene modification method of the present invention does not require marker-based selection. Mutations were introduced into all colonies whose sequences were determined.
[0053] [ Figure 9This diagram illustrates the simultaneous editing of multiple sites in the genome using the gene modification method according to the present invention. The upper diagram shows the nucleotide and amino acid sequences of the target sites of each gene, with arrows above the nucleotide sequences indicating the target nucleotide sequences. The numbers at the tails or arrowheads indicate the positions of the target nucleotide sequence ends on the ORF. The lower diagram shows the sequencing results of the target sites in each of the five clones from red (R) and white (W) colonies. Base transitions are shown in the nucleotides represented by letters in outline font. It should be noted that the reactivity to canavalialine (Canavalialine) is... R In this context, R indicates resistance and S indicates sensitivity.
[0054] [ Figure 10 The diagram shows how, according to the gene modification method of the present invention, mutations can be introduced simultaneously into two alleles on homologous chromosomes of a diploid genome. Figure 10 A shows the respective homologous mutation introduction efficiencies for the ade1 gene (above) and the can1 gene. Figure 10 B shows that a homologous mutation was actually introduced in the red colonies (bottom image). Additionally, heterologous mutations were also shown to have occurred in the white colonies (top image).
[0055] [ Figure 11 [Image showing a diagram illustrating genome editing of Escherichia coli, a prokaryotic cell, according to the gene modification method of the present invention.] Figure 11 A is a schematic illustration showing the plasmid used. Figure 11 B shows that the region within the galK gene can be used as a target to effectively introduce mutations (CAA→TAA). Figure 11 Figure C shows the sequencing results of two clones from each of the following media: one containing non-selective medium (none), one containing 25 μg / ml rifampicin (Rif 25), and one containing 50 μg / ml rifampicin (Rif 50). The introduction of a rifampicin-resistant mutation was confirmed (top figure). The estimated frequency of rifampicin-resistant strains is approximately 10% (bottom figure).
[0056] [ Figure 12 This diagram shows how the edited base sites are regulated based on the length of the guide RNA. Figure 12 A conceptual diagram showing the edited base sites for both 20-base and 24-base target nucleotide sequences. Figure 12B shows the results of editing the gsiA gene by altering the length of the target nucleotide sequence. Mutation sites are in bold; "T" and "A" indicate complete insertion of the mutation (C→T or G→A) into the clone; "t" indicates insertion of the mutation (C→T) into the clone at a rate of more than 50% (incomplete cloning); "c" indicates insertion efficiency of the mutation (C→T) into the clone is less than 50%.
[0057] [ Figure 13 [Illustrative illustration of the temperature-sensitive plasmid for mutation introduction used in Example 11]
[0058] [ Figure 14 [A diagram showing the mutation introduction scheme in Example 11.]
[0059] [ Figure 15 [Image showing the results of introducing mutations into the rpoB gene in Example 11]
[0060] [ Figure 16 [Image showing the results of introducing mutations into the galK gene in Example 11.] Detailed Implementation
[0061] This invention provides a method for modifying a target site of a double-stranded DNA without cleaving at least one strand of the DNA, by converting a target nucleotide sequence and nearby nucleotides in the double-stranded DNA into other nucleotides. The method includes a step of contacting the double-stranded DNA with a complex consisting of a base-converting enzyme and a nucleic acid sequence recognition module capable of specifically binding to the target nucleotide sequence in the double-stranded DNA, thereby converting the target site, i.e., the target nucleotide sequence and nearby nucleotides, into other nucleotides.
[0062] In this invention, "modification" of double-stranded DNA refers to the deletion or conversion of a certain nucleotide (e.g., dC) on the DNA strand into other nucleotides (e.g., dT, dA, or dG), or the insertion of nucleotides or nucleotide sequences between nucleotides on the DNA strand. Here, there is no particular limitation on the double-stranded DNA to be modified, but genomic DNA is preferred. Furthermore, the "target site" of the double-stranded DNA refers to all or part of the "target nucleotide sequence" that the nucleic acid sequence recognition module can specifically recognize and bind to, or to its vicinity (any one or both of the 5' upstream and 3' downstream sites). Depending on the purpose, this range can be appropriately adjusted between 1 base and several hundred bases in length.
[0063] In this invention, a "nucleic acid sequence recognition module" refers to a molecule or molecular complex that has the ability to specifically recognize and bind to a specific nucleotide sequence (i.e., target nucleotide sequence) on a DNA strand. The nucleic acid sequence recognition module can bind to the target nucleotide sequence, enabling the nucleic acid base-converting enzyme linked to the module to exert a specific effect at the target site of the double-stranded DNA.
[0064] In this invention, "nucleic acid base conversion enzyme" refers to an enzyme that can convert target nucleotides into other nucleotides by catalyzing the reaction of substituents on the purine or pyrimidine ring of DNA bases into other groups or atoms without cutting the DNA chain.
[0065] In this invention, a "nucleic acid modifying enzyme complex" refers to a complex comprising the aforementioned nucleic acid sequence recognition module and a nucleic acid base conversion enzyme, wherein the complex possesses nucleic acid base conversion enzyme activity and is endowed with specific nucleotide sequence recognition capabilities. The term "complex" used in this application includes not only forms composed of multiple molecules but also single molecules, such as fusion proteins, that possess both a nucleic acid sequence recognition module and a nucleic acid base conversion enzyme.
[0066] The nucleic acid base-converting enzyme used in this invention is not particularly limited as long as it can catalyze the above-described reaction. Examples include deaminases belonging to the nucleic acid / nucleotide deaminase superfamily that catalyze deamination reactions that convert amino groups to carbonyl groups. Preferred nucleic acid base-converting enzymes include cytidine deaminases that can convert cytosine or 5-methylcytosine to uracil or thymine, respectively; adenosine deaminases that can convert adenine to hypoxanthine; and guanine deaminases that can convert guanine to xanthine. More preferably, cytidine deaminases are activation-induced cytidine deaminases (hereinafter also referred to as AIDs) that are mutated by introducing a mutant enzyme into an immunoglobulin gene in acquired immunity in vertebrates.
[0067] There are no particular restrictions on the source of the nucleic acid base-converting enzyme; for example, PmCDA1 (Petromyzon marinus cytosine deaminase 1) from lampreys and AID (Activation-induced cytidine deaminase; AICDA) from mammals (e.g., humans, pigs, cattle, horses, monkeys, etc.) can be used. The base and amino acid sequences of the CDS of PmCDA1 are shown in sequence numbers 1 and 2, respectively, and the base and amino acid sequences of the CDS of human AID are shown in sequence numbers 3 and 4, respectively.
[0068] For the target nucleotide sequence in double-stranded DNA that is recognized by the nucleic acid sequence recognition module of the nucleic acid modifying enzyme complex of the present invention, there are no particular limitations as long as it can specifically bind to the module, and it can be any sequence in double-stranded DNA. Regarding the length of the target nucleotide sequence, it is sufficient to allow specific binding to the nucleic acid sequence recognition module; for example, when introducing mutations into specific sites in the genomic DNA of mammals, it is 12 nucleotides or more, preferably 15 nucleotides or more, and more preferably 17 nucleotides or more, depending on the genome size. There are no particular limitations on the upper limit of length, but it is preferably 25 nucleotides or less, and more preferably 22 nucleotides or less.
[0069] As the nucleic acid sequence recognition module of the nucleic acid modification enzyme complex of the present invention, it can use, for example, a CRISPR-Cas system (CRISPR-mutant Cas) with at least one DNA cleavage capability of Cas inactivated, zinc finger motifs, TAL effectors, and PPR motifs, etc., or fragments containing DNA-binding domains of proteins that can specifically bind to DNA, such as restriction enzymes, transcription factors, and RNA polymerases, but without DNA double-strand cleavage capability, etc., but are not limited to these. The module preferably includes: CRISPR-mutant Cas, zinc finger motifs, TAL effectors, PPR motifs, etc.
[0070] Zinc finger motifs are composed of 3–6 different Cys2His2 zinc finger units (each finger recognizes approximately 3 bases) linked together, and can recognize target nucleotide sequences of 9–18 bases. Zinc finger motifs can be fabricated using known methods such as the Modular assembly method (Nat Biotechnol (2002) 20:135-141), the OPEN method (Mol Cell (2008) 31:294-301), the CoDA method (Nat Methods (2011) 8:67-69), and the E. coli one-hybrid method (Nat Biotechnol (2008) 26:695-701). For details on the fabrication of zinc finger motifs, please refer to the aforementioned patent document 1.
[0071] For TAL effectors, there is a repeating structure with modules of approximately 34 amino acids as units. The binding stability and base specificity are determined using the 12th and 13th amino acid residues of each module (called RVD). Due to the high independence of each module, TAL effectors specific to the target nucleotide sequence can be created by simply connecting the modules together. TAL effectors can be constructed using open resource methods (REAL method (Curr Protoc Mol Biol (2012) Chapter 12: Unit 12.15), FLASH method (Nat Biotechnol (2012) 30: 460-465), Golden Gate method (Nucleic Acids Res (2011) 39: e82), etc.), making it relatively easy to design TAL effectors relative to the target nucleotide sequence. For details on the fabrication of TAL effectors, please refer to the aforementioned patent document 2.
[0072] For PPR motifs, the design is configured such that a specific nucleotide sequence is recognized by a single tandem 35-amino acid base containing the PPR motif, utilizing only the first, fourth, and penultimate amino acids of each motif to recognize the target base. Since it is independent of motif conformation and unaffected by interference from flanking motifs, similar to TAL effectors, PPR proteins specific to the target nucleotide sequence can be created by simply linking PPR motifs together. For details on the fabrication of PPR motifs, please refer to the aforementioned Patent Document 4.
[0073] In addition, when using fragments such as restriction enzymes, transcription factors, and RNA polymerases, since the DNA-binding domains of their proteins are well-known, fragments containing said domains and without DNA double-strand cutting ability can be easily designed and constructed.
[0074] The aforementioned nucleic acid sequence recognition module can also be provided as a fusion protein of the aforementioned nucleic acid base conversion enzyme. Alternatively, protein-binding domains such as the SH3 domain, PDZ domain, GK domain, and GB domain, along with their binding ligands, can be fused to the nucleic acid sequence recognition module and the nucleic acid base conversion enzyme, respectively, and provided as a protein complex through the interaction of these domains and their binding ligands. Alternatively, the nucleic acid sequence recognition module can be fused to the nucleic acid base conversion enzyme and to an intein, respectively, and the two can be linked by conjugation after the synthesis of each protein.
[0075] Regarding the contact between the nucleic acid-modifying enzyme complex of the present invention, which comprises a nucleic acid sequence recognition module and a nucleic acid base conversion enzyme (including fusion proteins), and double-stranded DNA, although it can be carried out in the form of an enzymatic reaction in a cell-free system, for the main purpose of the present invention, it is necessary to carry out contact by introducing the nucleic acid encoding the complex into a cell having double-stranded DNA of interest (e.g., genomic DNA).
[0076] Therefore, for the nucleic acid sequence recognition module and the nucleic acid base conversion enzyme, it is preferable to prepare them in the form of nucleic acids encoding their fusion proteins, or in the form of nucleic acids encoding them separately, such that after translation into proteins using binding domains, integrins, etc., they can form complexes within the host cell. Here, the nucleic acid can be either DNA or RNA. When it is DNA, double-stranded DNA is preferred, and it is provided in the form of an expression vector configured within the host cell under the control of a functional promoter. When it is RNA, single-stranded RNA is preferred.
[0077] Because the complex of the present invention, formed by the nucleic acid sequence recognition module and the nucleic acid base conversion enzyme, is not associated with DNA double-strand cutting (DSB), it enables genome editing with low toxicity. Therefore, the gene modification method of the present invention is applicable to a wide range of biological materials. Thus, the cells into which the nucleic acid sequence recognition module and / or nucleic acid-encoded nucleic acid are introduced can include cells of all biological species, from prokaryotes such as Escherichia coli and microorganisms such as yeast, to cells of higher eukaryotes such as insects and plants, and cells of vertebrates including mammals such as humans.
[0078] For DNA encoding nucleic acid sequence recognition modules such as zinc finger motifs, TAL effectors, and PPR motifs, each module can be obtained using any of the methods described above. For DNA encoding sequence recognition modules such as restriction enzymes, transcription factors, and RNA polymerases, cloning can be performed as follows: for example, based on their cDNA sequence information, oligoDNA primers are synthesized to cover the region encoding the desired portion of the protein (including the DNA-binding domain), and total RNA or mRNA fractions prepared from the cell that produces the protein are used as templates for amplification via RT-PCR.
[0079] DNA encoding nucleic acid base-converting enzymes can also be cloned in the same way: oligoDNA primers are synthesized based on the cDNA sequence information of the enzyme to be used, and total RNA or mRNA fractions prepared from the cell that produces the enzyme are used as templates for amplification by RT-PCR. For example, for DNA encoding PmCDA1 in the lamprey, appropriate primers can be designed upstream and downstream of the CDS based on the cDNA sequence registered in the NCBI database (accession No. EF094822), and cloned from the mRNA from the lamprey by RT-PCR. Similarly, for DNA encoding human AID, appropriate primers can be designed upstream and downstream of the CDS based on the cDNA sequence registered in the NCBI database (accession No. AB040431), and cloned from the mRNA from human lymph nodes, for example, by RT-PCR.
[0080] Cloned DNA can be directly digested with restriction enzymes as needed, or, after adding appropriate adapters and / or nuclear localization signals (or organelle transfer signals in the case of mitochondrial or chloroplast DNA of interest), conjugated with DNA encoding a nucleic acid sequence recognition module to prepare DNA encoding a fusion protein. Alternatively, DNA encoding a nucleic acid sequence recognition module and DNA encoding a nucleic acid base conversion enzyme can be fused separately with DNA encoding a binding domain or its binding partner, or both DNAs can be fused with DNA encoding a dissociative intapeptide, allowing the nucleic acid sequence recognition module and the nucleic acid base conversion enzyme to translate within the host cell, and then forming a complex. In these cases, adapters and / or nuclear localization signals can also be ligated at appropriate sites on one or both DNAs as needed.
[0081] For DNA encoding nucleic acid sequence recognition modules and DNA encoding nucleic acid base conversion enzymes, DNA chains can be chemically synthesized, or full-length DNA encoding these sequences can be constructed by ligating partially overlapping synthetic oligoDNA short chains using PCR or Gibson Assembly. The advantage of constructing full-length DNA using combinatorial chemical synthesis, PCR, or Gibson Assembly is that codons compatible with the host organism to which the DNA is to be introduced can be designed along the entire CDS length. By converting the DNA sequence to codons frequently used in the host organism during heterologous DNA expression, increased protein expression can be expected. Data on codon usage frequencies in the host organism can be obtained from, for example, the genetic code usage frequency database published on the homepage of the Kazusa DNA Institute (http: / / www.kazusa.or.jp / codon / index.html), or references to literature documenting codon usage frequencies in various hosts. By referring to the obtained data and the DNA sequence to be introduced, codons with low usage frequencies in the host organism within the DNA sequence can be converted to codons encoding the same amino acid with high usage frequencies.
[0082] DNA expression vectors containing a nucleic acid sequence recognition module and / or a nucleic acid base conversion enzyme can be manufactured, for example, by linking the DNA downstream of a suitable expression vector promoter.
[0083] As expression vectors, plasmids from *Escherichia coli* (e.g., pBR322, pBR325, pUC12, pUC13); plasmids from *Bacillus subtilis* (e.g., pUB110, pTP5, pC194); plasmids from yeast (e.g., pSH19, pSH15); insect cell expression plasmids (e.g., pFast-Bac); animal cell expression plasmids (e.g., pA1-11, pXT1, pRc / CMV, pRc / RSV, pcDNAI / Neo); bacteriophages such as λ phage; insect viral vectors such as baculoviruses (e.g., BmNPV, AcNPV); and animal viral vectors such as retroviruses, vaccinia viruses, and adenoviruses.
[0084] As a promoter, any appropriate promoter corresponding to the host of gene expression is acceptable. Due to the toxicity of conventional methods involving DSB, the survival rate of host cells is sometimes significantly reduced. Therefore, it is desirable to increase the number of cells up to the start of induction by using an inducible promoter. However, since sufficient cell proliferation can also be achieved by expressing the nuclease-modifying enzyme complex of the present invention, the promoter can be used without limitation.
[0085] For example, when the host is an animal cell, the SRα promoter, SV40 promoter, LTR promoter, CMV (cytomegalovirus) promoter, RSV (Rouse sarcoma virus) promoter, MoMuLV (Moloney mouse leukemia virus) LTR, HSV-TK (herpes simplex virus thymidine kinase) promoter, etc., can be used. Among them, the CMV promoter and SRα promoter are preferred.
[0086] When the host is Escherichia coli, the preferred promoters are trp, lac, recA, λPL, lpp, and T7.
[0087] When the host is Bacillus, the preferred promoters are SpO1, SpO2, and penP.
[0088] When the host is yeast, Gal1 / 10 promoter, PHO5 promoter, PGK promoter, GAP promoter, ADH promoter, etc. are preferred.
[0089] When the host is an insect cell, polyhedromic protein promoters, P10 promoters, etc. are preferred.
[0090] When the host is a plant cell, the CaMV35S promoter, CaMV19S promoter, NOS promoter, etc. are preferred.
[0091] In addition to the above, vectors containing enhancers, splicing signals, terminators, polyA addition signals, selection markers such as drug resistance genes and auxotrophic complementation genes, and origins of replication can be used as expression vectors as needed.
[0092] RNA encoding a nucleic acid sequence recognition module and / or a nucleic acid base conversion enzyme can be prepared by, for example, using a DNA-encoding vector encoding the aforementioned nucleic acid sequence recognition module and / or nucleic acid base conversion enzyme as a template, and transcribed into mRNA using a known in vitro transcription system.
[0093] A complex of a nucleic acid sequence recognition module and a nucleic acid base conversion enzyme can be expressed intracellularly by introducing a DNA expression vector containing a nucleic acid sequence recognition module and / or a nucleic acid base conversion enzyme into a host cell and culturing the host cell.
[0094] As a host, it can be used in the following ways: Escherichia coli, Bacillus spp., yeast, insect cells, insects, animal cells, etc.
[0095] For example, Escherichia coli K12? DH1 [Proc. Natl. Acad. Sci. USA, 60, 160 (1968)], Escherichia coli JM103 [Nucleic Acids Research, 9, 309 (1981)], Escherichia coli JA221 [Journal of Molecular Biology, 120, 517 (1978)], Escherichia coli HB101 [Journal of Molecular Biology, 41, 459 (1969)], Escherichia coli C600 [Genetics, 39, 440 (1954)], etc.
[0096] As a genus of Bacillus, examples include: Bacillus subtilis MI114 [Gene, 24, 255 (1983)], Bacillus subtilis 207-21 [Journal of Biochemistry, 95, 87 (1984)], etc.
[0097] As yeast, you can use, for example: Saccharomyces cerevisiae AH22, AH22R-, NA87-11A, DKD-5D, 20B-12, Schizosaccharomyces pombe NCYC1913, NCYC2036, Pichia pastoris KM71, etc.
[0098] As insect cells, for example, in the case of AcNPV, cell lines derived from the larvae of the cabbage looper (Spodoptera frugiperda cell; Sf cells), midgut MG1 cells from Trichoplusia ni, High Five™ cells from the eggs of Trichoplusia ni, cells from Mamestra brassicae, cells from Estigmena acrea, etc., can be used. In the case of BmNPV, cell lines derived from silkworms (Bombyx mori N cells; BmN cells), etc., can be used as insect cells. As Sf cells, for example, Sf9 cells (ATCCCRL1711), Sf21 cells [above, In Vivo, 13, 213-217 (1977)], etc., can be used.
[0099] As insects, for example, silkworm larvae, fruit flies, crickets, etc. can be used [Nature, 315, 592 (1985)].
[0100] As animal cells, the following can be used: monkey COS-7 cells, monkey Vero cells, Chinese hamster oocyst (CHO) cells, dhfr gene-deficient CHO cells, mouse L cells, mouse AtT-20 cells, mouse myeloma cells, rat GH3 cells, human FL cells, and other cell lines; pluripotent stem cells such as iPS cells and ES cells from humans and other mammals; and primary cultured cells prepared from various tissues. Furthermore, zebrafish embryos and Xenopus oocytes can also be used.
[0101] As plant cells, suspension culture cells, callus tissue, protoplasts, leaf sections, root sections, etc., prepared from various plants (e.g., grains such as rice, wheat, and corn; commercial crops such as tomatoes, cucumbers, and eggplants; horticultural plants such as carnations and lisianthus; experimental plants such as tobacco and Arabidopsis thaliana) can be used.
[0102] The aforementioned host cells can be haploid (monoploid) or polyploid (e.g., diploid, triploid, tetraploid, etc.). Conventional mutation introduction methods, as a principle, introduce mutations only into one homologous chromosome to form a heterologous genotype. Therefore, if it is not a dominant mutation, the desired trait will not be expressed, and achieving homology is time-consuming and laborious, often inconvenient. In contrast, according to the present invention, since mutations can be introduced into all alleles on homologous chromosomes within the genome, even if it is a recessive mutation, the desired trait can be expressed in that representative. Figure 10 From the perspective of overcoming the problems of conventional methods, it is extremely useful.
[0103] For the introduction of expression vectors, the method can be carried out according to the type of host, using known methods (e.g., lysozyme method, competent cell method, PEG method, CaCl2 coprecipitation method, electroporation method, microinjection method, particle gun method, lipid conversion method, Agrobacterium method, etc.).
[0104] For Escherichia coli, transformation can be performed according to methods described in, for example, Proc. Natl. Acad. Sci. USA, 69, 2110 (1972), Gene, 17, 107 (1982).
[0105] The vector can be introduced into Bacillus spp. using methods described, for example, in Molecular & General Genetics, 168, 111 (1979).
[0106] The vector can be introduced into yeast according to methods described in, for example, Methods in Enzymology, 194, 182-187 (1991), Proc. Natl. Acad. Sci. USA, 75, 1929 (1978).
[0107] The vector can be introduced into insect cells and insects according to methods described, for example, in Bio / Technology, 6, 47-55 (1988).
[0108] The vector can be introduced into animal cells according to, for example, the methods described in Cell Engineering Supplement 8 New Cell Engineering Experimental Protocols, 263-267 (1995) (published by Xiurun Company), Virology, 52, 456 (1973).
[0109] For the culture of cells that have been introduced with a vector, it can be carried out according to known methods, depending on the type of host.
[0110] For example, when culturing *Escherichia coli* or *Bacillus*, a liquid culture medium is preferred. Furthermore, the culture medium preferably contains carbon sources, nitrogen sources, and inorganic substances necessary for the growth of the transformant. Examples of carbon sources include glucose, dextrin, soluble starch, and sucrose; examples of nitrogen sources include ammonium salts, nitrates, corn steep liquor, peptone, casein, meat extract, soybean meal, and potato extract; examples of inorganic substances include calcium chloride, sodium dihydrogen phosphate, and magnesium chloride. Additionally, yeast extract, vitamins, and growth promoters may be added to the culture medium. The pH of the culture medium is preferably about 5 to about 8.
[0111] As a culture medium for culturing *E. coli*, M9 medium containing glucose and casein amino acids is preferred, for example [Journal of Experiments in Molecular Genetics, 431-433, Cold Spring Harbor Laboratory, New York 1972]. If necessary, to ensure the promoter functions effectively, reagents such as 3β-indolylic acid can be added to the medium. *E. coli* culture is typically carried out at approximately 15–43°C. Aeration and stirring can be performed as needed.
[0112] Bacillus cultures are typically carried out at approximately 30–40°C. Aeration and stirring may be performed as needed.
[0113] Examples of suitable culture media for yeast cultivation include Burkholder Minimal Medium [Proc. Natl. Acad. Sci. USA, 77, 4505 (1980)] and 0.5% SD Medium containing casein amino acids [Proc. Natl. Acad. Sci. USA, 81, 5330 (1984)]. The preferred pH of the medium is approximately 5 to approximately 8. Cultivation is typically carried out at approximately 20°C to approximately 35°C. Aeration and stirring can be performed as needed.
[0114] As a culture medium for culturing insect cells or insects, a medium such as Grace's Insect Medium [Nature, 195, 788 (1962)] with appropriate additions such as inactivated 10% bovine serum can be used. The pH of the medium is preferably about 6.2 to about 6.4. Incubation is usually carried out at about 27°C. Aeration and stirring can be performed as needed.
[0115] For culturing animal cells, the following media can be used: Minimal Essential Medium (MEM) containing approximately 5% to 20% fetal bovine serum [Science, 122, 501 (1952)], Dulbecco modified Eagle Medium (DMEM) [Virology, 8, 396 (1959)], RPMI 1640 Medium [The Journal of the American Medical Association, 199, 519 (1967)], and 199 Medium [Proceeding of the Society for the Biological Medicine, 73, 1 (1950)]. The pH of the medium is preferably approximately 6 to approximately 8. Culture is usually carried out at approximately 30°C to approximately 40°C. Aeration and stirring can be performed as needed.
[0116] MS medium, LS medium, B5 medium, etc., can be used as culture media for plant cells. The preferred pH of the medium is approximately 5 to approximately 8. Incubation is typically carried out at approximately 20°C to approximately 30°C. Aeration and stirring can be performed as needed.
[0117] As described above, a complex of nucleic acid sequence recognition module and nucleic acid base conversion enzyme, i.e., a nucleic acid modifying enzyme complex, can be expressed in cells.
[0118] The introduction of RNA encoding nucleic acid sequence recognition modules and / or nucleic acid base conversion enzymes into host cells can be performed via microinjection, lipid transfection, or other methods. RNA introduction can be performed once or multiple times at appropriate intervals (e.g., 2–5 times).
[0119] When an expression vector or RNA molecule introduced into a cell expresses a complex of a nucleic acid sequence recognition module and a nucleic acid base-converting enzyme, the nucleic acid sequence recognition module specifically recognizes and binds to the target nucleotide sequence within the target double-stranded DNA (e.g., genomic DNA). Utilizing the action of the nucleic acid base-converting enzyme linked to the nucleic acid sequence recognition module, a base transition occurs in the sense or antisense strand of the target site (which can be appropriately regulated within a range of several hundred bases containing all or part of the target nucleotide sequence or its vicinity), resulting in a mismatch within the double-stranded DNA (e.g., when cytidine deaminases such as PmCDA1 or AID are used as nucleic acid base-converting enzymes, cytosine on the sense or antisense strand of the target site is converted to uracil, producing a U:G or G:U mismatch). Various mutations are introduced as follows: the mismatch is not properly repaired, or repair results in the opposite strand pairing with the base of the transformed strand (TA or AT in the above examples), further substitution with other nucleotides during repair (e.g., U→A, G), or deletion or insertion of one to several tens of bases.
[0120] For zinc finger motifs, the creation efficiency of zinc fingers that specifically bind to target nucleotide sequences is low, and the selection of zinc fingers with high binding specificity is not easy. Therefore, it is not easy to create multiple functional zinc finger motifs. As for TAL effectors and PPR motifs, they have a higher degree of freedom in target nucleic acid sequence recognition than zinc finger motifs, but they require the design and construction of huge proteins based on the target nucleotide sequence each time, thus posing an efficiency problem.
[0121] In contrast, since the CRISPR-Cas system recognizes the sequence of double-stranded DNA of interest by using guide RNA complementary to the target nucleotide sequence, it can target any sequence by synthesizing only oligoDNAs that can form specific hybrids with the target nucleotide sequence.
[0122] Therefore, in a more preferred embodiment of the present invention, a CRISPR-Cas system (CRISPR-mutant Cas) with at least one DNA cutting capability of Cas is used as a nucleic acid sequence recognition module.
[0123] Figure 1 This is an illustrative description of the double-stranded DNA modification method of the present invention, which uses CRISPR-mutant Cas as a nucleic acid sequence recognition module.
[0124] The nucleic acid sequence recognition module of the present invention, which uses CRISPR-mutant Cas, is provided in the form of a complex of an RNA molecule and a mutant Cas protein, wherein the RNA molecule consists of a guide RNA complementary to the target nucleotide sequence and a tracrRNA essential for the recruitment of the mutant Cas protein.
[0125] The Cas protein used in this invention can belong to the CRISPR system without particular limitation, but Cas9 is preferred. Examples of Cas9 include, for example, Cas9 (SpCas9) from Streptococcus pyogenes and Cas9 (StCas9) from Streptococcus thermophilus, but are not limited to these. SpCas9 is preferred. As the mutant Cas used in this invention, Cas with cleavage ability to inactivate both strands of double-stranded DNA, or Cas with cleavage enzyme activity that inactivates only one strand, can be used. For example, in the case of SpCas9, the D10A mutant, which is converted from the As residue at position 10 to the Ala residue and lacks the ability to cleave the opposite strand that forms the strand complementary to the guide RNA, or the H840A mutant, which is converted from the His residue at position 840 to the Ala residue and lacks the ability to cleave the strand complementary to the guide RNA, can be used. Double mutants of these mutants can also be used, but other mutant Cas can also be used.
[0126] For nucleotide base-converting enzymes, they are provided as a complex of mutant Cas using the same method as the zinc fingers mentioned above. Alternatively, nucleotide base-converting enzymes and mutant Cas can be linked using an RNA scaffold composed of RNA aptamers such as MS2F6, PP7, etc., and their binding proteins. The guide RNA forms a complementary strand with the target nucleotide sequence, and then the tracrRNA recruits the mutant Cas, which recognizes the DNA cleavage site recognition sequence PAM (protospace radjacent motif) (when using SpCas9, PAM is NGG (N is any base) 3 bases; theoretically, any site on the genome can be targeted) without cleaving one or two DNA segments. The nucleotide base-converting enzyme linked to the mutant Cas causes a base shift at the target site (which can be appropriately regulated within a range of several hundred bases containing all or part of the target nucleotide sequence), resulting in a mismatch within the double-stranded DNA. Various mutations are introduced through the following methods: the mismatch is not properly repaired; repair results in the pairing of bases on the opposite strand with bases on the transformed strand; during repair, the mismatch is further transformed into other nucleotides; or a deletion or insertion of one to several dozen bases occurs (e.g., see reference). Figure 2 ).
[0127] In the case of using CRISPR-mutant Cas as a nucleic acid sequence recognition module, similarly to the case of using zinc fingers or the like as a nucleic acid sequence recognition module, it is preferable to introduce the nucleic acid sequence recognition module and the nucleic acid base conversion enzyme into the cell containing the double-stranded DNA of interest in the form of the nucleic acid encoding them.
[0128] The DNA encoding Cas can be cloned from the cell that produces the enzyme using the same method described above for the DNA encoding a nucleic acid base-transferase. Alternatively, mutant Cas can be obtained by introducing a mutation into the cloned Cas-encoding DNA using a site-specific mutation induction method known per se, by changing amino acid residues at sites important for DNA cleavage activity (e.g., in the case of Cas9, the Asp residue at position 10, the His residue at position 840, but not limited to these) to other amino acids.
[0129] Alternatively, for DNA encoding mutant Cas, DNA encoding nucleic acid sequence recognition modules or nucleic acid base-converting enzymes can also be constructed using the same methods and combinations of chemical synthesis, PCR, or Gibson Assembly to create a DNA form with codons suitable for expression in host cells. For example, the optimized CDS and amino acid sequences for expressing SpCas9 in eukaryotic cells are shown in sequences 5 and 6. If the "A" at base number 29 in the sequence shown in sequence 5 is changed to "C", DNA encoding the D10A mutant can be obtained; if the "CA" at bases 2518-2519 is changed to "GC", DNA encoding the H840A mutant can be obtained.
[0130] For the DNA encoding mutant CasDNA and the DNA encoding nucleic acid base conversion enzyme, they can be linked together to be expressed as a fusion protein, or they can be designed to be expressed separately using binding domains, integrins, etc., forming a complex in the host cell through protein-protein interactions and protein binding.
[0131] For the DNA encoding the mutant Cas and / or nucleic acid base conversion enzyme obtained, it can be inserted downstream of the promoter of the same expression vector as described above, depending on the host.
[0132] On the other hand, the DNA encoding guide RNA and tracrRNA can be designed as an oligoDNA sequence formed by linking a guide RNA sequence complementary to the target nucleotide sequence with a known tracrRNA sequence (e.g., gttttagagctagaaatagcaagttaaaataaggctagtccgttatcaacttgaaaaagtggcaccgagtcggtggtgctttt; sequence number 7), and then chemically synthesized using a DNA / RNA synthesizer.
[0133] The length of the guide RNA sequence is not particularly limited as long as it can specifically bind to the target nucleotide sequence; for example, it can be 15 to 30 nucleotides, preferably 18 to 24 nucleotides.
[0134] The DNA encoding guide RNA and tracrRNA can also be inserted into the same expression vector as described above, depending on the host, as a promoter. It is preferred to use a pol III type promoter (e.g., SNR6, SNR52, SCR1, RPR1, U6, H1 promoter, etc.) and a terminator (e.g., T6 sequence).
[0135] For RNA encoding mutant Cas and / or nucleic acid base conversion enzymes, it can be prepared, for example, by using a vector encoding DNA encoding the aforementioned mutant Cas and / or nucleic acid base conversion enzymes as a template, and transcribed into mRNA using an in vitro transcription system that is known per se.
[0136] The guide RNA (tracrRNA) can be designed as an oligoRNA sequence consisting of a sequence complementary to the target nucleotide sequence linked to a known tracrRNA sequence, and then chemically synthesized using a DNA / RNA synthesizer.
[0137] For DNA or RNA encoding mutant Cas and / or nucleic acid base conversion enzymes, guide RNA-tracrRNA or DNA encoding them, they can be introduced into host cells using the same methods described above, depending on the host.
[0138] In conventional artificial nucleases, due to the accompanying DNA double-strand cutting (DSB), random chromosomal cutting (off-target cutting) occurs when targeting sequences within the genome, leading to proliferation impairment and cell death. This effect is particularly lethal to many microorganisms and prokaryotes, hindering their applicability. In this invention, since mutations are introduced through a non-cutting reaction involving the transformation of substituents on DNA bases (especially deamination), a significant reduction in toxicity can be achieved. In fact, as shown in the examples below, in comparative experiments using budding yeast as a host, it was confirmed that when using Cas9 with conventional DSB activity, expression-induced cell survival decreased; in contrast, in the technique of this invention, which combines mutant Cas and a nucleoside transition enzyme, cells continued to proliferate, and cell survival increased. Figure 3 ).
[0139] It should be noted that the modification of the double-stranded DNA in this invention does not prevent the cleavage of the double-stranded DNA outside the target site (which can be appropriately regulated within a range of several hundred bases containing all or part of the target nucleotide sequence). However, one of the greatest advantages of this invention is the avoidance of toxicity caused by off-target cleavage. Considering its applicability to any biological species in principle, in a preferred embodiment, the modification of the double-stranded DNA in this invention is not only unrelated to DNA strand cleavage at the selected target site of the double-stranded DNA, but also at other sites.
[0140] Furthermore, as shown in the examples below, compared to the case where a mutant Cas that cannot cut either strand is used, in the case where a Cas with nickase activity that can only cut one strand of double-stranded DNA is used as the mutant Cas ( Figure 5 This increases mutation introduction efficiency. Therefore, for example, by further linking proteins with cleavage enzyme activity in addition to nucleic acid sequence recognition modules and nucleic acid base-converting enzymes, only one strand of DNA can be cleaved near the target nucleotide sequence, thereby improving mutation introduction efficiency while avoiding the strong toxicity of DSB-based mutations.
[0141] Furthermore, the effects of mutant Cas with two cleavage enzyme activities that cleave different chains were compared. The results showed that using one mutant Cas resulted in the accumulation of mutations near the center of the target nucleotide sequence, while using the other mutant Cas resulted in the random introduction of multiple mutations into a region extending from the target nucleotide sequence to several hundred bases. Figure 6Therefore, by selecting the strand to be cleaved by the cleavage enzyme, point mutations can be introduced into specific nucleotides or nucleotide regions, or multiple mutations can be randomly introduced into a wider range, depending on the purpose. For example, if the former technique is applied to iPS cells for genetic diseases, a cell transplantation therapy agent with a lower risk of rejection can be produced by differentiating iPS cells made from the patient's own cells into somatic cells of interest after repairing mutations in the pathogenic gene.
[0142] In Example 7 and the following examples, it is shown that mutations can be introduced almost at the target site in specific nucleotides. In this way, to introduce a desired nucleotide mutation at the target site, it is only necessary to set the target nucleotide sequence such that the positional relationship between the nucleotide to which the mutation is to be introduced and the target nucleotide sequence is regular. Using the CRIPR-Cas system as the nucleic acid sequence recognition module and AID as the nucleic acid base conversion enzyme, the target nucleotide sequence can be designed such that the C (or its opposite G) to which the mutation is to be introduced is located 2 to 5 nucleotides from the 5' end of the target nucleotide sequence. As mentioned above, the length of the guide RNA sequence can be appropriately set between 15 and 30 nucleotides, preferably between 18 and 24 nucleotides. Since the guide RNA sequence is complementary to the target nucleotide sequence, changing the length of the guide RNA sequence causes a change in the length of the target nucleotide sequence, but regardless of the nucleotide length, the regularity of the possible introduction of mutations at the C or G located 2 to 5 nucleotides from the 5' end is maintained. Figure 12 Therefore, by appropriately selecting the length of the target nucleotide sequence (as the guide RNA of its complementary strand), the site where the mutated base can be introduced can be moved. This also removes the restrictions imposed by the DNA cleavage site recognition sequence PAM (NGG), further increasing the freedom of mutation introduction.
[0143] As illustrated in the examples below, mutation introduction efficiency is significantly increased by creating sequence recognition modules relative to multiple neighboring target nucleotide sequences and using them simultaneously, compared to when a single nucleotide sequence is used as a target. Figure 7 The effect is that mutation induction is achieved whether the target nucleotide sequences are partially repeated or separated by about 600 bp. Furthermore, this is also true when the target nucleotide sequences are in the same orientation (on the same strand). Figure 7 ), and in the opposite direction (the target nucleotide sequence is present on each strand of the double-stranded DNA) Figure 4 It can occur in either of the following situations:
[0144] As shown in the examples below, with respect to the method for modifying the genome sequence of the present invention, if an appropriate target nucleotide sequence is selected, mutations can be introduced into almost all cells expressing the nuclease complex of the present invention. Figure 8 Therefore, the insertion and selection of selectable marker genes, which are necessary in conventional genome editing, are not required. This significantly simplifies gene manipulation and greatly broadens its applicability to crop breeding and other applications due to the absence of organisms with recombinant exogenous DNA.
[0145] Furthermore, because the mutation introduction efficiency is extremely high and no marker selection is required, the genome sequence modification method of this invention can target and modify multiple DNA regions at completely different locations. Figure 9 Therefore, in a preferred embodiment of the present invention, two or more nucleic acid sequence recognition modules that specifically bind to different target nucleotide sequences (which may be within one gene of interest or within two or more different genes of interest. These target genes may be located on the same chromosome or on different chromosomes) can be used. In this case, one of each of these nucleic acid sequence recognition modules is combined with a nucleotide base conversion enzyme to form a nucleotide editing enzyme complex. Here, a common enzyme can be used for the nucleotide editing enzyme. For example, when using the CRISPR-Cas system as the nucleic acid sequence recognition module, a common substance can be used for the complex of Cas protein and nucleotide base conversion enzyme (containing a fusion protein), and two or more chimeric RNAs can be prepared and used as guide RNAs - tracrRNAs, wherein the chimeric RNAs are formed by combining each of two or more guide RNAs that form complementary strands with different target nucleotide sequences with tracrRNAs. On the other hand, when using zinc finger motifs, TAL effectors, etc., as nucleic acid sequence recognition modules, for example, each nucleic acid sequence recognition module that specifically binds to different target nucleotides can be fused with a nucleotide base conversion enzyme.
[0146] To express the nucleic acid-modifying enzyme complex of the present invention in a host cell, as described above, an expression vector containing DNA encoding the nucleic acid-modifying enzyme complex and RNA encoding the nucleic acid-modifying enzyme complex are introduced into the host cell. However, for effective mutation introduction, it is preferable to maintain the expression of the nucleic acid-modifying enzyme complex at a specified time and at a specified level or higher. From this viewpoint, it is determined to introduce an expression vector (plasmid, etc.) capable of autonomous replication into the host cell. However, since the plasmid, etc., is exogenous DNA, it is preferable to remove it rapidly after successful mutation introduction. Therefore, although the method varies depending on the type of host cell, it is preferable, for example, to remove the introduced plasmid from the host cell using various plasmid removal methods known in the art 6 hours to 2 days after the expression vector is introduced.
[0147] Alternatively, wherever it is possible to express a nuclease complex sufficient for mutation introduction, it is also preferable to use an expression vector that does not have the ability to replicate autonomously in the host cell (e.g., a vector that lacks a replication origin and / or a gene encoding a protein essential for replication, etc.) or RNA, to introduce mutations into the double-stranded DNA of interest through transient expression.
[0148] Because the expression of the target gene is suppressed during the period when the nucleic acid modifying enzyme complex of the present invention is expressed in the host cell and the nucleic acid base conversion reaction occurs, it has been difficult to directly edit genes essential for the survival of the host cell as target genes (resulting in side effects such as host growth disorders, unstable mutation introduction efficiency, and mutations at sites different from the target). The present invention achieves successful and efficient direct editing of essential genes by transiently expressing the nucleic acid modifying enzyme complex of the present invention in the host cell at the desired time, during which the nucleic acid base conversion reaction occurs and the modification of the target site is stabilized. The time required for the nucleic acid base conversion reaction to occur and for the modification of the target site to be stabilized varies depending on the type of host cell, culture conditions, etc., but is generally considered to require 2 to 20 generations. For example, in the case of host cells such as yeast or bacteria (e.g., Escherichia coli), the expression of the nucleic acid modifying enzyme complex needs to be induced between 5 and 10 generations. Those skilled in the art can appropriately determine the suitable expression induction period based on the doubling time of the host cells under the culture conditions used. For example, when budding yeast is cultured in liquid medium with 0.02% galactose, the expression induction period can be 20 to 40 hours. Regarding the expression induction period of the nucleic acid encoding the nucleic acid-modifying enzyme complex of the present invention, it can also be extended beyond the aforementioned "period requiring stabilization of the target site modification" without causing side effects on the host cells.
[0149] As a method for transiently expressing the nucleic acid-modifying enzyme complex of the present invention at a desired time, a construct (expression vector) comprising a nucleic acid encoding the nucleic acid-modifying enzyme complex (in the CRISPR-Cas system, DNA encoding the guide RNA-tracrRNA and DNA encoding the mutant Cas and the nucleic acid base substitution enzyme) and introducing the construct into a host cell can be used. "In a form capable of controlling the expression period" specifically refers to a form in which the nucleic acid encoding the nucleic acid-modifying enzyme complex of the present invention is placed under the control of an inducible regulatory region. There is no particular limitation on the "inducible regulatory region," for example, in microbial cells such as bacteria (e.g., *Escherichia coli*), yeast, etc., temperature-sensitive (ts) mutation repressors and the operator's controller (operon) controlling them can be listed. Examples of ts mutation repressors include, for example, ts mutants of the cI repressor from λ phage, but are not limited to these. In the case of λ phage cI repressor (ts), it binds to the operator gene at temperatures below 30°C (e.g., 28°C) to inhibit downstream gene expression, but dissociates from the operator gene at temperatures above 37°C (e.g., 42°C), thus inducing gene expression. Figure 13 and 14 Therefore, by culturing host cells with nucleic acids encoding nucleic acid-modifying enzyme complexes, typically below 30°C, raising the temperature to above 37°C at an appropriate time and culturing for a certain period to induce nucleic acid base conversion, and then rapidly cooling the temperature back below 30°C after introducing mutations into the target gene, the period of suppression of target gene expression can be minimized. Even when targeting genes essential to the host cell, side effects can be suppressed and editing can be performed effectively. Figure 15 ).
[0150] In the case of utilizing temperature-sensitive mutations, for example, by including a temperature-sensitive mutant of the protein necessary for the autonomous replication of the vector in a DNA vector containing the nucleic acid-modifying enzyme complex of the present invention, the vector cannot rapidly replicate autonomously after the expression of the nucleic acid-modifying enzyme complex and will naturally detach during cell division. Examples of such temperature-sensitive mutant proteins include, but are not limited to, the temperature-sensitive mutant of Rep101 ori, which is necessary for the replication of pSC101 ori. While Rep101 ori(ts) can act on pSC101 ori and perform autonomous plasmid replication at temperatures below 30°C (e.g., 28°C), it loses function at temperatures above 37°C (e.g., 42°C), and the plasmid cannot replicate autonomously. Therefore, by combining it with the cI repressor(ts) of the aforementioned λ phage, the transient expression of the nucleic acid-modifying enzyme complex of the present invention and plasmid removal can be carried out simultaneously.
[0151] On the other hand, when higher eukaryotic cells such as animal cells, insect cells, and plant cells are used as host cells, DNA encoding the nucleic acid modifying enzyme complex of the present invention can be introduced into the host cell under the control of inducible promoters (e.g., metallothionein promoters (induced by heavy metal ions), heat shock protein promoters (induced by heat shock), Tet-ON / Tet-OFF promoters (induced by adding or removing tetracycline or its derivatives), steroid response promoters (induced by steroid hormones or their derivatives)). Inducing substances are added to the culture medium (or removed from the culture medium) at appropriate times to induce the expression of the nucleic acid modifying enzyme complex. After culturing for a certain period, a nucleic acid base conversion reaction is performed, which can achieve transient expression of the mutated nucleic acid modifying enzyme complex introduced into the target gene.
[0152] It should be noted that inducible promoters can also be used in prokaryotic cells such as E. coli. Examples of such inducible promoters include, but are not limited to, the lac promoter (inducible using IPTG), the cspA promoter (inducible using cold shock), and the araBAD promoter (inducible using arabinose).
[0153] Alternatively, the aforementioned inducible promoters can be used as a vector removal method when higher eukaryotic cells such as animal cells, insect cells, and plant cells are used as host cells. That is, by attaching the replication origin that functions in the host cell and the nucleic acid encoding the protein necessary for its replication (e.g., SV40 ori and large T antigen, oriP and EBNA-1, etc. in animal cells), the expression of the nucleic acid encoding the protein is controlled using the aforementioned inducible promoter. Although the vector can autonomously replicate in the presence of the inducing substance, it cannot replicate autonomously if the inducing substance is removed, and the vector will naturally detach during cell division (in Tet-OFF type vectors, conversely, the addition of tetracycline or doxycycline prevents autonomous replication).
[0154] The present invention will now be described with reference to embodiments. However, the present invention is not limited to these embodiments.
[0155] Example
[0156] In Examples 1 to 6 below, experiments were conducted as described below.
[0157] <Cell lines, culture, transformation, expression induction>
[0158] Saccharomyces cerevisiae BY4741 (requiring leucine and uracil) was cultured using standard YPDA medium or SD medium containing dropout components that met nutritional requirements. Culture was conducted between 25°C and 30°C using either static incubation on agar plates or shaking in liquid medium. For transformation, the lithium acetate method was used for selection using SD medium with appropriate nutritional requirements. For galactose-based expression induction, the cells were pre-cultured overnight in appropriate SD medium, then inoculated onto SR medium with 2% raffinose instead of 2% glucose as the carbon source and cultured overnight. Further inoculation was performed onto SGal medium with 0.2–2% galactose instead of 0.2% galactose as the carbon source for approximately 3 hours to two nights to induce expression.
[0159] For determining cell viability and Can1 mutation rate, the cell suspension was appropriately diluted and plated on SD agar plates and SD-Arg + 60 mg / L canavonine agar plates or SD + 300 mg / L canavonine agar plates. The number of colonies appearing after 3 days was counted as the cell viability. The number of surviving colonies on SD plates was taken as the total cell count, and the number of surviving colonies on canavonine agar plates was taken as the number of resistant mutant strains. The mutation rate was calculated and evaluated. For the mutation introduction site, DNA fragments containing the target gene region of each strain were amplified by colony PCR, followed by DNA sequencing. The sequences were compared and analyzed based on the sequences in the yeast genome database (http: / / www.yeastgenome.org / ) for identification.
[0160] <Nucleic Acid Procedures>
[0161] For DNA, processing and construction are performed using any of the following methods: PCR, restriction enzyme treatment, conjugation, Gibson assembly, or artificial chemical synthesis. For plasmids, leucine-selective pRS315 and uracil-selective pRS426 are used as the backbone, serving as a yeast-E. coli shuttle vector. The plasmids are amplified using E. coli strain XL-10gold or DH5α and introduced into yeast via the lithium acetate method.
[0162] <Constructor>
[0163] For induced expression, budding yeast pGal1 / 10 (sequence number 8), which has a galactose-inducible and bidirectional promoter, was used. The nuclear localization signal (ccc aag aag aag agg aaggtg; sequence number 9 (PKKKRV; coding sequence number 10)) was added to the ORF (sequence number 5) of the Cas9 gene from Streptococcus pyogenes, which had been optimized for eukaryotic expression. The added product was ligated downstream of the promoter and linked to the ORF (sequence number 1 or 3) of the deaminase gene (PmCDA1 from Petromyzon marinus or hAID from humans) via an adapter sequence, and expressed as a fusion protein. As a linker sequence, although the selection and combination of GS linker (ggt gga gga ggt tct; sequence number 11 (encoding GGGGS; sequence number 12) repetition), Flag tag (gac tat aag gac cacgac gga gac tac aag gat cat gatatt gat tac aaa gac gat gac gat aag; sequence number 13 (encoding DYKDHDGDYKDHDIDYKDDDDK; sequence number 14)), Strep-tag (tgg agc cac ccg cag ttc gaa aaa; sequence number 15 (encoding WSHPQFEK; sequence number 16)), and other structural domains are used, here, 2xGS, SH3 structural domains (sequence numbers 17 and 18), and Flag tag are specifically used. As a terminator, it is linked with the ADH1 terminator (sequence number 19) and Top2 terminator (sequence number 20) from budding yeast. In addition, regarding the domain-binding mechanism, the Cas9 gene ORF is linked to the SH3 domain via the 2xGS adapter and acts as a protein. The SH3 ligand sequences (sequence numbers 21 and 22) are added to the deaminase and act as another protein, linking it to both directions of the Gal1 / 10 promoter. Both are then expressed simultaneously. They are then assembled into the pRS315 plasmid.
[0164] To remove the cleavage ability of the DNA strands on each side, mutations were introduced into Cas9 that changed aspartic acid at position 10 to alanine (D10A, corresponding to DNA sequence mutation a29c) and histidine at position 840 to alanine (H840A, corresponding to DNA sequence mutation ca2518gc).
[0165] For the gRNA, it was configured as a chimeric structure with the tracrRNA (from Streptococcus pyogenes; sequence number 7) between the SNR52 promoter (sequence number 23) and the Sup4 terminator (sequence number 24) and assembled into the pRS426 plasmid. The target base sequences for the gRNA were the complementary strands of the CAN1 gene ORF: 187-206 (gatacgttctctatggagga; sequence number 25) (target1), 786-805 (ttggagaaacccaggtgcct; sequence number 26) (target3), 793-812 (aacccaggtgcctggggtcc; sequence number 27) (target4), 563-582 (ttggccaagtcattcaattt; sequence number 28) (target2), and 767-786 (ataacggaatccaactgggc; sequence number 29) (target5r). When expressing multiple targets simultaneously, sequences from promoter to terminator are used as one group, and multiple groups are assembled into the same plasmid. This plasmid, along with the Cas9-deaminase expression plasmid, is introduced into cells and expressed intracellularly, forming a gRNA-tracrRNA and Cas9-deaminase complex.
[0166] Example 1 Gene modification is achieved by linking the DNA sequence recognition capability of CRISPR-Cas with the deaminase PmCDA1. group sequence
[0167] To test the effectiveness of the genome sequence modification technology of this invention, which utilizes the nucleic acid sequence recognition capabilities of deaminases and CRISPR-Cas, a mutation was introduced into the CAN1 gene encoding the concanavalin A transporter, resulting in concanavalin A resistance due to gene defect. Using a sequence complementary to the 187-206 (target1) sequence of the CAN1 gene ORF as gRNA, an expression vector was constructed to express a chimeric RNA formed by linking this gRNA with a tracrRNA from *Streptococcus pyogenes*. A vector expressing dCas9 (a protein fused with mutations (D10A and H840A) from *Streptococcus pyogenes* Cas9 (SpCas9) to lose nuclease activity) and a protein derived from the deaminase PmCDA1 from *Gillardella* was also constructed and introduced into budding yeast using the lithium acetate method for co-expression. The results are shown in... Figure 2When cultured on SD plates containing canavanine, only cells expressing gRNA-tracrRNA and dCas9-PmCDA1 formed colonies. Resistant colonies were selected and the CAN1 gene region was sequenced, confirming the introduction of mutations in and around the target nucleotide sequence (target1).
[0168] Example 2 Significantly reduce side effects and toxicity
[0169] In conventional Cas9 and other artificial nucleases (ZFN, TALEN), targeting sequences within the genome produces proliferation disorders and cell death, believed to be caused by irregular chromosome cleavage. This effect is particularly lethal in many microorganisms and prokaryotes, hindering their applicability.
[0170] Therefore, to verify the safety and cytotoxicity of the genome sequence modification technology of this invention, a conventional CRISPR-Cas9 comparative experiment was performed. The colony-forming ability of surviving cells on SD plates was determined by using sequences within the CAN1 gene (targets 3 and 4) as gRNA targets, from immediately after galactose expression induction to 6 hours post-induction. The results are shown below. Figure 3 In conventional Cas9, cell death is induced due to inhibited proliferation, resulting in a reduced number of surviving cells. In contrast, in this technology (nCas9 D10A-PmCDA1), cells can proliferate continuously, and the number of surviving cells is greatly increased.
[0171] Example 3 Use of different connection types
[0172] The study examined whether mutational introduction into target genes could still occur if Cas9 and the deaminase were not in fusion protein form, but rather formed a nuclease complex through the binding of a domain and its ligand. dCas9, as used in Example 1, was used as Cas9, and human AID was used as the deaminase instead of PmCDA1. The former was fused with the SH3 domain, and the latter with its binding ligand to create… Figure 4 Various constructs are shown. Additionally, sequences within the CAN1 gene (target 4, 5r) were used as gRNA targets. These constructs were introduced into budding yeast. As a result, even when dCas9 was linked to the deaminase via a binding domain, mutations were effectively introduced into the CAN1 gene target site. Figure 4 By introducing multiple binding domains into dCas9, the mutation introduction efficiency was significantly improved. The main mutation introduction site was position 782 (g782c) of the ORF.
[0173] Example 4 Changes in efficiency and mutation patterns caused by the use of nickase
[0174] Instead of dCas9, the D10A mutant nCas9(D10A), which cleaves only the strand complementary to the gRNA, or the H840A mutant nCas9(H840A), which cleaves only the opposite strand complementary to the gRNA, was used. Otherwise, the procedure was the same as in Example 1, with mutations introduced into the CAN1 gene, and the sequence of the CAN1 gene region in colonies produced on SD plates containing canavanine was examined. The results showed that the former (nCas9(D10A)) was more efficient than dCas9. Figure 5 ), and the mutations are concentrated in the center of the target sequence ( Figure 6 Therefore, point mutation introduction can be performed according to this method. On the other hand, the results showed that the efficiency was higher than that of dCas9 in the latter (nCas9(H840A)). Figure 5 ), and multiple random mutations were introduced from the targeted nucleotide to a region of several hundred bases. Figure 6 ).
[0175] Even with alterations to the target nucleotide sequence, the same significant mutation introduction can be confirmed. In genome editing systems using the CRISPR-Cas9 system and cytidine deaminase, as shown in Table 1, cytosine preferentially deamination occurs in the range of approximately 2-5 bp from the 5' side of the target nucleotide sequence (20 bp). Therefore, by setting the target nucleotide sequence based on this regularity and further combining it with nCas9 (D10A), precise genome editing of one nucleotide unit is possible. On the other hand, if nCas9 (H840A) is used, multiple mutations can be inserted simultaneously within a range of approximately several hundred bp near the target nucleotide sequence. Furthermore, there is a possibility of further altering site specificity by changing the type of deaminase linkage.
[0176] These results suggest that the appropriate type of Cas9 protein can be used depending on the purpose.
[0177] [Table 1]
[0178]
[0179] Example 5 By targeting multiple neighboring DNA sequences, efficiency is synergistically increased.
[0180] Compared to using a single target, efficiency is significantly increased by using multiple neighboring targets simultaneously. Figure 7In fact, 10-20% of cells have concanavalin A resistance mutations (target3, 4). gRNA1 and gRNA2 in the figure target target3 and target4, respectively. PmCDA1 was used as the deaminase. It was confirmed that this effect occurred not only in cases of partial sequence repetition (target3, 4) but also in cases of separation of 600 bp (target1, 3). Additionally, the effect also occurred when the DNA sequences were in the same orientation (target3, 4) and opposite orientations (target4, 5). Figure 4 This effect occurs in both cases.
[0181] Example 6 Gene modifications that do not require selection labeling
[0182] For cells targeted at target3 and target4 in Example 5 (Target3, 4), 10 colonies grown on non-selective (canavanine-free) plates (SD plates without Leu and Ura) were randomly selected, and sequences of the CAN1 gene region were sequenced. The results showed that all examined colonies had a mutation introduced at the target site of the CAN1 gene. Figure 8 That is, according to the present invention, if an appropriate target sequence is selected, editing can be expected to occur in virtually all of the expressing cells. Therefore, the insertion and selection of marker genes, which are necessary in conventional gene manipulation, are not required. While significantly simplifying gene manipulation, its applicability to crop breeding and the like is greatly broadened because it is not a foreign DNA recombinant organism.
[0183] In the following embodiments, the experimental techniques common to Examples 1-6 are performed in the same manner as described above.
[0184] Example 7 Simultaneous editing of multiple sites (different genes)
[0185] In conventional gene manipulation methods, due to various limitations, mutations can typically only be performed at one site at a time. Therefore, this invention investigated whether the method could perform mutations at multiple sites simultaneously.
[0186] The first target nucleotide sequence was selected from positions 3-22 of the Ade1 gene ORF of budding yeast strain BY4741 (Ade1target5:GTCAATTACGAAGACTGAAC; sequence number 30), and the second target nucleotide sequence was selected from positions 767-786 (complementary strand) of the Can1 gene ORF (Can1 target8 (786-767;ATAACGGAATCCAACTGGGC; sequence number 29). Both DNA sequences containing chimeric RNAs encoding two gRNAs and a tracrRNA (sequence number 7) with complementary nucleotide sequences were loaded onto the same plasmid (pRS426) and then onto plasmid nCas9 containing nucleic acids encoding a fusion protein of mutant Cas9 and PmCDA1. D10A-PmCDA1 was introduced into BY4741 strain and expressed to verify the mutation introduction of the two genes. Cells were cultured on SD dropout medium (without uracil and leucine; SD-UL) to retain the plasmid. Cells were appropriately diluted and plated onto SD-UL and concanavalin A-supplemented medium to form colonies. After incubation at 28°C for 2 days, colonies were observed, and the incidence of red colonies caused by the ade1 mutation and the survival rate in concanavalin A-supplemented medium were counted. The results are shown in Table 2.
[0187] [Table 2]
[0188]
[0189] As a phenotype, the proportion of mutations introduced into both the Ade1 and Can1 genes was high, at approximately 31%.
[0190] Next, PCR was used to amplify and sequence the colonies on SD-UL medium. The ORF regions containing Ade1 and Can1 were broadened to obtain sequencing information for approximately 500 bytes surrounding the target sequence. Specifically, five red colonies and five white colonies were analyzed. The results showed that in all red colonies, the C at position 5 of the Ade1 gene ORF was changed to G, and in all white colonies, the C at position 5 was changed to T. Figure 9 Although the mutation rate of the target is 100%, since the mutation rate aimed at gene disruption requires changing the C at position 5 to G to become a stop codon, the expected mutation rate is considered to be 50%. Similarly, for the Can1 gene, although a G mutation at position 782 of the ORF has been confirmed in all clones (…), Figure 9 However, since the mutation that changes only to C provides resistance to concanavalin A, the expected mutation rate is 70%. Within the scope of the examination, the proportion of clones that obtain the desired mutation for both genes simultaneously is 40% (4 out of 10 clones), which is actually a high efficiency.
[0191] Example 8 Editing of polyploid genomes
[0192] Although many organisms possess diploid or polyploid genomes, conventional mutation introduction methods, as a matter of principle, only introduce mutations into one homologous chromosome to form a heterologous genotype. Therefore, if it is not a dominant mutation, the desired characteristic cannot be obtained, and making it homologous is time-consuming and labor-intensive. Therefore, the technique according to the present invention was tested to see if it is possible to introduce mutations into all target alleles on homologous chromosomes within the genome.
[0193] Specifically, the Ade1 and Can1 genes were simultaneously edited in the diploid budding yeast strain YPH501. Since the phenotypes resulting from these gene mutations (red colonies and concanavalin A resistance) are disadvantageous, these phenotypes will not be expressed unless mutations on both sides of the gene are introduced (homological mutations).
[0194] The first target nucleotide sequence was the 1173-1154 positions (complementary strand) of the Ade1 gene ORF (Ade1 target1: GTCAATAGGATCCCCTTTT; sequence number 31) or the 3-22 positions (Ade1 target5: GTCAATTACGAAGACTGAAC; sequence number 30). The second target nucleotide sequence was the 767-786 positions (complementary strand) of the Can1 gene ORF (Can1 target8: ATACGGAATCCAACTGGGC; sequence number 29). The DNA containing two types of gRNA and tracrRNA (sequence number 7) encoding their complementary nucleotide sequences was loaded onto the same plasmid (pRS426). This plasmid, along with the nCas9 D10A-PmCDA1 plasmid containing the nucleic acid encoding the fusion protein of mutant Cas9 and PmCDA1, was introduced into the BY4741 strain and expressed, verifying the introduction of mutations into each gene.
[0195] Colony counting results show that the characteristics of each phenotype can be obtained with a high probability (40%–70%). Figure 10 A).
[0196] Furthermore, to confirm the mutation, sequencing was performed on the Ade1 target regions of both white and red colonies. The results showed overlapping sequencing signals indicating a heterologous mutation at the target sites in the white colonies. Figure 10 In the upper part of image B, the signals of G and T overlap at the ↓ point. This confirms that the phenotype is not observed in the heterologous mutant colonies. On the other hand, in the red colonies, no overlapping signals were confirmed, indicating a homologous mutation. Figure 10 (See diagram B below, where the ↓ mark represents the signal of T).
[0197] Example 9 Genome editing in E. coli
[0198] In this embodiment, the effectiveness of this technology was demonstrated in *Escherichia coli*, a representative bacterial model organism. In particular, the superiority of this technology is highlighted because conventional nuclease-based genome editing techniques are lethal and difficult to apply in bacteria. Furthermore, its applicability to both prokaryotic and eukaryotic species was shown when used with yeast, a model cell for eukaryotes.
[0199] Amino acid mutations (dCas9) of D10A and H840A were introduced into the *Streptococcus pyogenes* Cas9 gene, which contains a bidirectional promoter region, to construct a fusion protein expressed with PmCDA1 via an adapter sequence. A plasmid containing a chimeric gRNA encoding sequences complementary to each target nucleotide sequence was further fabricated (full-length nucleotide sequence shown in sequence number 32). The n-nucleotide sequence was then used to construct the plasmid. 20 Partial import of sequences complementary to each target sequence) Figure 11 A).
[0200] First, a plasmid targeting positions 426-445 of the *E. coli* galK gene ORF (T CAA TGG GCT AAC TAC GTT C; sequence number 33) was introduced and transformed into various *E. coli* strains (XL10-gold, DH5a, MG1655, BW25113) using the calcium method or electroporation. After overnight recovery culture in SOC medium, cells loaded with the plasmid were selected using LB medium containing ampicillin to form colonies. The mutation introduction was verified by direct sequencing of the colony using PCR. Results are shown below. Figure 11 B.
[0201] Independent colonies (1–3) were randomly selected and sequenced. The results confirmed that there was a greater than 60% probability that the C at position 427 of the ORF was changed to T (clones 2 and 3), thereby disrupting the gene that produces the stop codon (TAA).
[0202] Below, using the complementary sequence (5'-GGTCCATAAACTGAGACAGC-3'; sequence number 34) of the 1530-1549 base region of the rpoB gene ORF (an essential gene), a specific point mutation was introduced using the same method described above to attempt to confer rifampicin resistance in *E. coli*. Sequencing analysis of colonies selected from media containing non-selective medium (none), 25 μg / ml rifampicin (Rif25), and 50 μg / ml rifampicin (Rif50) confirmed that rifampicin resistance was conferred by introducing an amino acid mutation that changes Asp(GAC) to Asn(AAC) by replacing G with A at position 1546 of the ORF. Figure 11 C, top figure). Ten-fold dilutions of the transformed cell suspension were spotted onto media containing non-selective medium (none), 25 μg / ml rifampicin (Rif25), and 50 μg / ml rifampicin (Rif50) and cultured. The results showed that rifampicin-resistant strains were obtained at a frequency of approximately 10%. Figure 11 C, as shown in the image below).
[0203] In this way, according to this technology, not only can genes be disrupted, but new functions can also be added through specific point mutations. Furthermore, this technology also demonstrates its superiority in directly editing essential genes.
[0204] Example 10 Using gRNA length to regulate the editing site
[0205] Previously, gRNAs targeting specific nucleotide sequences were based on a 20-bit length, with cytosine (or guanine on the opposite strand) located from their 5' end to 2–5 bits (15–19 bits upstream of the PAM sequence) as mutation targets. This study examined whether the target bases shifted if a nucleic acid that altered the gRNA length was expressed. Figure 12 A).
[0206] An example of an experiment using Escherichia coli is shown below. Figure 12 B. Multiple cytosine positions were searched in the *E. coli* genome using the gsiA gene, a putative ABC-transfer protein. The substitution of cytosine at target lengths of 24 bp, 22 bp, 20 bp, and 18 bp was examined. At the standard length of 20 bp, cytosine at positions 898 and 899 was replaced by thymine. It was found that when the target site was longer than 20 bp, cytosine at positions 896 and 897 was also substituted; and when the target site was shorter, cytosine at positions 900 and 901 could also be substituted. In fact, the target site can be shifted by changing the length of the gRNA.
[0207] Example 11 Development of temperature-dependent genome editing plasmids
[0208] The nucleic acid-modifying enzyme complex of this invention is designed as a plasmid for expression induction under high-temperature conditions. The aim is to optimize efficiency by controlling the expression state and to mitigate side effects (host growth impairment, unstable mutation introduction efficiency, and mutations at sites different from the target). Simultaneously, by combining a mechanism that utilizes high temperature to stop plasmid replication, the intention is to allow for simultaneous and easy removal of the plasmid after editing. Experimental details are shown below.
[0209] The temperature-sensitive plasmid pSC101-Rep101 (sequence of pSC101 ori shown in sequence number 35, and sequence of temperature-sensitive Rep101 shown in sequence number 36) was used as the backbone, and expression was induced using a temperature-sensitive λ repressor (cI857) system. A RecA-resistant G113E mutation was introduced into the λ repressor used for genome editing, enabling normal function even under SOS response (sequence number 37). dCas9-PmCDA1 (sequence number 38) was ligated downstream of the RightOperator (sequence number 39), and gRNA (sequence number 40) was ligated downstream of the Left Operator (sequence number 41) to control expression (all nucleotide sequences of the constructed expression vector are shown in sequence number 42). During culture at temperatures below 30°C, cells proliferated normally due to the inhibition of gRNA transcription and dCas9-PmCDA1 expression, respectively. When cultured above 37°C, it induces gRNA transcription and dCas9-PmCDA1 expression while simultaneously inhibiting plasmid replication. Therefore, a transient supply of the nuclease complex necessary for genome editing allows for easy removal of the plasmid after editing. Figure 13 ).
[0210] The specific scheme for base substitution is shown in Figure 14 .
[0211] The culture temperature for plasmid construction was set at approximately 28°C. First, *E. coli* colonies containing the desired plasmid were established. Then, the colonies were used directly, or, if the strain was changed, the plasmid was extracted and transformed again into the target strain, and the resulting colonies were used. The colonies were then incubated in liquid medium at 28°C overnight. Subsequently, the culture was diluted and induced at 42°C for approximately 1 hour to overnight. The cell suspension was appropriately diluted and spread or spotted onto a plate to obtain single colonies.
[0212] As a validation experiment, a point mutation was introduced into rpoB, an essential gene. If rpoB, a component of RNA polymerase, is missing or disabled, *E. coli* cannot survive. On the other hand, it is known that resistance to the antibiotic rifampin can be acquired by introducing point mutations into specific sites. Therefore, a target site was selected for introducing such a point mutation, and assays were performed.
[0213] The results are shown in Figure 15 The top left image shows an LB plate (with chloramphenicol added) on the left and an LB plate (with rifampin added) on the right. These plates were used to prepare samples with and without chloramphenicol, incubated at 28°C–42°C. Although the proportion of Rif resistance was lower at 28°C, rifampin resistance was obtained with extremely high efficiency at 42°C. Sequencing of eight colonies (non-selected) obtained from actual LB plates revealed that over 60% of the strains incubated at 42°C had guanine (g) at position 1546 replaced by adenine (A) (bottom left and top right images). This indicates that the bases were also completely substituted in the actual sequencing profile (bottom right image).
[0214] Similarly, base substitutions were performed on galK, one of the factors involved in galactose metabolism. Since galK is lethal to *E. coli* by metabolizing 2-deoxygalactose (2DOG), a galactose analogue, it was used as a selection method. Target sites were set such that a missense mutation was induced at target 8, and target 12 became a stop codon. Figure 16 (Bottom right)
[0215] The results are shown in Figure 16 In the top left and bottom left images, the left side shows LB plates (with chloramphenicol added), and the right side shows LB plates (with chloramphenicol added) with 2-DOG added. These plates were used to prepare samples with and without chloramphenicol, incubated at 28°C and 42°C. In target 8, only a slight colony formation occurred in the 2-DOG-added plates (top left image), but sequencing of the three colonies on the LB plate (red box) confirmed that cytosine (C) at position 61 was replaced by thymine (T) in all colonies (top right). It is speculated that this mutation is insufficient to cause galK inactivation. On the other hand, in target 12, colonies were obtained in the 2-DOG-added plates under either 28°C or 42°C incubation conditions (bottom left image). Sequencing of the three colonies on the LB plate confirmed that cytosine at position 271 was replaced by thymine in all colonies (bottom right). This shows that even in such different targets, mutations can be introduced more stably and efficiently.
[0216] The contents contained in all publications of the patents and patent application specifications mentioned herein are incorporated herein by reference in their entirety and to the same extent as stated herein.
[0217] This application is based on Japanese Patent Application No. 2014-43348 and Japanese Patent Application No. 2014-201859, filed in Japan on March 5, 2014 and September 30, 2014, respectively, the contents of which are incorporated herein by reference in their entirety.
[0218] Industrial applicability
[0219] According to the present invention, site-specific mutations can be safely introduced into any species of organism without the insertion of exogenous DNA or the cutting of the DNA double strand. Furthermore, the range of mutation introduction can be broadly defined, from a single base pinpoint to several hundred bases, making it extremely useful for localized evolutionary induction by randomly introducing mutations into specific, defined regions that have been virtually impossible to perform until now. SEQUENCE LISTING <110> Kobe University (National University Corporation) <120> Methods of modifying genomic sequences by specifically altering the nucleic acid bases of the target DNA sequence, and the molecular complexes used. converting nucleobase in targeted DNA sequence and molecular complex usedtherefor) <130> 092301 <150> JP 2014-043348 <151> 2014-03-05 <150> JP 2014–201859 <151> 2014-09-30 <160> 42 <170> PatentIn version 3.5 <210> 1 <211> 624 <212> DNA <213> Seven-gilled eel (Petromyzon marinus) <220> <221> CDS <222> (1)..(624) <400> 1 atg acc gac gct gag tac gtg aga atc cat gag aag ttg gac atc tac 48 Met Thr Asp With Glu Tyr Val Arg With His Glu Lys Leu Asp With Tyr 1 5 10 15 acg ttt aag aaa cag ttt ttc aac aac aaa aaa tcc gtg tcg cat aga 96 Thr Phe Lys Lys Gln Phe Phe Asn Asn Lys Lys Ser Val Ser His Arg 20 25 30 tgc tac gtt ctc ttt gaa tta aaa cga cgg ggt gaa cgt aga gcg tgt 144 Cys Tyr Val Phe Glu Leu Lys Arg Gly Glu Arg Arg Ala Cys 35 40 45 ttt tgg ggc tat gct gtg aat aaa cca cag agc ggg aca gaa cgt ggc 192 Phe Trp Gly Tyr Ala Val Asn Lys Pro Gln Ser Gly Thr Glu Arg Gly 50 55 60 att cac gcc gaa atc ttt agc att aga aaa gtc gaa gaa tac ctg cgc 240 I Has Only Glu Ile Phe Ser Ile Arg Lys Val Glu Glu Tyr Leu Arg 65 70 75 80 gac aac ccc gga caa ttc acg ata aat tgg tac tca tcc tgg agt cct 288 Asp Asn Pro Gly Gln Phe Thr Ile Asn Trp Tyr Ser Ser Trp Ser Pro 85 90 95 tgt gca gat tgc gct gaa aag atc tta gaa tgg tat aac cag gag ctg 336 Cys Ala Asp Cys Ala Glu Lys Ile Leu Glu Trp Tyr Asn Gln Glu Leu 100 105 110 cgg ggg aac ggc cac act ttg aaa atc tgg gct tgc aaa ctc tat tac 384 Arg Gly Asn Gly His Thr Leu Lys Ile Trp Ala Cys Lys Leu Tyr Tyr 115 120 125 gag aaa aat gcg agg aat caa att ggg ctg tgg aac ctc aga gat aac 432 Glu Lys Asn Ala Arg Asn Gln Ile Gly Leu Trp Asn Leu Arg Asp Asn 130 135 140 ggg gtt ggg ttg aat gta atg gta agt gaa cac tac caa tgt tgc agg 480 Gly Val Gly Leu Asn Val Met Val Ser Glu His Tyr Gln Cys Cys Arg 145 150 155 160 aaa ata ttc atc caa tcg tcg cac aat caa ttg aat gag aat aga tgg 528 Lys Ile Phe Ile Gln Ser Ser His Asn Gln Leu Asn Glu Asn Arg Trp 165 170 175 ctt gag aag act ttg aag cga gct gaa aaa cga cgg agc gag ttg tcc 576 Leu Glu Lys Thr Leu Lys Arg Ala Glu Lys Arg Arg Ser Glu Leu Ser 180 185 190 att atg att cag gta aaa ata ctc cac acc act aag agt cct gct gtt 624 Ile Met Ile Gln Val Lys Ile Leu His Thr Thr Lys Ser Pro Ala Val 195 200 205 <210> 2 <211> 208 <212> PRT <213> Lamprey <400> 2 Met Thr Asp Ala Glu Tyr Val Arg Ile His Glu Lys Leu Asp Ile Tyr 1 5 10 15 Thr Phe Lys Lys Gln Phe Phe Asn Asn Lys Lys Ser Val Ser His Arg 20 25 30 Cys Tyr Val Leu Phe Glu Leu Lys Arg Arg Gly Glu Arg Arg Ala Cys 35 40 45 Phe Trp Gly Tyr Ala Val Asn Lys Pro Gln Ser Gly Thr Glu Arg Gly 50 55 60 Ile His Ala Glu Ile Phe Ser Ile Arg Lys Val Glu Glu Tyr Leu Arg 65 70 75 80 Asp Asn Pro Gly Gln Phe Thr Ile Asn Trp Tyr Ser Ser Trp Ser Pro 85 90 95 Cys Ala Asp Cys Ala Glu Lys Ile Leu Glu Trp Tyr Asn Gln Glu Leu 100 105 110 Arg Gly Asn Gly His Thr Leu Lys Ile Trp Ala Cys Lys Leu Tyr Tyr 115 120 125 Glu Lys Asn Ala Arg Asn Gln Ile Gly Leu Trp Asn Leu Arg Asp Asn 130 135 140 Gly Val Gly Leu Asn Val Met Val Ser Glu His Tyr Gln Cys Cys Arg 145 150 155 160 Lys Ile Phe Ile Gln Ser Ser His Asn Gln Leu Asn Glu Asn Arg Trp 165 170 175 Leu Glu Lys Thr Leu Lys Arg Ala Glu Lys Arg Arg Ser Glu Leu Ser 180 185 190 Ile Met Ile Gln Val Lys Ile Leu His Thr Thr Lys Ser Pro Ala Val 195 200 205 <210> 3 <211> 600 <212> DNA <213> Homo sapiens <220> <221> CDS <222> (1)..(600) <400> 3 atg gac agc ctc ttg atg aac cgg agg aag ttt ctt tac caa ttc aaa 48 Met Asp Ser Leu Leu Met Asn Arg Arg Lys Phe Leu Tyr Gln Phe Lys 1 5 10 15 aat gtc cgc tgg gct aag ggt cgg cgt gag acc tac ctg tgc tac gta 96 Asn Val Arg Trp Ala Lys Gly Arg Arg Glu Thr Tyr Leu Cys Tyr Val 20 25 30 gtg aag agg cgt gac agt gct aca tcc ttt tca ctg gac ttt ggt tat 144 Val Lys Arg Arg Asp Ser Ala Thr Ser Phe Ser Leu Asp Phe Gly Tyr 35 40 45 ctt cgc aat aac ggc tgc cac gtg gaa ttg ctc ttc ctc cgc tac 192 Arg Asn Lys Asn Gly Cys His Val Glu Arg Phe Arg Tyr 50 55 60 atc tcg gac tgg gac cta gac cct ggc cgc tgc tac cgc gtc acc tgg 240 Serving Asp Trp Asp Leu Asp Pro Gly Arg Cys Tyr Arg Val Thr Trp 65 70 75 80 ttc acc tcc tgg agc ccc tgc tac gac tgt gcc cga cat gtg gcc gac 288 Phe Thr Ser Trp Ser Pro Cys Tyr Asp Cys Ala Arg His Val Ala Asp 85 90 95 ttt ctg cga ggg aac ccc tac ctc agt ctg agg atc ttc acc gcg cgc 336 Phe Leu Arg Gly Asn Pro Tyr Leu Ser Leu Arg Ile Phe Thr Ala Arg 100 105 110 ctc tac ttc tgt gag gac cgc aag gct gag ccc gag ggg ctg cgg cgg 384 Leu Tyr Phe Cys Glu Asp Arg Lys Ala Glu Pro Glu Gly Leu Arg Arg 115 120 125 ctg cac cgc gcc ggg gtg caa ata gcc atc atg acc ttc aaa gat tat 432 Leu His Arg Ala Gly Val Gln Ile Ala Ile Met Thr Phe Lys Asp Tyr 130 135 140 ttt tac tgc tgg aat act ttt gta gaa aac cat gaa aga act ttc aaa 480 Phe Tyr Cys Trp Asn Thr Phe Val Glu Asn His Glu Arg Thr Phe Lys 145 150 155 160 gcc tgg gaa ggg ctg cat gaa aat tca gtt cgt ctc tcc aga cag ctt 528 Ala Trp Glu Gly Leu His Glu Asn Ser Val Arg Leu Ser Arg Gln Leu 165 170 175 cgg cgc atc ctt ttg ccc ctg tat gag gtt gat gac tta cga gac gca 576 Arg Arg Ile Leu Leu Pro Leu Tyr Glu Val Asp Asp Leu Arg Asp Ala 180 185 190 ttt cgt act ttg gga ctt ctc gac 600 Phe Arg Thr Leu Gly Leu Leu Asp 195 200 <210> 4 <211> 200 <212> PRT <213> Human <400> 4 Met Asp Ser Leu Leu Met Asn Arg Arg Lys Phe Leu Tyr Gln Phe Lys 1 5 10 15 Asn Val Arg Trp Ala Lys Gly Arg Arg Glu Thr Tyr Leu Cys Tyr Val 20 25 30 Val Lys Arg Arg Asp Ser Ala Thr Ser Phe Ser Leu Asp Phe Gly Tyr 35 40 45 Leu Arg Asn Lys Asn Gly Cys His Val Glu Leu Leu Phe Leu Arg Tyr 50 55 60 Ile Ser Asp Trp Asp Leu Asp Pro Gly Arg Cys Tyr Arg Val Thr Trp 65 70 75 �0 Phe Thr Ser Trp Ser Pro Cys Tyr Asp Cys Ala Arg His Val Ala Asp 85 90 95 Phe Leu Arg Gly Asn Pro Tyr Leu Ser Leu Arg Ile Phe Thr Ala Arg 100 105 110 Leu Tyr Phe Cys Glu Asp Arg Lys Ala Glu Pro Glu Gly Leu Arg Arg 115 120 125 Leu His Arg Ala Gly Val Gln Ile Ala Ile Met Thr Phe Lys Asp Tyr 130 135 140 Phe Tyr Cys Trp Asn Thr Phe Val Glu Asn His Glu Arg Thr Phe Lys 145 150 155 160 Ala Trp Glu Gly Leu His Glu Asn Ser Val Arg Leu Ser Arg Gln Leu 165 170 175 Arg Arg Ile Leu Leu Pro Leu Tyr Glu Val Asp Asp Leu Arg Asp Ala 180 185 190 Phe Arg Thr Leu Gly Leu Leu Asp 195 200 <210> 5 <211> 4116 <212> DNA <213> Artificial sequence <220> <223> Cas9 CDS derived from Streptococcus pyogenes, optimized for eukaryotic expression. <220> <221> CDS <222> (1)..(4116) <400> 5 atg gac aag aag tac tcc att ggg ctc gat atc ggc aca aac agc gtc 48 Met Asp Lys Lys Tyr Ser Ile Gly Leu Asp Ile Gly Thr Asn Ser Val 1 5 10 15 ggt tgg gcc gtc att acg gac gag tac aag gtg ccg agc aaa aaa ttc 96 Gly Trp Ala Val Ile Thr Asp Glu Tyr Lys Val Pro Ser Lys Lys Phe 20 25 30 aaa gtt ctg ggc aat acc gat cgc cac agc ata aag aag aac ctc att 144 Lys Val Leu Gly Asn Thr Asp Arg His Ser Ile Lys Asn Leu Ile 35 40 45 ggc gcc ctc ctg ttc gac tcc ggg gag acg gcc gaa gcc acg cgg ctc 192 Gly Ala Leu Phe Asp Ser Gly Glu Thr Ala Glu Ala Thr Arg Leu 50 55 60 aaa aga aca gca cgg cgc aga tat acc cgc aga aag aat cgg atc tgc 240 Lys Arg Thr Ala Arg Arg Tyr Thr Arg Arg Lys Asn Arg With Cys 65 70 75 80 tac ctg cag gag atc ttt agt aat gag atg gct aag gtg gat gac tct 288 Tyr Leu Gln Glu Ile Phe Ser Asn Glu Met Ala Lys Val Asp Asp Ser 85 90 95 ttc ttc cat agg ctg gag gag tcc ttt ttg gtg gag gag gat aaa aag 336 Phe Phe His Arg Leu Glu Glu Ser Phe Leu Val Glu Glu Asp Lys Lys 100 105 110 cac gag cgc cac cca atc ttt ggc aat atc gtg gac gag gtg gcg tac 384 His Glu Arg His Pro Ile Phe Gly Asn Ile Val Asp Glu Val Ala Tyr 115 120 125 cat gaa aag tac cca acc ata tat cat ctg agg aag aag ctt gta gac 432 His Glu Lys Tyr Pro Thr Ile Tyr His Leu Arg Lys Lys Leu Val Asp 130 135 140 agt act gat aag gct gac ttg cgg ttg atc tat ctc gcg ctg gcg cat 480 Ser Thr Asp Lys Ala Asp Leu Arg Leu Ile Tyr Leu Ala Leu Ala His 145 150 155 160 atg atc aaa ttt cgg gga cac ttc ctc atc gag ggg gac ctg aac cca 528 Met Ile Lys Phe Arg Gly His Phe Leu Ile Glu Gly Asp Leu Asn Pro 165 170 175 gac aac agc gat gtc gac aaa ctc ttt atc caa ctg gtt cag act tac 576 Asp Asn Ser Asp Val Asp Lys Leu Phe Ile Gln Leu Val Gln Thr Tyr 180 185 190 aat cag ctt ttc gaa gag aac ccg atc aac gca tcc gga gtt gac gcc 624 Asn Gln Leu Phe Glu Glu Asn Pro Ile Asn Ala Ser Gly Val Asp Ala 195 200 205 aaa gca atc ctg agc gct agg ctg tcc aaa tcc cgg cgg ctc gaa aac 672 Lys Ala Ile Leu Ser Ala Arg Leu Ser Lys Ser Arg Arg Leu Glu Asn 210 215 220 ctc atc gca cag ctc cct ggg gag aag aag aac ggc ctg ttt ggt aat 720 Leu Ile Ala Gln Leu Pro Gly Glu Lys Lys Asn Gly Leu Phe Gly Asn 225 230 235 240 ctt atc gcc ctg tca ctc ggg ctg acc ccc aac ttt aaa tct aac ttc 768 Leu Ile Ala Leu Ser Leu Gly Leu Thr Pro Asn Phe Lys Ser Asn Phe 245 250 255 gac ctg gcc gaa gat gcc aag ctt caa ctg agc aaa gac acc tac gat 816 Asp Leu Ala Glu Asp Ala Lys Leu Gln Leu Ser Lys Asp Thr Tyr Asp 260 265 270 gat gat ctc gac aat ctg ctg gcc cag atc ggc gac cag tac gca gac 864 Asp Asp Leu Asp Asn Leu Leu Ala Gln Ile Gly Asp Gln Tyr Ala Asp 275 280 285 ctt ttt ttg gcg gca aag aac ctg tca gac gcc att ctg ctg agt gat 912 Leu Phe Leu Ala Ala Lys Asn Leu Ser Asp Ala Ile Leu Leu Ser Asp 290 295 300 to ctg cga gtg aac acg gag atc acc aaa gct ccg ctg agc gct agt 960 Ile Leu Arg Val Asn Thr Glu Ile Thr Lys Ala Pro Leu Ser Ala Ser 305 310 315 320 atg atc aag cgc tat gat gag cac cac caa gac ttg act ttg ctg aag 1008 Met Ile Lys Arg Tyr Asp Glu His His Gln Asp Leu Thr Leu Leu Lys 325 330 335 gcc ctt gtc aga cag caa ctg cct gag aag tac aag gaa to ttc ttc 1056 Ala Leu Val Arg Gln Gln Leu Pro Glu Lys Tyr Lys Glu Ile Phe Phe 340 345 350 gat cag tct aaa aat ggc tac gcc gga tac att gac ggc gga gca agc 1104 Asp Gln Ser Lys Asn Gly Tyr Ala Gly Tyr Ile Asp Gly Gly Ala Ser 355 360 365 cag gag gaa ttt tac aaa ttt att aag ccc atc ttg gaa aaa atg gac 1152 Gln Glu Glu Phe Tyr Lys Phe Ile Lys Pro Ile Leu Glu Lys Met Asp 370 375 380 ggc acc gag gag ctg ctg gta aag ctt aac aga gaa gat ctg ttg cgc 1200 Gly Thr Glu Glu Leu Leu Val Lys Leu Asn Arg Glu Asp Leu Leu Arg 385 390 395 400 aaa cag cgc act ttc gac aat gga agc atc ccc cac cag att cac ctg 1248 Lys Gln Arg Thr Phe Asp Asn Gly Ser Ile Pro His Gln Ile His Leu 405 410 415 ggc gaa ctg cac gct atc ctc agg cgg caa gag gat ttc tac ccc ttt 1296 Gly Glu Leu His Ala Ile Leu Arg Arg Gln Glu Asp Phe Tyr Pro Phe 420 425 430 ttg aaa gat aac agg gaa aag att gag aaa atc ctc aca ttt cgg ata 1344 Leu Lys Asp Asn Arg Glu Lys Ile Glu Lys Ile Leu Thr Phe Arg Ile 435 440 445 ccc tac tat gta ggc ccc ctc gcc cgg gga aat tcc aga ttc gcg tgg 1392 Pro Tyr Tyr Val Gly Pro Leu Ala Arg Gly Asn Ser Arg Phe Ala Trp 450 455 460 atg act cgc aaa tca gag acc atc act ccc tgg aac ttc gag gaa 1440 Met Thr Arg Lys Ser Glu Glu Thr Ile Thr Pro Trp Asn Phe Glu Glu 465 470 475 480 gtc gtg gat aag ggg gcc tct gcc cag tcc ttc atc gaa agg atg act 1488 Val Val Asp Lys Gly Ser Ala Gln Ser Phe Ile Glu Arg Met Thr 485,490,495 aac ttt gat aaa aat ctg cct aac gaa aag gtg ctt cct aaa cac tct Asn Phe Asp Lys Asn Leu Pro Asn Glu Lys Val Leu Pro Lys His Ser 500 505 510 ctg ctg tac gag tac ttc aca gtt tat aac gag ctc acc aag gtc aaa 1584 Leu Tyr Glu Tyr Phe Thr Val Tyr Asn Glu Thr Lys Val Lys 515,520,525 tac gtc aca gaa ggg atg aga aag cca gca ttc ctg tct gga gag cag 1632 Tyr Val Thr Glu Gly Met Arg Lys Pro Ala Phe Leu Ser Gly Glu Gln 530 535 540 aag aaa gct atc gtg gac ctc ctc ttc aag acg aac cgg aaa gtt acc 1680 Lys Lys Ala Ile Val Asp Leu Leu Phe Lys Thr Asn Arg Lys Val Thr 545 550 555 560 gtg aaa cag ctc aaa gaa gac tat ttc aaa aag att gaa tgt ttc gac 1728 Val Lys Gln Leu Lys Glu Asp Tyr Phe Lys Lys Ile Glu Cys Phe Asp 565 570 575 tct gtt gaa atc agc gga gtg gag gat cgc ttc aac gca tcc ctg gga 1776 Ser Val Glu Ile Ser Gly Val Glu Asp Arg Phe Asn Ala Ser Leu Gly 580 585 590 acg tat cac gat ctc ctg aaa atc att aaa gac aag gac ttc ctg gac 1824 Thr Tyr His Asp Leu Leu Lys Ile Ile Lys Asp Lys Asp Phe Leu Asp 595 600 605 aat gag gag aac gag gac att ctt gag gac att gtc ctc acc ctt acg 1872 Asn Glu Glu Asn Glu Asp Ile Leu Glu Asp Ile Val Leu Thr Leu Thr 610 615 620 ttg ttt gaa gat agg gag atg att gaa gaa cgc ttg aaa act tac gct 1920 Leu Phe Glu Asp Arg Glu Met Ile Glu Glu Arg Leu Lys Thr Tyr Ala 625 630 635 640 cat ctc ttc gac gac aaa gtc atg aaa cag ctc aag agg cgc cga tat 1968 His Leu Phe Asp Asp Lys Val Met Lys Gln Leu Lys Arg Arg Arg Tyr 645 650 655 aca gga tgg ggg cgg ctg tca aga aaa ctg atc aat ggg atc cga gac 2016 Thr Gly Trp Gly Arg Leu Ser Arg Lys Leu Ile Asn Gly Ile Arg Asp 660 665 670 aag cag agt gga aag aca atc ctg gat ttt ctt aag tcc gat gga ttt 2064 Lys Gln Ser Gly Lys Thr Ile Leu Asp Phe Leu Lys Ser Asp Gly Phe 675 680 685 gcc aac cgg aac ttc atg cag ttg atc cat gat gac tct ctc acc ttt 2112 Ala Asn Arg Asn Phe Met Gln Leu Ile His Asp Asp Ser Leu Thr Phe 690 695 700 aag gag gac atc cag aaa gca caa gtt tct ggc cag ggg gac agt ctt 2160 Lys Glu Asp With Gln Lys Ala Gln Val Ser Gly Gln Gly Asp Ser Leu 705 710 715 720 cac gag cac atc gct aat ctt gca ggt agc cca gct atc aaa aag gga 2208 His Glu His Ile Wing Asn Leu Wing Gly Ser Pro Wing Ile Lys Lys Gly 725 730 735 ata ctg cag acc gtt aag gtc gtg gat gaa ctc gtc aaa gta atg gga 2256 Ile Leu Gln Thr Val Lys Val Val Asp Glu Leu Val Lys Val Met Gly 740,745,750 agg cat aag ccc gag atc gtt atc gag atg gcc cga gag aac caa 2304 Arg His Pro Glu Asn Ile Val Ile Glu Met Ala Arg Glu Asn Gln 755,760,765 act acc cag aag gga cag aag aac agt agg gaa agg atg aag agg att 2352 Thr Thr Gln Lys Gly Gln Lys Asn Ser Arg Glu Arg Met Lys Arg Ile 770,775,780 gaa gag ggt ata aaa gaa ctg ggg tcc CA atc ctt aag gaa cac cca 2400 Glu Glu Gly Ile Lys Glu Leu Gly Ser Gln Ile Leu Lys Glu His Pro 785,790,795,800 gtt gaa aac acc cag ctt cag aat gag aag ctc tac ctg tac tac ctg 2448 Val Glu Asn Thr Gln Leu Gln Asn Glu Lys Leu Tyr Leu Tyr 805 810 815 cag aac ggc agg gac atg tac gtg gat cag gaa ctg gac atc atc at cgg 2496 Gln Asn Gly Arg Asp Met Tyr Val Asp Gln Glu Leu Asp With Asn Arg 820 825 830 ctc tcc gac tac gac gtg gat cat atc gtg ccc cag tct ttt ctc aaa 2544 Leu Ser Asp Tyr Asp Val Asp His Ile Val Pro Gln Ser Phe Leu Lys 835 840 845 gat gat tct att gat aat aaa gtg ttg aca tcc gat aaa at aga Asp Asp Ser Ile Asp Asn Lys Val Leu Thr Arg Ser Asp Lys Asn Arg 850 855 860 ggg aag agt gat aac gtc ccc tca gaa gaa gtt gtc aag aaa atg aaa 2640 Gly Lys Ser Asp Asn Val Pro Ser Glu Glu Val Lys Lys Met Lys 865 870 875 880 aat tat tgg cgg cg cctg aac gcc aaa ctg atc aca CA cgg aag 2688 Arg Tyr Asn To Lys Arg Asn With Thr Gln Arg Lys 885,890,895 ttc gat aat ctg act aag gct gaa cga ggt ggc ctg tct gag ttg gat 2736 Phe Asp Asp With Thr Lys Ala Glu Arg Gly Gly Leu Ser Glu With Asp 900 905 910 aaa gcc gcc ttc atc aaa agg cag ctt gtt gag aca cgc cag atc acc 2784 Lys Ala Gly Phe Ile Lys Arg Gln Leu Val Glu Thr Arg Gln Ile Thr 915,920,925 aag cac gtg gcc caa att ctc gat tca cgc atg aac acc aag tac gat 2832 Lys Tyr Asp Ser Arg Met Asn Thr Lys Tyr Asp 930,935,940 gaa aat gac aaa ctg att cga gag gtg aaa gtt att act ctg aag tct 2880 Glu Asn Asp Lys Leu Ile Arg Glu Val Lys Val Ile Thr Lys Ser 945 950 955 960 aag ctg gtc tca gat ttc aga aag gac ttt cag ttt tat aag gtg aga Lys Leu Val Ser Asp Phe Arg Lys Asp Phe Gln Phe Tyr Lys Val Arg 965,970,975 gag atc aac aat tac cac cat gcg cat gat gcc tac ctg aat gca gtg 2976 Glu Ile Asn Asn Tyr His Ala His Asp Ala Tyr Leu Asn Ala Val 980,985,990 gta ggc act gca ctt atc aaa aaa tat ccc aag ctt gaa tct gaa ttt 3024 Val Gly Thr Ala Leu Ile Lys Tyr Pro Lys Leu Glu Ser Glu Phe 995 1000 1005 gtt tac gga gac tat aaa gtg tac gat gtt agg aaa atg atc gca 3069 Val Tyr Gly Asp Tyr Lys Val Tyr Asp Val Arg Lys Met Ile Ala 1010 1015 1020 aag tct gag cag gaa ata ggc aag gcc acc gct aag tac ttc ttt 3114 Lys Ser Glu Gln Glu Ile Gly Lys Ala Thr Ala Lys Tyr Phe Phe 1025 1030 1035 tac agc aat att atg aat ttttc aag acc gag att aca ctg gcc 3159 Tyr Ser Asn Ile Met Asn Phe Phe Lys Thr Glu Ile Thr Leu Ala 1040 1045 1050 aat gga gag att cgg aag cga cca ctt atc gaa aca aac gga gaa 3204 Asn Gly Glu With Arg Lys Arg Pro Leu Glu Thr Asn Gly Glu 1055 1060 1065 aca gga gaa atc gtg tgg gac aag ggt agg gat ttc gcg aca gtc 3249 Thr Gly Glu Ile Val Trp Asp Lys Gly Arg Asp Phe Ala Thr Val 1070 1075 1080 cgg aag gtc ctg tcc atg cccg cag gtg aac atc gtt aaa aag acc 3294 Arg Lys Will Be Met Pro Gln Val Asn Ile Val Lys Thr 1085 1090 1095 gaa gta cag acc gga ggc ttc tcc aag gaa agt atc ctc ccg aaa Glu Val Gln Thr Gly Gly Phe Ser Lys Glu Ser Ile Leu Pro Lys 1100 1105 1110 agg aac agc gac aag ctg atc gca cgc aaa aaa gat tgg gac ccc 3384 Arg Asn Ser Asp Lys Leu Ile Ala Arg Lys Lys Asp Trp Asp Pro 1115 1120 1125 aag aaa tac ggc gga ttc gat tct cct aca gtc gct tac agt gta Lys Lys Tyr Gly Gly Phe Asp Ser Pro Thr Val Ala Tyr Ser Val 1130 1135 1140 ctg gtt gtg gcc aaa gtg gag aaa ggg aag tct aaa aaa ctc aaa 3474 Leu Val Val Ala Lys Val Glu Lys Gly Lys Ser Lys Lys Leu Lys 1145 1150 1155 agc gtc aag gaa ctg ctg ggc atc aca atc atc atg gag cga tca agc 3519 Ser Val Lys Glu Leu Leu Gly Ile Thr Ile Met Glu Arg Ser Ser 1160 1165 1170 ttc gaa aaa aac ccc atc gac ttt ctc gag gcg aaa gga tat aaa 3564 Phe Glu Lys Asn Pro Ile Asp Phe Leu Glu Ala Lys Gly Tyr Lys 1175 1180 1185 gag gtc aaa aaa gac ctc atc att aag ctt ccc aag tac tct ctc 3609 Glu Val Lys Asp Leu Ile Ile Leu Pro Lys Tyr Ser Leu 1190 1195 1200 ttt gag ctt gaa aac ggc cgg aaa cga atg ctc gct agt gcg ggc 3654 Phe Glu Leu Glu Asn Gly Arg Lys Arg Met Leu Ala Ser Ala Gly 1205 1210 1215 gag ctg cag aaa ggt aac gag ctg gca ctg ccc tct aaa tac gtt 3699 Leu Glu Gln Lys Gly Asn Leu Glu Wing Pro Ser Tyr Val 1220 1225 1230 aat ttc ttg tat ctg gcc agc cac tat gaa aag ctc aaa ggg tct 3744 Asn Phe Leu Tyr Leu Ala Ser is Tyr Glu Lys Leu Lys Gly Ser 1235 1240 1245 ccc GA gat at gag cag aag cag ctg ttc gtg GA CA cac aaa 3789 Pro Glu Asp Asn Glu Gln Lys Gln Leu Phe Val Glu Gln His Lys 1250 1255 1260 cac tac ctt gat gag atc atc gag caa ata agc gaa ttc tcc aaa 3834 His Tyr Leu Asp Glu Ile Ile Glu Gln Ile Ser Glu Phe Ser Lys 1265 1270 1275 aga gtg atc ctc gcc gac gct aac ctc gat aag gtg ctt tct gct 3879 Arg Valley Leu Valley Asp Ala Asn Leu Asp Lys Valley Leu Ser Ala 1280 1285 1290 Tac AAT AAG CAC Agg GAT AAG CCC ATC Agg GAG CAG GCA GA AAC 3924 Tyr Asn Lys His Arg Asp Lys Pro Ile Arg Glu Gln Ala Glu Asn 1295 1300 1305 atc cac ttg ttt act ctg acc aac ttg ggc gcg cct gca gcc 3969 Ile Ile His Leu Phe Thr Leu Thr Asn Leu Gly Ala Pro Ala Ala 1310 1315 1320 ttc aag tac ttc gac acc acc ata gac aga aag cgg tac acc tct 4014 Phe Lys Tyr Phe Asp Thr Thr Ile Asp Arg Lys Arg Tyr Thr Ser 1325 1330 1335 aca aag gag gtc ctg gac gcc aca ctg att cat cag tca att acg 4059 Thr Lys Glu Val Leu Asp Ala Thr Leu Ile His Gln Ser Ile Thr 1340 1345 1350 ggg ctc tat gaa aca aga atc gac ctc tct cag ctc ggt gga gac 4104 Gly Leu Tyr Glu Thr Arg Ile Asp Leu Ser Gln Leu Gly Gly Asp 1355 1360 1365 agc agg gct gac 4116 Looking Angry Ala Asp 1370 <210> 6 <211> 1372 <212> PRT <213> artificial sequence <220> <223> Synthetic Construct <400> 6 Met Asp Lys Lys Tyr Ser Ile Gly Leu Asp Ile Gly Thr Asn Ser Val 1 5 10 15 Gly Trp Ala Val Ile Thr Asp Glu Tyr Lys Val Pro Ser Lys Lys Phe 20 25 30 Lys Val Leu Gly Asn Thr Asp Arg His Ser Ile Lys Lys Asn Leu Ile 35 40 45 Gly Ala Leu Leu Phe Asp Ser Gly Glu Thr Ala Glu Ala Thr Arg Leu 50 55 60 Lys Arg Thr Ala Arg Arg Arg Tyr Thr Arg Arg Lys Asn Arg Ile Cys 65 70 75 80 Tyr Leu Gln Glu Ile Phe Ser Asn Glu Met Ala Lys Val Asp Asp Ser 85 90 95 Phe Phe His Arg Leu Glu Glu Ser Phe Leu Val Glu Glu Asp Lys Lys 100 105 110 His Glu Arg His Pro Ile Phe Gly Asn Ile Val Asp Glu Val Ala Tyr 115 120 125 His Glu Lys Tyr Pro Thr Ile Tyr His Leu Arg Lys Lys Leu Val Asp 130 135 140 Ser Thr Asp Lys Ala Asp Leu Arg Leu Ile Tyr Leu Ala Leu Ala His 145 150 155 160 Met Ile Lys Phe Arg Gly His Phe Leu Ile Glu Gly Asp Leu Asn Pro 165 170 175 Asp Asn Ser Asp Val Asp Lys Leu Phe Ile Gln Leu Val Gln Thr Tyr 180 185 190 Asn Gln Leu Phe Glu Glu Asn Pro Ile Asn Ala Ser Gly Val Asp Ala 195 200 205 Lys Ala Ile Leu Ser Ala Arg Leu Ser Lys Ser Arg Arg Leu Glu Asn 210 215 220 Leu Ile Ala Gln Leu Pro Gly Glu Lys Lys Asn Gly Leu Phe Gly Asn 225 230 235 240 Leu Ile Ala Leu Ser Leu Gly Leu Thr Pro Asn Phe Lys Ser Asn Phe 245 250 255 Asp Leu Ala Glu Asp Ala Lys Leu Gln Leu Ser Lys Asp Thr Tyr Asp 260 265 270 Asp Asp Leu Asp Asn Leu Leu Ala Gln Ile Gly Asp Gln Tyr Ala Asp 275 280 285 Leu Phe Leu Ala Ala Lys Asn Leu Ser Asp Ala Ile Leu Leu Ser Asp 290 295 300 Ile Leu Arg Val Asn Thr Glu Ile Thr Lys Ala Pro Leu Ser Ala Ser 305 310 315 320 Met Ile Lys Arg Tyr Asp Glu His His Gln Asp Leu Thr Leu Leu Lys 325 330 335 Ala Leu Val Arg Gln Gln Leu Pro Glu Lys Tyr Lys Glu Ile Phe Phe 340 345 350 Asp Gln Ser Lys Asn Gly Tyr Ala Gly Tyr Ile Asp Gly Gly Ala Ser 355 360 365 Gln Glu Glu Phe Tyr Lys Phe Ile Lys Pro Ile Leu Glu Lys Met Asp 370 375 380 Gly Thr Glu Glu Leu Leu Val Lys Leu Asn Arg Glu Asp Leu Leu Arg 385 390 395 400 Lys Gln Arg Thr Phe Asp Asn Gly Ser Ile Pro His Gln Ile His Leu 405 410 415 Gly Glu Leu His Ala Ile Leu Arg Arg Gln Glu Asp Phe Tyr Pro Phe 420 425 430 Leu Lys Asp Asn Arg Glu Lys Ile Glu Lys Ile Leu Thr Phe Arg Ile 435 440 445 Pro Tyr Tyr Val Gly Pro Leu Ala Arg Gly Asn Ser Arg Phe Ala Trp 450 455 460 Met Thr Arg Lys Ser Glu Glu Thr Ile Thr Pro Trp Asn Phe Glu Glu 465 470 475 480 Val Val Asp Lys Gly Ala Ser Ala Gln Ser Phe Ile Glu Arg Met Thr 485 490 495 Asn Phe Asp Lys Asn Leu Pro Asn Glu Lys Val Leu Pro Lys His Ser 500 505 510 Leu Leu Tyr Glu Tyr Phe Thr Val Tyr Asn Glu Leu Thr Lys Val Lys 515 520 525 Tyr Val Thr Glu Gly Met Arg Lys Pro Ala Phe Leu Ser Gly Glu Gln 530 535 540 Lys Lys Ala Ile Val Asp Leu Leu Phe Lys Thr Asn Arg Lys Val Thr 545 550 555 560 Val Lys Gln Leu Lys Glu Asp Tyr Phe Lys Lys Ile Glu Cys Phe Asp 565 570 575 Ser Val Glu Ile Ser Gly Val Glu Asp Arg Phe Asn Ala Ser Leu Gly 580 585 590 Thr Tyr His Asp Leu Leu Lys Ile Ile Lys Asp Lys Asp Phe Leu Asp 595 600 605 Asn Glu Glu Asn Glu Asp Ile Leu Glu Asp Ile Val Leu Thr Leu Thr 610 615 620 Leu Phe Glu Asp Arg Glu Met Ile Glu Glu Arg Leu Lys Thr Tyr Ala 625 630 635 640 His Leu Phe Asp Asp Lys Val Met Lys Gln Leu Lys Arg Arg Arg Tyr 645 650 655 Thr Gly Trp Gly Arg Leu Ser Arg Lys Leu Ile Asn Gly Ile Arg Asp 660 665 670 Lys Gln Ser Gly Lys Thr Ile Leu Asp Phe Leu Lys Ser Asp Gly Phe 675 680 685 Ala Asn Arg Asn Phe Met Gln Leu Ile His Asp Asp Ser Leu Thr Phe 690 695 700 Lys Glu Asp Ile Gln Lys Ala Gln Val Ser Gly Gln Gly Asp Ser Leu 705 710 715 720 His Glu His Ile Ala Asn Leu Ala Gly Ser Pro Ala Ile Lys Lys Gly 725 730 735 Ile Leu Gln Thr Val Lys Val Val Asp Glu Leu Val Lys Val Met Gly 740 745 750 Arg His Lys Pro Glu Asn Ile Val Ile Glu Met Ala Arg Glu Asn Gln 755 760 765 Thr Thr Gln Lys Gly Gln Lys Asn Ser Arg Glu Arg Met Lys Arg Ile 770 775 780 Glu Glu Gly Ile Lys Glu Leu Gly Ser Gln Ile Leu Lys Glu His Pro 785 790 795 800 Val Glu Asn Thr Gln Leu Gln Asn Glu Lys Leu Tyr Leu Tyr Tyr Leu 805 810 815 Gln Asn Gly Arg Asp Met Tyr Val Asp Gln Glu Leu Asp Ile Asn Arg 820 825 830 Leu Ser Asp Tyr Asp Val Asp His Ile Val Pro Gln Ser Phe Leu Lys 835 840 845 Asp Asp Ser Ile Asp Asn Lys Val Leu Thr Arg Ser Asp Lys Asn Arg 850 855 860 Gly Lys Ser Asp Asn Val Pro Ser Glu Glu Val Val Lys Lys Met Lys 865 870 875 880 Asn Tyr Trp Arg Gln Leu Leu Asn Ala Lys Leu Ile Thr Gln Arg Lys 885 890 895 Phe Asp Asn Leu Thr Lys Ala Glu Arg Gly Gly Leu Ser Glu Leu Asp 900 905 910 Lys Ala Gly Phe Ile Lys Arg Gln Leu Val Glu Thr Arg Gln Ile Thr 915 920 925 Lys Tyr Asp Ser Arg Met Asn Thr Lys Tyr Asp 930,935,940 Glu Asn Asp Lys Leu Ile Arg Glu Val Lys Val Ile Thr Lys Ser 945 950 955 960 Lys Leu Val Ser Asp Phe Arg Lys Asp Phe Gln Phe Tyr Lys Val Arg 965,970,975 Glu Ile Asn Asn Tyr His Ala His Asp Ala Tyr Leu Asn Ala Val 980,985,990 Val Gly Thr Ala Leu Ile Lys Tyr Pro Lys Leu Glu Ser Glu Phe 995 1000 1005 Val Tyr Gly Asp Tyr Lys Val Tyr Asp Val Arg Lys Met Ile Ala 1010 1015 1020 Lys Ser Glu Gln Glu Ile Gly Lys Ala Thr Ala Lys Tyr Phe Phe 1025 1030 1035 Tyr Ser Asn Ile Met Asn Phe Phe Lys Thr Glu Ile Thr Leu Ala 1040 1045 1050 Asn Gly Glu With Arg Lys Arg Pro Leu Glu Thr Asn Gly Glu 1055 1060 1065 Thr Gly Glu Ile Val Trp Asp Lys Gly Arg Asp Phe Ala Thr Val 1070 1075 1080 Arg Lys Val Leu Ser Met Pro Gln Val Asn Ile Val Lys Lys Thr 1085 1090 1095 Glu Val Gln Thr Gly Gly Phe Ser Lys Glu Ser Ile Leu Pro Lys 1100 1105 1110 Arg Asn Ser Asp Lys Leu Ile Ala Arg Lys Lys Asp Trp Asp Pro 1115 1120 1125 Light Light Tyr Gly Gly Phe Asp Ser Pro Thr Val Ala Tyr Ser Val 1130 1135 1140 Leu Val Val Ala Lys Val Glu Lys Gly Lys Ser Lys Lys Leu Lys 1145 1150 1155 Ser Val Lys Glu Leu Leu Gly Ile Thr Ile Met Glu Arg Ser Ser 1160 1165 1170 Phe Glu Lys Asn Pro Ile Asp Phe Leu Glu Ala Lys Gly Tyr Lys 1175 1180 1185 Glu Val Lys Lys Asp Leu Ile Ile Lys Leu Pro Lys Tyr Ser Leu 1190 1195 1200 Phe Glu Leu Glu Asn Gly Arg Lys Arg Met Leu Ala Ser Ala Gly 1205 1210 1215 Glu Leu Gln Lys Gly Asn Glu Leu Ala Leu Pro Ser Lys Tyr Val 1220 1225 1230 Asn Phe Leu Tyr Leu Ala Ser His Tyr Glu Lys Leu Lys Gly Ser 1235 1240 1245 Pro Glu Asp Asn Glu Gln Lys Gln Leu Phe Val Glu Gln His Lys 1250 1255 1260 His Tyr Leu Asp Glu Ile Ile Glu Gln Ile Ser Glu Phe Ser Lys 1265 1270 1275 Arg Val Ile Leu Ala Asp Ala Asn Leu Asp Lys Val Leu Ser Ala 1280 1285 1290 Tyr Asn Lys His Arg Asp Lys Pro Ile Arg Glu Gln Ala Glu Asn 1295 1300 1305 Ile Ile His Leu Phe Thr Leu Thr Asn Leu Gly Ala Pro Ala Ala 1310 1315 1320 Phe Lys Tyr Phe Asp Thr Thr Ile Asp Arg Lys Arg Tyr Thr Ser 1325 1330 1335 Thr Lys Glu Val Leu Asp Ala Thr Leu Ile His Gln Ser Ile Thr 1340 1345 1350 Gly Leu Tyr Glu Thr Arg Ile Asp Leu Ser Gln Leu Gly Gly Asp 1355 1360 1365 Ser Arg Ala Asp 1370 <210> 7 <211> 83 <212> DNA <213> Streptococcus pyogenes <400> 7 gttttagagc tagaaatagc aagttaaaat aaggctagtc cgttatcaac ttgaaaaagt 60 ggcaccgagt cggtggtgct ttt 83 <210> 8 <211> 665 <212> DNA <213> Saccharomyces cerevisiae <400> 8 tttcaaaaat tcttactttt tttttggatg gacgcaaaga agtttaataa tcatattaca 60 tggcattacc accatataca tatccatata catatccata tctaatctta cttatatgtt 120 gtggaaatgt aaagagcccc attatcttag cctaaaaaaa ccttctcttt ggaactttca 180 gtaatacgct taactgctca ttgctatatt gaagtacgga ttagaagccg ccgagcgggt 240 gacagccctc cgaaggaaga ctctcctccg tgcgtcctcg tcttcaccgg tcgcgttcct 300 gaaacgcaga tgtgcctcgc gccgcactgc tccgaacaat aaagattcta caatactagc 360 ttttatggtt atgaagagga aaaattggca gtaacctggc cccacaaacc ttcaaatgaa 420 cgaatcaaat taacaaccat aggatgataa tgcgattagt tttttagcct tatttctggg 480 gtaattaatc agcgaagcga tgatttttga tctattaaca gatatataaa tgcaaaaact 540 gcataaccac tttaactaat actttcaaca ttttcggttt gtattacttc ttattcaaat 600 gtaataaaag tatcaacaaa aaattgttaa tatacctcta tactttaacg tcaaggagaa 660 aaaac 665 <210> 9 <211> twenty one <212> DNA <213> Artificial sequence <220> <223> Nuclear transition signal. <220> <221> CDS <222> (1)..(21) <400> 9 ccc aag aag aag agg aag gtg 21 Pro Lys Lys Lys Arg Lys Val 1 5 <210> 10 <211> 7 <212> PRT <213> Artificial sequence <220> <223> Synthetic constructs <400> 10 Pro Lys Lys Lys Arg Lys Val 1 5 <210> 11 <211> 15 <212> DNA <213> Artificial sequence <220> <223> GS linker <220> <221> CDS <222> (1)..(15) <400> 11 ggt gga gga ggt tct 15 Gly Gly Gly Gly Ser 1 5 <210> 12 <211> 5 <212> PRT <213> Artificial sequence <220> <223> Synthetic constructs <400> 12 Gly Gly Gly Gly Ser 1 5 <210> 13 <211> 66 <212> DNA <213> Artificial sequence <220> <223> Flag tag <220> <221> CDS <222> (1)..(66) <400> 13 gac tat aag gac cac gac gga gac tac aag gat cat gat att gat tac 48 Asp Tyr Lys Asp His Asp Gly Asp Tyr Lys Asp His Asp Ile Asp Tyr 1 5 10 15 aaa gac gat gac gat aag 66 Lys Asp Asp Asp Asp Lys 20 <210> 14 <211> twenty two <212> PRT <213> Artificial sequence <220> <223> Synthetic constructs <400> 14 Asp Tyr Lys Asp His Asp Gly Asp Tyr Lys Asp His Asp Ile Asp Tyr 1 5 10 15 Lys Asp Asp Asp Asp Lys 20 <210> 15 <211> twenty four <212> DNA <213> Artificial sequence <220> <223> Strep-tag <220> <221> CDS <222> (1)..(24) <400> 15 tgg agc cac ccg cag ttc gaa aaa 24 Trp Ser His Pro Gln Phe Glu Lys 1 5 <210> 16 <211> 8 <212> PRT <213> Artificial sequence <220> <223> Synthetic constructs <400> 16 Trp Ser His Pro Gln Phe Glu Lys 1 5 <210> 17 <211> 171 <212> DNA <213> Artificial sequence <220> <223> SH3 structural domain <220> <221> CDS <222> (1)..(171) <400> 17 gca gag tat gtg cgg gcc ctc ttt gac ttt aat ggg aat gat gaa gaa 48 Ala Glu Tyr Val Arg Ala Leu Phe Asp Phe Asn Gly Asn Asp Glu Glu 1 5 10 15 gat ctt ccc ttt aag aaa gga gac atc ctg aga atc cgg gat aag cct 96 Asp Leu Pro Phe Lys Lys Gly Asp Ile Leu Arg Ile Arg Asp Lys Pro 20 25 30 gaa gag cag tgg tgg aat gca gag gac agc gaa gga aag agg ggg atg 144 Glu Glu Gln Trp Trp Asn Ala Glu Asp Ser Glu Gly Lys Arg Gly Met 35 40 45 att cct gtc cct tac gtg gag aag tat 171 Ile Pro Val Pro Tyr Val Glu Lys Tyr 50 55 <210> 18 <211> 57 <212> PRT <213> artificial sequence <220> <223> synthetic construction <400> 18 Ala Glu Tyr Val Arg Ala Leu Phe Asp Phe Asn Gly Asn Asp Glu Glu 1 5 10 15 Asp Leu Pro Phe Lys Lys Gly Asp Ile Leu Arg Ile Arg Asp Lys Pro 20 25 30 Glu Glu Gln Trp Trp Asn Ala Glu Asp Ser Glu Gly Lys Arg Gly Met 35 40 45 Ile Pro Val Pro Tyr Val Glu Lys Tyr 50 55 <210> 19 <211> 188 <212> DNA <213> Saccharomyces cerevisiae <400> 19 gcgaatttct tatgatttat gatttttatt attaaataag ttataaaaaa aataagtgta 60 tacaaatttt aaagtgactc ttaggtttta aaacgaaaat tcttattctt gagtaactct 120 ttcctgtagg tcaggttgct ttctcaggta tagcatgagg tcgctcttat tgaccacacc 180 tctaccgg 188 <210> 20 <211> 417 <212> DNA <213> Saccharomyces cerevisiae <400> 20 ataccaggca tggagcttat ctggtccgtt cgagttttcg acgagtttgg agacattctt 60 tatagatgtc ctttttttt aatgatattc gttaaagaac aaaaagtcaa agcagtttaa 120 cctaacacct gttgttgatg ctacttgaaa caaggcttct aggcgaatac ttaaaaaggt 180 aatttcaata gcggtttata tatctgtttg cttttcaaga tattatgtaa acgcacgatg 240 ttttcgccc aggctttatt ttttttgttg ttgttgtctt ctcgaagaat tttctcgggc 300 agatctttgt cggaatgtaa aaaagcgcgt aattaaactt tctattatgc tgactaaaat 360 ggaagtgatc accaaaggct atttctgatt atataatcta gtcattactc gctcgag 417 <210> twenty one <211> 33 <212> DNA <213> Artificial sequence <220> <223> SH3-binding ligands <220> <221> CDS <222> (1)..(33) <400> twenty one cct cca cct gct ctg cca cct aag aga agg aga 33 Pro Pro Pro Ala Leu Pro Pro Lys Arg Arg Arg 1 5 10 <210> twenty two <211> 11 <212> PRT <213> Artificial sequence <220> <223> Synthetic constructs <400> twenty two Pro Pro Pro Ala Leu Pro Pro Lys Arg Arg Arg 1 5 10 <210> twenty three <211> 269 <212> DNA <213> brewing yeast <400> twenty three tctttgaaaa gataatgtat gattatgctt tcactcatat ttatacagaa acttgatgtt 60 ttctttcgag tatatacaag gtgattacat gtacgtttga agtacaactc tagattttgt 120 agtgccctct tgggctagcg gtaaaggtgc gcattttttc acaccctaca atgttctgtt 180 caaaagattt tggtcaaacg ctgtagaagt gaaagttggt gcgcatgttt cggcgttcga 240 aacttctccg cagtgaaaga taaatgatc 269 <210> twenty four <211> 14 <212> DNA <213> brewing yeast <400> twenty four tgttttttat gtct 14 <210> 25 <211> 20 <212> DNA <213> brewing yeast <400> 25 gatacgttct ctatggagga 20 <210> 26 <211> 20 <212> DNA <213> brewing yeast <400> 26 ttggagaaac ccaggtgcct 20 <210> 27 <211> 20 <212> DNA <213> brewing yeast <400> 27 aacccaggtg cctggggtcc 20 <210> 28 <211> 20 <212> DNA <213> brewing yeast <400> 28 ttggccaagt cattcaattt 20 <210> 29 <211> 20 <212> DNA <213> brewing yeast <400> 29 ataacggaat ccaactgggc 20 <210> 30 <211> 20 <212> DNA <213> brewing yeast <400> 30 gtcaattacg aagactgaac 20 <210> 31 <211> 19 <212> DNA <213> brewing yeast <400> 31 gtcaatagga tcccctttt 19 <210> 32 <211> 10126 <212> DNA <213> Artificial sequence <220> <223> The plasmid carries the dCas9-PmCDA1 fusion protein and chimeric RNA. Targeting the galK gene in E. coli <220> <221> misc_feature <222> (5561)...(5580) <223> n is a, c, g, or t <400> 32 atcgccattc gccattcagg ctgcgcaact gttgggaagg gcgatcggtg cgggcctctt 60 cgctattacg ccagctggcg aaagggggat gtgctgcaag gcgattaagt tgggtaacgc 120 cagggttttc ccagtcacga cgttgtaaaa cgacggccag tgaattcgag ctcggtaccc 180 ggccgcaaac aacagataaa acgaaaggcc cagtctttcg actgagcctt tcgtttttt 240 300 aatggacaac tcgctccgtc gtttttcagc tcgcttcaaa gtttctcaa gccatctatt 360 ctcattcaat tgattgtgcg acgattggat gaatattttc ctgcaacatt ggtagtgttc 420 acttaccatt acattcaacc caaccccgtt atctctgagg ttccacagcc caatttgatt 480 cctcgcattt ttctcgtaat agagtttgca agcccagatt ttcaaagtgt ggccgttccc 540 ccgcagctcc tggttatacc attctaagat cttttcagcg caatctgcac areactcca 600 ggatgagtac caatttatcg tgaattgtcc ggggttgtcg cgcaggtatt cttcgacttt 660 tctaatgcta aagatttcgg cgtgaatgcc acgttctgtc ccgctctgtg gtttattcac 720 agcatagccc caaaaacacg ctctacgttc accccgtcgt tttaattcaa agagaacgta 780 gcatctatgc gacacggatt ttttgttgtt gaaaaactgt ttcttaaacg tgtagatgtc 840 caacttctca tggattctca cgtactcagc gtcggtcatc ctagacttat cgtcatcgtc 900 tttgtaatca atatcatgat ccttgtagtc tccgtcgtgg tccttatagt ctccggactc 960 gagcctagac ttatcgtcat cgtctttgta atcaatatca tgatccttgt agtctccgtc 1020 gtggtcctta tagtctccgg aatacttctc cacgtaaggg acaggaatca tccccctct 1080 tccttcgctg tcctctgcat tccaccactg ctcctcaggc ttatcccgga ttctcaggat 1140 gtctccttc ttaaagggaa gatcctcttc atcattccca ttaaagtcaa agagggctcg 1200 cacatactca gcagaacctc cacctccaga acctcctcca ccgtcacctc ctagctgact 1260 caaatcaatg cgtgtttcat aaagaccagt gatggattga tggataagag tggcatctaa 1320 aacttctttt gtagacgtat atcgtttacg atcaattgtt gtatcaaaat atttaaaagc 1380 agcgggagct ccaagattcg tcaacgtaaa taaatgaata atattttctg cttgttcacg 1440 tattggttg tctctatgtt tgttatatgc actaagaact ttatctaaat tggcatctgc 1500 taaaataaca cgcttagaaa attcactgat ttgctcaata atctcatcta aataatgctt 1560 atgctgctcc acaaacaatt gttttgttc gttatctttct ggactaccct tcaacttttc 1620 ataatgacta gctaaatata aaaaattcac atatttgctt ggcagagcca gctcattcc 1680 ttttgtaat tctccggcac tagccagcat ccgtttacga ccgttttcta actcaaaag 1740 actatattta ggtagtttta tgattaagtc tttttact tccttatatc ctttagcttc 1800 taaaaagtca atcggatttt ttcaagga acttctttcc atattgtga tcctagtaa 1860 ctctttaacg gatttttact tcttcgattt ccctttttcc accttagca ccactaggac 1920 tgaataagct accgttggac tatcaaacc accatatttt ttggatccc agtctttttt 1980 acgagcaata agcttgtccg aatttcttt tgtaaaatt gactccttgg agaatccgcc 2040 tgtctgtact tctgttttct tgacaatatt gacttggggc atggacaata ctttgcgcac 2100 tgtggcaaaa tctcgccctt tatcccagac aatttctcca gtttcccat tagtttcgat 2160 taggaggcgt ttgcgaatct ctccatttgc aagtgtaatt tctgttttga agaagttcat 2220 gatattagag taaaagaat atttgcggt tgctttgcct atttcttgct cagacttagc 2280 aatcatttta cgaacatcat aaactttata atcaccatag acaactccg attcaagtttt 2340 tggatatttc ttaatcaaag cagttccaac gacggcattt agatacgcat catgggcatg 2400 atggtaattg ttaatctcac gtactttata gaattggaaa tctttcgga agtcagaaac 2460 taatttagat tttaaggtaa tcactttaac ctctcgaata agtttatcat tttcatcgta 2520 tttagtattc atgcgactat ccaaaatttg tgccacatgc ttagtgattt ggcgagtttc 2580 aaccaattgg cgtttgataa aaccagcttt atcaagttca ctcaaacctc cacgttcagc 2640 tttcgttaaa tttcaaact tacgttgagt gattaacttg gcgtttagaa gttgtctcca 2700 atagtttc atctttttga ctacttcttc acttggaacg ttatccgatt taccacgatt 2760 tttatcagaa cgcgttaaga ccttattgtc tattgaatcg tctttaagga aactttgtgg 2820 aacaatggca tcgacatcat aatcacttaa acgattaata tctaattctt ggtccacata 2880 catgtctctt ccattttgga gataatagag atagagcttt tcattttgca attgagtatt 2940 ttcaacagga tgctctttaa gaatctgact tcctaattct ttgatacctt cttcgattcg 3000 tttcatacgc tctcgcgaat ttttctggcc cttttgagtt gtctgatttt cacgtgccat 3060 ttcaataacg atattttctg gcttatgccg ccccattact ttgaccaatt catcaacaac 3120 ttttacagtc tgtaaaatac cttttttaat agcagggcta ccagctaaat ttgcaatatg 3180 ttcatgtaaa ctatcgcctt gtccagacac ttgtgctttt tgaatgtctt cttaaatgt 3240 caaactatca tcatggatca gctgcataaa attgcgattg gcaaaaccat ctgatttcaa 3300 aaaatctaat attgttttgc cagattgctt atccctaata ccattaatca atttcgaga 3360 caaacgtccc caaccagtat aacggcgacg tttaagctgt ttcatcacct tatcatcaaa 3420 gaggtgagca tatgttttaa gtctttcctc aatcatctcc ctatcttcaa ataaggtcaa 3480 tgttaaaaca atacctcta agatatcttc attttcttca ttatccaaaa aatctttatc 3540 tttaataatt tttagcaaat catggtaggt acctaatgaa gcattaaatc tatcttcaac 3600 tcctgaaatt tcaactatt caaaacattc tatttttttg aaataatctt cttttaattg 3660 cttaacggtt acttttcgat ttgttttgaa gagtaaatca acaatggctt tcttctgttc 3720 acctgaaaga aatgctggtt ttcgcattcc ttcagtaaca tatttgacct ttgtcaattc 3780 gttataaacc gtaaaatact cataaagcaa actatgtttt ggtagtactt tttcatttgg 3840 aagatttta tcaaagtttg tcatgcgttc aataaatgat tgagctgaag cacctttatc 3900 gacaacttct tcaaaattcc atggggtaat tgtttcttca gacttccgag tcatccatgc 3960 aaaacgacta ttgccacgcg ccaatggacc aacataataa ggaattcgaa aagtcaagat 4020 tttttcaatc ttctcacgat tgtcttttaa aaatggataa aagtcttctt gtcttctcaa 4080 aatagcatgc agctcaccca agtgaatttg atggggaata gagccgttgt caaaggtccg 4140 ttgcttgcgc agcaaatctt cacgatttag tttcaccaat aattcctcag taccatccat 4200 tttttctaaa attggtttga taaatttata aaattcttct tggctagctc ccccatcaat 4260 ataacctgca tatccgtttt ttgattgatc aaaaaagatt tctttatact tttctggaag 4320 ttgttgtcga actaaagctt ttaaaagagt caagtcttga tgatgttcat cgtagcgttt 4380 aatcattgaa gctgataggg gagccttagt tatttcagta tttactctta ggatatctga 4440 aagtaaaaata gcatctgata aattcttagc tgccaaaaaac aaatcagcat attgatctcc 4500 aatttgcgcc aataaattat ctaaatcatc atcgtaagta tctttttgaaa gctgtaattt 4560 agcatcttct gccaaatcaa aatttgattt aaaattaggg gtcaaaccca atgacaaagc 4620 aatgagattc ccaaataagc cattttctt ctcaccgggg agctgagcaa tgagatttc 4680 taatcgtctt gatttactca atcgtgcaga aagaatcgct ttagcatcta ctccacttgc 4740 gttaataggg ttttcttcaa ataattgatt gtaggtttgt accaactgga taaatagttt 4800 gtccacatca ctattatcag gatttaaatc tccctcaatc aaaaaatgac cacgaaactt 4860 aatcatatgc gctaaggcca aatagattaa gcgcaaatcc gctttatcag tagaatctac 4920 caattttttt cgcagatgat agatagttgg atattctca tgataagcaa cttcatctac 4980 tatatttcca aaaataggat gacgttcatg cttcttgtct tcttccacca aaaaagactc 5040 ttcaagtcga tgaaagaaac tatcatctac tttcgccatc tcatttgaaa aaatctcctg 5100 tagataacaa atacgattct tccgacgtgt ataccttcta cgagctgtcc gtttgagacg 5160 agtcgcttc gctgtctctc cactgtcaaa taaaagagcc cctataagat ttttttgat 5220 actgtggcgg tctgtatttc ccagaacctt gaactttta gacgaacct tatattcatc 5280 agtgatcacc gcccatccga cgctatttgt gccgatagct aagcctattg agtatttctt 5340 atccatttt gcctcctaaa atgggccctt taaatttaat ccataatgag tttgatgatt 5400 tcaatatag ttttaatgac ctccgaaatt agtttatat gctttaattt ttctttttca 5460 aaatatctct tcaaaaata tcaccaata cttaata atagatt aacacaaat 5520 tcttttgaca agtagtttat tttgttataa ttctatagta nnnnnnnn nnnnnnnnn 5580 gttttagagc tagaatagc aagttaaaat aaggctagtc cgttatcaac ttgaaaagt 5640 ggcaccgagt cggtgcttttt ttgatactt ctattctact ctgactgcaa accaaaaaa 5700 caagcgcttt caaacgctt gttttatcat ttttaggaa attaatctct taatcctttt 5760 atcattctac atttaggcgc tgccatcttg ctaaacctac taagctccac aggatgatt 5820 cgtaatcccg caaggcccc ggcagtaccg gcataccaa gcctatgcct acagcatcca 5880 gggtgacggt gccgaggatg acgatgagcg cattgttaga ttcatacac ggtgcctgac 5940 tgcgttagca atttaactgt gataaactac cgcattaaag cttatcgatg ataagctgtc 6000 aaacatgaga attacaactt atatcgtatg gggctgactt caggtgctac atttgaagag 6060 ataaattgca ctgaaatcta gtcggatcct cgctcactga ctcgctgcgc tcggtcgttc 6120 ggctgcggcg agcggtatca gctcactcaa aggcggtaat acggttatcc acagaatcag 6180 gggataacgc aggaaagaac atgtgagcaa aaggccagca aaaggccagg aaccgtaaaa 6240 aggccgcgtt gctggcgttt ttccataggc tccgcccccc tgacgagcat cacaaaaaatc 6300 gacgctcaag tcagaggtgg cgaaacccga caggactata aagataccag gcgtttcccc 6360 ctggaagctc cctcgtgcgc tctcctgttc cgaccctgcc gcttaccgga tacctgtccg 6420 cctttctccc ttcgggaagc gtggcgcttt ctcatagctc acgctgtagg tatctcagtt 6480 cggtgtaggt cgttcgctcc aagctggggct gtgtgcacga accccccgtt cagcccgacc 6540 gctgcgcctt atccggtaac tatcgtcttg agtccaaccc ggtaagacac gacttatcgc 6600 cactggcagc agccactggt aacaggatta gcagagcgag gtatgtaggc ggtgctacag 6660 agttcttgaa gtggtggcct aactacggct acactagaag gacagtattt ggtatctgcg 6720 ctctgctgaa gccagttacc ttcggaaaaa gagttggtag ctcttgatcc ggcaaacaaa 6780 ccaccgctgg tagcggtggt ttttttgttt gcaagcagca gattacgcgc agaaaaaaag 6840 gatctcaaga agatccttg atcttttcta cggggtctga cgctcagtgg aacgaaaact 6900 cacgttaagg gattttggtc atgagattat caaaaaggat cttcacctag atccttttaa 6960 attaaaaatg aagttttaaa tcaatctaaa gtatatatga gtaaacttgg tctgacagtt 7020 accaatgctt aatcagtgag gcacctatct cagcgatctg tctatttcgt tcatccatag 7080 ttgcctgact ccccgtcgtg tagataacta cgatacggga gggcttacca tctggcccca 7140 gtgctgcaat gataccgcga gaaccacgct caccggctcc agatttatca gcaataaacc 7200 agccagccgg aagggccgag cgcagaagtg gtcctgcaac tttatccgcc tccatccagt 7260 ctattaattg ttgccgggaa gctagagtaa gtagttcgcc agttaatagt ttgcgcaacg 7320 ttgttgccat tgctgcaggc atcgtggtgt cacgctcgtc gtttggtatg gcttcattca 7380 gctccggttc ccaacgatca aggcgagtta catgatcccc catgttgtgc aaaaaagcgg ttagctcctt cggtcctccg atcgttgtca gaagtaagtt ggccgcagtg ttatcactca tggttatggc agcactgcat aattctctta ctgtcatgcc atccgtaaga tgcttttctg 7620. tgactggtga gtactcaacc aagtcattct gagaatagtg tatgcggcga ccgagttgct cttgcccggc gtcaacacgg cgccacatag cagaacttta aaagtgctca tcattggaaa acgttcttcg gggcgaaaac tctcaaggat cttaccgctg ttgagatcca gttcgatgta acccactcgt gcacccaact gatcttcagc atcttttact ttcaccagcg 7800 tttctgggtg agcaaaaca ggaaggcaaa atgccgcaaa aaagggaata agggcgacac ggaaatgttg aatactcata ctcttccttt ttcaatatta ttgaagcatt tatcagggtt attgtctcat gagcggatac attttgaat gtatttagaa aaataaacaa attaggggttc cgcgcacatt tccccgaaaa gtgccacctg acgtcaatgc cgagcgaaag cgagccgaag 8040 ggtagcattt acgttagata accccctgat atgctccgac gctttatata gaaaagaga ttcaactagg taaaatctta ataggttg agatgataag gtttataagg aatttgttg 8160 8220 acagttatga tagttaga atagttttaaa attaggtg agaaaaagat gaaagaaga 8280 tatggacag tcttaaagg ctctcagagg ctcatagacg aagaagtgg agagtcata 8340 gaggtagaca agttataccg taacaaacg tctggtaact tcgtaaaggc atatatagtg 8400 cattaata gtatgttaga tatgattggc ggaaaaaaac ttaaaatcgt taactatatc 8460 ctagataatg tccacttaag window atgatagcta storm atagcaaaa 8520 gctacaggaa caagtctaca aacagtaata aacactta aaatcttaga agaaggaat 8580 atataaaaa gaaaaactgg agtattaatg ttaaaccctg acactaat gagaggcgac 8640 gaccaaaaac aaaaataccct cttactcgaa ttgggaact ttgagcaga ggcaatgaa 8700 atagattgac ctcccaataa caccacgtag ttattgggag gtcaatctat gaatgcgat 8760 taagcttttt ctaattcaca taagcgtgca ggtttaaagt actaaaaa tataatgaaa 8820 aaaagcatca ttatactaac gttataccaa cattatactc tcattatact aattgcttat 8880 tccaatttcc tattggttgg aaccaacagg cgttagtgtg tgttgagtt ggtactttca 8940 tgggattaat cccatgaaac cccaccaa ctcgccaaag ctttggctaa cacacacgcc 9000 attccaacca atagttttct cggcataaag ccatgctctg acgcttaaat gcactaatgc 9060 cttaaaaaaa cattaaagtc taacacacta gacttattta cttcgtaatt aagtcgttaa 9120 accgtgtgct ctacgaccaa aagtataaaa cctttaagaa ctttctttt tcttgtaaaa 9180 aaagaaacta gataaatctc tcatatcttt tattcataa tcgcatcaga ttgcagtata 9240 aatttaacga tcactcatca tgttcatatt tatcagagct cgtgctataa ttatactaat 9300 tttataagga gaaaaaata aagaggggtta tgaaaataaaacacagtc 9360 aaaactttat tactcaaaa cattaatag aaaataat ghaaaata agattaatg 9420 aacatgataa tatctttgaa atcggctcag gaaaagggca ttttacctt gattagtac 9480 agaggtgtaa tttcgtaact gccattgaaa tagaccataa attatgcaa actacagaaa 9540 aaacttgt tgatcacgat aatttccaag ttttaacaa ggatatattg cagtttaaat 9600 ttcctaaaaa ccaatcctat aaaatattg gtaatatacc ttataacata agtacggata 9660 taatacgcaa aattgtttttt gatagtatag ctgatgagat ttatttaatc gtggaatacg 9720 ggtttgctaa agattatta atacaaac gctcattggc atttatttta atggcagaag 9780 ttgatatttc tatattagt atggttccaa gagaatattt tcatcctaaa cctaaagtga 9840 atagctcact tatcagatta atagaaaaa atcacact atcacacaa gataacaga 9900 agtataatta ttcgttatg aaatgggtta aaagaata caaaaaaa 9960 atcaatttaa cattcctta aaacatgcag gattgacga ttaacaat attagctttg 10020 aacattctt atctcttttc atagctata aattatta taagtagtt aagggatgca 10080 taaactgcat cccttaactt gttttcgtg tacctatttt tgtga 10126 <210> 33 <211> 20 <212> DNA <213> Escherichia coli <400> 33 tcaatgggct aactacgttc 20 <210> 34 <211> 20 <212> DNA <213> Escherichia coli <400> 34 ggtccataaa ctgagacagc 20 <210> 35 <211> 223 <212> DNA <213> Escherichia coli <400> 35 gagttataca cagggctggg atctattctt tttatctttt tttattcttt ctttattcta 60 taaattataa ccacttgaat ataaacaaaa aaaacacaca aaggtctagc ggaatttaca 120 gagggtctag cagaatttac aagttttcca gcaaaggtct agcagaattt acagataccc 180 acaactcaaa ggaaaaggac tagtaattat cattgactag ccc 223 <210> 36 <211> 951 <212> DNA <213> Escherichia coli <400> 36 atgtctgaat tagttgtttt caaagcaaat gaactagcga ttagtcgcta tgacttaacg 60 gagcatgaaa ccaagctaat tttatgctgt gtggcactac tcaaccccac gattgaaaac 120 cctacaatga aagaacggac ggtatcgttc acttataacc aatacgttca gatgatgaac 180 atcagtaggg aaaatgctta tggtgtatta gctaaagcaa ccagagagct gatgacgaga 240 actgtggaaa tcaggaatcc ttttgttaaa ggcttgaga tttccagtg ugaraactat 300 gccaagttct caagcgaaaa attagaatta gttttagtg aagagatatt gccttactt 360 ttccagttaa aaaaattcat aaaatataat ctggaacatg ttaagtcttt tgaaaacaaa 420 tactctatga ggatttag gtggttatta aagaactaa aactcacaag 480 gcaaatag agattagcct tgatgaattt aagttcatgt tatgcttga aaataactac 540 catgagttta aaggcttaa ccaatgggtt ttgaaaccaa taagtaaga tttaacact 600 tacagcaata tgaaattggt ggttgataag cgaggccgcc cgactgatac gttgattttc 660 caagttgaac tagatagaca aatggatctc gtaccgaac ttgagacaa ccagataaaa 720 atgaatggtg acaaaatacc aaaaccatt acatcagatt cctacctaca taacggacta 780 agaaaaacac tacacgatgc tttactgca aaaattcagc tcaccagtttt tgaggcaaa 840 ttttgagtg acatgcaag taagcatgat ctcaatggtt cgttctcatg gctcacgcaa 900 aaaacaacgaa ccacactaga gaacatactg gctaaatacg gaaggatctg a 951 <210> 37 <211> 714 <212> DNA <213> Bacteriophage lambda <400> 37 tcagccaaac gtctcttcag gccactgact agcgataact ttccccacaa cggaacaact 60 ctcattgcat gggatcattg ggtactgtgg gtttagtggt tgtaaaaaca cctgaccgct 120 atccctgatc agtttcttga aggtaaactc atcaccccca agtctggcta tgcagaaatc 180 acctggctca acagcctgct cagggtcaac gagaattaac attccgtcag gaaagcttgg 240 cttggagcct gttggtgcgg tcatggaatt accttcaacc tcaagccaga atgcagaatc 300 actggctttt ttggttgtgc ttacccatct ctccgcatca cctttggtaa aggttctaag 360 cttaggtgag aacatccctg cctgaacatg agaaaaaaca gggtactcat actcacttct 420 aagtgacggc tgcatactaa ccgcttcata catctcgtag atttctctgg cgattgaagg 480 gctaaattct tcaacgctaa ctttgagaat ttttgtaagc aatgcggcgt tataagcatt 540 taatgcattg atgccattaa ataaagcacc aacgcctgac tgccccatcc ccatcttgtc 600 tgcgacagat tcctgggata agccaagttc atttttcttt ttttcataaa ttgctttaag 660 gcgacgtgcg tcctcaagct gctcttgtgt taatggtttc ttttttgtgc tcat 714 <210> 38 <211> 5097 <212> DNA <213> Artificial Sequence <220> <223> dCas9-PmCDA1 <400> 38 atggataaga aatactcaat aggcttagct atcggcacaa atagcgtcgg atgggcggtg 60 atcactgatg aatataaggt tccgtctaaa aagttcaagg ttctgggaaa tacagaccgc 120 cacagtatca aaaaaaatct tataggggct cttttatttg acagtggaga gacagcggaa 180 gcgactcgtc tcaaacggac agctcgtaga aggtatacac gtcggaagaa tcgtatttgt 240 tatctacagg agattttttc aaatgagatg gcgaaagtag atgatagttt ctttcatcga 300 cttgaagagt cttttttggt ggaagaagac aagaagcatg aacgtcatcc tatttttgga 360 aatatagtag atgaagttgc ttatcatgag aaatatccaa ctatctatca tctgcgaaaa 420 aaattggtag attctactga taaagcggat ttgcgcttaa tctatttggc cttagcgcat 480 atgattaagt ttcgtggtca ttttttgatt gagggagatt taaatcctga taatagtgat 540 gtggacaaac tatttatcca gttggtacaa acctacaatc aattatttga agaaaaccct 600 attaacgcaa gtggagtaga tgctaaagcg attctttctg cacgattgag taaatcaaga 660 cgattagaaa atctcattgc tcagctcccc ggtgagaaga aaaatggctt atttgggaat 720 ctcattgctt tgtcattggg tttgacccct aattttaaat caaattttga tttggcagaa 780 gatgctaaat tacagctttc aaaagatact tacgatgatg atttagataa tttattggcg 840 caaattggag atcaatatgc tgatttgttt ttggcagcta agaatttatc agatgctatt 900 ttactttcag atatcctaag agtaaatact gaaataacta aggctcccct atcagcttca 960 atgattaaac gctacgatga acatcatcaa gacttgactc ttttaaaagc tttagttcga 1020 caacaacttc cagaaaagta taaagaaatc ttttttgatc aatcaaaaaa cggatatgca 1080 ggttatattg atgggggagc tagccaagaa gaattttata aatttatcaa accaatttta 1140 gaaaaaatgg atggtactga ggaattattg gtgaaactaa atcgtgaaga tttgctgcgc 1200 aagcaacgga cctttgacaa cggctctatt ccccatcaaa ttcacttggg tgagctgcat 1260 gctatttga gagacaaga agacttttat ccatttta aagacaatcg tgagagatt 1320 gaaaaaatct tgacttcg aattccttat tatgttgtc cattggcgcg tggcaatgt 1380 cgttttgcat ggatgactcg gaagtctgaa gaacaatta cccatggaa ttttgagaa 1440 gttgtcgata aaggtgcttc agctcaatca tttattgac gcatgacaaa ctttgataaa 1500 aatcttccaa atgaaaaagt actaccaaaa catagtttgc tttatgagta ttttacggtt 1560 tataacgaat tgacaaggt caataatgtt actgaaggaa tgcgaaacc agcattctt 1620 tcaggtgaac agaagaagc cattgttgat ttactctca aaaaaatcg aaagtaacc 1680 gttaagcaat t'aaagaga ttttcaa aaatagaat gttttgatag tgttgaatt 1740 tcaggagttg agatagatt taatgcttca ttaggtacct accatgatt gctaaaaatt 1800 Attaagatta Aagatttttt Ggaataatgaa Gaaatgaag Atacttaga Ggatattgtt 1860 ttaacattga ccttatttga agatagggag atgattgagg aagactta aacatatgct 1920 cacctctttg atgataggt gatgaacag cttaaacgtc gccgttatac tggttgggga 1980 cgtttgtctc gaaaattgat taatggtatt agggataagc aatctggcaa aacaatatta 2040 gattttttga aatcagatgg ttttgccaat cgcaatttta tgcagctgat ccatgatgat 2100 agtttgacat ttaaagaaga cattcaaaaa gcacaagtgt ctggacaagg cgatagttta 2160 catgaacata ttgcaaattt agctggtagc cctgctatta aaaaaggtat tttacagact 2220 gtaaaagttg ttgatgaatt ggtcaaagta atggggcggc ataagccaga aaatatcgtt 2280 attgaaatgg cacgtgaaaa tcagacaact caaaagggcc agaaaaattc gcgagagcgt 2340 atgaaacgaa tcgaagaagg tatcaaagaa ttaggaagtc agattcttaa agagcatcct 2400 gttgaaaata ctcaattgca aaatgaaaag ctctatctct attatctcca aaatggaaga 2460 gacatgtatg tggaccaaga attagatatt aatcgtttaa gtgattatga tgtcgatgcc 2520 attgttccac aaagtttcct taaagacgat tcaatagaca ataaggtctt aacgcgttct 2580 gataaaaatc gtggtaaatc ggataacgtt ccaagtgaag aagtagtcaa aaagatgaaa 2640 aactattgga gacaacttct aaacgccaag ttaatcactc aacgtaagtt tgataattta 2700 acgaaagctg aacgtggagg tttgagtgaa cttgataaag ctggttttat caaacgccaa 2760 ttggttgaaa ctcgccaaat cactaagcat gtggcacaaa ttttggatag tcgcatgaat 2820 actaaatacg atgaaaatga taaacttatt cgagaggtta aagtgattac cttaaaatct 2880 aaattagttt ctgacttccg aaaagatttc caattctata aagtacgtga gattaacaat 2940 taccatcatg cccatgatgc gtatctaaat gccgtcgttg gaactgcttt gattaagaaa 3000 tatccaaaac ttgaatcgga gtttgtctat ggtgattata aagtttatga tgttcgtaaa 3060 atgattgcta agtctgagca agaaataggc aaagcaaccg caaaatattt cttttactct 3120 aatatcatga acttcttcaa aacagaaatt acacttgcaa atggagagat tcgcaaacgc 3180 cctctaatcg aaactaatgg ggaaactgga gaaattgtct gggataaagg gcgagatttt 3240 gccacagtgc gcaaagtatt gtccatgccc caagtcaata ttgtcaagaa aacagaagta 3300 cagacaggcg gattctccaa ggagtcaatt ttaccaaaaa gaaattcgga caagcttatt 3360 gctcgtaaaa aagactggga tccaaaaaaa tatggtggtt ttgatagtcc aacggtagct 3420 tattcagtcc tagtggttgc taggtgga aaagggaat cgagaagtt aaatccgtt 3480 aaagagttac tagggatcac aattatggaa agaagttcct ttgaaaaaaa tccgattgac 3540 tttttagag ctaaggata taaggatt aaaaagact taatcattaa actacctaaa 3600 tatagtcttt ttgagttaga aaacggtcgt aaacggatgc tggctagtgc cggagaatta 3660 caaaaaggaa atgagctggc tctgccaagc aaattgtga attttta tttagctagt 3720 cattatgaaa agttgaaggg tagtccagaa gatacgaac aaaaacatt gtttgtggag 3780 cagcataagc attatttaga tgagattatt gagcaatca gtgaattttc taagcgtgtt 3840 attttagcag atgccaattt agataagtt cttagtgcat atacaaca tagagacaa 3900 ccaatacgtg aacaagcaga aaatattatt catttattta cgttgacgaa tcttggagct 3960 cccgctgctt ttaatattt tgatacaaca attgatcgta aacgatatac gtctacaaaa 4020 gaagttttag atgccactct tatccatca tccatcactg gtctttag aacacgcatt 4080 gatttgagtc agctaggagg tgacggtga gaggttctg gaggtggagg tctgagt 4140 tatgtgcgag ccctttga ctttaatggg aatgatgaag aggatctcc ctttaagaaa 4200 ggagacatcc tgagaatccg ggataagcct gaggagcagt ggtggaatgc agaggacagc 4260 gaaggaaga gggggat tcctgtccct tacgtggaga agtattccgg agactataag 4320 gaccacgacg gagactaca ggatcatgat attgattaca aagacgatga cgataagtct 4380 aggctcgagt ccgagacta taggaccac gacgagact ahagacta tgatattgat 4440 tacaagacg atgacgataa gtctaggatg accgacgctg agtacgtgag aatccatgag 4500 aagttggaca tctacacgtt taagaacag ttttcaca acaaaaatc cgtgtcgcat 4560 agatgctacg ttctctttga attaaacga cggggtgaac gtagagcgtg ttttggggc 4620 tatgctgtga ataaccaca gagcgggaca gaacgtggca ttcacgccga aatctttagc 4680 attagaaag tcgaagaata cctgcgcgac aaccccggac attcacgat aaattgtac 4740 tcatcctgga gtccttgtgc agatgcgct gaaagatct tegatgta taaccaggag 4800 ctgcggggga acggccacac tttgaaaatc tgggcttgca aactctatta cgagaaaaat 4860 gcgaggaatc aaattgggct gtggaacctc agagataacg gggttgggtt gaatgtaatg 4920 gtaagtgaac actaccaatg ttgcaggaaa atattcatcc aatcgtcgca caatcaattg 4980 aatgagaata gatggcttga gaagactttg aagcgagctg aaaaacgacg gagcgagttg 5040 tccattatga ttcaggtaaa aatactccac accactaaga gtcctgctgt tacttga 5097 <210> 39 <211> 105 <212> DNA <213> Escherichia coli <400> 39 acgttaaatc tatcaccgca agggataaat atctaacacc gtgcgtgttg actattttac 60 ctctggcggt gataatggtt gcagggccca ttttaggagg caaaa 105 <210> 40 <211> 247 <212> DNA <213> Artificial Sequence <220> <223> gRNA <400> 40 ggtttagcaa gatggcagcg cctaaatgta gaatgataaa aggattaaga gattaatttc 60 cctaaaaatg ataaaacaag cgttttgaaa gcgcttgttt ttttggtttg cagtcagagt 120 agaatagaag tatcaaaaaa agcaccgact cggtgccact ttttcaagtt gataacggac 180 tagccttatt ttaacttgct atttctagct ctaaaactga gaccatcccg ggtctctact 240 gcagaat 247 <210> 41 <211> 64 <212> DNA <213> Escherichia coli <400> 41 tatcaccgcc agtggtattt atgtcaacac cgccagagat aatttatcac cgcagatggt 60 tatc 64 <210> 42 <211> 10867 <212> DNA <213> Artificial Sequence <220> <223> Plasmid <400> 42 gtcggaactg actaaagtag tgagttatac acagggctgg gatctattct ttttatcttt 60 ttttattctt tctttattct ataaattata accacttgaa tataaacaaa aaaaacacac 120 aaaggtctag cggaatttac agagggtcta gcagaattta caagttttcc agcaaaggtc 180 tagcagaatt tacagatacc cacaactcaa aggaaaagga ctagtaatta tcattgacta 240 gcccatctca attggtatag tgattaaaat cacctagacc aattgagatg tatgtctgaa 300 ttagttgttt tcaaagcaaa tgaactagcg attagtcgct atgacttaac ggagcatgaa 360 accaagctaa ttttatgctg tgtggcacta ctcaacccca cgattgaaaa ccctacaatg 420 aaagaacgga cggtatcgtt cacttataac caatacgttc agatgatgaa catcagtagg 480 gaaaatgctt atggtgtatt agctaaagca accagagagc tgatgacgag aactgtggaa 540 atcaggaatc ctttggttaa aggctttgag attttccagt ggacaaacta tgccaagttc 600 tcaagcgaaa aattagaatt agtttttagt gaagagatat tgccttatct tttccagtta 660 aaaaaattca taaaatataa tctggaacat gttaagtctt ttgaaaacaa atactctatg 720 aggatttatg agtggttatt aaaagaacta acacaaaaga aaactcacaa ggcaaatata 780 gagattagcc ttgatgaatt taagttcatg ttaatgcttg aaaataacta ccatgagttt 840 aaaaggctta accaatgggt tttgaaacca ataagtaaag atttaaacac ttacagcaat 900 atgaaattgg tggttgataa gcgaggccgc ccgactgata cgttgatttt ccaagttgaa 960 ctagatagac aaatggatct cgtaaccgaa cttgagaaca accagataaa aatgaatggt 1020 gacaaaatac caacaaccat tacatcagat tcctacctac ataacggact aagaaaaaca 1080 ctacacgatg ctttaactgc aaaaattcag ctcaccagtt ttgaggcaaa atttttgagt 1140 gacatgcaaa gtaagcatga tctcaatggt tcgttctcat ggctcacgca aaaacaacga 1200 accacactag agaacatact ggctaaatac ggaaggatct gaggttctta tggctcttgt 1260 atctatcagt gaagcatcaa gactaacaaa caaaagtaga acaactgttc accgttacat 1320 atcaaaggga aaactgtcca tatgcacaga gataatctca tgaccaaaac cggtagctag 1380 aggggccgca ttaggcaccc caggctttac actttatgct tccggctcgt ataatgtgtg 1440 gattttgagt taggatccgg cgagatttc aggagctaag gaagctaaaa tggagaaaaa 1500 aatcactgga tataccaccg ttgatatatc ccaatggcat cgtaaagaac atttgaggc 1560 atttcagtca gttgctcaat gtacctataa ccagaccgtt cagctggata ttacggcctt 1620 tttaaagacc gtaaagaaaa ataagcacaa gttttatccg gcctttattc acattcttgc 1680 ccgcctgatg aatgctcatc cggaattccg tatggcaatg aaagacggtg agctggtgat 1740 atgggatagt gttcaccctt gttacaccgt tttccatgag caaactgaaa cgttttcatc 1800 gctctggagt gaataccacg acgatttccg gcagtttcta cacatatatt cgcaagatgt 1860 ggcgtgttac ggtgaaaacc tggcctattt ccctaaaggg tttattgaga atatgttttt 1920 cgtctcagcc aatccctggg tgagtttcac cagttttgat ttaaacgtgg ccaatatgga 1980 caacttcttc gccccgttt tcaccatggg caaatattat acgcaaggcg acaaggtgct 2040 gatgccgctg gcgattcagg ttcatcatgc cgtctgtgat ggcttccatg tcggcagaat 2100 gcttaatgaa ttacaacagt actgcgatga gtggcagggc ggggcgtaaa cgcgtggatc 2160 cggcttacta aaagccagat aacagtatgc gtatttgcgc gctgattttt gcggtctaga 2220 ggtttagcaa gatggcagcg cctaaatgta gaatgataaa aggattaaga gattaatttc 2280 cctaaaaatg ataaaacaag cgttttgaaa gcgcttgttt ttttggttg cagtcagagt 2340 agaatagaag tatcaaaaaa agcaccgact cggtgccact ttttcaagtt gataacggac 2400 tagccttatt ttaacttgct atttctagct ctaaaactga gaccatcccg ggtctctact 2460 gcagaattat caccgccagt ggtattatg tcaacaccgc cagagataat tttcaccgc 2520 agatggttat cgatgaagat tcttgctcaa ttgttatcag ctatgcgccg accagaacac 2580 cttgccgatc agccaaacgt ctcttcaggc cactgactag cgataacttt ccccacaacg 2640 gaacaactct cattgcatgg gatcattggg tactgtgggt ttagtggttg taaaaacacc 2700 tgaccgctat ccctgatcag tttcttgaag gtaaactcat cacccccaag tctggctatg 2760 cagaaatcac ctggctcaac agcctgctca gggtcaacga gaattaacat tccgtcagga 2820 aagcttggct tggagcctgt tggtgcggtc atggaattac cttcaacctc aagccagaat 2880 gcagaatcac tggctttttt ggttgtgctt acccatctct ccgcatcacc tttggtaaag 2940 gttctaagct taggtgagaa catccctgcc tgaacatgag aaaaaacagg gtactcatac 3000 tcacttctaa gtgacggctg catactaacc gcttcataca tctcgtagat ttctctggcg 3060 attgaagggc taaattcttc aacgctaact ttgagaattt ttgtaagcaa tgcggcgtta 3120 taagcattta atgcattgat gccattaaat aaagcaccaa cgcctgactg ccccatcccc 3180 atcttgtctg cgacagattc ctgggataag ccaagttcat ttttcttttt ttcataaatt 3240 gctttaaggc gacgtgcgtc ctcaagctgc tcttgtgtta atggttcttt ttttgctc 3300 atacgttaaa tctatcaccg caagggataa atatctaaca ccgtgcgtgt tgactatttt 3360 acctctggcg gtgataatgg ttgcagggcc cattttagga ggcaaaaatg gataagaaat 3420 actcaatagg cttagctatc ggcacaaata gcgtcggatg ggcggtgatc actgatgaat 3480 ataaggttcc gtctaaaaag ttcaaggttc tgggaaatac agaccgccac agtatcaaaa 3540 aaaatcttat aggggctctt ttatttgaca gtggagagac agcggaagcg actcgtctca 3600 aacggacagc tcgtagaagg tatacacgtc ggaagaatcg tatttgttat ctacaggaga 3660 ttttttcaaa tgagatggcg aaagtagatg atagtttctt tcatcgactt gaagagtctt 3720 ttttggtgga agaagacaag aagcatgaac gtcatcctat ttttggaaat atagtagatg 3780 aagttgctta tcatgagaaa tatccaacta tctatcatct gcgaaaaaaa ttggtagatt 3840 ctactgataa agcggatttg cgcttaatct atttggcctt agcgcatatg attaagtttc 3900 gtggtcattt tttgattgag ggagatttaa atcctgataa tagtgatgtg gacaaactat 3960 ttatccagtt ggtacaacc tacaatcaat tatttgaga aaaccctatt aacgcaagtg 4020 gagtagatgc taagcgatt ctttctgcac gattgagtaa atcaagacga ttagaaaatc 4080 tcattgctca gctccccggt gagagaaaa atggcttt tgggaatctc attgctttgt 4140 cattgggttt gacccctaat ttaatca atttgattt ggcagaagat gctaaattac 4200 agctttcaaa agatacttac gatgatgat agctattt attggcgcaa attggagatc 4260 atatgctga ttgttttg gcagctaaga atttatcaga tgctatttta ctttcagata 4320 tcctaagagt aaatactgaa ataactaagg ctcccctatc agctcaatg attaaacgct 4380 acgatgacaca tcatcagac ttgacttt taaaagcttt agttcgacaa caactccag 4440 aaagtataa agaatcttt ttgatcaat CAaaaacgg atatgcaggt tatattgatg 4500 ggggagctag ccaagaa ttttataat ttcaacc aattttagaa aaaatggatg 4560 gtactgagga attattgtg aaactaaatc gtgagattt gctgcgcaag caacggacct 4620 ttgacaacgg ctctattccc catcaattc acttgggtga gctgcatgct attttgagaa 4680 gacaagaaga cttttatcca ttttaaaag acaatcgtga gaagattgaa aaaatcttga 4740 cttttcgaat tccttattat gttggtccat tggcgcgtgg caatagtcgt tttgcatgga 4800 tgactcggaa gtctgaagaa acaattaccc catggaattt tgaagaagtt gtcgataaag 4860 gtgcttcagc tcaatcattt attgaacgca tgacaaactt tgataaaaat cttccaaatg 4920 aaaaagtact accaaaacat agtttgcttt atgagtattt tacggtttat aacgaattga 4980 caaaggtcaa atatgttact gaaggaatgc gaaaaccagc atttctttca ggtgaacaga 5040 agaaagccat tgttgattta ctcttcaaaa caaatcgaaa agtaaccgtt aagcaattaa 5100 aagaagatta tttcaaaaaa atagaatgtt ttgatagtgt tgaaatttca ggagttgaag 5160 atagatttaa tgcttcatta ggtacctacc atgatttgct aaaaattatt aaagataaag 5220 atttttgga taatgaagaa aatgaagata tcttagagga tattgtttta acattgacct 5280 tatttgaaga tagggagatg attgaggaaa gacttaaaac atatgctcac ctctttgatg 5340 ataaggtgat gaaacagctt aaacgtcgcc gttatactgg ttggggacgt ttgtctcgaa 5400 aattgattaa tggtattagg gataagcaat ctggcaaaac aatattagat tttttgaaat 5460 cagatggtt tgccaatcgc aattttatgc agctgatcca tgatgatagt ttgacattta 5520 aagaagacat tcaaaaagca caagtgtctg gacaaggcga tagtttacat gaacatattg 5580 caaatttagc tggtagccct gctattaaaa aaggtatttt acagactgta aaagttgttg 5640 atgaattggt caaagtaatg gggcggcata agccagaaaa tatcgttatt gaaatggcac 5700 gtgaaaatca gacaactcaa aagggccaga aaaattcgg agagcgtatg aaacgaatcg 5760 aagaaggtat caaagaatta ggaagtcaga ttcttaaaga gcatcctgtt gaaaatactc 5820 aattgcaaaa tgaaaagctc tatctctatt atctccaaaa tggaagac atgtatgtgg 5880 accaagaatt agatattaat cgtttaagtg attatgatgt cgatgccatt gttccacaaa 5940 gttccttaa agacgattca atagacaata aggtcttaac gcgttctgat aaaaatcgtg 6000 gtaaatcgga taacgttcca agtgaagaag tagtcaaaaa gatgaaaaac tattggagac 6060 aacttctaaa cgccaagtta atcactcaac gtaagtttga taatttaacg aaagctgaac 6120 gtggaggttt gagtgaactt gataaagctg gttttatcaa acgccaattg gttgaaactc 6180 6240 aaaatgataa acttattcga gaggttaaag tgattacctt aaaatctaaa ttagtttctg 6300 acttccgaaa agatttccaa ttctataaag tacgtgagat taacaattac catcatgccc 6360 atgatgcgta tctaaatgcc gtcgttggaa ctgctttgat taagaaatat ccaaaacttg 6420 aatcggagtt tgtctatggt gattaataaag tttatgatgt tcgtaaaatg attgctaagt 6480 ctgagcaaga aataggcaaa gcaaccgcaa aatatttctt ttactctaat atcatgaact 6540 tcttcaaac agaaattaca cttgcaaatg gagagattcg caaacgccct ctaatcgaaa 6600 ctaatggggga aactggagaa attgtctggg ataaagggcg agattttgcc acagtgcgca 6660 aagtattgtc catgccccaa gtcaatattg tcaagaaaac agaagtacag acaggcggat 6720 tctccaagga gtcaatttta ccaaaaagaa attcggacaa gcttattgct cgtaaaaaag 6780 actgggatcc aaaaaaaatat ggtggttttg atagtccaac ggtagcttat tcagtcctag 6840 tggttgctaa ggtggaaaaa gggaaatcga agagttaaa atccgttaaa gagttactag 6900 ggatcacaat tatggaaga agttccttg aaaaaaatcc gattgacttt ttagaagcta 6960 aaggatataa ggaagttaaaagacttaa tcattaact acctaata agtctttttg 7020 agttagaaaa cggtcgtaaa cggatgctgg ctagtgccgg agaattacaa aaggaaatg 7080 agctggctct gccaagcaaa tatgtgaattt tttatttt agctagtcat tatgaaagt 7140 tgaagggtag tccagagat aacaaaaaaatgtt tgtggagcag cataagcatt 7200 atttagatga gattattgag caatcagtg aattttctaa gcgtgttatt ttagcagatg 7260 ccaatttaga taaagttctt agtgcatata aacatag agacaacca atacgtgaac 7320 aagcagaaaa tattattcat tattacgt tgacgaatct tggagctccc gctgcttta 7380 atattttga tachaaaat gatcgtaaac gatatacgtc tachaaaga gttttagatg 7440 ccactcttat ccatcaatcc atcactggtc tttgaac acgcattgat ttgagtcagc 7500 taggaggtga cggtggagga gttctctggag gtggaggttc tgctgagtat gtgcgagccc 7560 tctttgactt taatgggaat gatgaagagg atcttccctt taagaaagga gacatcctga 7620 gaatccggga taagcctgag gagcagtggt ggaatgcaga ggacagcgaa ggaaagaggg 7680 ggatgattcc tgtcccttac gtggagaagt attccggaga ctataaggac cacgacggag 7740 actacaagga tcatgatat gattacaaag acgatgacga taagtctagg ctcgagtccg 7800 gagactataa ggaccacgac ggagactaca aggatcatga tattgattac aaagacgatg 7860 acgataagtc taggatgacc gacgctgagt acgtgagaat ccatgagaag ttggacatct 7920 acacgtttaa gaaacagttt ttcaacaaca aaaaatccgt gtcgcataga tgctacgttc 7980 tctttgaatt aaaacgacgg ggtgaacgta gagcgtgttt ttggggctat gctgtgaata 8040 aaccacagag cgggacagaa cgtggcattc acgccgaaat ctttagcatt agaaaagtcg 8100 aagaatacct gcgcgacaac cccggacaat tcacgataaa ttggtactca tcctggagtc 8160 cttgtgcaga ttgcgctgaa aagatcttag aatggtataa ccaggagctg cgggggaacg 8220 gccacacttt gaaaatctgg gcttgcaaac tctattacga gaaaaatgcg aggaatcaaa 8280 ttgggctgtg gaacctcaga gataacgggg ttgggttgaa tgtaatggta agtgaacact 8340 accaatgttg caggaaaata ttcatccaat cgtcgcacaa tcaattgaat gagaatagat 8400 ggcttgagaa gactttgaag cgagctgaaa aacgacggag cgagttgtcc attatgattc 8460 aggtaaaat actccacacc actaagagtc ctgctgttac ttgacaggca tcaaataaa 8520 cgaaaggctc agtcgaaaga ctgggccttt cgttttatct gttgtttgcg gccgggtacc 8580 gagctcgaat tcactggccg tcgttttaca acgtcgtgac tgggaaaacc ctggcgttac 8640 ccaacttaat cgccttgcag cacatcccccc tttcgccagc tggcgtaata gcgaagaggc 8700 ccgcaccgat cgcccttccc aacagttgcg cagcctgaat ggcgaatggc gattcacaaa 8760 aaataggtac acgaaaaaca agttaaggga tgcagtttat gcatccctta acttacttat 8820 taaataattt atagctattg aaaagagata agaattgttc aaagctaata ttgtttaaat 8880 cgtcaattcc tgcatgtttt aaagaattgt taaattgatt ttttgtaaat atttcttgt 8940 attctttgtt aacccatttc ataacgaaat aattatactt ctgtttatct ttgtgtgata 9000 ttcttgattt ttttctattt aatctgataa gtgagctatt cactttaggt ttaggatgaa 9060 aatattctct tggaaccata cttaatatag aaattcaac ttctgccatt aaaaataatg 9120 ccaatgagcg ttttgtattt aataatcttt tagcaaaccc gtattccacg attaaataaa 9180 tctcatcagc tatactatca aaaacaattt tgcgtattat atccgtactt atgttataag 9240 gtatattacc aaatatttta taggattggt ttttaggaaa tttaaactgc aatatatcct 9300 tgttaaaac ttggaaatta tcgtgatcaa caagtttatt ttctgtagtt ttgcataatt 9360 tatggtctat ttcaatggca gttacgaaat tacacctctg tactaattca agggtaaaat 9420 gccctttc tgagccgatt tcaaagatat tatcatgttc atttaatctt atatttgtca 9480 ttattttatc tatattatgt tttgaagtaa taaagttttg actgtgtttt atatttttct 9540 cgttcattat aaccctttt atttttcct ccttataaaa ttagtataat tatagcacga 9600 gctctgataa atatgaacat gatgagtgat cgttaaattt atactgcaat ctgatgcgat 9660 tattgaataa aagatatgag agatttatct agtttcttttt tttacaagaa aaaagaaagt 9720 tcttaaaggt tttatacttt tggtcgtaga gcacacggtt taacgactta attacgaagt 9780 aaataagtct agtgtgttag actttaatgt ttttttaagg cattagtgca tttaagcgtc 9840 agagcatggc tttatgccga gaaaactatt ggttggaatg gcgtgtgtgt tagccaaagc 9900 tttggcgagt tggttggggg tttcatggga ttaatcccat gaaagtacca actcaacaac 9960 acactaacgc ctgttggttc caaccaatag gaaattggaa taagcaatta gtataatgag 10020 agtataatgt tggtataacg ttagtataat gatgcttttt ttcattatat tttttatgta 10080 ctttaaacct gcacgcttat gtgaattaga aaaagcttaa tcgcatttca tagattgacc 10140 tcccaataac tacgtggtgt tattgggagg tcaatctatt tcatttgcct cttgctcaaa 10200 gttcccaaat tcgagtaaga ggtatttttg tttttggtcg tcgcctctca ttagtagttc 10260 agggtttaac attaatactc cagtttttct ttttataata tttccttctt ctaagatttt 10320 aagtgttgtt attactgttt gtagacttgt tcctgtagct tttgctattt ctcttgttgt 10380 agctatcatt gtattgttac ttaagtggac attatctagg atatagttaa cgattttaag 10440 tttttttccg ccaatcatat ctaacatact tattaattgc actatatatg cctttacgaa gttaccagac gtttgtttac ggtataactt gtctacctct atgacttctc cactttcttc gtctatgagc ctctgagagc ctttatagac tgttccatat ctttctttca tctttttctc 10620 actccttatt ttaactatt ctaactatat cataactgtt ctaaaaaaaa aagaacattt gttaaaaga attagaacaa aatgagtgaa aaattagaac aaacaaattc cttataaacc ttatcatctc aacctatatt aagattttc ctagttgaat cttcttttct fatheragcg tcggagcata tcagggggtt atctaacgta aatgctaccc ttcggctcgc tttcgctcgg 10860. cattgac 10867
Claims
1. A method of modifying a targeted site of double-stranded DNA for a non-disease treatment purpose, comprising: a process of contacting a complex formed by linking a nucleic acid base converting enzyme and a nucleic acid sequence recognition module that specifically binds to a target nucleotide sequence in a double-stranded DNA with the double-stranded DNA, and causing deletion of one or more nucleotides at the target site or conversion of the one or more nucleotides at the target site to other one or more nucleotides, or insertion of one or more nucleotides into the target site, wherein the nucleic acid sequence recognition module is a CRISPR-Cas system, and wherein the CRISPR-Cas system comprises Cas, wherein the method uses two or more nucleic acid sequence recognition modules that specifically bind to different target nucleotide sequences, respectively; wherein the nucleic acid sequence recognition module is a CRISPR-Cas system in which Cas has only one DNA cleavage ability inactivated, wherein the nucleic acid base converting enzyme is a deaminase; and wherein the Cas is Cas9.
2. The method of claim 1, wherein, The different target nucleotide sequences exist in different genes.
3. The method of claim 1, wherein, The deaminase is AID (AICDA).
4. The method of claim 1, wherein, The contacting of the double-stranded DNA with the complex is performed by introducing a nucleic acid encoding the complex into a cell having the double-stranded DNA.
5. The method of claim 4, wherein, The cell is a prokaryotic cell.
6. The method of claim 4, wherein, The cell is a eukaryotic cell.
7. The method of claim 4, wherein, The cell is a microbial cell.
8. The method of claim 4, wherein, The cell is a plant cell.
9. The method of claim 4, wherein, The cell is an insect cell.
10. The method of claim 4, wherein, The cell is an animal cell.
11. The method of claim 4, wherein, The cell is a vertebrate cell.
12. The method of claim 4, wherein, The cell is a mammalian cell.
13. The method of claim 6, wherein, The cell is a polyploid cell, and the modification is performed on a site within all targeted alleles on homologous chromosomes.
14. The method of claim 4, comprising: a process of introducing into the cell an expression vector containing a nucleic acid encoding the complex in a form that enables control of the expression period, and a step of inducing expression of the nucleic acid at a period in which it is desired to stabilize the modification of the target site of the double-stranded DNA.
15. The method of claim 14, wherein, The target nucleotide sequence in the double-stranded DNA exists in a gene essential for the cell.
16. A nucleic acid modifying enzyme complex, which is linked to a nucleic acid sequence recognition module that specifically binds to a target nucleotide sequence in a double-stranded DNA, and a nucleic acid base converting enzyme, which causes deletion of one or more nucleotides from the target site or conversion of the one or more nucleotides to other one or more nucleotides, or insertion of one or more nucleotides into the target site, wherein, The nucleic acid sequence recognition module is a CRISPR-Cas system, and wherein the CRISPR-Cas system comprises Cas, and wherein at least one DNA cleavage ability of Cas is inactivated, wherein the complex comprises two or more nucleic acid sequence recognition modules that specifically bind to different target nucleotide sequences, respectively; wherein the nucleic acid sequence recognition module is a CRISPR-Cas system in which Cas has only one DNA cleavage ability inactivated or in which Cas has two DNA cleavage abilities inactivated; wherein the nucleic acid base converting enzyme is a deaminase; and wherein the Cas is Cas9.
17. A nucleic acid encoding the nucleic acid modifying enzyme complex of claim 16.
18. The nucleic acid modifying enzyme complex of claim 16, wherein, The deaminase is AID (AICDA).
19. The nucleic acid modifying enzyme complex of claim 16, wherein, The nucleic acid sequence recognition module is a CRISPR-Cas system in which Cas has only one DNA cleavage ability inactivated.
20. The nucleic acid modifying enzyme complex of claim 16, wherein, The nucleic acid sequence recognition module is a CRISPR-Cas system in which Cas has two DNA cleavage abilities inactivated.
Citation Information
Patent Citations
JP0923000001S
JP1974068498A
Cultures with improved phage resistance
JP2010519929A
Method for modifying RNA-binding protein using PPR motif
JP2013128413A
DNA modification mediated by TAL effectors
JP2013513389A