A method for site-directed mutagenesis using a dCas9-p450 system
By using the dCas9-p450 system fusion protein and CYP3A4 enzyme, and utilizing aflatoxin B1 to induce DNA damage, the problem of introducing random mutations in specific genomic regions in existing technologies has been solved, achieving highly efficient gene editing results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUAZHONG AGRI UNIV
- Filing Date
- 2021-08-10
- Publication Date
- 2026-04-28
AI Technical Summary
Existing technologies struggle to efficiently introduce random mutations into specific genomic regions, particularly for aflatoxin B1-induced DNA damage, and there is a lack of effective gene editing tools.
Using the dCas9-p450 system, by fusing the dCas9 protein with the CYP3A4 enzyme, aflatoxin B1 induces the formation of AFB1-8,9-epoxides in the DNA at the target site, leading to DNA damage and introducing random mutations.
It enables the efficient introduction of random mutations in specific genomic regions, providing a flexible and powerful gene editing tool for gene function research, and is able to introduce mutations in specific nucleotide sequences at target sites.
Smart Images

Figure BDA0003204712330000051 
Figure BDA0003204712330000061 
Figure BDA0003204712330000062
Abstract
Description
Technical Field
[0001] This invention belongs to the field of animal gene editing engineering and relates to a method for site-directed mutagenesis using the dCas9-p450 system. Background Technology
[0002] Aflatoxins (AFTs) are a class of chemically similar compounds, primarily produced by Aspergillus flavus and Aspergillus parasiticus. Aflatoxin B1 (AFB1) is the most common type found in naturally contaminated foods. In fact, AFB1 itself does not possess toxic, carcinogenic, or mutagenic effects. Its potent carcinogenicity stems from cytochrome P450 oxidase (CYP), which converts AFB1 into the highly reactive and unstable AFB1-8,9-epoxide (AFBO). AFBO can covalently bind to various nucleophilic centers of cellular macromolecules, such as DNA, RNA, or proteins, forming complexes. This triggers DNA mismatch repair mechanisms, leading to a series of genetic mutations, including base damage, DNA single- or double-strand breaks, DNA oxidative modification, and increased sister chromosome crossing frequency, ultimately resulting in carcinogenic effects. The most crucial enzyme in the CYP450 family members mediating AFB1 metabolism to form AFBO is cytochrome P450 3A4 (CYP3A4 for short). CYP3A4 is a heme protein that oxidizes exogenous small organic molecules, such as toxins or drugs, to facilitate their excretion. CYP3A4 protein is a key factor in the carcinogenic process induced by low-dose AFB1 exposure.
[0003] The CRISPR / Cas system is an adaptive immune defense system developed by bacteria and archaea over long periods of evolution to combat invading viruses and foreign DNA. The CRISPR / Cas9 gene editing system consists of a Cas9 protein with endonuclease activity and a single-stranded guide RNA (sgRNA). The sgRNA binds to the Cas9 protein and guides it to the target site for cleavage. The two cleavage domains of Cas9, RuvC and HNH, cleave the DNA double strand, resulting in double-strand breaks (DSBs). The body then initiates DNA repair mechanisms: non-homologous end joining (NHEJ) and homology-directed repair (HR). NHEJ is a mismatch repair mechanism where random insertions and deletions (indels) occur in the DNA double strand. HR is a precise repair mechanism where, in the presence of a homologous donor, the foreign gene fragment from the donor integrates into the target site via homology recombination.
[0004] The RuvC and HNH domains of the Cas9 protein are responsible for cleaving the two strands of the DNA double helix, respectively, determining the nuclease cleavage activity of Cas9. Point mutations in these two domains can cause the loss of Cas9 cleavage activity. When both the RuvC (D10A mutation) and HNH (H840A mutation) domains are simultaneously inactivated (RuvC...), the Cas9 protein loses its cleavage activity. - &HNH -Cas9, lacking nuclease activity, is called dCas9 (dead Cas9). Although dCas9 cannot cleave DNA, it can still bind to specific DNA sequences under the guidance of sgRNA. Studies have found that if sgRNA is designed into the promoter or enhancer region of a target gene, dCas9, after fusing with other proteins, can act as a transcription factor, promoting or inhibiting gene expression. The dCas9-VPR system refers to dCas9 fused with certain transcription activators (herpes simplex virus protein VP64, NF-κB subunit p65, and Epstein-Barr virus R transactivator Rta5), which can target promoter and enhancer regions to regulate gene upregulation. When used in conjunction with an sgRNA library, this system can also support high-throughput genome-wide functional activation screening. The dCas9-KRAB system refers to the fusion of dCas9 with the KRAB (Krüppel-associated box) of Kox1. Relying on KRAB, it recruits various histone modifiers and reversibly inhibits gene expression by forming heterochromatin. It can highly specifically reduce endogenous gene expression by 60-80%, and dCas9-KRAB has no effect on cell growth, making it a non-toxic gene silencing method. Therefore, dCas9 fusion proteins have gradually become a powerful tool for studying biological processes and pathways. The base editor (BE) system refers to the fusion of dCas9 protein with cytosine deaminase or adenine deaminase, thereby achieving C>T or A>G mutations within the editing window. BE has been widely used in life science research. Furthermore, the dCas9-p300 system, which fuses the catalytic core of histone acetyltransferase p300 with dCas9, can directly alter the chromatin state near the target gene. When targeting coding regions or promoter regions, this system can successfully induce high gene expression. Summary of the Invention
[0005] The purpose of this invention is to provide a DNA site-directed mutagenesis system that can introduce random mutations at DNA target sites. The technical problem to be solved is not limited to the technical subject matter described herein; other technical subject matter not mentioned herein will be clearly understood by those skilled in the art through the following description.
[0006] To achieve the above objectives, the present invention first provides a fusion protein named PdCas9-p450, which includes dCas9 protein and CYP3A4 protein.
[0007] The CYP3A4 protein is cytochrome P450 3A4 enzyme (CYP3A4 for short). CYP3A4 is a heme protein that can oxidize exogenous small organic molecules, such as toxins or drugs, so that they can be excreted from the body.
[0008] Further, the dCas9 protein may be either A1) or A2) as follows: A1) the amino acid sequence is the protein of SEQ ID No. 1; A2) a protein that has more than 80% identity with and has the same function as the protein shown in A1) obtained by substituting and / or deleting and / or adding amino acid residues of the amino acid sequence shown in SEQ ID No. 1.
[0009] And / or, the CYP3A4 protein may be B1) or B2 as follows: B1) the amino acid sequence is the protein of SEQ ID No. 2; B2) a protein that has more than 80% identity with and has the same function as the protein shown in B1) obtained by substituting and / or deleting and / or adding amino acid residues of the amino acid sequence shown in SEQ ID No. 2.
[0010] Furthermore, the CYP3A4 protein can be linked to the C-terminus of the dCas9 protein, or the CYP3A4 protein and the C-terminus of the dCas9 protein can be linked via a linker. The linker is used to connect the CYP3A4 protein and the dCas9 protein.
[0011] In one embodiment of the invention, the amino acid sequence of the linker is shown as positions 1369-1401 of SEQ ID No. 3. Further, the amino acid sequence of the fusion protein PdCas9-p450 may be as shown in SEQ ID No. 3.
[0012] Those skilled in the art can readily mutate the nucleotide sequence encoding the fusion protein PdCas9-p450 of this invention using known methods, such as directed evolution or point mutation. Nucleotides that are artificially modified and possess 75% or more identity with the nucleotide sequence of the fusion protein PdCas9-p450 isolated in this invention, as long as they encode and function the fusion protein PdCas9-p450, are derived from and equivalent to the sequence of this invention. The aforementioned 75% or more identity can be 80%, 85%, 90%, or 95% or more. In this document, identity refers to the identity of the amino acid sequence or nucleotide sequence. The identity of the amino acid sequence can be determined using homology search sites on the Internet, such as the BLAST page on the NCBI homepage. For example, in advanced BLAST 2.1, by using blastp as the procedure, setting the Expect value to 10, setting all filters to OFF, using BLOSUM62 as the matrix, setting the Gap existence cost, Per residue gap cost, and Lambda ratio to 11, 1, and 0.85 (default values) respectively, and performing an identity calculation for a pair of amino acid sequences, the identity value (%) can be obtained. In this document, the identity of 80% or more can be at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity.
[0013] The present invention also provides biological materials, which may be any of the following: C1) a nucleic acid molecule encoding the fusion protein PdCas9-p450; C2) an expression cassette containing the nucleic acid molecule of C1); C3) a recombinant vector containing the nucleic acid molecule of C1, or a recombinant vector containing the expression cassette of C2); C4) a recombinant microorganism containing the nucleic acid molecule of C1, or a recombinant microorganism containing the expression cassette of C2, or a recombinant microorganism containing the recombinant vector of C3); C5) a recombinant cell containing the nucleic acid molecule of C1, or a recombinant cell containing the expression cassette of C2, or a recombinant cell containing the recombinant vector of C3); C6) a nucleic acid molecule encoding the dCas9 protein; C7) a nucleic acid molecule encoding the CYP3A4 protein.
[0014] In the above-mentioned biological materials, the nucleic acid molecule may be any of the following:
[0015] D1) A DNA molecule whose coding sequence is shown in positions 1009-5112 of SEQ ID No. 4; D2) A DNA molecule whose coding sequence is shown in positions 5212-6720 of SEQ ID No. 4; D3) A DNA molecule whose coding sequence is shown in positions 1009-6720 of SEQ ID No. 4; D4) A DNA molecule having 75% or more identity with the nucleotide sequence defined in D1) and encoding the dCas9 protein; D5) A DNA molecule having 75% or more identity with the nucleotide sequence defined in D2) and encoding the CYP3A4 protein; D6) A DNA molecule having 75% or more identity with the nucleotide sequence defined in D3) and encoding the fusion protein PdCas9-p450.
[0016] Specifically, the DNA molecule shown in positions 1009-6720 of SEQ ID No. 4 encodes the fusion protein PdCas9-p450; the DNA molecule shown in positions 5212-6720 of SEQ ID No. 4 encodes the optimized CYP3A4 protein; the DNA molecule shown in positions 1009-5112 of SEQ ID No. 4 encodes the dCas9 protein; and the DNA molecule shown in positions 5113-5211 of SEQ ID No. 4 encodes the linker.
[0017] In the aforementioned biological materials, the recombinant vector may further include, but is not limited to, the following operably linked elements: a CMV enhancer, a CMV promoter, and a bGH poly(A) termination signal. Further, the recombinant vector may also include, but is not limited to, the following operably linked elements: an ampicillin resistance gene and a Neo resistance gene.
[0018] The term "operable linkage" refers to the linkage of a regulatory element (e.g., but not limited to, promoters, transcription terminators, etc.) to a nucleic acid (e.g., coding sequences or open reading frames) such that the transcription of nucleotides is controlled and regulated by the transcriptional regulatory element. Techniques for operably linking regulatory elements to nucleic acid molecules are known in the art.
[0019] The nucleotide sequence of the CMV enhancer is shown in positions 235-614 of SEQ ID No. 4; the nucleotide sequence of the CMV promoter is shown in positions 615-818 of SEQ ID No. 4; and the nucleotide sequence of the bGH poly(A) termination signal is shown in positions 6801-7025 of SEQ ID No. 4.
[0020] In the above-mentioned biological materials, the nucleotide sequence of the recombinant vector may be SEQ ID No. 4.
[0021] The method for constructing the recombinant vector includes codon optimization of the coding sequence (CDS) of the CYP3A4 protein and its ligation to the C-terminus of the dCas9 protein.
[0022] The recombinant vector is the expression vector for the fusion protein PdCas9-p450.
[0023] In the aforementioned biological materials, the carrier may be a plasmid, a granule, a bacteriophage, or a viral vector.
[0024] In the aforementioned biological materials, the microorganisms may be yeast, bacteria, algae, or fungi. Specifically, the bacteria may be derived from species such as *Escherichia*, *Erwinia*, *Agrobacterium*, *Flavobacterium*, *Alcaligenes*, *Pseudomonas*, and *Bacillus*. The cells in the aforementioned biological materials may be animal cells, specifically HEK293T cells. The recombinant vector may specifically be the recombinant vector dCas9-p450, whose nucleotide sequence is shown in SEQ ID No. 4, and its map is shown in... Figure 2 As shown.
[0025] The present invention also provides a gene editing system (dCas9-p450 system) (composition), the gene editing system comprising the recombinant vector dCas9-p450, the sgRNA expression vector and aflatoxin B1 (AFB1).
[0026] This invention also provides a method for site-directed mutagenesis using the gene editing system (dCas9-p450 system), the method comprising the following steps: transfecting host cells with a recombinant vector expressing the fusion protein PdCas9-p450 and a target-site-specific sgRNA expression vector; and inducing nucleotide mutations at the target site using aflatoxin B1 (AFB1); wherein the target sequence of the sgRNA is 5′-N. 19-20 PAM-3′, the N 19-20 There are 19-20 N's, and the PAM is NGG; the N's can be A, G, C, or T. The mutation can be a random mutation, such as G mutating into A, C, or T, or A mutating into G, C, or T.
[0027] In one embodiment of the present invention, the method for site-directed mutagenesis using the gene editing system (dCas9-p450 system) includes the following steps:
[0028] (1) Construct the recombinant vector dCas9-p450; (2) Construct the sgRNA expression vector according to the target site of DNA;
[0029] (3) The recombinant vector dCas9-p450 constructed in step (1) and the sgRNA expression vector constructed in step (2) were co-transfected into host cells; (4) Aflatoxin B1 (AFB1) was used to induce random mutations in the nucleotides at the target site.
[0030] The present invention also provides the fusion protein PdCas9-p450, and / or the biomaterial, and / or the application of the gene editing system in gene editing and / or in the mutation of DNA at target sites.
[0031] Furthermore, the mutation can be a random mutation.
[0032] In one embodiment of the present invention, the method for site-directed mutagenesis using the gene editing system may specifically include the following steps: (1) Constructing a dCas9-p450 recombinant expression vector (referred to as recombinant vector dCas9-p450 or dCas9-p450): The synthesized codon-optimized CYP3A4 protein-coding DNA is linked to dCas9 (D10A&H840A) protein-coding DNA through a linker, and then linked with CMV enhancer and CMV promoter, Flag tag protein, SV40NLS nuclear localization signal and bGH. (1) Poly(A) termination signal, construct dCas9-p450 recombinant expression vector; (2) Construct sgRNA expression vector: design sgRNA oligonucleotide primers according to target site, connect them to pGL3-U6-EGFP vector digested with BsaI, and construct pGL3-U6-sgRNA-EGFP vector; (3) Transfect HEK293T cells with the dCas9-p450 recombinant expression vector constructed in step (1) and the sgRNA expression vector constructed in step (2); (4) After transfection for 4-6 h, change the medium with fresh medium containing 4 μg / mL aflatoxin B1 (AFB1), continue to culture the cells for 72 h, and detect the nucleotide mutation status of the target site.
[0033] This invention discloses an in vivo DNA site-directed mutagenesis system. Specifically, dCas9 protein and p450 protein are fused to form a fusion protein PdCas9-p450. sgRNA guides the fusion protein PdCas9-p450 to bind to a target site (target region). Cytochrome p450 oxidase is expressed at the target site. Upon addition of aflatoxin B1 (AFB1), p450 oxidase metabolizes AFB1 into AFB1-8,9-epoxide (AFB1-8,9-epoxide, AFBO). AFBO forms a complex with DNA bases, triggering the body's repair mechanism, thereby introducing random mutations at the target site. Figure 1 As shown.
[0034] This invention synthesizes codon-optimized CYP3A4 protein-coding DNA, ligates a CMV enhancer, CMV promoter, Flag tag protein, NLS nuclear localization signal, dCas9 (D10A & H840A) protein-coding DNA, and CYP3A4 protein-coding DNA, constructs it into the backbone vector pCMV-BE3, and names it dCas9-p450. The recombinant vector dCas9-p450 map is shown below. Figure 2 As shown, the recombinant vector dCas9-p450 sequence was synthesized by the company, transformed into E. coli, amplified, and the plasmid was extracted and stored at -20℃ for later use.
[0035] In one embodiment of the present invention, three pairs of sgRNA oligonucleotide primers are designed according to the desired mutation site, and BsaI restriction sticky ends are added to the 5' end of the primers. After annealing, the primers are ligated into the BsaI-digested pGL3-U6-EGFP vector, transformed into Escherichia coli, single colonies are picked, expanded cultured and sequenced, and the successfully constructed vector is named pGL3-U6-sgRNA-EGFP.
[0036] Furthermore, HEK293T cells were co-transfected with the recombinant vector dCas9-p450 and the sgRNA expression vector using the liposome transfection method (jetPRIME). After 4-6 hours of transfection, the medium was replaced with fresh medium containing 4 μg / mL AFB1. At the same time, control groups were cultured with only plasmid transfection without AFB1, control groups with only AFB1 medium but no plasmid transfection, and wild-type control groups without treatment. Genomic DNA was extracted from all cells at the same time 72 hours after transfection.
[0037] Furthermore, corresponding primers were designed to amplify the target site regions of the cell genome, and after purification, high-throughput sequencing (deep sequence) was performed to analyze the target regions and obtain mutation types and mutation frequencies.
[0038] The beneficial effects of this invention are as follows: This invention provides a method for high expression of the p450 enzyme in a specific genomic region using a CRISPR / Cas9 system, thereby causing AFB1 in that region to metabolize into AFBO. AFBO then forms a complex with target site DNA, allowing for the targeted introduction of random mutations. This method can mutate specific nucleotide sequences, providing a feasible approach and a more powerful and flexible gene editing tool for gene function research. Attached Figure Description
[0039] Figure 1 This is a schematic diagram illustrating the principle of site-directed mutagenesis of dCas9-p450 in this invention.
[0040] Figure 2This is a spectrum of the dCas9-p450 carrier element synthesized according to the sequence of this invention.
[0041] Figure 3 The plasmid map of the sgRNA expression vector pGL3-U6-sgRNA-EGFP constructed for this invention.
[0042] Figure 4 This is a diagram showing the high-throughput sequencing (deep sequence) results of the target regions of the three gene experimental groups in this embodiment of the invention.
[0043] Figure 5 This is a diagram showing the high-throughput sequencing (deep sequence) results of the target regions for the three gene control groups in this embodiment of the invention. Detailed Implementation
[0044] The present invention will now be described in further detail with reference to specific embodiments. The given embodiments are merely illustrative of the invention and not intended to limit its scope. The embodiments provided below can serve as a guide for further improvements by those skilled in the art and do not constitute a limitation on the invention in any way.
[0045] Unless otherwise specified, the experimental methods used in the following examples are conventional methods, performed according to the techniques or conditions described in the literature in this field or according to the product instructions. Unless otherwise specified, the materials and reagents used in the following examples are commercially available.
[0046] Example 1: Construction of the recombinant vector dCas9-p450
[0047] A codon-optimized CYP3A4 sequence was synthesized and linked with the CMV enhancer, CMV promoter, Flag tag protein, NLS nuclear localization signal, dCas9 (D10A & H840A) protein sequence, and the CDS sequence of CYP3A4. This sequence was then constructed into the backbone vector pCMV-BE3 (purchased from Addgene) and named dCas9-p450, which is the recombinant vector dCas9-p450. The recombinant vector dCas9-p450 is the expression vector for the fusion protein PdCas9-p450. The map of the recombinant vector dCas9-p450 is shown below. Figure 2As shown in SEQ ID NO.4, the nucleotide sequence of the recombinant vector dCas9-p450 was synthesized by Suzhou Genewise Biotechnology Co., Ltd., transformed into Escherichia coli, amplified, and the plasmid was extracted and stored at -20℃ for later use. Wherein: the DNA molecule shown in positions 1009-5112 of SEQ ID No. 4 encodes the dCas9 protein; the DNA molecule shown in positions 5212-6720 of SEQ ID No. 4 (a DNA molecule whose encoding DNA of the CYP3A4 protein has been codon-optimized) encodes the CYP3A4 protein; the DNA molecule shown in positions 5113-5211 of SEQ ID No. 4 encodes the linker; the DNA molecule shown in positions 1009-6720 of SEQ ID No. 4 encodes the fusion protein PdCas9-p450; the nucleotide sequence of the CMV enhancer is shown in positions 235-614 of SEQ ID No. 4; the nucleotide sequence of the CMV promoter is shown in positions 615-818 of SEQ ID No. 4; the nucleotide sequence of the Flag tag protein is shown in positions 907-972 of SEQ ID No. 4; and the nucleotide sequence of the NLS nuclear localization signal is shown in positions 979-999 and 6721-6741 of SEQ ID No. 4.
[0048] Example 2: Construction of sgRNA expression vector
[0049] Three sgRNAs were randomly selected from the human genome at GRIN2B, DYRK1A, and PDCD1 sites, and named GRIN2B-sgRNA, DYRK1A-sgRNA, and PDCD1-sgRNA, respectively. The upstream and downstream sequences of each sgRNA guide sequence were modified by adding corresponding BsaI restriction enzyme sticky-terminal bases to their 5' ends before being sent to the company for synthesis. The oligonucleotide sequences to be synthesized are shown in Table 1, with the underlined portions representing sticky-terminal bases. After centrifuging the synthesized sgRNA guide sequence oligonucleotide powder, it was dissolved and diluted with ddH2O to a final concentration of 10 μM. Then, 5 μL of each of the corresponding upstream and downstream sgRNA primers were added to a PCR tube, vortexed, briefly centrifuged, and annealed to form double-stranded oligos. The annealing program was 95℃ for 10 min; 65℃ for 30 min.
[0050] Table 1. Guide primer sequences for sgRNA of 3 genes
[0051]
[0052]
[0053] The annealed product was reacted with Ligation Mix at 16°C in a metal bath for 1 hour and then ligated into the backbone of the pGL3-U6-EGFP vector (Addgene#107721) after BsaI digestion and recovery, resulting in a 3-site pGL3-U6-sgRNA-EGFP expression vector. Figure 3 The three expression vectors were GRIN2B-sgRNA, DYRK1A-sgRNA, and PDCD1-sgRNA, respectively. After transformation into competent *E. coli* cells, single clones were picked, and Sanger sequencing was performed to verify successful sgRNA ligation. Following successful sgRNA ligation, the cells were expanded using bacterial culture, and plasmids were extracted using an endotoxin-free plasmid extraction kit for transfection.
[0054] Example 3: Validation of Mutation Efficiency of Gene Editing System (dCas9-p450 System)
[0055] The gene editing system (dCas9-p450 system) includes the recombinant vector dCas9-p450 from Example 1, the sgRNA expression vector (pGL3-U6-sgRNA-EGFP expression vector) from Example 2, and aflatoxin B1 (AFB1). The dCas9-p450 system was used to edit HEK293T cells to verify the mutation efficiency.
[0056] Three gene experimental groups were set up: Before transfection, HEK293T cells were plated in 12-well cell culture dishes, and transfection was performed when the cell confluence reached 80-90%. HEK293T cells were co-transfected with 500 ng each of the sgRNA expression vectors for the three sites constructed in Example 2 and the recombinant vector dCas9-p450 constructed in Example 1 using a liposome transfection method (jetPRIME). 4-6 hours after transfection, the culture medium was changed, and the cells were cultured in fresh medium containing 4 μg / mL AFB1 (10% FBS + 90% DMEM). Cell luminescence was observed under a fluorescence microscope 24 hours after transfection.
[0057] Meanwhile, three control groups of cells were set up: a group with only plasmid transfection without AFB1 culture medium (the only difference from the three gene experimental groups was that the culture medium was replaced with fresh medium without 4 μg / mL AFB1 4-6 hours after plasmid transfection), a group with only culture medium containing 4 μg / mL AFB1 without plasmid transfection (the only difference from the three gene experimental groups was that plasmid transfection was not performed), and a wild-type group without any treatment (the only difference from the three gene experimental groups was that plasmid transfection was not performed, and the cells were cultured in fresh medium without 4 μg / mL AFB1). After culturing for 48 hours, all cells were collected, and genomic DNA was extracted from the cells. Corresponding primers were designed to amplify the target fragments with target sites (Table 4).
[0058] The PCR amplification procedure and amplification system are as follows:
[0059] Table 2 PCR amplification system
[0060]
[0061] Table 3 PCR amplification program
[0062]
[0063] Table 4 Primer sequence list for amplifying the target sequences of the three genes.
[0064] Gene name Upstream primer sequence (5'-3') Downstream primer sequence (5'-3') GRIN2B GGTTTGGTGCTCAATGAAAGG CCACCTCGTCGGAAGTGC DYRK1A ACCTCACTTATCTTCTTGTAGGAGG ACTGCCATTCCAATAGTCATTTCTG PDCD1 ACAGTTTCCCTTCCGCTCAC GGACTGAGGGTGGAAGGTCC
[0065] The PCR products were extracted using a gel extraction kit and then analyzed by high-throughput sequencing (deepsequence). The sequencing results of the experimental group and the control group were compared and analyzed to determine the mutation type and mutation frequency. Figure 4 , Figure 5 ).
[0066] The results showed that no mutations were found at the target sites of any genes in the untransfected plasmid group and the wild-type group without any treatment when cultured in medium containing 4 μg / mL AFB1. Only one mutation type was found in the experimental group transfected with the GRIN2B-sgRNA expression vector: the G at position 9 in the 5′ to 3′ direction of the target sequence was replaced with A. This mutation occurred in 1029 out of a total of 910,549 reads, with a mutation frequency of 0.1%.
[0067] The experimental group transfected only with the DYRK1A-sgRNA expression vector had only one type of mutation: the G at position 19 in the 5′ to 3′ direction of the target sequence was replaced with A. This mutation occurred in 1081 out of a total of 994,452 reads, with a mutation frequency of 0.1%.
[0068] The experimental group transfected only with the PDCD1-sgRNA expression vector had only one type of mutation: the C at position 4 in the 5′ to 3′ direction of the target sequence was replaced with T. This mutation occurred in 1081 out of a total of 994,452 reads, with a mutation frequency of 0.1%.
[0069] In summary, the gene editing system (dCas9-p450 system) constructed in this invention can achieve gene editing by random point mutations in specific genomic regions.
[0070] The present invention has been described in detail above. For those skilled in the art, the invention can be practiced in a wide range of ways with equivalent parameters, concentrations, and conditions without departing from its spirit and scope, and without requiring unnecessary experiments. Although specific embodiments have been given, it should be understood that further modifications can be made to the invention. In summary, according to the principles of the invention, this application is intended to include any changes, uses, or improvements to the invention, including changes made using conventional techniques known in the art that depart from the scope disclosed herein. Some of the essential features can be applied within the scope of the following appended claims. SEQUENCE LISTING <110> Huazhong Agricultural University <120> A method for site-directed mutagenesis using the dCas9-p450 system <160> 4 <170> PatentIn version 3.5 <210> 1 <211> 1368 <212> PRT <213> Artificial sequence <400> 1 Met Asp Lys Lys Tyr Ser Ile Gly Leu Ala Ile Gly Thr Asn Ser Val 1 5 10 15 Gly Trp Ala Val Ile Thr Asp Glu Tyr Lys Val Pro Ser Lys Lys Phe 20 25 30 Lys Val Leu Gly Asn Thr Asp Arg His Ser Ile Lys Lys Asn Leu Ile 35 40 45 Gly Ala Leu Leu Phe Asp Ser Gly Glu Thr Ala Glu Ala Thr Arg Leu 50 55 60 Lys Arg Thr Ala Arg Arg Arg Tyr Thr Arg Arg Lys Asn Arg Ile Cys 65 70 75 80 Tyr Leu Gln Glu Ile Phe Ser Asn Glu Met Ala Lys Val Asp Asp Ser 85 90 95 Phe Phe His Arg Leu Glu Glu Ser Phe Leu Val Glu Glu Asp Lys Lys 100 105 110 His Glu Arg His Pro Ile Phe Gly Asn Ile Val Asp Glu Val Ala Tyr 115 120 125 His Glu Lys Tyr Pro Thr Ile Tyr His Leu Arg Lys Lys Leu Val Asp 130 135 140 Ser Thr Asp Lys Ala Asp Leu Arg Leu Ile Tyr Leu Ala Leu Ala His 145 150 155 160 Met Ile Lys Phe Arg Gly His Phe Leu Ile Glu Gly Asp Leu Asn Pro 165 170 175 Asp Asn Ser Asp Val Asp Lys Leu Phe Ile Gln Leu Val Gln Thr Tyr 180 185 190 Asn Gln Leu Phe Glu Glu Asn Pro Ile Asn Ala Ser Gly Val Asp Ala 195 200 205 Lys Ala Ile Leu Ser Ala Arg Leu Ser Lys Ser Arg Arg Leu Glu Asn 210 215 220 Leu Ile Ala Gln Leu Pro Gly Glu Lys Lys Asn Gly Leu Phe Gly Asn 225 230 235 240 Leu Ile Ala Leu Ser Leu Gly Leu Thr Pro Asn Phe Lys Ser Asn Phe 245 250 255 Asp Leu Ala Glu Asp Ala Lys Leu Gln Leu Ser Lys Asp Thr Tyr Asp 260 265 270 Asp Asp Leu Asp Asn Leu Leu Ala Gln Ile Gly Asp Gln Tyr Ala Asp 275 280 285 Leu Phe Leu Ala Ala Lys Asn Leu Ser Asp Ala Ile Leu Leu Ser Asp 290 295 300 Ile Leu Arg Val Asn Thr Glu Ile Thr Lys Ala Pro Leu Ser Ala Ser 305 310 315 320 Met Ile Lys Arg Tyr Asp Glu His His Gln Asp Leu Thr Leu Leu Lys 325 330 335 Ala Leu Val Arg Gln Gln Leu Pro Glu Lys Tyr Lys Glu Ile Phe Phe 340 345 350 Asp Gln Ser Lys Asn Gly Tyr Ala Gly Tyr Ile Asp Gly Gly Ala Ser 355 360 365 Gln Glu Glu Phe Tyr Lys Phe Ile Lys Pro Ile Leu Glu Lys Met Asp 370 375 380 Gly Thr Glu Glu Leu Leu Val Lys Leu Asn Arg Glu Asp Leu Leu Arg 385 390 395 400 Lys Gln Arg Thr Phe Asp Asn Gly Ser Ile Pro His Gln Ile His Leu 405 410 415 Gly Glu Leu His Ala Ile Leu Arg Arg Gln Glu Asp Phe Tyr Pro Phe 420 425 430 Leu Lys Asp Asn Arg Glu Lys Ile Glu Lys Ile Leu Thr Phe Arg Ile 435 440 445 Pro Tyr Tyr Val Gly Pro Leu Ala Arg Gly Asn Ser Arg Phe Ala Trp 450 455 460 Met Thr Arg Lys Ser Glu Glu Thr Ile Thr Pro Trp Asn Phe Glu Glu 465 470 475 480 Val Val Asp Lys Gly Ala Ser Ala Gln Ser Phe Ile Glu Arg Met Thr 485 490 495 Asn Phe Asp Lys Asn Leu Pro Asn Glu Lys Val Leu Pro Lys His Ser 500 505 510 Leu Leu Tyr Glu Tyr Phe Thr Val Tyr Asn Glu Leu Thr Lys Val Lys 515 520 525 Tyr Val Thr Glu Gly Met Arg Lys Pro Ala Phe Leu Ser Gly Glu Gln 530 535 540 Lys Lys Ala Ile Val Asp Leu Leu Phe Lys Thr Asn Arg Lys Val Thr 545 550 555 560 Val Lys Gln Leu Lys Glu Asp Tyr Phe Lys Lys Ile Glu Cys Phe Asp 565 570 575 Ser Val Glu Ile Ser Gly Val Glu Asp Arg Phe Asn Ala Ser Leu Gly 580 585 590 Thr Tyr His Asp Leu Leu Lys Ile Ile Lys Asp Lys Asp Phe Leu Asp 595 600 605 Asn Glu Glu Asn Glu Asp Ile Leu Glu Asp Ile Val Leu Thr Leu Thr 610 615 620 Leu Phe Glu Asp Arg Glu Met Ile Glu Glu Arg Leu Lys Thr Tyr Ala 625 630 635 640 His Leu Phe Asp Asp Lys Val Met Lys Gln Leu Lys Arg Arg Arg Tyr 645 650 655 Thr Gly Trp Gly Arg Leu Ser Arg Lys Leu Ile Asn Gly Ile Arg Asp 660 665 670 Lys Gln Ser Gly Lys Thr Ile Leu Asp Phe Leu Lys Ser Asp Gly Phe 675 680 685 Ala Asn Arg Asn Phe Met Gln Leu Ile His Asp Asp Ser Leu Thr Phe 690 695 700 Lys Glu Asp Ile Gln Lys Ala Gln Val Ser Gly Gln Gly Asp Ser Leu 705 710 715 720 His Glu His Ile Ala Asn Leu Ala Gly Ser Pro Ala Ile Lys Lys Gly 725 730 735 Ile Leu Gln Thr Val Lys Val Val Asp Glu Leu Val Lys Val Met Gly 740 745 750 Arg His Lys Pro Glu Asn Ile Val Ile Glu Met Ala Arg Glu Asn Gln 755 760 765 Thr Thr Gln Lys Gly Gln Lys Asn Ser Arg Glu Arg Met Lys Arg Ile 770 775 780 Glu Glu Gly Ile Lys Glu Leu Gly Ser Gln Ile Leu Lys Glu His Pro 785 790 795 800 Val Glu Asn Thr Gln Leu Gln Asn Glu Lys Leu Tyr Leu Tyr Tyr Leu 805 810 815 Gln Asn Gly Arg Asp Met Tyr Val Asp Gln Glu Leu Asp Ile Asn Arg 820 825 830 Leu Ser Asp Tyr Asp Val Asp Ala Ile Val Pro Gln Ser Phe Leu Lys 835 840 845 Asp Asp Ser Ile Asp Asn Lys Val Leu Thr Arg Ser Asp Lys Asn Arg 850 855 860 Gly Lys Ser Asp Asn Val Pro Ser Glu Glu Val Val Lys Lys Met Lys 865 870 875 880 Asn Tyr Trp Arg Gln Leu Leu Asn Ala Lys Leu Ile Thr Gln Arg Lys 885 890 895 Phe Asp Asn Leu Thr Lys Ala Glu Arg Gly Gly Leu Ser Glu Leu Asp 900 905 910 Lys Ala Gly Phe Ile Lys Arg Gln Leu Val Glu Thr Arg Gln Ile Thr 915 920 925 Lys His Val Ala Gln Ile Leu Asp Ser Arg Met Asn Thr Lys Tyr Asp 930 935 940 Glu Asn Asp Lys Leu Ile Arg Glu Val Lys Val Ile Thr Leu Lys Ser 945 950 955 960 Lys Leu Val Ser Asp Phe Arg Lys Asp Phe Gln Phe Tyr Lys Val Arg 965 970 975 Glu Ile Asn Asn Tyr His His Ala His Asp Ala Tyr Leu Asn Ala Val 980 985 990 Val Gly Thr Ala Leu Ile Lys Lys Tyr Pro Lys Leu Glu Ser Glu Phe 995 1000 1005 Val Tyr Gly Asp Tyr Lys Val Tyr Asp Val Arg Lys Met Ile Ala 1010 1015 1020 Lys Ser Glu Gln Glu Ile Gly Lys Ala Thr Ala Lys Tyr Phe Phe 1025 1030 1035 Tyr Ser Asn Ile Met Asn Phe Phe Lys Thr Glu Ile Thr Leu Ala 1040 1045 1050 Asn Gly Glu Ile Arg Lys Arg Pro Leu Ile Glu Thr Asn Gly Glu 1055 1060 1065 Thr Gly Glu Ile Val Trp Asp Lys Gly Arg Asp Phe Ala Thr Val 1070 1075 1080 Arg Lys Val Leu Ser Met Pro Gln Val Asn Ile Val Lys Lys Thr 1085 1090 1095 Glu Val Gln Thr Gly Gly Phe Ser Lys Glu Ser Ile Leu Pro Lys 1100 1105 1110 Arg Asn Ser Asp Lys Leu Ile Ala Arg Lys Lys Asp Trp Asp Pro 1115 1120 1125 Light Light Tyr Gly Gly Phe Asp Ser Pro Thr Val Ala Tyr Ser Val 1130 1135 1140 Leu Val Val Ala Lys Val Glu Lys Gly Lys Ser Lys Lys Leu Lys 1145 1150 1155 Ser Val Lys Glu Leu Leu Gly Ile Thr Ile Met Glu Arg Ser Ser 1160 1165 1170 Phe Glu Lys Asn Pro Ile Asp Phe Leu Glu Ala Lys Gly Tyr Lys 1175 1180 1185 Glu Val Lys Lys Asp Leu Ile Ile Lys Leu Pro Lys Tyr Ser Leu 1190 1195 1200 Phe Glu Leu Glu Asn Gly Arg Lys Arg Met Leu Ala Ser Ala Gly 1205 1210 1215 Glu Leu Gln Lys Gly Asn Glu Leu Ala Leu Pro Ser Lys Tyr Val 1220 1225 1230 Asn Phe Leu Tyr Leu Ala Ser His Tyr Glu Lys Leu Lys Gly Ser 1235 1240 1245 Pro Glu Asp Asn Glu Gln Lys Gln Leu Phe Val Glu Gln His Lys 1250 1255 1260 His Tyr Leu Asp Glu Ile Ile Glu Gln Ile Ser Glu Phe Ser Lys 1265 1270 1275 Arg Val Ile Leu Ala Asp Ala Asn Leu Asp Lys Val Leu Ser Ala 1280 1285 1290 Tyr Asn Lys His Arg Asp Lys Pro Ile Arg Glu Gln Ala Glu Asn 1295 1300 1305 Ile Ile His Leu Phe Thr Leu Thr Asn Leu Gly Ala Pro Ala Ala 1310 1315 1320 Phe Lys Tyr Phe Asp Thr Thr Ile Asp Arg Lys Arg Tyr Thr Ser 1325 1330 1335 Thr Lys Glu Val Leu Asp Ala Thr Leu Ile His Gln Ser Ile Thr 1340 1345 1350 Gly Leu Tyr Glu Thr Arg Ile Asp Leu Ser Gln Leu Gly Gly Asp 1355 1360 1365 <210> 2 <211> 503 <212> PRT <213> Artificial sequence <400> 2 Met Ala Leu Ile Pro Asp Leu Ala Met Glu Thr Trp Leu Leu Leu Ala 1 5 10 15 Val Ser Leu Val Leu Leu Tyr Leu Tyr Gly Thr His Ser His Gly Leu 20 25 30 Phe Lys Lys Leu Gly Ile Pro Gly Pro Thr Pro Leu Pro Phe Leu Gly 35 40 45 Asn Ile Leu Ser Tyr His Lys Gly Phe Cys Met Phe Asp Met Glu Cys 50 55 60 His Lys Lys Tyr Gly Lys Val Trp Gly Phe Tyr Asp Gly Gln Gln Pro 65 70 75 80 Val Leu Ala Ile Thr Asp Pro Asp Met Ile Lys Thr Val Leu Val Lys 85 90 95 Glu Cys Tyr Ser Val Phe Thr Asn Arg Arg Pro Phe Gly Pro Val Gly 100 105 110 Phe Met Lys Ser Ala Ile Ser Ile Ala Glu Asp Glu Glu Trp Lys Arg 115 120 125 Leu Arg Ser Leu Leu Ser Pro Thr Phe Thr Ser Gly Lys Leu Lys Glu 130 135 140 Met Val Pro Ile Ile Ala Gln Tyr Gly Asp Val Leu Val Arg Asn Leu 145 150 155 160 Arg Arg Glu Ala Glu Thr Gly Lys Pro Val Thr Leu Lys Asp Val Phe 165 170 175 Gly Ala Tyr Ser Met Asp Val Ile Thr Ser Thr Ser Phe Gly Val Asn 180 185 190 Ile Asp Ser Leu Asn Asn Pro Gln Asp Pro Phe Val Glu Asn Thr Lys 195 200 205 Lys Leu Leu Arg Phe Asp Phe Leu Asp Pro Phe Phe Leu Ser Ile Thr 210 215 220 Val Phe Pro Phe Leu Ile Pro Ile Leu Glu Val Leu Asn Ile Cys Val 225 230 235 240 Phe Pro Arg Glu Val Thr Asn Phe Leu Arg Lys Ser Val Lys Arg Met 245 250 255 Lys Glu Ser Arg Leu Glu Asp Thr Gln Lys His Arg Val Asp Phe Leu 260 265 270 Gln Leu Met Ile Asp Ser Gln Asn Ser Lys Glu Thr Glu Ser His Lys 275 280 285 Ala Leu Ser Asp Leu Glu Leu Val Ala Gln Ser Ile Ile Phe Ile Phe 290 295 300 Ala Gly Tyr Glu Thr Thr Ser Ser Val Leu Ser Phe Ile Met Tyr Glu 305 310 315 320 Leu Ala Thr His Pro Asp Val Gln Gln Lys Leu Gln Glu Glu Ile Asp 325 330 335 Ala Val Leu Pro Asn Lys Ala Pro Pro Thr Tyr Asp Thr Val Leu Gln 340 345 350 Met Glu Tyr Leu Asp Met Val Val Asn Glu Thr Leu Arg Leu Phe Pro 355 360 365 Ile Ala Met Arg Leu Glu Arg Val Cys Lys Lys Asp Val Glu Ile Asn 370 375 380 Gly Met Phe Ile Pro Lys Gly Val Val Val Met Ile Pro Ser Tyr Ala 385 390 395 400 Leu His Arg Asp Pro Lys Tyr Trp Thr Glu Pro Glu Lys Phe Leu Pro 405 410 415 Glu Arg Phe Ser Lys Lys Asn Lys Asp Asn Ile Asp Pro Tyr Ile Tyr 420 425 430 Thr Pro Phe Gly Ser Gly Pro Arg Asn Cys Ile Gly Met Arg Phe Ala 435 440 445 Leu Met Asn Met Lys Leu Ala Leu Ile Arg Val Leu Gln Asn Phe Ser 450 455 460 Phe Lys Pro Cys Lys Glu Thr Gln Ile Pro Leu Lys Leu Ser Leu Gly 465 470 475 480 Gly Leu Leu Gln Pro Glu Lys Pro Val Val Leu Lys Val Glu Ser Arg 485 490 495 Asp Gly Thr Val Ser Gly Ala 500 <210> 3 <211> 1904 <212> PRT <213> Artificial sequence <400> 3 Met Asp Lys Lys Tyr Ser Ile Gly Leu Ala Ile Gly Thr Asn Ser Val 1 5 10 15 Gly Trp Ala Val Ile Thr Asp Glu Tyr Lys Val Pro Ser Lys Lys Phe 20 25 30 Lys Val Leu Gly Asn Thr Asp Arg His Ser Ile Lys Asn Leu Ile 35 40 45 Gly Ala Leu Phe Asp Ser Gly Glu Thr Ala Glu Ala Thr Arg Leu 50 55 60 Lys Arg Thr Ala Arg Arg Tyr Thr Arg Arg Lys Asn Arg With Cys 65 70 75 80 Tyr Leu Gln Glu Ile Phe Ser Asn Glu Met Ala Asp Ser 85 90 95 Phe Phe His Arg Leu Glu Glu Ser Phe Leu Val Glu Glu Asp Lys Lys 100 105 110 His Glu Arg His Pro Ile Phe Gly Asn Ile Val Asp Glu Val Ala Tyr 115 120 125 His Glu Lys Tyr Pro Thr Ile Tyr His Leu Arg Lys Leu Val Asp 130 135 140 Ser Thr Asp Lys Ala Asp Leu Arg Leu Ile Tyr Leu Ala Leu Ala His 145 150 155 160 Met Ile Lys Phe Arg Gly His Phe Leu Ile Glu Gly Asp Leu Asn Pro 165 170 175 Asp Asn Ser Asp Val Asp Lys Leu Phe Ile Gln Leu Val Gln Thr Tyr 180 185 190 Asn Gln Leu Phe Glu Glu Asn Pro Ile Asn Ala Ser Gly Val Asp Ala 195 200 205 Lys Ala Ile Leu Ser Ala Arg Leu Ser Lys Ser Arg Arg Leu Glu Asn 210 215 220 Leu Ile Ala Gln Leu Pro Gly Glu Lys Lys Asn Gly Leu Phe Gly Asn 225 230 235 240 Leu Ile Ala Leu Ser Leu Gly Leu Thr Pro Asn Phe Lys Ser Asn Phe 245 250 255 Asp Leu Ala Glu Asp Ala Lys Leu Gln Leu Ser Lys Asp Thr Tyr Asp 260 265 270 Asp Asp Leu Asp Asn Leu Leu Ala Gln Ile Gly Asp Gln Tyr Ala Asp 275 280 285 Leu Phe Leu Ala Ala Lys Asn Leu Ser Asp Ala Ile Leu Leu Ser Asp 290 295 300 Ile Leu Arg Val Asn Thr Glu Ile Thr Lys Ala Pro Leu Ser Ala Ser 305 310 315 320 Met Ile Lys Arg Tyr Asp Glu His His Gln Asp Leu Thr Leu Leu Lys 325 330 335 Ala Leu Val Arg Gln Gln Leu Pro Glu Lys Tyr Lys Glu Ile Phe Phe 340 345 350 Asp Gln Ser Lys Asn Gly Tyr Ala Gly Tyr Ile Asp Gly Gly Ala Ser 355 360 365 Gln Glu Glu Phe Tyr Lys Phe Ile Lys Pro Ile Leu Glu Lys Met Asp 370 375 380 Gly Thr Glu Glu Leu Leu Val Lys Leu Asn Arg Glu Asp Leu Leu Arg 385 390 395 400 Lys Gln Arg Thr Phe Asp Asn Gly Ser Ile Pro His Gln Ile His Leu 405 410 415 Gly Glu Leu His Ala Ile Leu Arg Arg Gln Glu Asp Phe Tyr Pro Phe 420 425 430 Leu Lys Asp Asn Arg Glu Lys Ile Glu Lys Ile Leu Thr Phe Arg Ile 435 440 445 Pro Tyr Tyr Val Gly Pro Leu Ala Arg Gly Asn Ser Arg Phe Ala Trp 450 455 460 Met Thr Arg Lys Ser Glu Glu Thr Ile Thr Pro Trp Asn Phe Glu Glu 465 470 475 480 Val Val Asp Lys Gly Ala Ser Ala Gln Ser Phe Ile Glu Arg Met Thr 485 490 495 Asn Phe Asp Lys Asn Leu Pro Asn Glu Lys Val Leu Pro Lys His Ser 500 505 510 Leu Leu Tyr Glu Tyr Phe Thr Val Tyr Asn Glu Leu Thr Lys Val Lys 515 520 525 Tyr Val Thr Glu Gly Met Arg Lys Pro Ala Phe Leu Ser Gly Glu Gln 530 535 540 Lys Lys Ala Ile Val Asp Leu Leu Phe Lys Thr Asn Arg Lys Val Thr 545 550 555 560 Val Lys Gln Leu Lys Glu Asp Tyr Phe Lys Lys Ile Glu Cys Phe Asp 565 570 575 Ser Val Glu Ile Ser Gly Val Glu Asp Arg Phe Asn Ala Ser Leu Gly 580 585 590 Thr Tyr His Asp Leu Leu Lys Ile Ile Lys Asp Lys Asp Phe Leu Asp 595 600 605 Asn Glu Glu Asn Glu Asp Ile Leu Glu Asp Ile Val Leu Thr Leu Thr 610 615 620 Leu Phe Glu Asp Arg Glu Met Ile Glu Glu Arg Leu Lys Thr Tyr Ala 625 630 635 640 His Leu Phe Asp Asp Lys Val Met Lys Gln Leu Lys Arg Arg Arg Tyr 645 650 655 Thr Gly Trp Gly Arg Leu Ser Arg Lys Leu Ile Asn Gly Ile Arg Asp 660 665 670 Lys Gln Ser Gly Lys Thr Ile Leu Asp Phe Leu Lys Ser Asp Gly Phe 675 680 685 Ala Asn Arg Asn Phe Met Gln Leu Ile His Asp Asp Ser Leu Thr Phe 690 695 700 Lys Glu Asp Ile Gln Lys Ala Gln Val Ser Gly Gln Gly Asp Ser Leu 705 710 715 720 His Glu His Ile Ala Asn Leu Ala Gly Ser Pro Ala Ile Lys Lys Gly 725 730 735 Ile Leu Gln Thr Val Lys Val Val Asp Glu Leu Val Lys Val Met Gly 740 745 750 Arg His Lys Pro Glu Asn Ile Val Ile Glu Met Ala Arg Glu Asn Gln 755 760 765 Thr Thr Gln Lys Gly Gln Lys Asn Ser Arg Glu Arg Met Lys Arg Ile 770 775 780 Glu Glu Gly Ile Lys Glu Leu Gly Ser Gln Ile Leu Lys Glu His Pro 785 790 795 800 Val Glu Asn Thr Gln Leu Gln Asn Glu Lys Leu Tyr Leu Tyr Tyr Leu 805 810 815 Gln Asn Gly Arg Asp Met Tyr Val Asp Gln Glu Leu Asp Ile Asn Arg 820 825 830 Leu Ser Asp Tyr Asp Val Asp Ala Ile Val Pro Gln Ser Phe Leu Lys 835 840 845 Asp Asp Ser Ile Asp Asn Lys Val Leu Thr Arg Ser Asp Lys Asn Arg 850 855 860 Gly Lys Ser Asp Asn Val Pro Ser Glu Glu Val Val Lys Lys Met Lys 865 870 875 880 Asn Tyr Trp Arg Gln Leu Leu Asn Ala Lys Leu Ile Thr Gln Arg Lys 885 890 895 Phe Asp Asn Leu Thr Lys Ala Glu Arg Gly Gly Leu Ser Glu Leu Asp 900 905 910 Lys Ala Gly Phe Ile Lys Arg Gln Leu Val Glu Thr Arg Gln Ile Thr 915 920 925 Lys His Val Ala Gln Ile Leu Asp Ser Arg Met Asn Thr Lys Tyr Asp 930 935 940 Glu Asn Asp Lys Leu Ile Arg Glu Val Lys Val Ile Thr Lys Ser 945 950 955 960 Lys Leu Val Ser Asp Phe Arg Lys Asp Phe Gln Phe Tyr Lys Val Arg 965,970,975 Glu Ile Asn Asn Tyr His Ala His Asp Ala Tyr Leu Asn Ala Val 980,985,990 Val Gly Thr Ala Leu Ile Lys Tyr Pro Lys Leu Glu Ser Glu Phe 995 1000 1005 Val Tyr Gly Asp Tyr Lys Val Tyr Asp Val Arg Lys Met Ile Ala 1010 1015 1020 Lys Ser Glu Gln Glu Ile Gly Lys Ala Thr Ala Lys Tyr Phe Phe 1025 1030 1035 Tyr Ser Asn Ile Met Asn Phe Phe Lys Thr Glu Ile Thr Leu Ala 1040 1045 1050 Asn Gly Glu With Arg Lys Arg Pro Leu Glu Thr Asn Gly Glu 1055 1060 1065 Thr Gly Glu Ile Val Trp Asp Lys Gly Arg Asp Phe Ala Thr Val 1070 1075 1080 Arg Lys Will Be Met Pro Gln Val Asn Ile Val Lys Thr 1085 1090 1095 Glu Val Gln Thr Gly Gly Phe Ser Lys Glu Ser Ile Leu Pro Lys 1100 1105 1110 Arg Asn Ser Asp Lys Leu Ile Ala Arg Lys Lys Asp Trp Asp Pro 1115 1120 1125 Light Light Tyr Gly Gly Phe Asp Ser Pro Thr Val Ala Tyr Ser Val 1130 1135 1140 Leu Val Val Ala Lys Val Glu Lys Gly Lys Ser Lys Lys Leu Lys 1145 1150 1155 Ser Val Lys Glu Leu Leu Gly Ile Thr Ile Met Glu Arg Ser Ser 1160 1165 1170 Phe Glu Lys Asn Pro Ile Asp Phe Leu Glu Ala Lys Gly Tyr Lys 1175 1180 1185 Glu Val Lys Lys Asp Leu Ile Ile Lys Leu Pro Lys Tyr Ser Leu 1190 1195 1200 Phe Glu Leu Glu Asn Gly Arg Lys Arg Met Leu Ala Ser Ala Gly 1205 1210 1215 Glu Leu Gln Lys Gly Asn Glu Leu Ala Leu Pro Ser Lys Tyr Val 1220 1225 1230 Asn Phe Leu Tyr Leu Ala Ser His Tyr Glu Lys Leu Lys Gly Ser 1235 1240 1245 Pro Glu Asp Asn Glu Gln Lys Gln Leu Phe Val Glu Gln His Lys 1250 1255 1260 His Tyr Leu Asp Glu Ile Ile Glu Gln Ile Ser Glu Phe Ser Lys 1265 1270 1275 Arg Val Ile Leu Ala Asp Ala Asn Leu Asp Lys Val Leu Ser Ala 1280 1285 1290 Tyr Asn Lys His Arg Asp Lys Pro Ile Arg Glu Gln Ala Glu Asn 1295 1300 1305 Ile Ile His Leu Phe Thr Leu Thr Asn Leu Gly Ala Pro Ala Ala 1310 1315 1320 Phe Lys Tyr Phe Asp Thr Thr Ile Asp Arg Lys Arg Tyr Thr Ser 1325 1330 1335 Thr Lys Glu Val Leu Asp Ala Thr Leu Ile His Gln Ser Ile Thr 1340 1345 1350 Gly Leu Tyr Glu Thr Arg Ile Asp Leu Ser Gln Leu Gly Gly Asp 1355 1360 1365 Ser Gly Gly Ser Ser Gly Gly Ser Ser Gly Ser Glu Thr Pro Gly 1370 1375 1380 Thr Ser Glu Ser Ala Thr Pro Glu Ser Ser Gly Gly Ser Ser Gly 1385 1390 1395 Gly Ser Ser Met Ala Leu Ile Pro Asp Leu Ala Met Glu Thr Trp 1400 1405 1410 Leu Leu Leu Ala Val Ser Leu Val Leu Leu Tyr Leu Tyr Gly Thr 1415 1420 1425 His Ser His Gly Leu Phe Lys Lys Leu Gly Ile Pro Gly Pro Thr 1430 1435 1440 Pro Leu Pro Phe Leu Gly Asn Ile Leu Ser Tyr His Lys Gly Phe 1445 1450 1455 Cys Met Phe Asp Met Glu Cys His Lys Lys Tyr Gly Lys Val Trp 1460 1465 1470 Gly Phe Tyr Asp Gly Gln Gln Pro Val Leu Ala Ile Thr Asp Pro 1475 1480 1485 Asp Met Ile Lys Thr Val Leu Val Lys Glu Cys Tyr Ser Val Phe 1490 1495 1500 Thr Asn Arg Arg Pro Phe Gly Pro Val Gly Phe Met Lys Ser Ala 1505 1510 1515 Ile Ser Ile Ala Glu Asp Glu Glu Trp Lys Arg Leu Arg Ser Leu 1520 1525 1530 Leu Ser Pro Thr Phe Thr Ser Gly Lys Leu Lys Glu Met Val Pro 1535 1540 1545 Ile Ile Ala Gln Tyr Gly Asp Val Leu Val Arg Asn Leu Arg Arg 1550 1555 1560 Glu Ala Glu Thr Gly Lys Pro Val Thr Leu Lys Asp Val Phe Gly 1565 1570 1575 Ala Tyr Ser Met Asp Val Ile Thr Ser Thr Ser Phe Gly Val Asn 1580 1585 1590 Ile Asp Ser Leu Asn Asn Pro Gln Asp Pro Phe Val Glu Asn Thr 1595 1600 1605 Lys Lys Leu Leu Arg Phe Asp Phe Leu Asp Pro Phe Phe Leu Ser 1610 1615 1620 Ile Thr Val Phe Pro Phe Leu Ile Pro Ile Leu Glu Val Leu Asn 1625 1630 1635 Ile Cys Val Phe Pro Arg Glu Val Thr Asn Phe Leu Arg Lys Ser 1640 1645 1650 Val Lys Arg Met Lys Glu Ser Arg Leu Glu Asp Thr Gln Lys His 1655 1660 1665 Arg Val Asp Phe Leu Gln Leu Met Ile Asp Ser Gln Asn Ser Lys 1670 1675 1680 Glu Thr Glu Ser His Lys Ala Leu Ser Asp Leu Glu Leu Val Ala 1685 1690 1695 Gln Ser Ile Ile Phe Ile Phe Ala Gly Tyr Glu Thr Thr Ser Ser 1700 1705 1710 Val Leu Ser Phe Ile Met Tyr Glu Leu Ala Thr His Pro Asp Val 1715 1720 1725 Gln Gln Lys Leu Gln Glu Glu Ile Asp Ala Val Leu Pro Asn Lys 1730 1735 1740 Ala Pro Pro Thr Tyr Asp Thr Val Leu Gln Met Glu Tyr Leu Asp 1745 1750 1755 Met Val Val Asn Glu Thr Leu Arg Leu Phe Pro Ile Ala Met Arg 1760 1765 1770 Leu Glu Arg Val Cys Lys Lys Asp Val Glu Ile Asn Gly Met Phe 1775 1780 1785 Ile Pro Lys Gly Val Val Val Met Ile Pro Ser Tyr Ala Leu His 1790 1795 1800 Arg Asp Pro Lys Tyr Trp Thr Glu Pro Glu Lys Phe Leu Pro Glu 1805 1810 1815 Arg Phe Ser Lys Lys Asn Lys Asp Asn Ile Asp Pro Tyr Ile Tyr 1820 1825 1830 Thr Pro Phe Gly Ser Gly Pro Arg Asn Cys Ile Gly Met Arg Phe 1835 1840 1845 Ala Leu Met Asn Met Lys Leu Ala Leu Ile Arg Val Leu Gln Asn 1850 1855 1860 Phe Ser Phe Lys Pro Cys Lys Glu Thr Gln Ile Pro Leu Lys Leu 1865 1870 1875 Ser Leu Gly Gly Leu Leu Gln Pro Glu Lys Pro Val Val Leu Lys 1880 1885 1890 Val Glu Ser Arg Asp Gly Thr Val Ser Gly Ala 1895 1900 <210> 4 <211> 11201 <212> DNA <213> Artificial sequence <400> 4 gacggatcgg gagatctccc gatcccctat ggtgcactct cagtacaatc tgctctgatg 60 ccgcatagtt aagccagtat ctgctccctg cttgtgtgtt ggaggtcgct gagtagtgcg 120 cgagcaaaat ttaagctaca acaaggcaag gcttgaccga caattgcatg aagaatctgc 180 ttagggttag gcgttttgcg ctgcttcgg atgtacgggc cagatatacg cgttgacatt 240 gattattgac tagttattaa tagtaatcaa ttacggggtc attagttcat agcccatata 300 tggagttccg cgttacataa cttacggtaa atggcccgcc tggctgaccg cccaacgacc 360 cccgcccatt gacgtcaata atgacgtatg ttcccatagt aacgccaata gggactttcc 420 attgacgtca atgggtggag tatttacggt aaactgccca cttggcagta catcaagtgt 480 atcatatgcc aagtacgcccc cctattgacg tcaatgacgg taaatggccc gcctggcatt 540 atgcccagta catgacctta tgggactttc ctacttggca gtacatctac gtattagtca 600 tcgctattac catggtgatg cggttttggc agtacatcaa tgggcgtgga tagcggtttg 660 actcacgggg atttccaagt ctccacccca ttgacgtcaa tgggagtttg ttttggcacc 720 aaaatcaacg ggactttcca aaatgtcgta acaactccgc cccattgacg caaatgggcg 780 gtaggcgtgt acggtgggag gtctatataa gcagagctct ctggctaact agagaaccca 840 ctgcttactg gcttatcgaa attaatacga ctcactatag ggagaccacaa gctggctagc 900 gccatggact acaaagacca tgacggtgat tataaagatc atgacatcga ttacaaggat 960 gacgatgaca agatggcccc caagagaag aggaaggtgg gccgcggaat ggataagaaa 1020 tactcaatag gcttagctat cggcacaaat agcgtcggat gggcggtgat cactgatgaa 1080 tataaggttc cgtctaaaaa gttcaaggtt ctgggaaata cagaccgcca cagtatcaaa 1140 aaaaatctta taggggctct tttatttgac agtggagaga cagcggaagc gactcgtctc 1200 aaacggacag ctcgtagaag gtatacacgt cggaagaatc gtatttgtta tctacaggag 1260 atttttcaa atgagatggc gaaagtagat gatagttct ttcatcgact tgaagagtct 1320 ttttggtgg aagaagacaa gaagcatgaa cgtcatccta ttttggaaa tatagtagat 1380 gaagttgctt atcatgagaa atatccaact atctatcatc tgcgaaaaaa attggtagat 1440 tctactgata aagcggattt gcgcttaatc tatttggcct tagcgcatat gattaagtttt 1500 cgtggtcatt ttttgattga gggagattta aatcctgata atagtgatgt ggacaaacta 1560 tttatccagt tggtacaaac ctacaatcaa ttatttgaag aaaaccctat taacgcaagt 1620 ggagtagatg ctaaagcgat tctttctgca cgattgagta aatcaagacg attagaaaat 1680 ctcattgctc agctccccgg tgagaagaaa aatggcttat ttgggaatct cattgctttg 1740 tcattgggtt tgacccctaa ttttaaatca aatttgatt tggcagaaga tgctaaatta 1800 cagctttcaa aagatactta cgatgatgat ttagataatt tattggcgca aattggagat 1860 caatatgctg atttgtttt ggcagctaag aatttatcag atgctatttt actttcagat 1920 atcctaagag taaatactga aataactaag gctcccctat cagcttcaat gattaaacgc 1980 tacgatgaac atcatcaaga cttgactctt ttaaaagctt tagttcgaca acaacttcca 2040 gaaaagtata aagaaatctt ttttgatcaa tcaaaaaacg gatatgcagg ttatattgat 2100 gggggagcta gccaagaaga attttataaa tttatcaaac caattttaga aaaaatggat 2160 ggtactgagg aattatggt gaaactaaat cgtgaagatt tgctgcgcaa gcaacggacc 2220 tttgacaacg gctctattcc ccatcaaatt cacttgggtg agctgcatgc tattttgaga 2280 agacaagaag acttttatcc attttaaaaa gacaatcgtg agaagattga aaaaatcttg 2340 acttttcgaa ttccttatta tgttggtcca ttggcgcgtg gcaatagtcg ttttgcatgg 2400 atgactcgga agtctgaaga aacaattacc ccatggaatt ttgaagaagt tgtcgataaa 2460 ggtgcttcag ctcaatcatt tattgaacgc atgacaaact ttgataaaaa tcttccaaat 2520 gaaaaagtac taccaaaaca tagtttgctt tatgagtatt ttacggttta taacgaattg 2580 acaaaggtca aatatgttac tgaaggaatg cgaaaaccag catttctttc aggtgaacag 2640 aagaaagcca ttgttgattt actcttcaaa acaaatcgaa aagtaaccgt taagcaatta 2700 aaagaagatt atttcaaaaa aatagaatgt tttgatagtg ttgaaatttc aggagttgaa 2760 gatagattta atgcttcatt aggtacctac catgatttgc taaaaattat taaagataaa 2820 gattttttgg ataatgaaga aaatgaagat atcttagagg atattgtttt aacattgacc 2880 ttatttgaag atagggagat gattgaggaa agacttaaaa catatgctca cctctttgat 2940 gataaggtga tgaaacagct taaacgtcgc cgttatactg gttggggacg tttgtctcga 3000 aaattgatta atggtattag ggataagcaa tctggcaaaa caatattaga ttttttgaaa 3060 tcagatggtt ttgccaatcg caattttatg cagctgatcc atgatgatag tttgacattt 3120 aaagaagaca ttcaaaaagc acaagtgtct ggacaaggcg atagtttaca tgaacatatt 3180 gcaaatttag ctggtagccc tgctattaaa aaaggtattt tacagactgt aaaagttgtt 3240 gatgaattgg tcaaagtaat ggggcggcat aagccagaaa atatcgttat tgaaatggca 3300 cgtgaaaatc agacaactca aaagggccag aaaaattcgc gagagcgtat gaaacgaatc 3360 gaagaaggta tcaaagaatt aggaagtcag attcttaaag agcatcctgt tgaaaatact 3420 caattgcaaa atgaaaagct ctatctctat tatctccaaa atggaagaga catgtatgtg 3480 gaccaagaat tagatattaa tcgtttaagt gattatgatg tcgatgccat tgttccacaa 3540 agtttcctta aagacgattc aatagacaat aaggtcttaa cgcgttctga taaaaatcgt 3600 ggtaaatcgg ataacgttcc aagtgaagaa gtagtcaaaa agatgaaaaa ctattggaga 3660 caacttctaa acgccaagtt aatcactcaa cgtaagtttg ataatttaac gaaagctgaa 3720 cgtggaggtt tgagtgaact tgataaagct ggtttatca aacgccaatt ggttgaaact 3780 cgccaaatca ctaagcatgt ggcacaaatt ttggatagtc gcatgaatac taatacgat 3840 gaaaatgata aacttattcg agaggttaaa gtgattacct taaaatctaa attagttct 3900 gacttccgaa aagatttcca attctataaa gtacgtgaga ttaacaatta ccatcatgcc 3960 catgatgcgt atctaaatgc cgtcgttgga actgctttga ttaagaaata tccaaaactt 4020 gaatcggagt ttgtctatgg tgattataaa gtttatgatg ttcgtaaaat gattgctaag 4080 tctgagcaag aaataggcaa agcaaccgca aaatatttct tttactctaa tatcatgaac 4140 ttcttcaaaa cagaaattac acttgcaaat ggagagattc gcaaacgccc tctaatcgaa 4200 actaatgggg aaactggaga aattgtctgg gataaagggc gagattttgc cacagtgcgc 4260 aaagtattgt ccatgcccca agtcaatatt gtcaagaaaa cagaagtaca gacaggcgga 4320 ttctccaagg agtcaatttt accaaaaaga aattcggaca agcttattgc tcgtaaaaaa 4380 gactgggatc caaaaaaaata tggtggtttt gatagtccaa cggtagctta ttcagtccta 4440 gtggttgcta aggtggaaaa agggaaatcg aagaagttaa aatccgttaa agagttacta 4500 gggatcacaa ttatggaaag aagttccttt gaaaaaaaatc cgattgactt tttagaagct 4560 aaaggatata aggaagttaa aaaagactta atcattaaac tacctaaata tagtcttttt 4620 gagttagaaa acggtcgtaa acggatgctg gctagtgccg gagaattaca aaaaggaaat 4680 gagctggctc tgccaagcaa atatgtgaat tttttatatt tagctagtca ttgaaaag 4740 ttgaagggta gtccagaaga taacgaacaa aaacaattgt ttgtggagca gcataagcat 4800 tatttagatg agattattga gcaaatcagt gaattttcta agcgtgttat tttagcagat 4860 gccaatttag ataaagttct tagtgcatat aacaaacata gagacaaacc ataacgtgaa 4920 caagcagaaa atattattca tttatttacg ttgacgaatc ttggagctcc cgctgctttt 4980 aaatattttg atacaacaat tgatcgtaaa cgatatacgt ctacaaaaga agttttagat 5040 gccactctta tccatcaatc catcactggt ctttatgaaa cacgcattga tttgagtcag 5100 ctaggaggtg actctggagg atctagcgga ggatcctctg gcagcgagac accaggaaca 5160 agcgagtcag caacaccaga gagcagtggc ggcagcagcg gcggcagcag catggctctc 5220 atcccagact tggccatgga aacctggctt ctcctggctg tcagcctggt gctcctctat 5280 ctatatggaa cccattcaca tggacttttt aagaagcttg gaattccagg gcccacacct 5340 ctgcctttt tgggaaatat tttgtcctac cataagggct tttgtatgtt tgacatggaa 5400 tgtcataaaa agtatggaaa agtgtggggc ttttatgatg gtcaacagcc tgtgctggct 5460 atcacagatc ctgacatgat caaaacagtg ctagtgaaag aatgttattc tgtcttcaca 5520 aaccggaggc cttttggtcc agtgggattt atgaaaagtg ccatctctat agctgaggat 5580 gaagaatgga agagattacg atcattgctg tctccaacct tcaccagtgg aaaactcaag 5640 gagatggtcc ctatcattgc ccagtatgga gatgtgttgg tgagaaatct gaggcgggaa 5700 gcagagacag gcaagcctgt caccttgaaa gacgtctttg gggcctacag catggatgtg 5760 atcactagca catcatttgg agtgaacatc gactctctca acaatccaca agaccccttt 5820 gtggaaaaca ccaagaagct tttaagattt gattttttgg atccattctt tctctcaata 5880 acagtctttc cattcctcat cccaattctt gaagtattaa atatctgtgt gtttccaaga 5940 gaagttacaa attttttaag aaaatctgta aaaaggatga aagaaagtcg cctcgaagat 6000 acacaaaagc accgagtgga tttccttcag ctgatgattg actctcagaa ttcaaaagaa 6060 actgagtccc acaaagctct gtccgatctg gagctcgtgg cccaatcaat tatctttatt 6120 tttgctggct atgaaaccac gagcagtgtt ctctccttca tttgtatga actggccact 6180 caccctgatg tccagcagaa actgcaggag gaaattgatg cagttttacc caataaggca 6240 ccacccacct atgatactgt gctacagatg gagtatctttg acatggtggt gaatgaaacg 6300 ctcagattat tcccaattgc tatgagactt gagagggtct gcaaaaaaga tgttgagatc 6360 aatgggatgt tcattcccaa aggggtggtg gtgatgattc caagctatgc tcttcaccgt 6420 gacccaaagt actggacaga gcctgagaag ttcctccctg aaagattcag caagaagaac 6480 aaggacaaca tagatcctta catatacaca ccctttggaa gtggacccag aaactgcatt 6540 ggcatgaggt ttgctctcat gaacatgaaa cttgctctaa tcagagtcct tcagaacttc 6600 tccttcaaac cttgtaaaga aacacagatc cccctgaaat taagcttagg aggacttctt 6660 caaccagaaa aacccgttgt tctaaaggtt gagtcaaggg atggcaccgt aagtggagcc 6720 cccaagaaga agaggaaagt ctgaatcggt aggaattcgc ggccgtctag acttaagtttt 6780 aaaccgctga tcagcctcga ctgtgccttc tagttgccag ccatctgttg tttgcccctc 6840 ccccgtgcct tccttgaccc tggaaggtgc cactccact gtcctttcct aataaaatga 6900 ggaaattgca tcgcattgtc tgagtaggtg tcattctatt ctggggggtg gggtggggca 6960 ggacagcaag ggggaggatt gggaagacaa tagcaggcat gctgggggatg cggtgggctc 7020 tatggcttct gaggcggaaa gaaccagctg gggctctagg gggtatcccc acgcgccctg 7080 tagcggcgca ttaagcgcgg cgggtgtggt ggttacgcgc agcgtgaccg ctacacttgc 7140 cagcgcccta gcgcccgctc ctttcgcttt cttcccttcc tttctcgcca cgttcgccgg 7200 ctttccccgt caagctctaa atcggggct ccctttaggg ttccgattta gtgctttacg 7260 gcacctgac cccaaaaaac ttgattaggg tgatggttca cgtagtgggc catcgccctg 7320 atagacggtt tttcgccctt tgacgttgga gtccacgttc tttaatagtg gactcttgtt 7380 ccaaactgga acaacactca accctatctc ggtctattct tttgatttat aagggatttt 7440 gccgatttcg gcctattggt taaaaaatga gctgatttaa caaaaattta acgcgaatta 7500 attctgtgga atgtgtgtca gttagggtgt ggaaagtccc caggctcccc agcaggcaga 7560 agtatgcaaa gcatgcatct caattagtca gcaaccaggt gtggaaagtc cccaggctcc 7620 ccagcaggca gaagtatgca aagcatgcat ctcaattagt cagcaaccat agtcccgccc 7680 ctaactccgc ccatcccgcc cctaactccg cccagttccg cccattctcc gccccatggc 7740 tgactaattt tttttattta tgcagaggcc gaggccgcct ctgcctctga gctattccag 7800 aagtagtgag gaggcttttt tggaggccta ggcttttgca aaaagctccc gggagcttgt 7860 atatccattt tcggatctga tcaagagaca ggatgaggat cgtttcgcat gattgaacaa 7920 gatggattgc acgcaggttc tccggccgct tgggtggaga ggctattcgg ctatgactgg 7980 gcacaacaga caatcggctg ctctgatgcc gccgtgttcc ggctgtcagc gcaggggcgc 8040 ccggttcttt ttgtcaagac cgacctgtcc ggtgccctga atgaactgca ggacgaggca 8100 gcgcggctat cgtggctggc cacgacgggc gttccttgcg cagctgtgct cgacgttgtc 8160 actgaagcgg gaagggactg gctgctattg ggcgaagtgc cggggcagga tctcctgtca 8220 tctcaccttg ctcctgccga gaaagtatcc atcatggctg atgcaatgcg gcggctgcat 8280 acgcttgatc cggctacctg cccattcgac caccaagcga aacatcgcat cgagcgagca 8340 cgtactcgga tggaagccgg tcttgtcgat caggatgatc tggacgaaga gcatcagggg 8400 ctcgcgccag ccgaactgtt cgccaggctc aaggcgcgca tgcccgacgg cgaggatctc 8460 gtcgtgaccc atggcgatgc ctgcttgccg aatatcatgg tggaaaatgg ccgcttttct 8520 ggattcatcg actgtggccg gctgggtgtg gcggaccgct atcaggacat agcgttggct 8580 acccgtgata ttgctgaaga gcttggcggc gaatgggctg accgcttcct cgtgctttac 8640 ggtatcgccg ctcccgattc gcagcgcatc gccttctatc gccttcttga cgagttcttc 8700 tgagcgggac tctggggttc gaaatgaccg accaagcgac gcccaacctg ccatcacgag 8760 atttcgattc caccgccgcc ttctatgaaa ggttgggctt cggaatcgtt ttccgggacg 8820 ccggctggat gatcctccag cgcggggatc tcatgctgga gttcttcgcc caccccaact 8880 tgtttattgc agcttataat ggttacaaat aaagcaatag catcacaaat ttcacaaata 8940 aagcattttt ttcactgcat tctagttgtg gtttgtccaa actcatcaat gtatcttatc 9000 atgtctgtat accgtcgacc tctagctaga gcttggcgta atcatggtca tagctgtttc 9060 ctgtgtgaaa ttgttatccg ctcacaattc cacacaacat acgagccgga agcataaagt 9120 gtaaagcctg gggtgcctaa tgagtgagct aactcacatt aattgcgttg cgctcactgc 9180 ccgctttcca gtcgggaaac ctgtcgtgcc agctgcatta atgaatcggc caacgcgcgg 9240 ggagaggcgg tttgcgtatt gggcgctctt ccgcttcctc gctcactgac tcgctgcgct 9300 cggtcgttcg gctgcggcga gcggtatcag ctcactcaaa ggcggtaata cggttatcca 9360 cagaatcagg ggataacgca ggaaagaaca tgtgagcaaa aggccagcaa aaggccagga 9420 accgtaaaaa ggccgcgttg ctggcgtttt tccataggct ccgcccccct gacgagcatc 9480 acaaaaatcg acgctcaagt cagaggtggc gaaacccgac aggactataa agataccagg 9540 cgtttccccc tggaagctcc ctcgtgcgct ctcctgttcc gaccctgccg cttaccggat 9600 acctgtccgc ctttctccct tcgggaagcg tggcgctttc tcatagctca cgctgtaggt 9660 atctcagttc ggtgtaggtc gttcgctcca agctgggctg tgtgcacgaa ccccccgttc 9720 agcccgaccg ctgcgcctta tccggtaact atcgtcttga gtccaacccg gtaagacacg 9780 acttatcgcc actggcagca gccactggta acaggattag cagagcgagg tatgtaggcg 9840 gtgctacaga gttcttgaag tggtggccta actacggcta cactagaaga acagtatttg 9900 gtatctgcgc tctgctgaag ccagttacct tcggaaaaag agttggtagc tcttgatccg 9960 gcaaacaaac caccgctggt agcggttttt ttgtttgcaa gcagcagatt acgcgcagaa 10020 aaaaaggatc tcaagaagat cctttgatct tttctacggg gtctgacgct cagtggaacg 10080 aaaactcacg ttaagggatt ttggtcatga gattatcaaa aaggatcttc acctagatcc 10140 ttttaaatta aaaatgaagt tttaaatcaa tctaaagtat atatgagtaa acttggtctg 10200 acagttacca atgcttaatc agtgaggcac ctatctcagc gatctgtcta tttcgttcat 10260 ccatagttgc ctgactcccc gtcgtgtaga taactacgat acgggagggc ttaccatctg 10320 gccccagtgc tgcaatgata ccgcgagacc cacgctcacc ggctccagat ttatcagcaa 10380 taaaccagcc agccggaagg gccgagcgca gaagtggtcc tgcaacttta tccgcctcca 10440 tccagtctat taattgttgc cgggaagcta gagtaagtag ttcgccagtt aatagtttgc 10500 gcaacgttgt tgccattgct acaggcatcg tggtgtcacg ctcgtcgttt ggtatggctt 10560 cattcagctc cggttcccaa cgatcaaggc gagttacatg atcccccatg ttgtgcaaaa 10620 aagcggttag ctccttcggt cctccgatcg ttgtcagaag taagttggcc gcagtgttat 10680 cactcatggt tatggcagca ctgcataatt ctcttactgt catgccatcc gtaagatgct 10740 tttctgtgac tggtgagtac tcaaccaagt cattctgaga atagtgtatg cggcgaccga 10800 gttgctcttg cccggcgtca atacgggata ataccgcgcc acatagcaga actttaaaag 10860 tgctcatcat tggaaaacgt tcttcggggc gaaaactctc aaggatctta ccgctgttga 10920 gatccagttc gatgtaaccc actcgtgcac ccaactgatc ttcagcatct tttactttca 10980 ccagcgtttc tgggtgagca aaaacaggaa ggcaaaatgc cgcaaaaaag ggaataaggg 11040 cgacacggaa atgttgaata ctcatactct tcctttttca atattattga agcatttatc 11100 agggttattg tctcatgagc ggatacatat ttgaatgtat ttagaaaaat aaacaaatag 11160 gggttccgcg cacatttccc cgaaaagtgc cacctgacgt c 11201
Claims
1. A site-directed mutagenesis system for DNA, characterized in that, The DNA site-directed mutagenesis system comprises: a recombinant vector containing a nucleic acid molecule encoding the fusion protein PdCas9-p450, an sgRNA expression vector, and aflatoxin B1; the amino acid sequence of the fusion protein PdCas9-p450 is shown in SEQ ID No.
3.
2. The DNA site-directed mutagenesis system according to claim 1, characterized in that, The nucleotide sequence of the nucleic acid molecule encoding the fusion protein PdCas9-p450 is shown in positions 1009-6720 of SEQ ID No.
4.
3. The DNA site-directed mutagenesis system according to claim 1 or 2, characterized in that, The nucleotide sequence of the recombinant vector is SEQ ID No.
4.
4. A method for performing site-directed mutagenesis using any one of the DNA site-directed mutagenesis systems according to claims 1-3, characterized in that, The method includes the following steps: transfecting host cells with any of the recombinant vectors described in claims 1-3 and a target site-specific sgRNA expression vector, and inducing nucleotide mutations at the target site with aflatoxin B1.
5. The method according to claim 4, characterized in that, The method includes the following steps: (1) Construct a recombinant vector with the nucleotide sequence SEQ ID No. 4; (2) Construct sgRNA expression vectors based on DNA target sites; (3) The recombinant vector constructed in step (1) and the sgRNA expression vector constructed in step (2) are co-transfected into host cells; (4) Aflatoxin B1 (AFB1) induces random mutations in nucleotides at the target site.
6. The application of the DNA site-directed mutagenesis system according to any one of claims 1-3 in the mutation of DNA at target sites.
Citation Information
Patent Citations
A:t to c:g base editors and uses thereof
WO2020181180A1