A high-precision adenine base editor and its use
By fusing the e18 protein with the ABEs base editing system, the problems of large and off-target effects of the ABEs editing window are solved, and higher accuracy and lower cost gene editing effects are achieved.
Patent Information
- Application Number
- CN202210538473.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-17
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2042-05-17
AI Technical Summary
The existing adenine base editors (ABEs) have large editing windows and obvious off-target effects, high cost of synthesis of RNPs and difficult to preserve, which limits their application prospects.
A high-precision ABEs base editor is constructed to fuse e18 protein with ABEs base editing system, and use e18 to accelerate the degradation of nCas9-TadA, reduce half-life, reduce off-target effects, and simplify operations and reduce costs.
It achieves higher editing accuracy and lower off-target effects, reduces the need for RNP synthesis, improves operation ease and storage ease.
Smart Images

Figure CN115093482B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of biotechnology, and specifically relates to a high-precision adenine base editor and its use. Background Art
[0002] A single-base editor is a gene editing system composed of a fusion of the Cas9 protein and deaminases (primarily adenine deaminase and cytosine deaminase). This system can precisely and irreversibly alter one base pair to another without introducing double-strand breaks (DSBs) or exogenous repair templates. Currently, single-base editors primarily include adenine base editors (ABEs), cytosine base editors (CBEs), guanine base editors (GCBEs), and DdCBEs and TALEDs, which can precisely edit the mitochondrial genome.
[0003] Adenine base editors (ABEs) are fusion proteins composed of nCas9 (D10A) and a modified adenine deaminase. Guided by a sgRNA, these proteins specifically recognize and bind to target sequences, deaminating adenine (A) to form creatinine (I), which is then converted to G, achieving an A-to-G conversion. Due to their high efficiency and ease of use, various versions of ABEs (primarily including ABE7.10, ABEmax, and ABE8e) have been widely adopted, significantly promoting advancements in life sciences fields such as gene therapy, human disease modeling, biopharmaceuticals, and disease pathogenesis. However, as research continues, researchers have discovered that ABEs and CBEs suffer from issues such as large editing windows and genome-wide off-target effects (both Cas9-dependent and Cas9-independent).
[0004] Therefore, developing base editors with higher precision is of great significance for the advancement of life sciences. Studies have shown that DNA base editing ribonucleoprotein complexes (RNPs) composed of sgRNA and ABE fusion proteins can be rapidly degraded by proteases in cells, significantly improving the purity of the editing product and reducing off-target effects. However, in practical applications, the high cost of RNP synthesis and the difficulty of preservation significantly limit its application prospects. The ubiquitin-proteasome system (UPS), as the primary pathway for protein degradation, is the most important mechanism for controlling protein levels. Ubiquitination involves three main steps: activation, conjugation, and ligation. The protein degradation process primarily involves the activation of ubiquitin by an E1 ubiquitin activating enzyme, followed by the conjugation of ubiquitin transferred by the E1 ubiquitin activating enzyme by an E2 ubiquitin conjugating enzyme. Finally, the E3 ligase selectively attaches ubiquitin to lysine, serine, threonine, or cysteine residues on the target protein. The E3 ligase directly binds to the substrate and determines the specificity of the ubiquitin system. Rad18 is a RING-type E3 ubiquitin ligase that plays a crucial role in DNA damage repair. The e18 protein is a Rad18 variant protein with the SAP domain removed. Summary of the Invention
[0005] The purpose of the present invention is to construct a highly precise ABEs base editor by fusing the e18 protein with the ABEs base editing system. The ABEs base editor utilizes the e18 fused to it to accelerate the degradation of its own nCas9-TadA, reducing its half-life and thus increasing its accuracy. Compared with the DNA base editing RNP complex, it does not require in vitro synthesis of RNP, and has the advantages of being cheaper, simpler to operate, and easier to store.
[0006] The purpose of the present invention is achieved through the following technical solutions:
[0007] A fusion protein comprising a first region and a second region from the N-terminus to the C-terminus, wherein the first region comprises ABEs, wherein the ABEs comprise adenine deaminase or an enzymatically active component thereof and nCas9; the second region comprises e18 protein; and the fusion protein optionally comprises one or more linker amino acid sequences located in the first region of the fusion protein and between the first region and the second region.
[0008] As a preferred technical solution of the present invention: the ABEs is one of the base editors ABE7.10, ABEmax and ABE8e.
[0009] As a preferred technical solution of the present invention: the fusion protein further includes a nuclear localization signal segment.
[0010] As a preferred technical solution of the present invention: when the ABEs is ABEmax, the amino acid sequence of the fusion protein is shown as SEQ ID NO. 40; when the ABEs is ABE8e, the amino acid sequence of the fusion protein is shown as SEQ ID NO. 41.
[0011] Another object of the present invention is to provide a polynucleotide encoding the fusion protein described above. The polynucleotide sequence comprises a first region and a second region, wherein the first region encodes ABEs and the second region encodes e18 protein. The polynucleotide may optionally comprise one or more linker amino acid sequences encoding the amino acid sequence of the first region and between the first and second regions of the fusion protein.
[0012] As a preferred technical solution of the present invention: when the ABEs is ABEmax, the polynucleotide sequence is constructed as follows: the ABEmax gene sequence before gene modification is shown in SEQ ID NO. 2; the upstream primer sequence F1 that can effectively amplify the target e18 gene fragment 1 is shown in SEQ ID NO. 3, and the downstream primer sequence R1 is shown in SEQ ID NO. 4; the upstream primer sequence F2 that can effectively amplify the target e18 gene fragment 2 is shown in SEQ ID NO. 5, and the downstream primer sequence R2 is shown in SEQ ID NO. 6; the upstream primer sequence F3 that can effectively amplify the e18 fragment gene sequence with the restriction enzyme cleavage site is shown in SEQ ID NO.
[0013] 7, the downstream primer sequence R3 is shown in SEQ ID NO. 8, and the e18 gene sequence with double enzyme cutting sites is shown in SEQ ID NO. 9; the fragment after enzyme cutting SEQ ID NO. 9 is separated from the fragment after enzyme cutting SEQ ID NO.
[0014] The sequence obtained by connecting the fragments after 2 is the polynucleotide sequence shown in SEQ ID NO. 1.
[0015] As a preferred technical solution of the present invention: when the ABEs is ABE8e, the polynucleotide sequence shown in SEQ ID NO. 10 is constructed as follows; the ABE8e gene sequence before gene modification is shown in SEQ ID NO. 11; the e18 gene sequence with double enzyme cleavage sites is shown in SEQ ID NO. 9, and the fragment SEQ ID NO. 9 is connected with the fragment after enzyme cleavage of SEQ ID NO. 11 to obtain the polynucleotide sequence shown in SEQ ID NO. 10.
[0016] Another object of the present invention is to provide a construct comprising the polynucleotide. The construct can be constructed by inserting the polynucleotide into a suitable expression vector. The expression vector may include, but is not limited to, a pCMV expression vector, a pSV2 expression vector, or the like.
[0017] Another object of the present invention is to provide an expression system comprising the construct or the exogenous polynucleotide integrated into the genome. The expression system can be a host cell that can express the fusion protein described above, which can be combined with an sgRNA to localize the fusion protein to the target region, thereby achieving base editing in the target region.
[0018] Another object of the present invention is to provide a use, specifically the use of the fusion protein, the polynucleotide, the construct or the expression system in gene editing, wherein the gene editing is to convert base A to G.
[0019] Another object of the present invention is to provide a base editing system, comprising the fusion protein and sgRNA, wherein the fusion protein cooperates with the sgRNA to locate the fusion protein to the target area.
[0020] Another object of the present invention is to provide a gene editing method, comprising using the fusion protein or the base editing system to perform gene editing, wherein the gene editing is to convert base A to G.
[0021] The beneficial effects are as follows:
[0022] The highly precise adenine base editor described in this invention fuses the exogenous protein e18 to an existing adenine base editor. The gene sequences constituting the plasmids are shown in SEQ ID NO. 1 and SEQ ID NO. 10, namely ABEmax-e18 and ABE8e-e18. The editing windows are both highly precise, at positions 5-7 and 1-9, respectively. PCSK9-edited positive clones significantly increased LDL uptake, demonstrating its potential for gene therapy. This provides more possible directions for improvement in the subsequent development of precise base editors while increasing the number of editor tools.
[0023] The present invention constructs an ABEs expression plasmid that can shorten the half-life of ABEs protein and accelerate its degradation in cells, thereby increasing the accuracy of ABEs and reducing its off-target effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1a is a schematic diagram of the ABEmax-e18 plasmid vector of the present invention; Figure 1b is a schematic diagram of the ABE8e-e18 plasmid vector of the present invention;
[0025] FIG2 is a sequencing statistical graph showing the sgRNA base editing efficiency of ABEmax and ABEmax-e18 of the present invention at multiple endogenous sites;
[0026] FIG3 is a sequencing diagram of a positive PCSK9 editing clone of the present invention;
[0027] FIG4 is a comparison of LDL uptake by wild-type cells and positive PCSK9 edited clone cells of the present invention; Figure 5 This is a sequencing statistical graph of the sgRNA base editing efficiency of ABE8e and ABE8e-e18 of the present invention at multiple endogenous sites. DETAILED DESCRIPTION
[0028] In a first aspect, the present invention provides a fusion protein comprising, from the N-terminus to the C-terminus, a first region and a second region, wherein the first region comprises ABEs, which comprise adenine deaminase or an enzymatically active component thereof and nCas9; the second region comprises e18 protein; and the fusion protein optionally comprises one or more linker amino acid sequences located within the first region and between the first and second regions of the fusion protein. The fusion protein performs base editing at a targeted site under the guidance of sgRNA. The ABEs fragment is an existing ABEs base editor, which utilizes the e18 sequence fused to it to accelerate the degradation of its own nCas9-TadA, reducing its half-life and thereby increasing its accuracy.
[0029] In the fusion protein provided by the present invention, the ABEs is one of the base editors ABE7.10, ABEmax and ABE8e.
[0030] In the fusion protein provided by the present invention, the amino acid sequence of the fusion protein is as shown in SEQ ID NO. 40 or SEQ ID NO. 41; or an amino acid sequence having a sequence similarity of more than 80% with SEQ ID NO. 40 or SEQ ID NO. 41, and having the function of the amino acid sequence defined by SEQ ID NO. 40 or SEQ ID NO. 41. Specifically, the amino acid sequence having 80% or more sequence similarity to SEQ ID NO. 40 or SEQ ID NO. 41 specifically refers to: an amino acid sequence as shown in SEQ ID NO. 40 or SEQ ID NO. 41 obtained by substitution, deletion or addition of one or more (specifically 1-50, 1-30, 1-20, 1-10, 1-5, 1-3, 1, 2, or 3) amino acids, or a polypeptide fragment obtained by adding one or more (specifically 1-50, 1-30, 1-20, 1-10, 1-5, 1-3, 1, 2, or 3) amino acids to the N-terminus and / or C-terminus, and having the function of the polypeptide fragment as shown in SEQ ID NO. 40 or SEQ ID NO. 41. The sequence similarity generally refers to the percentage of identical amino acid residues in the compared sequences. The similarity of two or more sequences can be calculated using calculation software known in the art, for example, software from NCBI.
[0031] In the fusion proteins provided herein, the substitutions, deletions, or additions may be conservative amino acid substitutions. "Conservative amino acid substitutions" specifically refer to situations where an amino acid residue is replaced by another amino acid residue having a similar side chain. Families of amino acid residues having similar side chains are known to those skilled in the art.
[0032] In the fusion protein provided by the present invention, the fusion protein may further include a nuclear localization signal segment. The nuclear localization signal segment can usually interact with a nuclear import vector, thereby enabling the protein to be transported into the cell nucleus.
[0033] The second aspect of the present invention provides a polynucleotide encoding the above-mentioned fusion protein.
[0034] The polynucleotide sequence provided by the present invention is shown as SEQ ID NO. 1 or SEQ ID NO. 10. Among the polynucleotides provided by the present invention, the polynucleotide sequence shown in SEQ ID NO. 1 is constructed as follows: the ABEmax gene sequence before gene modification is shown in SEQ ID NO. 2; the upstream primer sequence F1 capable of effectively amplifying the target e18 gene fragment 1 is shown in SEQ ID NO. 3, and the downstream primer sequence R1 is shown in SEQ ID NO. 4; the upstream primer sequence F2 capable of effectively amplifying the target e18 gene fragment 2 is shown in SEQ ID NO. 5, and the downstream primer sequence R2 is shown in SEQ ID NO. 6; the upstream primer sequence F3 capable of effectively amplifying the e18 fragment gene sequence with restriction enzyme cleavage sites is shown in SEQ ID NO. 7, and the downstream primer sequence R3 is shown in SEQ ID NO. 8; the e18 fragment with double restriction enzyme cleavage sites is shown in SEQ ID NO. 9; and the sequence obtained by ligating the fragment after restriction enzyme cleavage of SEQ ID NO. 9 with the fragment after restriction enzyme cleavage of SEQ ID NO. 2 is the polynucleotide sequence shown in SEQ ID NO. 1.
[0035] Among the polynucleotides provided by the present invention, the polynucleotide sequence shown in SEQ ID NO. 10 is constructed as follows; the ABE8e gene sequence before genetic modification is shown in SEQ ID NO. 11; and the sequence obtained by ligating the fragment SEQ ID NO. 9 with the fragment after enzyme digestion of SEQ ID NO. 11 is shown in SEQ ID NO. 10.
[0036] A third aspect of the present invention is to provide a construct comprising the polynucleotide. The construct can be constructed by inserting the polynucleotide into a suitable expression vector. A person skilled in the art can select a suitable expression vector, for example, the expression vector can include but is not limited to a pCMV expression vector, a pSV2 expression vector, or the like.
[0037] A fourth aspect of the present invention is to provide an expression system, wherein the expression system contains the construct or the above-mentioned polynucleotide integrated with an exogenous source in the genome. The expression system can be a host cell, and the host cell can express the fusion protein as described above, and the fusion protein can be coordinated with the sgRNA, so that the fusion protein can be located in the target region to achieve base editing of the target region. In another specific embodiment of the present invention, the host cell can be a eukaryotic cell and / or a prokaryotic cell, more specifically a mouse cell, a human cell, etc., more specifically a mouse brain neuroma cell, a human embryonic kidney cell, a human cervical cancer cell, a human colon cancer cell, a human osteosarcoma cell, etc.
[0038] A fifth aspect of the present invention provides a use of the fusion protein, polynucleotide, construct, or expression system in gene editing. Preferably, the use is in gene editing of eukaryotic organisms, specifically metazoans, including but not limited to humans and mice. Specifically, the use may include, but is not limited to, base editing from A to G. These base editing techniques can be applied to edit splice acceptor / donor sites to regulate RNA splicing, construct models (e.g., disease models, cell models, animal models, etc.), or treat human diseases. In a specific embodiment of the present invention, the edited object may be an embryo, a cell, or the like.
[0039] A sixth aspect of the present invention provides a base editing system comprising the fusion protein and sgRNA. Those skilled in the art can select an appropriate sgRNA targeting a specific site based on the target gene region. For example, the sgRNA sequence is typically at least partially complementary to the target region, thereby cooperating with the fusion protein to localize the fusion protein to the target region and achieve base editing within the target region, wherein the gene editing is to convert the base A to a G.
[0040] A seventh aspect of the present invention provides a gene editing method, comprising using the fusion protein or base editing system to perform gene editing, wherein the gene editing converts the base A to G. For example, the gene editing method may include culturing the expression system provided in the fourth aspect of the present invention under appropriate conditions to express the fusion protein. The fusion protein can perform base editing in the target region in the presence of a complementary sgRNA targeting the target region. Methods for providing the conditions for the presence of the sgRNA should be known to those skilled in the art. For example, the method may include culturing an expression system capable of expressing the sgRNA under appropriate conditions. The expression system may be a host cell containing an expression vector containing a polynucleotide encoding the sgRNA, or a host cell having the polynucleotide encoding the sgRNA integrated into its chromosome. In one embodiment of the present invention, the sgRNA and the fusion protein can be expressed in the same host cell, which may be a target cell. In another embodiment of the present invention, the gene editing is in vitro gene editing.
[0041] The following describes the embodiments of the present invention through specific examples. Those skilled in the art will readily understand the other advantages and benefits of the present invention from the disclosure herein. The present invention may also be implemented or applied through various other specific embodiments, and the details in this specification may be modified or altered based on different viewpoints and applications without departing from the spirit of the present invention.
[0042] Before further describing the specific embodiments of the present invention, it should be understood that the scope of protection of the present invention is not limited to the specific specific embodiments described below; it should also be understood that the terms used in the examples of the present invention are for describing specific specific embodiments rather than for limiting the scope of protection of the present invention; in the present specification and claims, unless otherwise expressly stated herein, the singular forms "a", "an" and "the" include plural forms.
[0043] When the embodiments provide numerical ranges, it should be understood that, unless otherwise specified in the present invention, both endpoints of each numerical range and any numerical value between the two endpoints may be selected. Unless otherwise defined, all technical and scientific terms used in the present invention have the same meaning as those generally understood by those skilled in the art. In addition to the specific methods, equipment, and materials used in the embodiments, according to the understanding of the prior art by those skilled in the art and the description of the present invention, any methods, equipment, and materials of the prior art similar or equivalent to the methods, equipment, and materials described in the embodiments of the present invention can also be used to implement the present invention.
[0044] Unless otherwise specified, the experimental methods, detection methods, and preparation methods disclosed in the present invention all adopt conventional techniques in molecular biology, biochemistry, chromatin structure and analysis, analytical chemistry, cell culture, recombinant DNA technology, and related fields in the art.
[0045] Example 1 Construction of ABEmax-e18 and ABE8e-e18 Plasmids
[0046] Primers targeting the e18 sequence were designed and synthesized. PCR amplification was performed in vitro to generate an e18 fragment with restriction enzyme cleavage sites. This fragment was then digested and ligated with the similarly digested ABEmax and ABE8e fragments, respectively, to create the ABEmax-e18 and ABE8e-e18 plasmids. The main components of this vector are, in order: adenine deaminase, nCas9, e18, and a nuclear localization signal. This editor plasmid, when used in conjunction with an sgRNA plasmid, can reveal its editing activity and other characteristic features. After verification by sequencing, the constructed plasmid was extracted and ethanol precipitated. The purified expression vector was then concentrated to a desired concentration. The complete plasmid map is shown in Figure 1.
[0047] The present invention constructs an ABEs expression plasmid that can shorten the half-life of ABEs protein and accelerate its degradation in cells, thereby increasing the accuracy of ABEs and reducing its off-target effects. The fusion protein comprises deaminase, nCas9, and e18 fragments in sequence.
[0048] Example 2 Design of sgRNA sequences and plasmid construction.
[0049] sgRNA sequences suitable for use in human cells were designed and synthesized. The designed sgRNA sequences were synthesized; the single-stranded sgRNA DNA sequences were annealed to form multiple oligonucleotide chains targeting different sites; these oligonucleotides were then ligated into the sgRNA backbone plasmid vector. The sequences of these multiple sgRNAs are:
[0050] SgRNA-1 sequence: 5-GAATACTAAGCATAGACTCC -3 SgRNA-2 sequence: 5-GTAAACAAAGCATAGACTGA -3 SgRNA-3 sequence: 5-GAACACAAAGCATAGACTGC -3 SgRNA-4 sequence: 5-GATGAGATAATGATGAGTCA -3 SgRNA-5 sequence: 5-GACAAACCAGAAGCCGCTCC -3 SgRNA-6 sequence: 5-GGGAATAAATCATAGAATCC -3 SgRNA-7 sequence: 5-GGAACACAAAGCATAGACTG -3 SgRNA-8 sequence: 5-GCACCTACCTCGGGAGCTGA -3 SgRNA-9 sequence: 5-GGAATCCCTTCTGCAGCACC -3 SgRNA-10 sequence: 5-TCAGAAAGTGGTGGCTGGTG -3 SgRNA-11 sequence: 5-GGCCCAGACTGAGCACGTGA -3SgRNA-12 sequence: 5-ATATTTTGCATTGAGATAGTG -3SgRNA-13 sequence: 5-GTCATCTTAGTCATTACCTG -3SgRNA-14 sequence: 5-GAAGATAGGAATAGACTGC -3.
[0051] After sequencing and verifying the constructed sgRNA expression vector, the target plasmid was extracted and ethanol precipitation was performed, and the purified sgRNA expression vector was purified to a certain concentration.
[0052] Example 3 Co-transfection of ABEs plasmid and sgRNA expression vector.
[0053] and co-transfection of ABEmax-e18 plasmid and sgRNA expression vector.
[0054] HEK293T cells were plated and, when the cells reached a density of approximately 80%, ABEmax and ABEmax-e18 and sgRNA expression vectors were introduced into the cells by lipofectamine transfection. 72 hours after transfection, the genome of each group of cells was extracted, and PCR reactions were performed using specific primers for detecting mutation efficiency. The obtained PCR products were sent for sequencing. The editing efficiency of the sgRNA site was evaluated by analyzing the sequencing peak graph and the editing window of the editor was edited. The results in Figure 2 show the sgRNA1 sequences corresponding to SEQ ID NO. 12 and SEQ ID NO. 13; the sgRNA2 sequences corresponding to SEQ ID NO. 14 and SEQ ID NO. 15; the sgRNA3 sequences corresponding to SEQ ID NO. 16 and SEQ ID NO. 17; the sgRNA4 sequences corresponding to SEQ ID NO. 18 and SEQ ID NO. 19; and the sgRNA1 sequences corresponding to SEQ ID NO. 20 and SEQ ID NO. 21.
[0055] The corresponding sgRNA5 sequence is SEQ ID NO. 22; the corresponding sgRNA6 sequence is SEQ ID NO. 24 and SEQ ID NO. 25; the corresponding sgRNA7 sequence is SEQ ID NO. 25. The sgRNA sequence can effectively guide the cas9 protein to edit the target site in the cell and obtain the characteristics of its editing window. The sites within the editable window of ABEmax range from 3 to 8, while the sites within the editable window of ABEmax-e18 range from 5 to 7.
[0056] 3-2 Co-transfection of ABE8e and ABE8e-e18 plasmids and sgRNA expression vector.
[0057] HEK293T cells were plated in the same manner. When the cells reached approximately 80% density, ABE8e and ABE8e-e18 and sgRNA expression vectors were introduced into the cells via lipofectamine transfection. 72 hours after transfection, the genomes of each group of cells were extracted, and PCR reactions were performed using specific primers for detecting mutation efficiency. The resulting PCR products were sequenced to evaluate the editing efficiency of the sgRNA site and define the editing window of the editor by analyzing the sequencing peaks. Figure 5 The results showed that the SEQ ID NO. 12 and SEQ ID NO.
[0058] 13 corresponding to the sgRNA1 sequence; SEQ ID NO. 14 and SEQ ID NO. 15 corresponding to the sgRNA2 sequence; SEQ ID NO. 16 and SEQ ID NO. 17 corresponding to the sgRNA3 sequence; SEQ ID NO. 18 and SEQ ID NO. 19 corresponding to the sgRNA4 sequence; SEQ ID NO. 22 and SEQ ID NO. 23 corresponding to the sgRNA6 sequence; SEQ ID NO. 28 and SEQ ID NO. 29 corresponding to the sgRNA9 sequence; SEQ ID NO. 30 and SEQ ID NO. 31 corresponding to the sgRNA10 sequence; SEQ ID NO. 32 and SEQ ID NO. 33 corresponding to the sgRNA11 sequence; SEQ ID NO. 34 and SEQ ID NO. 35 corresponding to the sgRNA12 sequence; SEQ ID NO. 36 and SEQ ID NO. 37 corresponding to the sgRNA13 sequence; SEQ ID NO. The sgRNA14 sequences corresponding to SEQ ID NO. 38 and SEQ ID NO. 39 effectively guide Cas9 protein editing at the target site in cells, characterizing its editing window. The editable window for ABE8e extends from positions 1 to 14, while the editable window for ABE8e-e18 extends from positions 1 to 9. See Figure 2 for the editable window site range and efficiency range.
[0059] 3-3 Co-transfection of ABEmax-e18 plasmid and PCSK9-sgRNA expression vector.
[0060] Resuscitate HepG2 cells and, when they are nearly confluent, wash them 2-3 times with PBS, discard the supernatant, add electroporation buffer, and then add the ABEmax-e18 plasmid and sgRNA expression plasmid to the cells and buffer in proportion. Gently mix with a pipette and gently aspirate the mixture into a pipette with a dedicated tip. Insert the pipette into the pipette holder with an electrode cup, add electroporation buffer to the electrode cup, set the program on the device, and press start. After the electroporation is completed, let it stand for 2 minutes, and then transfer the mixture in the electroporation gun to the cells.
[0061] The cell culture dish was then placed in a 37°C CO2 incubator. After 12 hours of culture, the medium was changed. The results in Figure 3 demonstrate that the sgRNA8 sequences corresponding to SEQ ID NO. 26 and SEQ ID NO. 27 obtained above can effectively guide Cas9 protein editing of the target site in HepG2 cells (see Figure 3).
[0062] Example 4. Preparation of PCSK9-edited positive clone HepG2 cells.
[0063] 72 hours after electroporation, HepG2 cells were trypsinized and plated onto 100mm culture dishes using limiting dilution. The culture medium was changed every 2-3 days. After 8-10 days, once cell clones had established, they were uniformly labeled under a fluorescence microscope and then transferred to 24-well culture plates for further culture. After 2-3 days, once the cells in the 24-well plates had reached confluence, half of the cells were lysed with NP40 lysis buffer and further verified by PCR sequencing for PCSK9 site-directed base editing events (see Figure 4).
[0064] Experimental Example 5: In vitro LDL uptake assay in HepG2 cells with positive PCSK9 site-directed base editing.
[0065] The HepG2 cells obtained above were cultured in DMEM supplemented with 5% FBS, 1% double-antibody, 1 mM sodium pyruvate, 1% glutamine, and 1% non-essential amino acids. The cells were seeded at a specific density in 6-well cell culture plates. When the cell density reached 70%, Dil-LDL was diluted 1:100 and incubated with the cells for 3 hours. The supernatant was then discarded, the cells were washed three times with PBS, and fixed in the culture plate with 4% paraformaldehyde for 2 hours. The samples were then washed three times with PBS for 5 minutes each. The cells were permeabilized with 0.5% triton X-100 in PBS for 10 minutes, washed three times with PBS for 5 minutes each, and stained with 0.5 μg / ml DAPI in PBS for 10 minutes. The cells were washed three times with PBS and observed under a fluorescence microscope. Compared to the fluorescence intensity of the unedited control cells, the site-directed edited cells were able to take up a greater amount of LDL (see Figure 5).
[0066] In summary, the present invention realizes the use of exogenous proteins to transform gene sequences such as SEQ ID NO.
[0067] The e18 (rad18 gene with the SAP domain removed) gene sequence shown in Figure 9 was fused to the commonly used ABEmax and the ABE8e adenine base editor with the highest editing activity, resulting in an editor with a more precise editing window and lower editing activity, providing a new direction for further improvement of the editor.
[0068] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the present invention. Anyone skilled in the art may modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by one of ordinary skill in the art without departing from the spirit and technical principles disclosed herein are intended to be covered by the claims of the present invention. Sequence Listing <110> Chongqing Jitang Biotechnology Research Institute Co., Ltd. <120> A high-precision adenine base editor and its use <160> 41 <170> SIPOSequenceListing 1.0 <210> 1 <211> 10224 <212> DNA <213> Artificial sequence <400> 1 atatgccaag tacgccccct attgacgtca atgacggtaa atggcccgcc tggcattatg 60 cccagtacat gaccttatgg gactttccta cttggcagta catctacgta ttagtcatcg 120 ctattaccat ggtgatgcgg ttttggcagt acatcaatgg gcgtggatag cggtttgact 180 cacggggatt tccaagtctc caccccattg acgtcaatgg gagtttgttt tggcaccaaa 240 atcaacggga ctttccaaaa tgtcgtaaca actccgcccc attgacgcaa atgggcggta 300 ggcgtgtacg gtgggaggtc tatataagca gagctggttt agtgaaccgt cagatccgct 360 agagatccgc ggccgctaat acgactcact ataggagag ccgccaccat gaaacggaca 420 gccgacggaa gcgagttcga gtcaccaaag aagaagcgga aagtctctga agtcgagttt 480 agccacgagt attggatgag gcacgcactg accctggcaa agcgagcatg ggatgaaaga 540 gaagtccccg tgggcgccgt gctggtgcac aacaatagag tgatcggaga gggatggaac 600 aggccaatcg gccgccacga ccctaccgca cacgcagaga tcatggcact gaggcaggga 660 ggcctggtca tgcagaatta ccgcctgatc gatgccaccc tgtatgtgac actggagcca 720 tgcgtgatgt gcgcaggagc aatgatccac agcaggatcg gaagagtggt gttcggagca 780 cgggacgcca agaccggcgc agcaggctcc ctgatggatg tgctgcacca ccccggcatg 840 aaccaccggg tggagatcac agagggaatc ctggcagacg agtgcgccgc cctgctgagc 900 gatttcttta gaatgcggag acaggagatc aaggcccaga agaaggcaca gagctccacc 960 gactctggag gatctagcgg aggatcctct ggaagcgaga caccaggcac aagcgagtcc 1020 gccacaccag agagctccgg cggctcctcc ggaggatcct ctgaggtgga gttttcccac 1080 gagtactgga tgagacatgc cctgaccctg gccaagaggg cacgcgatga gagggaggtg 1140 cctgtgggag ccgtgctggt gctgaacaat agagtgatcg gcgagggctg gaacagagcc 1200 atcggcctgc acgacccaac agcccatgcc gaaattatgg ccctgagaca gggcggcctg 1260 gtcatgcaga actacagact gattgacgcc accctgtacg tgacattcga gccttgcgtg 1320 atgtgcgccg gcgccatgat ccactctagg atcggccgcg tggtgtttgg cgtgaggaac 1380 gcaaaaaccg gcgccgcagg ctccctgatg gacgtgctgc actaccccgg catgaatcac 1440 cgcgtcgaaa ttaccgaggg aatcctggca gatgaatgtg ccgccctgct gtgctatttc 1500 tttcggatgc ctagacaggt gttcaatgct cagaagaagg cccagagctc caccgactcc 1560 ggaggatcta gcggaggctc ctctggctct gagacacctg gcacaagcga gagcgcaaca 1620 cctgaaagca gcgggggcag cagcgggggg tcagacaaga agtacagcat cggcctggcc 1680 atcggcacca actctgtggg ctgggccgtg atcaccgacg agtacaaggt gcccagcaag 1740 aaattcaagg tgctggggcaa caccgaccgg cacagcatca agaagaacct gatcggagcc 1800 ctgctgttcg acagcggcga aacagccgag gccaccggc tgaagagaac cgccagaaga 1860 1920 gccaaggtgg acgacagctt cttccacaga ctggaagagt ccttcctggt ggaagaggat 1980 aagaagcacg agcggcaccc catcttcggc aacatcgtgg acgaggtggc ctaccacgag 2040 aagtacccca ccatctacca cctgagaaag aaactggtgg acagcaccga caaggccgac 2100 ctgcggctga tctatctggc cctggcccac atgatcaagt tccggggcca cttcctgatc 2160 gagggcgacc tgaaccccga caacgcgac gtggacaagc tgttcatcca gctggtgcag 2220 acctacaacc agctgttcga ggaaaacccc atcaacgcca gcggcgtgga cgccaaggcc 2280 atcctgtctg ccagactgag caagagcaga cggctggaaa atctgatcgc ccagctgccc 2340 ggcgaagaagaatggcct gttcgggaaac ctgattgccc tgagcctggg cctgaccccc 2400 aacttcaaga gcaacttcga cctggccgag gatgccaaac tgcagctgag caaggacacc 2460 tacgacgacg acctggacaa cctgctggcc cagatcggcg accagtacgc cgacctgttt 2520 ctggccgcca agaacctgtc cgacgccatc ctgctgagcg acatcctgag agtgaacacc 2580 gagatcacca aggcccccct gagcgcctct atgatcaaga gatacgacga gcaccaccag 2640 gacctgaccc tgctgaaagc tctcgtgcgg cagcagctgc ctgagaagta caaagagatt 2700 ttcttcgacc agagcaagaa cggctacgcc ggctacattg acggcggagc cagccaggaa 2760 gagttctaca agttcatcaa gcccatcctg gaaaagatgg acggcaccga ggaactgctc 2820 gtgaagctga acagagagga cctgctgcgg aagcagcgga ccttcgacaa cggcagcatc 2880 ccccaccaga tccacctggg agagctgcac gccattctgc ggcggcagga agatttttac 2940 ccattcctga aggacaaccg ggaaaagatc gagaagatcc tgaccttccg catcccctac 3000 tacgtgggcc ctctggccag gggaaacagc agattcgcct ggatgaccag aaagagcgag 3060 gaaaccatca ccccctggaa cttcgaggaa gtggtggaca agggcgcttc cgcccagagc 3120 ttcatcgagc ggatgaccaa cttcgataag aacctgccca acgagaaggt gctgcccaag 3180 cacagcctgc tgtacgagta cttcaccgtg tataacgagc tgaccaaagt gaaatacgtg 3240 accgagggaa tgagaaagcc cgccttcctg agcggcgagc agaaaaaggc catcgtggac 3300 ctgctgttca agaccaaccg gaaagtgacc gtgaagcagc tgaaagagga ctacttcaag 3360 aaaatcgagt gcttcgactc cgtggaaatc tccggcgtgg aagatcggtt caacgcctcc 3420 ctgggcacat accacgatct gctgaaaatt atcaaggaca aggacttcct ggacaatgag 3480 gaaaacgagg acattctgga agatatcgtg ctgaccctga cactgtttga ggacagagag 3540 atgatcgagg aacggctgaa aacctatgcc cacctgttcg acgacaaagt gatgaagcag 3600 ctgaagcggc ggagatacac cggctggggc aggctgagcc ggaagctgat caacggcatc 3660 cgggacaagc agtccggcaa gacaatcctg gatttcctga agtccgacgg cttcgccaac 3720 agaaacttca tgcagctgat ccacgacgac agcctgacct ttaaagagga catccagaaa 3780 gcccaggtgt ccggccaggg cgatagcctg cacgagcaca ttgccaatct ggccggcagc 3840 cccgccatta agaagggcat cctgcagaca gtgaaggtgg tggacgagct cgtgaaagtg 3900 atgggccggc acaagcccga gaacatcgtg atcgaaatgg ccagagagaa ccagaccacc 3960 cagaagggac agagaacag ccgcgagaga atgaagcgga tcgaagaggg catcaaagag 4020 ctgggcagcc agatcctgaa agaacacccc gtggaaaaca cccagctgca gaacgagaag 4080 ctgtacctgt actacctgca gaatgggcgg gatatgtacg tggaccagga actggacatc 4140 aaccggctgt ccgactacga tgtggaccat atcgtgcctc agagctttct gaaggacgac 4200 tccatcgaca acaaggtgct gaccagaagc gacaagaacc ggggcaagag cgacaacgtg 4260 ccctccgaag aggtcgtgaa gaagatgaag aactactggc ggcagctgct gaacgccaag 4320 ctgattaccc agaaagtt cgacaatctg accaaggccg agaggcgg cctgagcgaa 4380 ctggataagg ccggcttcat caagacag ctggtggaaa cccggcagat cacaaagcac 4440 4500 cgggaagtga aagtgatcac cctgaagtcc aagctgggtgt ccgatttccg gaaggatttc 4560 cagttttaca aagtgcgcga gatcaac taccaccacg cccacgacgc ctacctgaac 4620 gccgtcgtgg gaaccgccct gatcaaaaag tacctaagc tggaaagcga gttcgtgtac 4680 4740 aaggctaccg ccaagtactt cttctacagc aacatcatga actttttcaa gaccgagatt 4800 accctggcca acggcgagat ccggaagcgg cctctgatcg agaaacgg cgaaaccggg 4860 gagatcgtgt gggataaggg ccgggatttt gccaccgtgc ggaaagtgct gagcatgccc 4920 gaagtgaata tcgtgaaaaa gaccgaggtg cagacaggcg gcttcagcaa agagtctatc 4980 cggcccaaga ggaacagcga tagctgatc gccagaaaga agactggga ccctaagaag 5040 tacggcggct tcgtgagccc caccgtggcc tattctgtgc tggtggtggc caaagtggaa 5100 aagggcaagt ccaagaaact gaaggtgtg aaagagctgc tgggatcac catcatggaa 5160 agaagcagct tcgagaagaa tcccatcgac tttctggaag ccaagggcta caaagaagtg 5220 aaaaaggacc tgatcatcaa gctgcctaag tactccctgt tcgagctgga aaacggccgg 5280 aagagaatgc tggcctctgc cagattcctg cagaagggaa acgaactggc cctgccctcc 5340 aaatatgtga acttcctgta cctggccagc cactatgaga agctgaaggg ctcccccgag 5400 gataatgagc agaaacagct gttgtgga cagcacaagc actacctgga cgagatcatc 5460 gagcagatca gcgagttctc caagagagtg atcctggccg acgctaatct ggacaagtg 5520 ctgtccgcct acaacaagca ccgggataag cccatcagag agcaggccga gatatcatc 5580 cacctgttta cctgaccaa tctgggagcc cctcgggcct tcaagtactt tgacaccacc 5640 atcgaccgga aggtgtaccg gagcaccaaa gaggtgctgg acgccaccct gatccaccag 5700 agcatcaccg gcctgtacga gaacggatc gacctgctc agctgggagg tgactctggc 5760 ggctcaaaaa gaaccgccga cggcagcgaa ttcagcacag ggagcatggg aatggactcc 5820 ctggccgagt ctcggtggcc tccggcctg gcagtcatga agacaataga tgattgctg 5880 cggtgtggaa ttgcttcga gtatttcaac attgcaatga taatacctca gtgttcacat 5940 aactactgct ctctctgtat aagaaaattt ctgtcctata aactcagtg tccaacttgc 6000 tgtgtgactg tcacagagcc ggatctgaaa aaaccgca tattagatga actggtaaaa 6060 agcttgaatt ttgcacggaa tcatctgctg cagtttgctt tegtcacc agccaaatct 6120 cctgctctt cctctcaa gatcttgct gtcaaagtat atactcctgt agcctccaga 6180 cagtctttaa agcagggag caggttaatg gataatttct tgatcagaga atgagtggt 6240 tctacatcag agttgttgat aaagaaaat aaagcaat tcagccctca aaagaggcg 6300 agccctgctg caagaccaa agagacacgt tctgtagaag agatcgctcc agatccctca 6360 gaggctaagc gtcctgagcc accctcgaca tccacttga aacaagttac taagtggat 6420 tgtcctgttt gcggggttaa cattccagaa agtcacatta ataagcatttt agacagctgt 6480 ttcacgcg aagagaa ggaagccctc agaagttctg ttcacaaag gaagccgcac 6540 atgtacaatg cccaatgcga tgctttgcat cctaaatcag ctgctgaaat agttcgagaa 6600 atcgaaaata tagagac taggagcgt cttgaagcta gtaaactca tgaaagtgta 6660 atggtttttta caaggacca aagaaaag gaatagatg aaatccacag taataatcgt 6720 aaaaaacata agagtgaatt tcagctctg gtggatcagg ctagaaaagg atacaagaaa 6780 attgctggaa tgtcacaaa aacagtaca atacaaag aagatgaatc tacagaaag 6840 ctatcttctg tatgcatggg acaggaagat atatgacct cagtaacaa ccactttct 6900 caatcaagc tggactcccc agaggaattg gaacctgaca gagagagga ttctcttagc 6960 tgtattgata tcagagt tctttcttca tcagaatcag attcatgcaa tagttccagt 7020 tcagacatca tagagatct tttagagaa gaggaagcct gggaagcatc actaaaaac 7080 gatcttcaag acacagaaat aagtccaag cagaatcgcc gcacaagagc cgctgaagt 7140 gctgagattg aaccaagaa caagcgtaat aggaatgaa aaagaccgc cgacggcagc 7200 gagttcgagc ccagagaa gaggaagtc caaccggtca tcatcaccat caccattgag 7260 tttaaacccg ctgatcagcc tcgactgtgc cttctagttg ccagccatct gttgtttgcc 7320 cctcccccgt gccttccttg accctggaag gtgccactcc cactgtcctt tcctataaa 7380 atgaggaaat tgcatcgcat tgtctgagta gtgtcattc tattctgggg ggtggggtgg 7440 ggcaggacag cagggggag gattgggag acatagcag gcatgctggg gatgcggtgg 7500 gctctatggc ttctgaggcg gaaagaacca gctggggctc gataccgtcg acctctagct 7560 agagcttggc gtaatcatgg tcatagctgt ttcctgtgtg aaattgttat ccgctcacaa 7620 ttccacacaa catacgagcc ggaagcataa agtgtaaagc ctagggtgcc taatgagtga 7680 gctaactcac attaattgcg ttgcgctcac tgcccgcttt ccagtcggga aacctgtcgt 7740 gccagctgca ttaatgaatc ggccaacgcg cggggagagg cggtttgcgt attgggcgct 7800 cttccgcttc ctcgctcact gactcgctgc gctcggtcgt tcggctgcgg cgagcggtat 7860 cagctcactc aaaggcggta atacggttat ccacagaatc aggggataac gcaggaaaga 7920 acatgtgagc aaaaggccag caaaaggcca ggaaccgtaa aaaggccgcg ttgctggcgt 7980 ttttccatag gctccgcccc cctgacgagc atcacaaaaa tcgacgctca agtcagaggt 8040 ggcgaaaccc gacaggacta taaagatacc aggcgtttcc ccctggaagc tccctcgtgc 8100 gctctcctgt tccgaccctg ccgcttaccg gatacctgtc cgcctttctc ccttcgggaa 8160 gcgtggcgct ttctcatagc tcacgctgta ggtatctcag ttcggtgtag gtcgttcgct 8220 ccaagctggg ctgtgtgcac gaaccccccg ttcagcccga ccgctgcgcc ttatccggta 8280 actatcgtct tgagtccaac ccggtaagac acgacttatc gccactggca gcagccactg 8340 gtaacaggat tagcagagcg aggtatgtag gcggtgctac agagttcttg aagtggtggc 8400 ctaactacgg ctacactaga agaacagtat ttggtatctg cgctctgctg aagccagtta 8460 ccttcggaaa aagagttggt agctcttgat ccggcaaaca aaccaccgct ggtagcggtg 8520 gtttttttgt ttgcaagcag cagattacgc gcagaaaaaa aggatctcaa gaagatcctt 8580 tgatcttttc tacggggtct gacactcagt ggaacgaaaa ctcacgttaa gggattttgg 8640 tcatgagatt atcaaaaagg atcttcacct agatcctttt aaattaaaaa tgaagtttta 8700 aatcaatcta aagtatatat gagtaaactt ggtctgacag ttaccaatgc ttaatcagtg 8760 aggcacctat ctcagcgatc tgtctatttc gttcatccat agttgcctga ctccccgtcg 8820 tgtagataac tacgatacgg gagggcttac catctggccc cagtgctgca atgataccgc 8880 gagacccacg ctcaccggct ccagatttat cagcaataaa ccagccagcc ggaagggccg 8940 agcgcagaag tggtcctgca actttatccg cctccatcca gtctattaat tgttgccggg 9000 aagctagagt aagtagttcg ccagttaata gtttgcgcaa cgttgttgcc attgctacag 9060 gcatcgtggt gtcacgctcg tcgtttggta tggcttcatt cagctccggt tcccaacgat 9120 caaggcgagt tacatgatcc cccatgttgt gcaaaaaagc ggttagctcc ttcggtcctc 9180 cgatcgttgt cagaagtaag ttggccgcag tgttatcact catggttatg gcagcactgc 9240 ataattctct tactgtcatg ccatccgtaa gatgcttttc tgtgactggt gagtactcaa 9300 ccaagtcatt ctgagaatag tgtatgcggc gaccgagttg ctcttgcccg gcgtcaatac 9360 gggataatac cgcgccacat agcagaactt taaaagtgct catcattgga aaacgttctt 9420 cggggcgaaa actctcaagg atcttaccgc tgttgagatc cagttcgatg taacccactc 9480 gtgcacccaa ctgatcttca gcatctttta ctttcaccag cgtttctggg tgagcaaaaa 9540 caggaaggca aaatgccgca aaaaagggaa taagggcgac acggaaatgt tgaatactca 9600 tactcttcct ttttcaatat tattgaagca tttatcaggg ttattgtctc atgagcggat 9660 acatatttga atgtatttag aaaaataaac aaataggggt tccgcgcaca tttccccgaa 9720 aagtgccacc tgacgtcgac ggatcgggag atcgatctcc cgatccccta gggtcgactc 9780 tcagtacaat ctgctctgat gccgcatagt taagccagta tctgctccct gcttgtgtgt 9840 tggaggtcgc tgagtagtgc gcgagcaaaa tttaagctac aacaaggcaa ggcttgaccg 9900 acaattgcat gaagaatctg cttagggtta ggcgttttgc gctgcttcgc gatgtacggg 9960 ccagatatac gcgttgacat tgattattga ctagttatta atagtaatca attacggggt 10020 cattagttca tagcccatat atggagttcc gcgttacata acttacggta aatggcccgc 10080 ctggctgacc gcccaacgac ccccgcccat tgacgtcaat aatgacgtat gttcccatag 10140 taacgccaat agggactttc cattgacgtc aatgggtgga gtatttacgg taaactgccc 10200 acttggcagt acatcaagtg tatc 10224 <210> 2 <211> 8811 <212> DNA <213> Artificial Sequence <400> 2 atatgccaag tacgccccct attgacgtca atgacggtaa atggcccgcc tggcattatg 60 cccagtacat gaccttatgg gactttccta cttggcagta catctacgta ttagtcatcg 120 ctattaccat ggtgatgcgg ttttggcagt acatcaatgg gcgtggatag cggtttgact 180 cacggggatt tccaagtctc caccccattg acgtcaatgg gagtttgttt tggcaccaaa 240 atcaacggga ctttccaaaa tgtcgtaaca actccgcccc attgacgcaa atgggcggta 300 ggcgtgtacg gtgggaggtc tatataagca gagctggttt agtgaaccgt cagatccgct 360 agagatccgc ggccgctaat acgactcact ataggagag ccgccaccat gaaacggaca 420 gccgacggaa gcgagttcga gtcaccaaag aagaagcgga aagtctctga agtcgagttt 480 agccacgagt attggatgag gcacgcactg accctggcaa agcgagcatg ggatgaaaga 540 gaagtccccg tgggcgccgt gctggtgcac aacaatagag tgatcggaga gggatggaac 600 aggccaatcg gccgccacga ccctaccgca cacgcagaga tcatggcact gaggcaggga 660 ggcctggtca tgcagaatta ccgcctgatc gatgccaccc tgtatgtgac actggagcca 720 tgcgtgatgt gcgcaggagc aatgatccac agcaggatcg gaagagtggt gttcggagca 780 cgggacgcca agaccggcgc agcaggctcc ctgatggatg tgctgcacca ccccggcatg 840 aaccaccggg tggagatcac agagggaatc ctggcagacg agtgcgccgc cctgctgagc 900 gatttcttta gaatgcggag acaggagatc aaggcccaga agaaggcaca gagctccacc 960 gactctggag gatctagcgg aggatcctct ggaagcgaga caccaggcac aagcgagtcc 1020 gccacaccag agagctccgg cggctcctcc ggaggatcct ctgaggtgga gttttcccac 1080 gagtactgga tgagacatgc cctgaccctg gccaagaggg cacgcgatga gagggaggtg 1140 cctgtgggag ccgtgctggt gctgaacaat agagtgatcg gcgagggctg gaacagagcc 1200 atcggcctgc acgacccaac agcccatgcc gaaattatgg ccctgagaca gggcggcctg 1260 gtcatgcaga actacagact gattgacgcc accctgtacg tgacattcga gccttgcgtg 1320 atgtgcgccg gcgccatgat ccactctagg atcggccgcg tggtgtttgg cgtgaggaac 1380 gcaaaaaccg gcgccgcagg ctccctgatg gacgtgctgc actaccccgg catgaatcac 1440 cgcgtcgaaa ttaccgaggg aatcctggca gatgaatgtg ccgccctgct gtgctatttc 1500 ttcggatgc ctagacaggt gttcaatgct cagaagaagg cccagagctc caccgactcc 1560 ggaggatcta gcggaggctc ctctggctct gagacacctg gcacaagcga gagcgcaaca 1620 cctgaaagca gcgggggcag cagcgggggg tcagacaaga agtacagcat cggcctggcc 1680 atcggcacca actctgtggg ctgggccgtg atcaccgacg agtacaaggt gcccagcaag 1740 aaattcaagg tgctgggcaa caccgaccgg cacagcatca agaagaacct gatcggagcc 1800 ctgctgttcg acagcggcga aacagccgag gccacccggc tgaagagaac cgccagaaga 1860 agatacacca gacggaagaa ccggatctgc tatctgcaag agatcttcag caacgagatg 1920 gccaaggtgg acgacagctt cttccacaga ctggaagagt ccttcctggt ggaagaggat 1980 aagaagcacg agcggcaccc catcttcggc aacatcgtgg acgaggtggc ctaccacgag 2040 aagtacccca ccatctacca cctgagaaag aaactggtgg acagcaccga caaggccgac 2100 ctgcggctga tctatctggc cctggcccac atgatcaagt tccggggcca cttcctgatc 2160 gagggcgacc tgaaccccga caacagcgac gtggacaagc tgttcatcca gctggtgcag 2220 acctacaacc agctgttcga ggaaaacccc atcaacgcca gcggcgtgga cgccaaggcc 2280 atcctgtctg ccagactgag caagagcaga cggctggaaa atctgatcgc ccagctgccc 2340 ggcgaagaagaatggcct gttcgggaaac ctgattgccc tgagcctggg cctgaccccc 2400 aacttcaaga gcaacttcga cctggccgag gatgccaaac tgcagctgag caaggacacc 2460 tacgacgacg acctggacaa cctgctggcc cagatcggcg accagtacgc cgacctgttt 2520 ctggccgcca agaacctgtc cgacgccatc ctgctgagcg acatcctgag agtgaacacc 2580 gagatcacca aggcccccct gagcgcctct atgatcaaga gatacgacga gcaccaccag 2640 gacctgaccc tgctgaaagc tctcgtgcgg cagcagctgc ctgagagta caaagagatt 2700 2760 gagttctaca agttcatcaa gcccatcctg gaaaagatgg acggcaccga ggaactgctc 2820 gtgaagctga agagagga cctgctgcgg aagcagcgga ccttgcaa cggcagcatc 2880 ccccaccaga tccacctggg agagctgcac gccattctgc ggcggcagga agatttttac 2940 ccattcctga aggacaaccg ggaaaagatc gagaagatcc tgaccttccg catcccctac 3000 tacgtgggcc ctctggccag gggaaacagc agattcgcct ggatgaccag aaagagcgag 3060 gaaaccatca ccccctggaa cttcgaggaa gtggtggaca agggcgcttc cgcccagagc 3120 ttcatcgagc ggatgaccaa cttcgataag aacctgccca acgagaaggt gctgcccaag 3180 cacagcctgc tgtacgagta cttcaccgtg tataacgagc tgaccaaagt gaaatacgtg 3240 accgagggaa tgagaaagcc cgccttcctg agcggcgagc agaaaaaggc catcgtggac 3300 ctgctgttca agaccaaccg gaaagtgacc gtgaagcagc tgaaagagga ctacttcaag 3360 aaaatcgagt gcttcgactc cgtggaaatc tccggcgtgg aagatcggtt caacgcctcc 3420 ctgggcacat accacgatct gctgaaaatt atcaaggaca aggacttcct ggacaatgag 3480 gaaaacgagg acattctgga agatatcgtg ctgaccctga cactgtttga ggacagagag 3540 atgatcgagg aacggctgaa aacctatgcc cacctgttcg acgacaaagt gatgaagcag 3600 ctgaagcggc ggagatacac cggctggggc aggctgagcc ggaagctgat caacggcatc 3660 cgggacaagc agtccggcaa gatacctg gatttcctga agtccgacgg cttcgccac 3720 agaaacttca tgcagctgat ccacgacgac agcctgacct ttaagagga catccagaaa 3780 gcccaggtgt ccggccagggg cgatagcctg cacgagcaca ttgccaatct ggccggcagc 3840 cccgccatta agaagggcat cctgcagaca gtgaaggtgg tggacgagct cgtgaaagtg 3900 atggggccggc acaagcccga gaacatcgtg atcgaatgg ccagagaa ccagaccacc 3960 cagaaggac agahaagg ccgcgagaga agaagcgg tcgaagggg catchaagag 4020 ctgggcagcc agatcctgaa agacacccc gtggaaaaca cccagctgca gaacgagaag 4080 ctgtacctgt actacctgca gatgggcgg gatatgtacg tggaccagga actggacatc 4140 aaccggctgt ccgactacga tgtggaccat atcgtcctc agagctttct gaaggacgac 4200 tccatcgaca acaagtgct gaccagaagc vakaagaacc ggggcaagg cgacaacgtg 4260 ccctccgaag aggtcgtgaa gagatgaag aactactggc ggcagctgct gaacgccaag 4320 ctgattaccc agagaagtt cgacaatctg accaagccg agagaggcgg cctgagcgaa 4380 ctggataagg ccggcttcat caagacag ctggtggaaa cccggcagat cacaaagcac 4440 4500 cgggaagtga aagtgatcac cctgaagtcc aagctgggtgt ccgatttccg gaaggatttc 4560 cagttttaca aagtgcgcga gatcaac taccaccacg cccacgacgc ctacctgaac 4620 gccgtcgtgg gaaccgccct gatcaaaaag tacctaagc tggaaagcga gttcgtgtac 4680 4740 aaggctaccg ccaagtactt cttctacagc aacatcatga actttttcaa gaccgagatt 4800 accctggcca acggcgagat ccggaagcgg cctctgatcg agaaacgg cgaaaccggg 4860 gagatcgtgt gggataaggg ccgggatttt gccaccgtgc ggaaagtgct gagcatgccc 4920 gaagtgaata tcgtgaaaaa gaccgaggtg cagacaggcg gcttcagcaa agagtctatc 4980 cggcccaaga ggaacagcga tagctgatc gccagaaaga agactggga ccctaagaag 5040 tacggcggct tcgtgagccc caccgtggcc tattctgtgc tggtggtggc caaagtggaa 5100 aagggcaagt ccaagaact gagagtgtg aaagagctgc tgggatcac catcatggaa 5160 agaagcagct tcgagagaa tcccatcgac ttctggaag ccaagggcta caagaagtg 5220 aaaaaggacc tgatcatca gctgcctaag tactccctgt tcgagctgga aaacggccgg 5280 aagagaatgc tggcctctgc cagattcctg cagaagggaa acgaactggc cctgccctcc 5340 aaatatgtga acttcctgta cctggccagc cactatgaga agctgaaggg ctcccccgag 5400 gataatgagc agaaacagct gttgtgga cagcacaagc actacctgga cgagatcatc 5460 gagcagatca gcgagttctc caagagagtg atcctggccg acgctaatct ggacaagtg 5520 ctgtccgcct acaacaagca ccgggataag cccatcagag agcaggccga gatatcatc 5580 cacctgttta cctgaccaa tctgggagcc cctcgggcct tcaagtactt tgacaccacc 5640 atcgaccgga aggtgtaccg gagcaccaaa gaggtgctgg acgccaccct gatccaccag 5700 agcatcaccg gcctgtacga gaacggatc gacctgctc agctgggagg tgactctggc 5760 ggctcaaaaa gaaccgccga cggcagcgaa ttcgagcccca agaagag gaaagtctaa 5820 ccggtcatca tcaccatcac cattgagttt aaacccgctg atcagcctcg actgtgcctt 5880 ctagttgcca gccatctgtt gtttgcccct cccccgtgcc ttccttgacc ctggaaggtg 5940 ccactcccac tgtcctttcc taataaaatg aggaaattgc atcgcattgt ctgagtaggt 6000 gtcattctat tctggggggt ggggtggggc aggacagcaa gggggaggat tgggaagaca 6060 atagcaggca tgctggggat gcggtgggct ctatggcttc tgaggcggaa agaaccagct 6120 ggggctcgat accgtcgacc tctagctaga gcttggcgta atcatggtca tagctgtttc 6180 ctgtgtgaaa ttgttatccg ctcacaattc cacacaacat acgagccgga agcataaagt 6240 gtaaagccta gggtgcctaa tgagtgagct aactcacatt aattgcgttg cgctcactgc 6300 ccgctttcca gtcgggaaac ctgtcgtgcc agctgcatta atgaatcggc caacgcgcgg 6360 ggagaggcgg tttgcgtatt gggcgctctt ccgcttcctc gctcactgac tcgctgcgct 6420 cggtcgttcg gctgcggcga gcggtatcag ctcactcaaa ggcggtaata cggttatcca 6480 cagaatcagg ggataacgca ggaaagaaca tgtgagcaaa aggccagcaa aaggccagga 6540 accgtaaaaa ggccgcgttg ctggcgtttt tccataggct ccgcccccct gacgagcatc 6600 acaaaaatcg acgctcaagt cagaggtggc gaaacccgac aggactataa agataccagg 6660 cgtttccccc tggaagctcc ctcgtgcgct ctcctgttcc gaccctgccg cttaccggat 6720 acctgtccgc ctttctccct tcgggaagcg tggcgctttc tcatagctca cgctgtaggt 6780 atctcagttc ggtgtaggtc gttcgctcca agctgggctg tgtgcacgaa ccccccgttc 6840 agcccgaccg ctgcgcctta tccggtaact atcgtcttga gtccaacccg gtaagacacg 6900 acttatcgcc actggcagca gccactggta acaggattag cagagcgagg tatgtaggcg 6960 gtgctacaga gttcttgaag tggtggccta actacggcta cactagaaga acagtatttg 7020 gtatctgcgc tctgctgaag ccagttacct tcggaaaaag agttggtagc tcttgatccg 7080 gcaaacaaac caccgctggt agcggtggtt tttttgtttg caagcagcag attacgcgca 7140 gaaaaaaagg atctcaagaa gatcctttga tcttttctac ggggtctgac actcagtgga 7200 acgaaaactc acgttaaggg attttggtca tgagattatc aaaaaggatc ttcacctaga 7260 tccttttaaa ttaaaaatga agttttaaat caatctaaag tatatatgag taaacttggt 7320 ctgacagtta ccaatgctta atcagtgagg cacctatctc agcgatctgt ctatttcgtt 7380 catccatagt tgcctgactc cccgtcgtgt agataactac gatacgggag ggcttaccat 7440 ctggccccag tgctgcaatg ataccgcgag acccacgctc accggctcca gatttatcag 7500 caataaacca gccagccgga agggccgagc gcagaagtgg tcctgcaact ttatccgcct 7560 ccatccagtc tattaattgt tgccgggaag ctagagtaag tagttcgcca gttaatagtt 7620 tgcgcaacgt tgttgccatt gctacaggca tcgtggtgtc acgctcgtcg tttggtatgg 7680 cttcattcag ctccggttcc caacgatcaa ggcgagttac atgatcccccc atgttgtgca 7740 aaaaagcggt tagctccttc ggtcctccga tcgttgtcag aagtaagttg gccgcagtgt 7800 tatcactcat ggttatggca gcactgcata attctcttac tgtcatgcca tccgtaagat 7860 gcttttctgt gactggtgag tactcaacca agtcattctg agaatagtgt atgcggcgac 7920 cgagttgctc ttgcccggcg tcaatacggg samaaccgc gccacatagc agaactttaa 7980 aagtgctcat cattggaaaa cgttcttcgg ggcgaaaact ctcaaggatc ttaccgctgt 8040 tgagatccag ttcgatgtaa cccactcgtg cacccaactg atcttcagca tcttttactt 8100 tcaccagcgt ttctgggtga gcaaaaacag gaaggcaaaa tgccgcaaaa aagggaataa 8160 gggcgacacg gaaatgttga atactcatac tcttcctttt tcaatattat tgaagcattt 8220 atcagggtta ttgtctcatg agcggataca tatttgaatg tatttagaaa aataaacaaa 8280 taggggttcc gcgcacattt ccccgaaaag tgccacctga cgtcgacgga tcgggagatc 8340 gatctcccga tcccctaggg tcgactctca gtacaatctg ctctgatgcc gcatagttaa 8400 gccagtatct gctccctgct tgtgtgttgg aggtcgctga gtagtgcgcg agcaaaattt 8460 aagctacaac aaggcaaggc ttgaccgaca attgcatgaa gaatctgctt agggttaggc 8520 gttttgcgct gcttcgcgat gtacgggcca gatatacgcg ttgacattga ttattgacta 8580 gttattaata gtaatcaatt acggggtcat tagttcatag cccatatatg gagttccgcg 8640 ttacataact tacggtaaat ggcccgcctg gctgaccgcc caacgacccc cgcccattga 8700 cgtcaataat gacgtatgtt cccatagtaa cgccaatagg gactttccat tgacgtcaat 8760 gggtggagta tttacggtaa actgcccact tggcagtaca tcaagtgtat c 8811 <210> 3 <211> twenty one <212> DNA <213> Artificial sequence <400> 3 atggactccc tggccgagtc t 21 <210> 4 <211> 20 <212> DNA <213> Artificial sequence <400> 4 cggcttcctt ttgtgaacag 20 <210> 5 <211> 42 <212> DNA <213> Artificial sequence <400> 5 ctgttcacaa aaggaagccg cacatgtaca atgcccaatg cg 42 <210> 6 <211> twenty four <212> DNA <213> Artificial sequence <400> 6 attcctatta cgcttgtttc ttgg 24 <210> 7 <211> 45 <212> DNA <213> Artificial sequence <400> 7 gaattcagca cagggagcat gggaatggac tccctggccg agtct 45 <210> 8 <211> 86 <212> DNA <213> Artificial sequence <400> 8 accggttgga ctttcctctt cttcttgggc tcgaactcgc tgccgtcggc ggttcttttt 60 tcattcctat tacgcttgtt tcttgg 86 <210> 9 <211> 1451 <212> DNA <213> Artificial sequence <400> 9 gaattcagca cagggagcat gggaatggac tccctggccg agtctcggtg gcctccgggc 60 ctggcagtca tgaagacaat agatgatttg ctgcggtgtg gaatttgctt cgagtatttc 120 aacattgcaa tgataatacc tcagtgttca cataactact gctctctctg tataagaaaa 180 tttctgtcct ataaaactca gtgtccaact tgctgtgtga ctgtcacaga gccggatctg 240 aaaaataacc gcatattaga tgaactggta aaaagcttga attttgcacg gaatcatctg 300 ctgcagtttg ctttagagtc accagccaaa tctcctgctt cttcctcttc aaagaatctt 360 gctgtcaaag tatatactcc tgtagcctcc agacagtctt taaagcaggg gagcaggtta 420 atggataatt tcttgatcag agaaatgagt ggttctacat cagagttgtt gataaaagaa 480 aataaaagca aattcagccc tcaaaaagag gcgagccctg ctgcaaagac caaagagaca 540 cgttctgtag aagagatcgc tccagatccc tcagaggcta agcgtcctga gccaccctcg 600 acatccactt tgaacaagt tactaagtg gattgtcctg tttgcggggt taacattcca 660 gaaagtcaca ttaataagca tttagacagc tgtttatcac gcgaagagaa gaaggaagc 720 ctcagaagtt ctgttcacaa aaggaagccg cacatgtaca atgcccaatg cgatgctttg 780 catcctaaat cagctgctga atagttcga gaatcgaaa atagagaa gactaggatg 840 cgtcttgaag ctagtaaact caatgaaagt gtaatggttt ttacaagga caacagaa 900 aaggaaatag atgaatcca cagtaaatat cgtaaaaaac atagagtga atttcagctt 960 ctggtggatc aggctagaaa aggatacaag aaattgctg gatgtcaca aaaaacagta 1020 acaatacaaaagaatga atctacagaa aagctatctt ctgtagcat gggacaggaa 1080 gataatatga cctcagtaac aaaccactt tctcaatca agctggactc cccagaggaa 1140 ttggaacctg acagagaga ggattcttct agctgtattg atttcaaga agttctttct 1200 tcatcagaat cagattcatg catagttcc agttcagaca tcatagaga tctttagaa 1260 gaagaggaag cctgggaagc atcacataaa aacgatcttc aagacacaga aataagtcca 1320 agacagaatc gccgcacaag agccgctgaa agtgctgaga ttgaaccaag aaacaagcgt 1380 aataggaatg aaaaaagaac cgccgacggc agcgagttcg agcccaagaa gaagaggaaa 1440 gtccaaccgg t 1451 <210> 10 <211> 9630 <212> DNA <213> Artificial Sequence <400> 10 atatgccaag tacgccccct attgacgtca atgacggtaa atggcccgcc tggcattatg 60 cccagtacat gaccttatgg gactttccta cttggcagta catctacgta ttagtcatcg 120 ctattaccat ggtgatgcgg ttttggcagt acatcaatgg gcgtggatag cggtttgact 180 cacggggatt tccaagtctc caccccattg acgtcaatgg gagtttgttt tggcaccaaa 240 atcaacggga ctttccaaaa tgtcgtaaca actccgcccc attgacgcaa atgggcggta 300 ggcgtgtacg gtgggaggtc tatataagca gagctggttt agtgaaccgt cagatccgct 360 agagatccgc ggccgctaat acgactcact atagggagag ccgccaccat gaaacggaca 420 gccgacggaa gcgagttcga gtcaccaaag aagaagcgga aagtctctga ggtggagttt 480 tcccacgagt actggatgag acatgccctg accctggcca agagggcacg ggatgagagg 540 gaggtgcctg tgggagccgt gctggtgctg aacaatagag tgatcggcga gggctggaac 600 agagccatcg gcctgcacga cccaacagcc catgccgaaa ttatggccct gagacagggc 660 ggcctggtca tgcagaacta cagactgatt gacgccaccc tgtacgtgac attcgagcct 720 tgcgtgatgt gcgccggcgc catgatccac tctaggatcg gccgcgtggt gtttggcgtg 780 aggaactcaa aaagaggcgc cgcaggctcc ctgatgaacg tgctgaacta ccccggcatg 840 aatcaccgcg tcgaaattac cgagggaatc ctggcagatg aatgtgccgc cctgctgtgc 900 gatttctatc ggatgcctag acaggtgttc aatgctcaga agaaggccca gagctccatc 960 aactccggag gatctagcgg aggctcctct ggctctgaga cacctggcac aagcgagagc 1020 gcaacacctg aaagcagcgg gggcagcagc ggggggtcag acaagaagta cagcatcggc 1080 ctggccatcg gcaccaactc tgtgggctgg gccgtgatca ccgacgagta caaggtgccc 1140 agcaagaaat tcaaggtgct gggcaacacc gaccggcaca gcatcaagaa gaacctgatc 1200 ggagccctgc tgttcgacag cggcgaaaca gccgaggcca cccggctgaa gagaaccgcc 1260 agaagaagat acaccagacg gaagaaccgg atctgctatc tgcaagagat cttcagcaac 1320 gagatggcca aggtggacga cagcttcttc cagagactgg aagagtcctt cctggtggaa 1380 1440 cacgagaagt acccccaccat ctaccacctg agaagaaac tggtggacag caccgacaag 1500 gccgacctgc ggctgatcta tctggccctg gcccacatga tcaagttccg gggccacttc 1560 ctgatcgagg gcgacctgaa ccccgacaac agcgacgtgg acaagctgtt catccagctg 1620 gtgcagacct aaaccagct gttcgaggaa aaccccatca acgccagcgg cgtggacgcc 1680 aaggccatcc tgtctgccag actgagcaag agcagacggc tggaaaatct gatcgcccag 1740 1800 acccccaact tcaagagcaa cttcgacctg gccgaggatg ccaaactgca gctgagcaag 1860 gabacctacg acgacgacct ggacaacctg ctggcccaga tcggcgacca gtacgccgac 1920 ctgttctctgg cctgtccgac gccatccctgc tgagcgacat cctgagagtg 1980 aacaccgaga tcaccaaggc ccccctgagc gcctctatga tcaagagata cgacgagcac 2040 caccaggacc tgaccctgct gaagctctc gtgcggcagc agctgcctga gaagtacaaa 2100 gagattttct tcgaccagag cagaacggc tacgccggct acatgacgg cggagccagc 2160 caggaagt tctacaagtt catcaagccc atcctggaaa agatggacgg caccgaggaa 2220 ctgctcgtga agctgaacag agaggacctg ctgcggaagc agcggacctt cgacaacggc 2280 agcatccccc accagatcca cctgggag ctgcacgcca ttctcggcg gcaggaagat 2340 ttttacccat tcctgaagga caacgggaa aagatcgaga agatcctgac cttccgcatc 2400 ccctactacg tggccctct ggccagggga aacagcagat tcgcctggat gaccagaaag 2460 agcgaggaaa ccatcacccc ctggaacttc gaggaagtgg tggacaaggg cgcttccgcc 2520 cagagcttca tcgagcggat gaccaacttc gataagaacc tgcccacga gaggtgctg 2580 cccaagcaca gcctgctgta cgagtacttc accgtgtata acgagctgac caagtgaaa 2640 tacgtgaccg agggaatgag aaagcccgcc ttcctgagcg gcagagcagaa aaggccac 2700 gtggacctgc tgttcagac caaccggaaa gtgaccgtga agcagctgaa agaggactac 2760 ttcaagaaaa tcgagtgctt cgactccgtg gaatctccg gcgtggaaga tcggttcaac 2820 gcctccctgg gcacatacca cgatctgctg aaaattatca aggacagga cttcctggac 2880 aatgaggaaa acgaggacat tctggagat atcgtgctga ccctgacact gtttgaggac 2940 agagagatga tcgaggaacg gctgaaaacc tatgcccacc tgttcgacga caagtgatg 3000 aagcagctga agcggcggag atacaccggc tggggcaggc tgagccgaa gctgatcac 3060 ggcatccggg acagcagtc cggcaagaca atcctggatt tcctgaagtc cgacggctc 3120 gccaacagaa acttcatgca gctgatccac gacgacagcc tgacctttaa agaggacatc 3180 cagaaagcccc aggtgtccgg ccaggggcgat agcctgcacg agcacattgc caatctggcc 3240 ggcagccccg ccattaagaa gggcatcctg cagacagtga aggtggtgga cgagctcgtg 3300 aaagtgatgg gccggcacaa gcccgagaac atcgtgatcg aaatggccag agagaaccag 3360 acccccaga aggagagaa gaacagccgc gagagaatga agcggatcga agagggcatc 3420 aaagagctgg gcagccagat cctgaaagaa caccccgtgg aaaacaccca gctgcagaac 3480 gagaagctgt acctgtacta cctgcagaat gggcgggata tgtacgtgga ccaggaactg 3540 gacatcaacc ggctgtccga ctacgatgtg gaccatatcg tgcctcagag ctttctgaag 3600 gacgactcca tcgacaacaa ggtgctgacc agaagcgaca agaaccgggg caagagcgac 3660 aacgtgccct ccgaagaggt cgtgaagaag atgaagaact actggcggca gctgctgaac 3720 gccaagctga ttacccagag aaagttcgac aatctgacca aggccgagag aggcggcctg 3780 agcgaactgg ataaggccgg cttcatcaag agacagctgg tggaaacccg gcagatcaca 3840 aagcacgtgg cacagatcct ggactcccgg atgaacacta agtacgacga gaatgacaag 3900 ctgatccggg aagtgaaagt gatcaccctg aagtccaagc tggtgtccga tttccggaag 3960 gatttccagt tttacaaagt gcgcgagatc aacaactacc accacgccca cgacgcctac 4020 ctgaacgccg tcgtgggaac cgccctgatc aaaaagtacc ctaagctgga aagcgagttc 4080 gtgtacggcg actacaaggt gtacgacgtg cggagaatga tcgccaagag cgagcaaggaa 4140 atcggcaagg ctaccgccaa gtacttcttc tacagcaaca tcatgaactt tttcaagacc 4200 gagattaccc tggccaacgg cgagatccgg aagcggcctc tgatcgagac aaacggcgaa 4260 accggggaga tcgtgtggga taagggccgg gattttgcca ccgtgcggaa agtgctgagc 4320 atgccccaag tgaatatcgt gaaaaagacc gaggtgcaga caggcggctt cagcaaagag 4380 tctatcctgc ccaaggaa cagcgataag ctgatcgcca gaagaagga ctgggaccct 4440 aagaagtcg gcggcttcga cagcccccacc gtggcctatt ctgtgctggt ggtggccaaa 4500 gtggaaaagg gcaagtccaa gaactgaag agtgtgaaag agctgctggg gatcaccatc 4560 atggaaagaa gcagcttcga gaagaatccc atcgactttc tggaagccaa gggctacaaa 4620 gaagtgaaaa aggacctgat catcaagctg cctaagtact ccctgttcga gctggaaaac 4680 ggccggagagaatgctggc ctctgccggc gaactgcaga agggaagaga actggccctg 4740 ccctccaaat atgtgaactt cctgtacctg gccagccact atgagaagct gaagggctcc 4800 cccgaggata atgagcagaa acagctgttt gtggaacagc acaagcacta cctggacgag 4860 atcatcgagc agatcagcga gttctccaag agagtgatcc tggccgacgc taatctggac 4920 aaagtgctgt ccgcctacaa caagcaccgg gataagccca tcagagagca ggccgagaat 4980 atcatccacc tgtttaccct gaccaatctg ggagcccctg ccgccttcaa gtactttgac 5040 accaccatcg accggaagag gtacaccagc accaaagagg tgctggacgc caccctgatc 5100 caccagagca tcaccggcct gtacgagaca cggatcgacc tgtctcagct gggaggtgac 5160 tctggcggct caaaaagaac cgccgacggc agcgaattca gcacagggag catgggaatg 5220 gactccctgg ccgagtctcg gtggcctccg ggcctggcag tcatgaagac aatagatgat 5280 ttgctgcggt gtggaatttg cttcgagtat ttcaacattg caatgataat acctcagtgt 5340 tcacataact actgctctct ctgtataaga aaatttctgt cctataaaac tcagtgtcca 5400 acttgctgtg tgactgtcac agagccggat ctgaaaaata accgcatatt agatgaactg 5460 gtaaaaaagct tgaattttgc acggaatcat ctgctgcagt ttgctttaga gtcaccagcc 5520 aaatctcctg cttcttctc ttcaaagaat cttgctgtca aagtatatac tcctgtagcc 5580 tccagacagt ctttaaagca ggggagcagg ttaatggata atttcttgat cagagaaatg 5640 agtggttcta catcagagtt gttgataaaa gaaataaaaa gcaaattcag ccctcaaaaa 5700 gaggcgagcc ctgctgcaaa gaccaaagag acacgttctg tagagaagat cgctccagat 5760 ccctcagagg ctaagcgtcc tgagccaccc tcgacatcca ctttgaaaca agttactaaa 5820 gtggattgtc ctgtttgcgg ggttaacatt ccagaaagtc acattaataa gcatttagac 5880 agctctttat cacgcgaaga gaagaaggaa agcctcagaa gttctgttca caaaaggaag 5940 ccgcacatgt acaatgccca atgcgatgct ttgcatccta aatcagctgc tgaaatagtt 6000 cgagaaatcg aaaatataga gaagactagg atgcgtcttg aagctagtaa actcaatgaa 6060 agtgtaatgg tttttacaaa ggaccaaaca gaaaaggaaa tagatgaaat ccacagtaaa 6120 tatcgtaaaa aacataagag tgaatttcag cttctggtgg atcaggctag aaaaggatac 6180 aagaaaattg ctggaatgtc acaaaaaaca gtaacaataa caaaaagaaga tgaatctaca 6240 gaaaagctat cttctgtatg catgggacag gaataata tgacctcagt aaaaaccac 6300 ttttctcaat caaagctgga ctccccagag gaattggaac ctgacagaga agaggattct 6360 tctagctgta ttgatattca agaagttctt tcttcatcag aatcagattc atgcaatagt 6420 tccagttcag acatcataag agatctttta gaagaagagg aagcctggga agcatcacat 6480 aaaaacgatc ttcaagacac agaataagt ccaagacaga atcgccgcac aagagccgct 6540 gaaagtgctg agattgaacc aagaaacaag cgtatatagga atgaaaaaag aaccgccgac 6600 ggcagcgagt tcgagcccaa gaagaagagg aaagtccaac cggtcatcat caccatcacc 6660 attgagttta aacccgctga tcagcctcga ctgtgccttc tagttgccag ccatctgttg 6720 tttgcccctc ccccgtgcct tccttgaccc tggaaggtgc cactccact gtcctttcct 6780 ataaatga ggaaattgca tcgcattgtc tgagtaggtg tcattctatt ctggggggtg 6840 gggtggggca ggacagcaag gggaggatt gggaagacaa tagcaggcat gctggggatg 6900 cggtgggctc tatggcttct gaggcggaaa gaaccagctg gggctcgata ccgtcgacct 6960 ctagctagag cttggcgtaa tcatggtcat agctgtttcc tgtgtgaaat tgttatccgc 7020 tcacaattcc acacaacata cgagccggaa gcataaagtg taaagcctag ggtgcctaat 7080 gagtgagcta actcacatta attgcgttgc gctcactgcc cgctttccag tcgggaaacc 7140 tgtcgtgcca gctgcattaa tgaatcggcc aacgcgcggg gagaggcggt ttgcgtattg 7200 ggcgctcttc cgcttcctcg ctcactgact cgctgcgctc ggtcgttcgg ctgcggcgag 7260 cggtatcagc tcactcaaag gcggtaatac ggttatccac agaatcaggg gataacgcag 7320 gaaagaacat gtgagcaaaa ggccagcaaa aggccaggaa ccgtaaaaag gccgcgttgc 7380 tggcgttttt ccataggctc cgcccccctg acgagcatca caaaaatcga cgctcaagtc 7440 agaggtggcg aaacccgaca ggactataaa gataccaggc gtttccccct ggaagctccc 7500 tcgtgcgctc tcctgttccg accctgccgc ttaccggata cctgtccgcc tttctccctt 7560 cgggaagcgt ggcgctttct catagctcac gctgtaggta tctcagttcg gtgtaggtcg 7620 ttcgctccaa gctgggctgt gtgcacgaac cccccgttca gcccgaccgc tgcgccttat 7680 ccggtaacta tcgtcttgag tccaacccgg taagacacga cttatcgcca ctggcagcag 7740 ccactggtaa caggattagc agagcgaggt atgtaggcgg tgctacagag ttcttgaagt 7800 ggtggcctaa ctacggctac actagaagaa cagtatttgg tatctgcgct ctgctgaagc 7860 cagttacctt cggaaaaaga gttggtagct cttgatccgg caaacaaacc accgctggta 7920 gcggtggtt ttttgtttgc aagcagcaga ttacgcgcag aaaaaaagga tctcaagaag 7980 atcctttgat cttttctacg gggtctgaca ctcagtggaa cgaaaactca cgttaaggga 8040 ttttggtcat gagattatca aaaaggatct tcacctagat ccttttaaat taaaaatgaa 8100 gttttaaatc aatctaaagt atatatgagt aaacttggtc tgacagttac caatgcttaa 8160 tcagtgaggc acctatctca gcgatctgtc tatttcgttc atccatagtt gcctgactcc 8220 ccgtcgtgta gataactacg atacgggagg gcttaccatc tggccccagt gctgcaatga 8280 taccgcgaga cccacgctca ccggctccag attatcagc aataaaccag ccagccggaa 8340 gggccgagcg cagaagtggt cctgcaactt tatccgcctc catccagtct atttattgtt 8400 gccgggaagc tagatagt agttcgccag ttaatagttt gcgcaacgtt gttgccattg 8460 ctacaggcat cgtggtgtca cgctcgtcgt ttggtatggc ttcattcagc tccggttccc 8520 aacgatcaag gcgagttaca tgatccccca tgttgtgcaa aaagcggtt agctccttcg 8580 gtcctccgat cgttgtcaga agtaagttgg ccgcagtgtt atcactcatg gttatggcag 8640 cactgcataa ttctcttact gtcatgccat ccgtaagatg cttttctgtg actggtgagt 8700 actcaaccaa gtcattctga gatagtgta tgcggcgacc gagttgctt tgcccggcgt 8760 caatacggga taataccgcg ccacatagca gaactttaaa agtgctcatc attggaaac 8820 gttcttcggg gcgaaaactc tcaggatct taccgctgtt gagatccagt tcgatgtaac 8880 ccactcgtgc acccaactga tcttcagcat cttttactt caccagcgtt tctgggtgag 8940 caaaaacagg aaggcaaat gccgcaaaaa agggaataag ggcgacacgg aaatgttgaa 9000 tactcatact cttccttttt caatattatt gaagcattta tcaggttat tgtctcatga 9060 gcggatacat atttgaatgt atttagaaaa ataaacaaat aggggttccg cgcacatttc cccgaaaagt gccacctgac gtcgacggat cgggagatcg atctcccgat cccctagggt 9180. cgactctcag tacaatctgc tctgatgccg catagttaag ccagtatctg ctccctgctt gtgtgttgga ggtcgctgag taggcgcga gcaaaattta agctacaaca aggcaaggct tgaccgacaa ttgcatgaag aatctgctta gggttaggcg ttttgcgctg cttcgcgatg tacgggccag fathercgcgt tgacattgat tattgactag ttattaatag father cggggtcatt agttcatagc ccatatattg agttccgcgt tacataactt acggtaaatg gcccgcctgg ctgaccgccc aacgacccc gcccattgac gtcaataatg acgtatgttc 9540 ccatagtaac gccaataggg actttccatt gacgtcaatg ggtggagtat ttacggtaaa ctgcccactt ggcagtacat caagtgtatc 9630 <210> 11 <211> 8217 <212> DNA <213> The snowstorm <400> 11 atatgccaag tacgccccct attgacgtca atgacggtaa atggcccgcc tggcattatg cccagtacat gaccttatgg gactttccta cttggcagta catctacgta ttagtcatcg 120 ctattaccat ggtgatgcgg ttttggcagt acatcaatgg gcgtggatag cggtttgact 180 cacggggatt tccaagtctc caccccattg acgtcaatgg gagtttgttt tggcaccaaa 240 atcaacggga ctttccaaaa tgtcgtaaca actccgcccc attgacgcaa atgggcggta 300 ggcgtgtacg gtgggaggtc tatataagca gagctggttt agtgaaccgt cagatccgct 360 agagatccgc ggccgctaat acgactcact ataggagag ccgccaccat gaaacggaca 420 gccgacggaa gcgagttcga gtcaccaaag aagaagcgga aagtctctga ggtggagttt 480 tcccacgagt actggatgag acatgccctg accctggcca agagggcacg ggatgagagg 540 gaggtgcctg tgggagccgt gctggtgctg aacaatagag tgatcggcga gggctggaac 600 agagccatcg gcctgcacga cccaacagcc catgccgaaa ttatggccct gagacagggc 660 ggcctggtca tgcagaacta cagactgatt gacgccaccc tgtacgtgac attcgagcct 720 tgcgtgatgt gcgccggcgc catgatccac tctaggatcg gccgcgtggt gtttggcgtg 780 aggaactcaa aaagaggcgc cgcaggctcc ctgatgaacg tgctgaacta ccccggcatg 840 aatcaccgcg tcgaaattac cgagggaatc ctggcagatg aatgtgccgc cctgctgtgc 900 gatttctatc ggatgcctag acaggtgttc aatgctcaga agaaggccca gagctccatc 960 aactccggag gatctagcgg aggctcctct ggctctgaga cacctggcac aagcgagagc 1020 gcaacacctg aaagcagcgg gggcagcagc gggggtcag acaagaagta cagcatcggc 1080 ctggccatcg gcaccaactc tgtgggctgg gccgtgatca ccgacgagta caaggtgccc 1140 agcaagaaat tcaaggtgct gggcaacacc gaccggcaca gcatcaagaa gaacctgatc 1200 ggagccctgc tgttcgacag cggcgaaaca gccgaggcca cccggctgaa gagaaccgcc 1260 agaagaagat acaccagacg gaagaaccgg atctgctatc tgcaagagat cttcagcaac 1320 gagatggcca aggtggacga cagcttcttc cagagactgg aagagtcctt cctggtggaa 1380 1440 cacgagaagt acccccaccat ctaccacctg agaagaaac tggtggacag caccgacaag 1500 gccgacctgc ggctgatcta tctggccctg gcccacatga tcaagttccg gggccacttc 1560 ctgatcgagg gcgacctgaa ccccgacaac agcgacgtgg acaagctgtt catccagctg 1620 gtgcagacct acaaccagct gttcgaggaa aaccccatca acgccagcgg cgtggacgcc 1680 aaggccatcc tgtctgccag actgagcaag agcagacggc tggaaaatct gatcgcccag 1740 ctgcccggcg agaagaagaa tggcctgttc ggaaacctga ttgccctgag cctgggcctg 1800 acccccaact tcaagagcaa cttcgacctg gccgaggatg ccaaactgca gctgagcaag 1860 gacacctacg acgacgacct ggacaacctg ctggcccaga tcggcgacca gtacgccgac 1920 ctgtttctgg ccgccaagaa cctgtccgac gccatcctgc tgagcgacat cctgagagtg 1980 aacaccgaga tcaccaaggc ccccctgagc gcctctatga tcaagagata cgaccagcac 2040 caccaggacc tgaccctgct gaaagctctc gtgcggcagc agctgcctga gaagtacaaa 2100 gagattttct tcgaccagag caagaacggc tacgccggct acattgacgg cggagccagc 2160 caggaagagt tctacaagtt catcaagccc atcctggaaa agatggacgg caccgaggaa 2220 ctgctcgtga agctgaacag agaggacctg ctgcggaagc agcggacctt cgacaacggc 2280 agcatccccc accagatcca cctgggag ctgcacgcca ttctcggcg gcaggaagat 2340 ttttacccat tcctgaagga caacgggaa aagatcgaga agatcctgac cttccgcatc 2400 ccctactacg tggccctct ggccagggga aacagcagat tcgcctggat gaccagaaag 2460 agcgaggaaa ccatcacccc ctggaacttc gaggaagtgg tggacaaggg cgcttccgcc 2520 cagagcttca tcgagcggat gaccaacttc gataagaacc tgcccacga gaggtgctg 2580 cccaagcaca gcctgctgta cgagtacttc accgtgtata acgagctgac caagtgaaa 2640 tacgtgaccg agggaatgag aaagcccgcc ttcctgagcg gcagagcagaa aaggccac 2700 gtggacctgc tgttcagac caaccggaaa gtgaccgtga agcagctgaa agaggactac 2760 ttcaagaaaa tcgagtgctt cgactccgtg gaatctccg gcgtggaaga tcggttcaac 2820 gcctccctgg gcacatacca cgatctgctg aaaattatca aggacagga cttcctggac 2880 aatgaggaaa acgaggacat tctggagat atcgtgctga ccctgacact gtttgaggac 2940 agagagatga tcgaggaacg gctgaaaacc tatgcccacc tgttcgacga caagtgatg 3000 aagcagctga agcggcggag atacaccggc tggggcaggc tgagccgaa gctgatcac 3060 ggcatccggg acagcagtc cggcaagaca atcctggatt tcctgaagtc cgacggctc 3120 gccaacagaa acttcatgca gctgatccac gacgacagcc tgacctttaa agaggacatc 3180 cagaaagcccc aggtgtccgg ccaggggcgat agcctgcacg agcacattgc caatctggcc 3240 ggcagccccg ccattaagaa gggcatcctg cagacagtga aggtggtgga cgagctcgtg 3300 aaagtgatgg gccggcacaa gcccgagaac atcgtgatcg aaatggccag aggaaccag 3360 accacccaga agggagaa gaacagccgc gagagaatga agcggatcga agggggcatc 3420 aaagagctgg gcagccagat cctgaaagaa caccccgtgg aaaacaccca gctgcagaac 3480 gagaagctgt acctgtacta cctgcagaat gggcgggata tgtacgtgga ccaggaactg 3540 gatacaacc ggctgtccga ctacgatgtg gaccatatcg tgcctcagag ctttctgaag 3600 gacgactcca tcgacaacaa ggtgctgacc agaagcgaca agaaccgggg gaagagcgac 3660 aacgtgccct ccgaagaggt cgtgaagaag atgaagaact actggcggca gctgctgaac 3720 gccaagctga ttacccagag aaagttcgac aatctgacca aggccgagag aggcggcctg 3780 agcgaactgg ataaggccgg cttcatcaag agacagctgg tggaaacccg gcagatcaca 3840 aagcacgtgg cacagatcct ggactcccgg atgaacacta agtacgacga gaatgacaag 3900 ctgatccggg aagtgaaagt gatcaccctg aagtccaagc tggtgtccga tttccggaag 3960 gatttccagt tttacaaagt gcgcgagatc aacaactacc accacgccca cgacgcctac 4020 ctgaacgccg tcgtgggaac cgccctgatc aaaaagtacc ctaagctgga aagcgagttc 4080 gtgtacggcg actacaaggt gtacgacgtg cggaagatga tcgccaagag cgagcaggaa 4140 atcggcaagg ctaccgccaa gtacttcttc tacagcaaca tcatgaactt tttcaagacc 4200 gagattaccc tggccaacgg cgagatccgg aagcggcctc tgatcgagac aaacggcgaa 4260 accggggaga tcgtgtggga taagggccgg gattttgcca ccgtgcggaa agtgctgagc 4320 atgccccaag tgaatatcgt gaaaaagacc gaggtgcaga caggcggctt cagcaaagag 4380 tctatcctgc ccaagaggaa cagcgataag ctgatcgcca gaaagaagga ctgggaccct 4440 aagaagtacg gcggcttcga cagccccacc gtggcctatt ctgtgctggt ggtggccaaa 4500 gtggaaaagg gcaagtccaa gaaactgaag agtgtgaaag agctgctggg gatcaccatc 4560 atggaaagaa gcagcttcga gaagaatccc atcgactttc tggaagccaa gggctacaaa 4620 gaagtgaaaa aggacctgat catcaagctg cctaagtact ccctgttcga gctggaaaac 4680 ggccggaaga gaatgctggc ctctgccggc gaactgcaga agggaaacga actggccctg 4740 ccctccaaat atgtgaactt cctgtacctg gccagccact atgagaagct gaagggctcc 4800 cccgaggata atgagcagaa acagctgttt gtggaacagc acaagcacta cctggacgag 4860 atcatcgagc agatcagcga gttctccaag agagtgatcc tggccgacgc taatctggac 4920 aaagtgctgt ccgcctacaa caagcaccgg gataagccca tcagagagca ggccgagaat 4980 atcatccacc tgtttaccct gaccaatctg ggagcccctg ccgccttcaa gtactttgac 5040 accaccatcg accggaagag gtacaccagc accaaagagg tgctggacgc caccctgatc 5100 caccagagca tcaccggcct gtacgagaca cggatcgacc tgtctcagct gggaggtgac 5160 tctggcggct caaaaagaac cgccgacggc agcgaattcg agcccaagaa gaagaggaaa 5220 gtctaaccgg tcatcatcac catcaccatt gagtttaaac ccgctgatca gcctcgactg 5280 tgccttctag ttgccagcca tctgttgttt gcccctcccc cgtgccttcc ttgaccctgg 5340 aaggtgccac tcccactgtc ctttcctaat aaaatgagga aattgcatcg cattgtctga 5400 gtaggtgtca ttctattctg gggggtgggg tggggcagga cagcaagggg gaggattggg 5460 aagacaatag caggcatgct ggggatgcgg tgggctctat ggcttctgag gcggaaagaa 5520 ccagctgggg ctcgataccg tcgacctcta gctagagctt ggcgtaatca tggtcatagc 5580 tgtttcctgt gtgaaattgt tatccgctca caattccaca caacatacga gccggaagca 5640 taaagtgtaa agcctagggt gcctaatgag tgagctaact cacattaatt gcgttgcgct 5700 cactgcccgc tttccagtcg ggaaacctgt cgtgccagct gcattaatga atcggccaac 5760 gcgcggggag aggcggtttg cgtattgggc gctcttccgc ttcctcgctc actgactcgc 5820 tgcgctcggt cgtcggctg cggcgagcgg tatcagctca ctcaaaggcg gtaatacggt 5880 tatccacaga atcaggggat aacgcaggaa agaacatgtg agcaaaaggc cagcaaaagg 5940 ccaggaaccg taaaaggcc gcgttgctgg cgtttttcca taggctccgc ccccctgacg 6000 agcatcacaa aaatcgacgc tcaagtcaga ggtggcgaaa cccgacagga ctataaagat 6060 accaggcgtt tccccctgga agctccctcg tgcgctctcc tgttccgacc ctgccgctta 6120 ccggatacct gtccgccttt ccccttcgg gaagcgtggc gctttctcat agctcacgct 6180 gtaggtatct cagttcggtg tagtcgttc gctccaagct gggctgtgtg cacgaacccc 6240 ccgttcagcc cgaccgctgc gccttatccg gtaactatcg tcttgagtcc aacccggtaa 6300 gacacgactt atcgccactg gcagcagcca ctggtaacag gattagcaga gcgaggtatg 6360 taggcggtgc tacagagttc ttgaagtggt ggcctaacta cggctacact agaagaacag 6420 tatttggtat ctgcgctctg ctgaagccag ttaccttcgg aaaaagagtt ggtagctctt 6480 gatccggcaa acaaaccacc gctggtagcg gtggtttttt tgtttgcaag cagcagatta 6540 cgcgcagaaa aaaaggatct caagaagatc ctttgatctt ttctacgggg tctgacactc 6600 agtgggaacga aaactcacgt taagggattt tggtcatgag attatcaaaa aggatcttca 6660 6720 cttggtctga cagttaccaa tgcttaatca gtgaggcacc tatctcagcg atctgtctat 6780 ttcgttcatc catagttgcc tgactccccg tcgtgtagat aactacgata cgggaggct 6840 taccatctgg ccccagtgct gcaatgatac cgcgagaccc acgctcaccg gctccagatt 6900 tatcagcaat aaaccagcca gccggaaggg ccgagcgcag aagtggtcct gcaactttat 6960 ccgcctccat ccagtctatt aattgttgcc gggaagctag agtaagtagt tcgccagtta 7020 atagtttgcg caacgttgtt gccattgcta caggcatcgt ggtgtcacgc tcgtcgtttg 7080 gtatggcttc attcagctcc ggttcccaac gatcaaggcg agttacatga tcccccatgt 7140 tgtgcaaaaa agcggttagc tccttcggtc ctccgatcgt tgtcagaagt aagttggccg 7200 cagtgttatc actcatggtt atggcagcac tgcataattc tcttactgtc atgccatccg 7260 taagatgctt ttctgtgact ggtgagtact caaccaagtc attctgagaa tagtgtatgc 7320 ggcgaccgag ttgctcttgc ccggcgtcaa tacgggataa taccgcgcca catagcagaa 7380 ctttaaaagt gctcatcatt ggaaaacgtt cttcggggcg aaaactctca aggatcttac 7440 cgctgttgag atccagttcg atgtaaccca ctcgtgcacc caactgatct tcagcatctt 7500 ttactttcac cagcgtttct gggtgagcaa aaacaggaag gcaaaatgcc gcaaaaaaagg 7560 gaataagggc gacacggaaa tgttgaatac tcatactctt cctttttcaa tattattgaa 7620 gcatttatca gggttatgt ctcatgagcg gatacatatt tgaatgtatt tagaaaaata 7680 aacaaatagg ggttccgcgc acatttcccc gaaaagtgcc acctgacgtc gacggatcgg 7740 gagatcgatc tcccgatccc ctagggtcga ctctcagtac aatctgctct gatgccgcat 7800 agttaagcca gtatctgctc cctgcttgtg tgttggaggt cgctgagtag tgcgcgagca 7860 aaatttaagc tacaacaagg caaggcttga ccgacaattg catgaagaat ctgcttaggg 7920 ttaggcgttt tgcgctgctt cgcgatgtac gggccagata tacgcgttga cattgattat 7980 tgactagtta ttaatagtaa tcaattacgg ggtcattagt tcatagccca tatattgagt 8040 tccgcgttac ataacttacg gtaaatggcc cgcctggctg accgcccaac gacccccgcc 8100 cattgacgtc aataatgacg tatgttccca tagtaacgcc aatagggact ttccattgac 8160 gtcaatgggt ggagtattta cggtaaactg cccacttggc agtacatcaa gtgtatc 8217 <210> 12 <211> 20 <212> DNA <213> Artificial sequence <400> 12 gaatactaag catagactcc 20 <210> 13 <211> 20 <212> DNA <213> Artificial sequence <400> 13 ggagtctatg cttagtattc 20 <210> 14 <211> 20 <212> DNA <213> Artificial sequence <400> 14 gtaaacaaag catagactga 20 <210> 15 <211> 20 <212> DNA <213> Artificial sequence <400> 15 tcagtctatg ctttgtttac 20 <210> 16 <211> 20 <212> DNA <213> Artificial sequence <400> 16 gaacacaaag catagactgc 20 <210> 17 <211> 20 <212> DNA <213> Artificial sequence <400> 17 gcagtctatg ctttgtgttc 20 <210> 18 <211> 20 <212> DNA <213> Artificial sequence <400> 18 gatgagataa tgatgagtca 20 <210> 19 <211> 20 <212> DNA <213> Artificial sequence <400> 19 tgactcatca ttatctcatc 20 <210> 20 <211> 20 <212> DNA <213> Artificial sequence <400> 20 gacaaaccag aagccgctcc 20 <210> twenty one <211> 20 <212> DNA <213> Artificial sequence <400> twenty one ggagcggctt ctggtttgtc 20 <210> twenty two <211> 20 <212> DNA <213> Artificial sequence <400> twenty two gggaataaat catagaatcc 20 <210> twenty three <211> 20 <212> DNA <213> Artificial sequence <400> twenty three ggattctatg atttattccc 20 <210> twenty four <211> 20 <212> DNA <213> Artificial sequence <400> twenty four ggaacacaaa gcatagactg 20 <210> 25 <211> 20 <212> DNA <213> Artificial sequence <400> 25 cagtctatgc tttgtgttcc 20 <210> 26 <211> 20 <212> DNA <213> Artificial sequence <400> 26 gcacctacct cgggagctga 20 <210> 27 <211> 20 <212> DNA <213> Artificial sequence <400> 27 tcagctcccg aggtaggtgc 20 <210> 28 <211> 20 <212> DNA <213> Artificial sequence <400> 28 ggaatccctt ctgcagcacc 20 <210> 29 <211> 20 <212> DNA <213> Artificial sequence <400> 29 ggtgctgcag aagggattcc 20 <210> 30 <211> 20 <212> DNA <213> Artificial sequence <400> 30 tcagaaagtg gtggctggtg 20 <210> 31 <211> 20 <212> DNA <213> Artificial sequence <400> 31 caccagccac cactttctga 20 <210> 32 <211> 20 <212> DNA <213> Artificial sequence <400> 32 ggcccagact gagcacgtga 20 <210> 33 <211> 20 <212> DNA <213> Artificial sequence <400> 33 tcacgtgctc agtctgggcc 20 <210> 34 <211> 20 <212> DNA <213> Artificial sequence <400> 34 atatttgcat tgagatagtg 20 <210> 35 <211> 20 <212> DNA <213> Artificial sequence <400> 35 cactatctca atgcaaatat 20 <210> 36 <211> 20 <212> DNA <213> Artificial sequence <400> 36 gtcatcttag tcattacctg 20 <210> 37 <211> 20 <212> DNA <213> Artificial sequence <400> 37 caggtaatga ctaagatgac 20 <210> 38 <211> 20 <212> DNA <213> Artificial sequence <400> 38 gaagatagag aatagactgc 20 <210> 39 <211> 20 <212> DNA <213> Artificial sequence <400> 39 gcagtctatt ctctatcttc 20 <210> 40 <211> 1791 <212> PRT <213> Artificial sequence <400> 40 Pro Lys Lys Lys Arg Lys Val Ser Glu Val Glu Phe Ser His Glu Tyr 1 5 10 15 Trp Met Arg His Ala Leu Thr Leu Ala Lys Arg Ala Trp Asp Glu Arg 20 25 30 Glu Val Pro Val Gly Ala Val Leu Val His Asn Asn Arg Val Ile Gly 35 40 45 Glu Gly Trp Asn Arg Pro Ile Gly Arg His Asp Pro Thr Ala His Ala 50 55 60 Glu Ile Met Ala Leu Arg Gln Gly Gly Leu Val Met Gln Asn Tyr Arg 65 70 75 80 Leu Ile Asp Ala Thr Leu Tyr Val Thr Leu Glu Pro Cys Val Met Cys 85 90 95 Ala Gly Ala Met Ile His Ser Arg Ile Gly Arg Val Val Phe Gly Ala 100 105 110 Arg Asp Ala Lys Thr Gly Ala Ala Gly Ser Leu Met Asp Val Leu His 115 120 125 His Pro Gly Met Asn His Arg Val Glu Ile Thr Glu Gly Ile Leu Ala 130 135 140 Asp Glu Cys Ala Ala Leu Leu Ser Asp Phe Phe Arg Met Arg Arg Gln 145 150 155 160 Glu Ile Lys Ala Gln Lys Lys Ala Gln Ser Ser Thr Asp Ser Gly Gly 165 170 175 Ser Ser Gly Gly Ser Ser Gly Ser Glu Thr Pro Gly Thr Ser Glu Ser 180 185 190 Ala Thr Pro Glu Ser Ser Gly Gly Ser Ser Gly Gly Ser Ser Glu Val 195 200 205 Glu Phe Ser His Glu Tyr Trp Met Arg His Ala Leu Thr Leu Ala Lys 210 215 220 Arg Ala Arg Asp Glu Arg Glu Val Pro Val Gly Ala Val Leu Val Leu 225 230 235 240 Asn Asn Arg Val Ile Gly Glu Gly Trp Asn Arg Ala Ile Gly Leu His 245 250 255 Asp Pro Thr Ala His Ala Glu Ile Met Ala Leu Arg Gln Gly Gly Leu 260 265 270 Val Met Gln Asn Tyr Arg Leu Ile Asp Ala Thr Leu Tyr Val Thr Phe 275 280 285 Glu Pro Cys Val Met Cys Ala Gly Ala Met Ile His Ser Arg Ile Gly 290 295 300 Arg Val Val Phe Gly Val Arg Asn Ala Lys Thr Gly Ala Ala Gly Ser 305 310 315 320 Leu Met Asp Val Leu His Tyr Pro Gly Met Asn His Arg Val Glu Ile 325 330 335 Thr Glu Gly Ile Leu Ala Asp Glu Cys Ala Ala Leu Leu Cys Tyr Phe 340 345 350 Phe Arg Met Pro Arg Gln Val Phe Asn Ala Gln Lys Lys Ala Gln Ser 355 360 365 Ser Thr Asp Ser Gly Gly Ser Ser Gly Gly Ser Ser Gly Ser Glu Thr 370 375 380 Pro Gly Thr Ser Glu Ser Ala Thr Pro Glu Ser Ser Gly Gly Ser Ser 385 390 395 400 Gly Gly Ser Asp Lys Lys Tyr Ser Ile Gly Leu Ala Ile Gly Thr Asn 405 410 415 Ser Val Gly Trp Ala Val Ile Thr Asp Glu Tyr Lys Val Pro Ser Lys 420 425 430 Lys Phe Lys Val Leu Gly Asn Thr Asp Arg His Ser Ile Lys Lys Asn 435 440 445 Leu Ile Gly Ala Leu Leu Phe Asp Ser Gly Glu Thr Ala Glu Ala Thr 450 455 460 Arg Leu Lys Arg Thr Ala Arg Arg Arg Tyr Thr Arg Arg Lys Asn Arg 465 470 475 480 Ile Cys Tyr Leu Gln Glu Ile Phe Ser Asn Glu Met Ala Lys Val Asp 485 490 495 Asp Ser Phe Phe His Arg Leu Glu Glu Ser Phe Leu Val Glu Glu Asp 500 505 510 Lys Lys His Glu Arg His Pro Ile Phe Gly Asn Ile Val Asp Glu Val 515 520 525 Ala Tyr His Glu Lys Tyr Pro Thr Ile Tyr His Leu Arg Lys Lys Leu 530 535 540 Val Asp Ser Thr Asp Lys Ala Asp Leu Arg Leu Ile Tyr Leu Ala Leu 545 550 555 560 Ala His Met Ile Lys Phe Arg Gly His Phe Leu Ile Glu Gly Asp Leu 565 570 575 Asn Pro Asp Asn Ser Asp Val Asp Lys Leu Phe Ile Gln Leu Val Gln 580 585 590 Thr Tyr Asn Gln Leu Phe Glu Glu Asn Pro Ile Asn Ala Ser Gly Val 595 600 605 Asp Ala Lys Ala Ile Leu Ser Ala Arg Leu Ser Lys Ser Arg Arg Leu 610 615 620 Glu Asn Leu Ile Ala Gln Leu Pro Gly Glu Lys Lys Asn Gly Leu Phe 625 630 635 640 Gly Asn Leu Ile Ala Leu Ser Leu Gly Leu Thr Pro Asn Phe Lys Ser 645 650 655 Asn Phe Asp Leu Ala Glu Asp Ala Lys Leu Gln Leu Ser Lys Asp Thr 660 665 670 Tyr Asp Asp Asp Leu Asp Asn Leu Leu Ala Gln Ile Gly Asp Gln Tyr 675 680 685 Ala Asp Leu Phe Leu Ala Ala Lys Asn Leu Ser Asp Ala Ile Leu Leu 690 695 700 Ser Asp Ile Leu Arg Val Asn Thr Glu Ile Thr Lys Ala Pro Leu Ser 705 710 715 720 Ala Ser Met Ile Lys Arg Tyr Asp Glu His His Gln Asp Leu Thr Leu 725 730 735 Leu Lys Ala Leu Val Arg Gln Gln Leu Pro Glu Lys Tyr Lys Glu Ile 740 745 750 Phe Phe Asp Gln Ser Lys Asn Gly Tyr Ala Gly Tyr Ile Asp Gly Gly 755 760 765 Ala Ser Gln Glu Glu Phe Tyr Lys Phe Ile Lys Pro Ile Leu Glu Lys 770 775 780 Met Asp Gly Thr Glu Glu Leu Leu Val Lys Leu Asn Arg Glu Asp Leu 785 790 795 800 Leu Arg Lys Gln Arg Thr Phe Asp Asn Gly Ser Ile Pro His Gln Ile 805 810 815 His Leu Gly Glu Leu His Ala Ile Leu Arg Arg Gln Glu Asp Phe Tyr 820 825 830 Pro Phe Leu Lys Asp Asn Arg Glu Lys Ile Glu Lys Ile Leu Thr Phe 835 840 845 Arg Ile Pro Tyr Tyr Val Gly Pro Leu Ala Arg Gly Asn Ser Arg Phe 850 855 860 Ala Trp Met Thr Arg Lys Ser Glu Glu Thr Ile Thr Pro Trp Asn Phe 865 870 875 880 Glu Glu Val Val Asp Lys Gly Ala Ser Ala Gln Ser Phe Ile Glu Arg 885 890 895 Met Thr Asn Phe Asp Lys Asn Leu Pro Asn Glu Lys Val Leu Pro Lys 900 905 910 His Ser Leu Leu Tyr Glu Tyr Phe Thr Val Tyr Asn Glu Leu Thr Lys 915 920 925 Val Lys Tyr Val Thr Glu Gly Met Arg Lys Pro Ala Phe Leu Ser Gly 930 935 940 Glu Gln Lys Lys Ala Ile Val Asp Leu Leu Phe Lys Thr Asn Arg Lys 945 950 955 960 Val Thr Val Lys Gln Leu Lys Glu Asp Tyr Phe Lys Lys Ile Glu Cys 965 970 975 Phe Asp Ser Val Glu Ile Ser Gly Val Glu Asp Arg Phe Asn Ala Ser 980 985 990 Leu Gly Thr Tyr His Asp Leu Leu Lys Ile Ile Lys Asp Lys Asp Phe 995 1000 1005 Leu Asp Asn Glu Glu Asn Glu Asp Ile Leu Glu Asp Ile Val Leu Thr 1010 1015 1020 Leu Thr Leu Phe Glu Asp Arg Glu Met Ile Glu Glu Arg Leu Lys Thr 1025 1030 1035 1040 Tyr Ala His Leu Phe Asp Asp Lys Val Met Lys Gln Leu Lys Arg Arg 1045 1050 1055 Arg Tyr Thr Gly Trp Gly Arg Leu Ser Arg Lys Leu Ile Asn Gly Ile 1060 1065 1070 Arg Asp Lys Gln Ser Gly Lys Thr Ile Leu Asp Phe Leu Lys Ser Asp 1075 1080 1085 Gly Phe Ala Asn Arg Asn Phe Met Gln Leu Ile His Asp Asp Ser Leu 1090 1095 1100 Thr Phe Lys Glu Asp Ile Gln Lys Ala Gln Val Ser Gly Gln Gly Asp 1105 1110 1115 1120 Ser Leu His Glu His Ile Ala Asn Leu Ala Gly Ser Pro Ala Ile Lys 1125 1130 1135 Lys Gly Ile Leu Gln Thr Val Lys Val Val Asp Glu Leu Val Lys Val 1140 1145 1150 Met Gly Arg His Lys Pro Glu Asn Ile Val Ile Glu Met Ala Arg Glu 1155 1160 1165 Asn Gln Thr Thr Gln Lys Gly Gln Lys Asn Ser Arg Glu Arg Met Lys 1170 1175 1180 Arg Ile Glu Glu Gly Ile Lys Glu Leu Gly Ser Gln Ile Leu Lys Glu 1185 1190 1195 1200 His Pro Val Glu Asn Thr Gln Leu Gln Asn Glu Lys Leu Tyr Leu Tyr 1205 1210 1215 Tyr Leu Gln Asn Gly Arg Asp Met Tyr Val Asp Gln Glu Leu Asp Ile 1220 1225 1230 Asn Arg Leu Ser Asp Tyr Asp Val Asp His Ile Val Pro Gln Ser Phe 1235 1240 1245 Leu Lys Asp Asp Ser Ile Asp Asn Lys Val Leu Thr Arg Ser Asp Lys 1250 1255 1260 Asn Arg Gly Lys Ser Asp Asn Val Pro Ser Glu Glu Val Val Lys Lys 1265 1270 1275 1280 Met Lys Asn Tyr Trp Arg Gln Leu Leu Asn Ala Lys Leu Ile Thr Gln 1285 1290 1295 Arg Lys Phe Asp Asn Leu Thr Lys Ala Glu Arg Gly Gly Leu Ser Glu 1300 1305 1310 Leu Asp Lys Ala Gly Phe Ile Lys Arg Gln Leu Val Glu Thr Arg Gln 1315 1320 1325 Thr Lys His Val Ala Gln Ile Leu Asp Ser Arg Met Asn Thr Lys 1330 1335 1340 Tyr Asp Glu Asn Asp Lys Ile Arg Glu Val Lys Val Ile Thr Leu 1345 1350 1355 1360 Lys Ser Lys Leu Val Ser Asp Phe Arg Lys Asp Phe Gln Phe Tyr Lys 1365 1370 1375 Val Arg Glu Ile Asn Asn Tyr His Ala His Asp Ala Tyr Leu Asn 1380 1385 1390 Ala Val Val Gly Thr Ala Leu Ile Lys Lys Tyr Pro Lys Leu Glu Ser 1395 1400 1405 Glu Phe Val Tyr Gly Asp Tyr Lys Val Tyr Asp Val Arg Lys Met Ile 1410 1415 1420 Ala Lys Ser Glu Gln Glu Ile Gly Lys Ala Thr Ala Lys Tyr Phe Phe 1425 1430 1435 1440 Tyr Ser Asn With Asn Phe Phe Lys Thr Glu Ile Thr Leu Ala Asn 1445 1450 1455 Gly Glu Ile Arg Lys Arg Pro Leu Ile Glu Thr Asn Gly Glu Thr Gly 1460 1465 1470 Glu Ile Val Trp Asp Lys Gly Arg Asp Phe Ala Thr Val Arg Lys Val 1475 1480 1485 Leu Ser Met Pro Gln Val Asn Ile Val Lys Lys Thr Glu Val Gln Thr 1490 1495 1500 Gly Gly Phe Ser Lys Glu Ser Ile Arg Pro Lys Arg Asn Ser Asp Lys 1505 1510 1515 1520 Leu Ile Ala Arg Lys Lys Asp Trp Asp Pro Lys Lys Tyr Gly Gly Phe 1525 1530 1535 Val Ser Pro Thr Val Ala Tyr Ser Val Leu Val Val Ala Lys Val Glu 1540 1545 1550 Lys Gly Lys Ser Lys Lys Leu Lys Ser Val Lys Glu Leu Leu Gly Ile 1555 1560 1565 Thr Ile Met Glu Arg Ser Ser Phe Glu Lys Asn Pro Ile Asp Phe Leu 1570 1575 1580 Glu Ala Lys Gly Tyr Lys Glu Val Lys Lys Asp Leu Ile Ile Lys Leu 1585 1590 1595 1600 Pro Lys Tyr Ser Leu Phe Glu Leu Glu Asn Gly Arg Lys Arg Met Leu 1605 1610 1615 Ala Ser Ala Arg Phe Leu Gln Lys Gly Asn Glu Leu Ala Leu Pro Ser 1620 1625 1630 Lys Tyr Val Asn Phe Leu Tyr Leu Ala Ser His Tyr Glu Lys Leu Lys 1635 1640 1645 Gly Ser Pro Glu Asp Asn Glu Gln Lys Gln Leu Phe Val Glu Gln His 1650 1655 1660 Lys His Tyr Leu Asp Glu Ile Ile Glu Gln Ile Ser Glu Phe Ser Lys 1665 1670 1675 1680 Arg Val Ile Leu Ala Asp Ala Asn Leu Asp Lys Val Leu Ser Ala Tyr 1685 1690 1695 Asn Lys His Arg Asp Lys Pro Ile Arg Glu Gln Ala Glu Asn Ile Ile 1700 1705 1710 His Leu Phe Thr Leu Thr Asn Leu Gly Ala Pro Arg Ala Phe Lys Tyr 1715 1720 1725 Phe Asp Thr Thr Ile Asp Arg Lys Val Tyr Arg Ser Thr Lys Glu Val 1730 1735 1740 Leu Asp Ala Thr Leu Ile His Gln Ser Ile Thr Gly Leu Tyr Glu Thr 1745 1750 1755 1760 Arg Ile Asp Leu Ser Gln Leu Gly Gly Asp Ser Gly Gly Ser Lys Arg 1765 1770 1775 Thr Ala Asp Gly Ser Glu Phe Glu Pro Lys Lys Lys Arg Lys Val 1780 1785 1790 <210> 41 <211> 1605 <212> PRT <213> Artificial sequence <400> 41 Met Lys Arg Thr Ala Asp Gly Ser Glu Phe Glu Ser Pro Lys Lys Lys 1 5 10 15 Arg Lys Val Ser Glu Val Glu Phe Ser His Glu Tyr Trp Met Arg His 20 25 30 Ala Leu Thr Leu Ala Lys Arg Ala Arg Asp Glu Arg Glu Val Pro Val 35 40 45 Gly Ala Val Leu Val Leu Asn Asn Arg Val Ile Gly Glu Gly Trp Asn 50 55 60 Arg Ala Ile Gly Leu His Asp Pro Thr Ala His Ala Glu Ile Met Ala 65 70 75 80 Leu Arg Gln Gly Gly Leu Val Met Gln Asn Tyr Arg Leu Ile Asp Ala 85 90 95 Thr Leu Tyr Val Thr Phe Glu Pro Cys Val Met Cys Ala Gly Ala Met 100 105 110 Ile His Ser Arg Ile Gly Arg Val Val Phe Gly Val Arg Asn Ser Lys 115 120 125 Arg Gly Ala Ala Gly Ser Leu Met Asn Val Leu Asn Tyr Pro Gly Met 130 135 140 Asn His Arg Val Glu Ile Thr Glu Gly Ile Leu Ala Asp Glu Cys Ala 145 150 155 160 Ala Leu Leu Cys Asp Phe Tyr Arg Met Pro Arg Gln Val Phe Asn Ala 165 170 175 Gln Lys Lys Ala Gln Ser Ser Ile Asn Ser Gly Gly Ser Ser Gly Gly 180 185 190 Ser Ser Gly Ser Glu Thr Pro Gly Thr Ser Glu Ser Ala Thr Pro Glu 195 200 205 Ser Ser Gly Gly Ser Ser Gly Gly Ser Asp Lys Lys Tyr Ser Ile Gly 210 215 220 Leu Ala Ile Gly Thr Asn Ser Val Gly Trp Ala Val Ile Thr Asp Glu 225 230 235 240 Tyr Lys Val Pro Ser Lys Lys Phe Lys Val Leu Gly Asn Thr Asp Arg 245 250 255 His Ser Ile Lys Lys Asn Leu Ile Gly Ala Leu Leu Phe Asp Ser Gly 260 265 270 Glu Thr Ala Glu Ala Thr Arg Leu Lys Arg Thr Ala Arg Arg Arg Tyr 275 280 285 Thr Arg Arg Lys Asn Arg Ile Cys Tyr Leu Gln Glu Ile Phe Ser Asn 290 295 300 Glu Met Ala Lys Val Asp Asp Ser Phe Phe His Arg Leu Glu Glu Ser 305 310 315 320 Phe Leu Val Glu Glu Asp Lys Lys His Glu Arg His Pro Ile Phe Gly 325 330 335 Asn Ile Val Asp Glu Val Ala Tyr His Glu Lys Tyr Pro Thr Ile Tyr 340 345 350 His Leu Arg Lys Lys Leu Val Asp Ser Thr Asp Lys Ala Asp Leu Arg 355 360 365 Leu Ile Tyr Leu Ala Leu Ala His Met Ile Lys Phe Arg Gly His Phe 370 375 380 Leu Ile Glu Gly Asp Leu Asn Pro Asp Asn Ser Asp Val Asp Lys Leu 385 390 395 400 Phe Ile Gln Leu Val Gln Thr Tyr Asn Gln Leu Phe Glu Glu Asn Pro 405 410 415 Ile Asn Ala Ser Gly Val Asp Ala Lys Ala Ile Leu Ser Ala Arg Leu 420 425 430 Ser Lys Ser Arg Arg Leu Glu Asn Leu Ile Ala Gln Leu Pro Gly Glu 435 440 445 Lys Lys Asn Gly Leu Phe Gly Asn Leu Ile Ala Leu Ser Leu Gly Leu 450 455 460 Thr Pro Asn Phe Lys Ser Asn Phe Asp Leu Ala Glu Asp Ala Lys Leu 465 470 475 480 Gln Leu Ser Lys Asp Thr Tyr Asp Asp Asp Leu Asp Asn Leu Leu Ala 485 490 495 Gln Ile Gly Asp Gln Tyr Ala Asp Leu Phe Leu Ala Ala Lys Asn Leu 500 505 510 Ser Asp Ala Ile Leu Leu Ser Asp Ile Leu Arg Val Asn Thr Glu Ile 515 520 525 Thr Lys Ala Pro Leu Ser Ala Ser Met Ile Lys Arg Tyr Asp Glu His 530 535 540 His Gln Asp Leu Thr Leu Leu Lys Ala Leu Val Arg Gln Gln Leu Pro 545 550 555 560 Glu Lys Tyr Lys Glu Ile Phe Phe Asp Gln Ser Lys Asn Gly Tyr Ala 565 570 575 Gly Tyr Ile Asp Gly Gly Ala Ser Gln Glu Glu Phe Tyr Lys Phe Ile 580 585 590 Lys Pro Ile Leu Glu Lys Met Asp Gly Thr Glu Glu Leu Leu Val Lys 595 600 605 Leu Asn Arg Glu Asp Leu Leu Arg Lys Gln Arg Thr Phe Asp Asn Gly 610 615 620 Ser Ile Pro His Gln Ile His Leu Gly Glu Leu His Ala Ile Leu Arg 625 630 635 640 Arg Gln Glu Asp Phe Tyr Pro Phe Leu Lys Asp Asn Arg Glu Lys Ile 645 650 655 Glu Lys Ile Leu Thr Phe Arg Ile Pro Tyr Tyr Val Gly Pro Leu Ala 660 665 670 Arg Gly Asn Ser Arg Phe Ala Trp Met Thr Arg Lys Ser Glu Glu Thr 675 680 685 Ile Thr Pro Trp Asn Phe Glu Glu Val Val Asp Lys Gly Ala Ser Ala 690 695 700 Gln Ser Phe Ile Glu Arg Met Thr Asn Phe Asp Lys Asn Leu Pro Asn 705 710 715 720 Glu Lys Val Leu Pro Lys His Ser Leu Leu Tyr Glu Tyr Phe Thr Val 725 730 735 Tyr Asn Glu Leu Thr Lys Val Lys Tyr Val Thr Glu Gly Met Arg Lys 740 745 750 Pro Ala Phe Leu Ser Gly Glu Gln Lys Lys Ala Ile Val Asp Leu Leu 755 760 765 Phe Lys Thr Asn Arg Lys Val Thr Val Lys Gln Leu Lys Glu Asp Tyr 770 775 780 Phe Lys Lys Ile Glu Cys Phe Asp Ser Val Glu Ile Ser Gly Val Glu 785 790 795 800 Asp Arg Phe Asn Ala Ser Leu Gly Thr Tyr His Asp Leu Leu Lys Ile 805 810 815 Ile Lys Asp Lys Asp Phe Leu Asp Asn Glu Glu Asn Glu Asp Ile Leu 820 825 830 Glu Asp Ile Val Leu Thr Leu Thr Leu Phe Glu Asp Arg Glu Met Ile 835 840 845 Glu Glu Arg Leu Lys Thr Tyr Ala His Leu Phe Asp Asp Lys Val Met 850 855 860 Lys Gln Leu Lys Arg Arg Arg Tyr Thr Gly Trp Gly Arg Leu Ser Arg 865 870 875 880 Lys Leu Ile Asn Gly Ile Arg Asp Lys Gln Ser Gly Lys Thr Ile Leu 885 890 895 Asp Phe Leu Lys Ser Asp Gly Phe Ala Asn Arg Asn Phe Met Gln Leu 900 905 910 Ile His Asp Asp Ser Leu Thr Phe Lys Glu Asp Ile Gln Lys Ala Gln 915 920 925 Val Ser Gly Gln Gly Asp Ser Leu His Glu His Ile Ala Asn Leu Ala 930 935 940 Gly Ser Pro Ala Ile Lys Lys Gly Ile Leu Gln Thr Val Lys Val Val 945 950 955 960 Asp Glu Leu Val Lys Val Met Gly Arg His Lys Pro Glu Asn Ile Val 965 970 975 Ile Glu Met Ala Arg Glu Asn Gln Thr Thr Gln Lys Gly Gln Lys Asn 980 985 990 Ser Arg Glu Arg Met Lys Arg Ile Glu Glu Gly Ile Lys Glu Leu Gly 995 1000 1005 Ser Gln Ile Leu Lys Glu His Pro Val Glu Asn Thr Gln Leu Gln Asn 1010 1015 1020 Glu Lys Leu Tyr Leu Tyr Tyr Leu Gln Asn Gly Arg Asp Met Tyr Val 1025 1030 1035 1040 Asp Gln Glu Leu Asp Ile Asn Arg Leu Ser Asp Tyr Asp Val Asp His 1045 1050 1055 Ile Val Pro Gln Ser Phe Leu Lys Asp Asp Ser Ile Asp Asn Lys Val 1060 1065 1070 Leu Thr Arg Ser Asp Lys Asn Arg Gly Lys Ser Asp Asn Val Pro Ser 1075 1080 1085 Glu Glu Val Val Lys Lys Met Lys Asn Tyr Trp Arg Gln Leu Leu Asn 1090 1095 1100 Ala Lys Leu Ile Thr Gln Arg Lys Phe Asp Asn Leu Thr Lys Ala Glu 1105 1110 1115 1120 Arg Gly Gly Leu Ser Glu Leu Asp Lys Ala Gly Phe Ile Lys Arg Gln 1125 1130 1135 Leu Val Glu Thr Arg Gln Ile Thr Lys His Val Ala Gln Ile Leu Asp 1140 1145 1150 Ser Arg Met Asn Thr Lys Tyr Asp Glu Asn Asp Lys Leu Ile Arg Glu 1155 1160 1165 Val Lys Val Ile Thr Leu Lys Ser Lys Leu Val Ser Asp Phe Arg Lys 1170 1175 1180 Asp Phe Gln Phe Tyr Lys Val Arg Glu Ile Asn Asn Tyr His His Ala 1185 1190 1195 1200 His Asp Ala Tyr Leu Asn Ala Val Val Gly Thr Ala Leu Ile Lys Lys 1205 1210 1215 Tyr Pro Lys Leu Glu Ser Glu Phe Val Tyr Gly Asp Tyr Lys Val Tyr 1220 1225 1230 Asp Val Arg Lys Met Ile Ala Lys Ser Glu Gln Glu Ile Gly Lys Ala 1235 1240 1245 Thr Ala Lys Tyr Phe Phe Tyr Ser Asn Ile Met Asn Phe Phe Lys Thr 1250 1255 1260 Glu Ile Thr Leu Ala Asn Gly Glu Ile Arg Lys Arg Pro Leu Ile Glu 1265 1270 1275 1280 Thr Asn Gly Glu Thr Gly Glu Ile Val Trp Asp Lys Gly Arg Asp Phe 1285 1290 1295 Ala Thr Val Arg Lys Val Leu Ser Met Pro Gln Val Asn Ile Val Lys 1300 1305 1310 Lys Thr Glu Val Gln Thr Gly Gly Phe Ser Lys Glu Ser Ile Leu Pro 1315 1320 1325 Lys Arg Asn Ser Asp Lys Leu Ile Ala Arg Lys Lys Asp Trp Asp Pro 1330 1335 1340 Lys Lys Tyr Gly Gly Phe Asp Ser Pro Thr Val Ala Tyr Ser Val Leu 1345 1350 1355 1360 Val Val Ala Lys Val Glu Lys Gly Lys Ser Lys Lys Leu Lys Ser Val 1365 1370 1375 Lys Glu Leu Leu Gly Ile Thr Ile Met Glu Arg Ser Ser Phe Glu Lys 1380 1385 1390 Asn Pro Ile Asp Phe Leu Glu Ala Lys Gly Tyr Lys Glu Val Lys Lys 1395 1400 1405 Asp Leu Ile Ile Lys Leu Pro Lys Tyr Ser Leu Phe Glu Leu Glu Asn 1410 1415 1420 Gly Arg Lys Arg Met Leu Ala Ser Ala Gly Glu Leu Gln Lys Gly Asn 1425 1430 1435 1440 Glu Leu Ala Leu Pro Ser Lys Tyr Val Asn Phe Leu Tyr Leu Ala Ser 1445 1450 1455 His Tyr Glu Lys Leu Lys Gly Ser Pro Glu Asp Asn Glu Gln Lys Gln 1460 1465 1470 Leu Phe Val Glu Gln His Lys His Tyr Leu Asp Glu Ile Ile Glu Gln 1475 1480 1485 Ile Ser Glu Phe Ser Lys Arg Val Ile Leu Ala Asp Ala Asn Leu Asp 1490 1495 1500 Lys Val Leu Ser Ala Tyr Asn Lys His Arg Asp Lys Pro Ile Arg Glu 1505 1510 1515 1520 Gln Ala Glu Asn Ile Ile His Leu Phe Thr Leu Thr Asn Leu Gly Ala 1525 1530 1535 Pro Ala Ala Phe Lys Tyr Phe Asp Thr Thr Ile Asp Arg Lys Arg Tyr 1540 1545 1550 Thr Ser Thr Lys Glu Val Leu Asp Ala Thr Leu Ile His Gln Ser Ile 1555 1560 1565 Thr Gly Leu Tyr Glu Thr Arg Ile Asp Leu Ser Gln Leu Gly Gly Asp 1570 1575 1580 Ser Gly Gly Ser Lys Arg Thr Ala Asp Gly Ser Glu Phe Glu Pro Lys 1585 1590 1595 1600 Lys Lys Arg Lys Val 1605
Claims
1. A fusion protein, characterized in that: The fusion protein consists of a first region and a second region from the N-terminus to the C-terminus, wherein the first region is ABEs, wherein the ABEs are ABEmax or ABE8e, wherein the ABEs are adenine deaminase or an enzymatically active component thereof and nCas9; the second region is the e18 protein; the amino acid sequence of the ABEmax is shown in SEQ ID NO. 40; the amino acid sequence of the ABE8e is shown in SEQ ID NO. 41; and the nucleotide sequence of the e18 protein is shown in SEQ ID NO.
9.
2. An isolated polynucleotide, characterized in that: Encodes the fusion protein as claimed in claim 1.
3. A construct, characterized in that: The construct contains the polynucleotide according to claim 2.
4. An expression system, characterized in that: The expression system comprises the construct according to claim 3 or the exogenous polynucleotide according to claim 2 integrated into the genome.
5. A use, characterized in that: Use of the fusion protein according to claim 1, the polynucleotide according to claim 2, the construct according to claim 3, or the expression system according to claim 4 in gene editing, wherein the gene editing is to convert base A to G.
6. A base editing system, characterized in that: Comprising the fusion protein and sgRNA as claimed in claim 1.
7. A gene editing method, characterized in that: The method comprises performing gene editing using the fusion protein as described in claim 1 or the base editing system as described in claim 6, wherein the gene editing is to convert the base A into G.
Citation Information
Patent Citations
Adenine base editing tool and use thereof
CN110029096A