A method for analyzing DNA integration information in transgenic plants and its application

By screening and grouping the genome sequencing data of transgenic plants and cutting out the unmatched SCR parts, the problem of difficult efficient and accurate analysis of transgenic plant insertion situations and genome changes in existing technologies is solved, and efficient and accurate evaluation and detection are achieved.

CN114743595BActive Publication Date: 2025-09-23INST OF PLANT PROTECTION CHINESE ACAD OF AGRI SCI
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210384790.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-13
Publication Date
2025-09-23
Estimated Expiration
2042-04-13

Smart Images

  • Figure CN114743595B_ABST
    Figure CN114743595B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of biological information processing and discloses a method and use for analyzing DNA integration information in transgenic plants. The method comprises the following steps: comparing the genome sequence of the transgenic plant with the plant genome sequence and the transgenic vector sequence, respectively, and screening for SCRs that align to the plant genome and transgenic vector sequences; grouping the SCRs according to their SCR alignment positions; arranging the unaligned sequence portions of each SCR in each SCR group from longest to shortest according to sequence length; if a short sequence appears within a long sequence, the short sequence and the long sequence belong to the same SCR subgroup; accurately aligning the longest sequence portion in each SCR subgroup that does not align to the transgenic vector sequence with the plant genome and transgenic vector sequences, and obtaining transgenic insertion site information based on the SCR subgroup's alignment information to the transgenic vector and plant genome. The method of the present invention can evaluate the molecular characteristics of inserted fragments, including various insertion scenarios, in transgenic plants.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of biological information processing, and relates to a method and application for analyzing DNA integration information in transgenic plants. Background Art

[0002] Plant transgenic systems include Agrobacterium-mediated T-DNA transformation, gene gun-mediated transformation, protoplast fusion, pollen tube pathway transformation, etc.

[0003] Inserts in transgenic plants often occur in a complex pattern, such as truncated or concatenated fragments. Furthermore, insertions in the plant genome can manifest as small deletions, insertions, duplications, and chromosomal rearrangements. Consequently, conventional analysis methods struggle to analyze the various chromosomal changes induced by transgenes.

[0004] Agrobacterium-mediated T-DNA transformation is a core technology in plant genetic engineering, widely used in basic scientific research and genetic breeding. The insertion site and number of T-DNA fragments in the genome are random, and inserted T-DNA fragments often break through the left boundary (LB) and right boundary (RB) restrictions. Transgenic plants generated using other methods, including gene gun-mediated transformation, protoplast fusion, and pollen tube transformation, exhibit similar characteristics. Molecular characterization of transgenic plants (including T-DNA insertion site, insert copy number, insert sequence, and genomic changes) is a key component and core of transgenic plant safety assessment and an important indicator of transgenic plant performance. Traditionally, southern hybridization and fluorescent quantitative real-time PCR are used to assess T-DNA copy number, while PCR-based chromosome walking techniques (TAIL-PCR and Adapter-PCR) are often used to isolate T-DNA insertion sites. However, these methods are only suitable for transgenic plants with relatively simple T-DNA inserts and are complex and time-consuming.

[0005] In recent years, with the development of high-throughput sequencing technology and the reduction in sequencing time and cost, high-throughput sequencing has begun to be applied to evaluate the molecular characteristics of inserted T-DNA in transgenic plants. During the sequencing process, when reads spanning deletion sites and splice sites are mapped back to the genome, a single read is split into two segments, matching to different regions. Such reads are called soft-clipped reads (SCRs), and they are important for identifying chromosomal structural variations and foreign sequence integration. Mapping SCRs to plant genomes requires sequence alignment. Sequence alignment methods include exact and inexact alignment. Inexact alignment has the advantage of high speed, but the disadvantage is that it cannot find all matching positions in all cases. Exact alignment has the advantage of finding all matching positions in all cases, but the disadvantage is slow operation, which limits its practical application. In particular, when plant genome sequencing generates a large number of SCRs, locating SCRs through exact alignment is impractical.

[0006] Although high-throughput genome sequencing data from transgenic plants contains a wealth of molecular signatures of T-DNA insertions, accurately and efficiently extracting this information relies on mature bioinformatics algorithms. Currently, the only developed pipeline, TDNAscan, cannot accurately assess various T-DNA insertion scenarios and cannot evaluate changes in the plant genome caused by transgenics. Summary of the Invention

[0007] In view of this, the present invention proposes a method for evaluating DNA integration information characteristics such as various insertion situations of transgenic plants and changes caused by transgenic plant genomes.

[0008] According to one aspect of the present invention, a method for analyzing DNA integration information in transgenic plants is provided, comprising the following steps:

[0009] Step 1: Comparing the genome sequence of the transgenic plant with the genome sequence of the plant and the transgenic vector sequence, and then screening the SCRs that align to the plant genome and the transgenic vector sequence;

[0010] Step 2: SCRs are grouped according to their positions relative to the plant genome and the transgenic vector; the unaligned sequence portions of each SCR in each SCR group are arranged from longest to shortest according to sequence length; if a short sequence appears within a long sequence, the short sequence and the long sequence belong to the same SCR subgroup; otherwise, the short sequence becomes a new SCR subgroup;

[0011] Step 3: The longest sequence portion that is not aligned with the transgenic vector in each SCR subgroup in the SCR grouping that is aligned with the transgenic vector sequence is accurately aligned with the plant genome, and the longest sequence portion that is not aligned with the plant genome in each SCR subgroup in the SCR grouping that is aligned with the plant genome is accurately aligned with the transgenic vector sequence, and the transgenic insertion site information is obtained based on the information of the SCR subgroups aligned with the transgenic vector and the plant genome.

[0012] The present invention groups SCRs according to the positions of the SCRs compared to the plant genome and the transgenic vector, and further divides them into SCR subgroups according to the inclusion relationship of the unaligned sequence portion of each SCR in each SCR group, so that accurate alignment and conventional inaccurate matching (such as blast, bwa, bowtie, etc.) of the longest unaligned sequence portion of each SCR subgroup with the genome or transgenic vector becomes possible, greatly reducing the time for accurate alignment, thereby enabling the evaluation of characteristics such as various insertion situations of transgenic plants and changes in the genome of transgenic plants.

[0013] Preferably, in step 3, the longest sequence portion not mapped to the plant genome in each SCR subgroup within the SCR grouping mapped to the plant genome is accurately aligned to the plant genome sequence; if the longest sequence portion not mapped to the plant genome in the SCR subgroup is accurately aligned to the plant genome, information on the chromosomal rearrangement of the transgenic plant caused by the transgene is obtained based on the SCR subgroup information and the transgene insertion site information.

[0014] Preferably, in step 3, the longest sequence portion that is not aligned with the transgenic vector in each SCR subgroup in the SCR grouping aligned with the transgenic vector sequence is accurately aligned with the transgenic vector sequence; if the longest sequence portion that is not aligned with the transgenic vector in the SCR subgroup is accurately aligned with the transgenic vector, the transgenic vector fragment tandem insertion information is obtained based on the SCR subgroup information and the insertion site information.

[0015] Preferably, in step 1, SCRs containing two or more distinct segments are screened for subsequent analysis. Preferably, in step 1, non-transgenic SCRs are used as controls to screen for SCR groups unique to the transgenic plant. This screening step can significantly reduce the amount of data analyzed and significantly improve the accuracy of the analysis.

[0016] Preferably, in step 3, if the longest sequence portion of each SCR subgroup in the SCR grouping that is aligned to the transgenic vector sequence that is not aligned to the transgenic vector is accurately aligned to the plant genome, the full-length read of the SCR subgroup is aligned to the plant genome sequence, and if the full-length read of the SCR subgroup matches the transgenic vector sequence, the SCR subgroup is deleted; if the longest sequence portion of each SCR subgroup in the SCR grouping that is aligned to the genome that is not aligned to the plant group is accurately aligned to the transgenic vector, the full-length read of the SCR subgroup is aligned to the transgenic vector sequence, and if the full-length read of the SCR subgroup matches the transgenic vector sequence, the SCR subgroup is deleted. By searching for the full-length SCR, false positives caused by the common sequence of the plant genome and the transgenic vector can be deleted.

[0017] Preferably, in step 3, if a complete match is not achieved, the matching conditions allow for a mismatch of 1 to 5 nt; if a match is still not achieved, a 6 to 110 nt sequence is removed from the end of the unmatched sequence and the match is repeated. Therefore, the method of the present invention can detect single nucleotide variants (SNVs), small insertion-deletion mutations, and small DNA insertions at T-DNA insertion sites in plant genomes. The present invention addresses small DNA insertions during transgenic development by shearing the unmatched portion of the SCR.

[0018] Preferably, the precise sequence alignment is performed using a program such as ShortRead or Perl that can perform precise character / sequence queries.

[0019] Preferably, in step 1, genome alignment software such as bwa or bowtie2 that can use Local parameters is used.

[0020] Preferably, the transgenic plant is produced by Agrobacterium-mediated transformation and other transgenic methods such as gene gun-mediated transformation, protoplast fusion, pollen tube pathway transformation, etc.

[0021] Preferably, the transgenic vector is a vector required for transformation via Agrobacterium-mediated transformation and other transgenic methods such as gene gun-mediated transformation, protoplast fusion, pollen tube pathway transformation, etc.

[0022] Preferably, the transgenic insertion site information includes: T-DNA insertion position, T-DNA insertion information, SCR pattern, T-DNA insertion site chimeric sequence, and transgenic plant T-DNA insertion site sequencing coverage.

[0023] Preferably, the copy number of T-DNA is determined based on the coverage of T-DNA between the left and right boundaries of the SCR subgroup in the T-DNA insertion site information, the plant genome coverage, and the homozygous information of the T-DNA insertion in the transgenic plant.

[0024] Preferably, the transgenic plant is a plant whose genome has been sequenced, including but not limited to rice, Arabidopsis, rapeseed, tomato, soybean, cotton, cucumber, corn, wheat, etc.

[0025] According to another aspect of the present invention, the use of the method for identifying the genotype of transgenic plants is provided.

[0026] According to the method of the present invention, high-throughput sequencing reads are aligned to the plant genome and the transgenic vector, and SCRs are selected; then, when processing the unaligned portions of the SCRs, exact matching in the ShortRead program is used instead of conventional non-exact matching alignment software (such as blast, bwa, bowtie, etc.), and the time required for the program is reduced by grouping the SCRs; during the insertion of T-DNA into the genome, some small DNA fragments may be introduced, and the present invention handles this situation by shearing the unaligned portions of the SCRs; by searching the full-length SCRs, false positives caused by the common sequences of the plant genome and the transgenic vector are deleted; and by combining the SCRs belonging to the REF-REF grouping, the chromosomal rearrangement caused by the T-DNA insertion can be determined; the present invention outputs the genomic coverage of the T-DNA insertion site to confirm changes in the plant genome insertion site.

[0027] Other features and advantages of the present invention will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention, and the exemplary embodiments of the present invention and their descriptions are used to explain the present invention. In the accompanying drawings:

[0029] Figure 1 Flowchart of the method of the present invention.

[0030] Figure 2 Detailed information of the SCR sequence corresponding to one T-DNA insertion site is shown.

[0031] Figure 3 The T-DNA insertion site in the plant genome and the T-DNA fusion 2 kb sequence are shown.

[0032] Figure 4 The plant genome coverage at the T-DNA insertion site is shown.

[0033] Figure 5 The results of PCR detection of T-DNA insertion sites (TIS) are shown.

[0034] Figure 6 The diagram shows various insertion situations of T-DNA in transgenic rice detected by the present invention.

[0035] Figure 7 Describes the chromosomal rearrangements caused by T-DNA insertion in transgenic rice. A, Schematic diagram of chromosomal rearrangements at the T-DNA insertion site; B, SCR schematic diagram; C, Nanopore long-range sequencing verification.

[0036] Figure 8 Figure 3. T-DNA insertions leading to plant genome segment duplication. A, Plant genome segment duplication pattern and sequencing coverage; B, SCR pattern.

[0037] Figure 9 Showing perfect repair of the T-DNA insertion site in the plant genome. A, SCR pattern; B, Sequencing read distribution at the T-DNA insertion site.

[0038] Figure 10 A comparison of the analysis results of the method of the present invention (T-LOC) and TDNAscan is shown. DETAILED DESCRIPTION

[0039] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in each embodiment can be combined with each other.

[0040] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0041] The present invention proposes a method for analyzing DNA integration information in transgenic plants. Figure 1 The method of the present invention is described.

[0042] Example 1

[0043] The method for analyzing DNA integration information in transgenic plants according to the present invention comprises the following steps:

[0044] Step 1: Sequencing the genome of the transgenic plant;

[0045] Leaves or calli from wild-type and transgenic Kitaake rice T0 generations were thoroughly ground and lysed in 600 μL of 2× CTAB extract at 65°C for 45 min. Mixing was repeated 2-3 times, followed by extraction with 500 μL of chloroform and centrifugation at 14,000 rpm for 10 min. 400 μL of the supernatant was added to 400 μL of isopropanol and mixed again at 14,000 rpm for 10 min. The supernatant was discarded, resulting in a white DNA precipitate. The precipitate was washed with 70% ethanol, air-dried in a 37°C oven, and dissolved in 30 μL of ddH₂O. One μg of genomic DNA was then sent to Novogene Sequencing (https: / / cn.novogene.com / ) for whole-genome sequencing to generate high-throughput sequencing data (https: / / www.ncbi.nlm.nih.gov / geo / query / acc.cgi-acc=GSE185495).

[0046] Step 2: Download and install ShortRead 1.52.0 (https: / / bioconductor.org / packages / release / bioc / html / ShortRead.html) and bwa 0.7.17 (https: / / sourceforge.net / projects / bio-bwa / files / ) on Linux.

[0047] Step 3: Prepare the rice genome sequence (https: / / phytozome-next.jgi.doe.gov / info / OsativaKitaake_v3_1) and binary vector sequence (Sequence No. 1 in the sequence listing);

[0048] Step 4: Call bwa to align the genome sequencing data to the plant genome and vector (using default parameters), and then filter the SCRs aligned to the plant genome and vector;

[0049] Step 5: Remove redundant SCRs and group them according to their locations in the plant genome and vector. Remove groups with only a single SCR to avoid false positives, so that each group contains two or more SCRs with different fragments.

[0050] Steps 6-7: Processing the SCR groups from the alignment to the vector;

[0051] Step 6: Extract the sequences that are not aligned to the vector in each SCR group and arrange them from long to short according to sequence length. If a short sequence appears in a long sequence, the short sequence and the long sequence belong to the same SCR subgroup. Otherwise, the short sequence becomes a new SCR subgroup.

[0052] Step 7: Use ShortRead to search for a completely matching position in the plant genome and the vector for the longest unmatched vector sequence of each SCR subgroup (using default parameters). If not found, relax the matching conditions to allow 1-5nt mismatches. If still not found, remove the 6-110nt sequence at the end of the unmatched sequence and match again. If a matching position is found in the plant genome, the full-length sequencing read is searched for a matching position in the genome. If a match is found, the SCR subgroup is removed. If no match is found, the SCR subgroup is assigned to the TDNA-REF category, and the positions aligned to the vector and matched to the plant genome are marked. If a matching position is found in the vector, the SCR subgroup is assigned to the TDNA-TDNA category, and the positions aligned to the vector and matched to the vector are marked respectively; if no match is found in either the plant genome or the vector sequence, the SCR subgroup is marked as the SCR_T_N category;

[0053] Steps 8-10: Processing of SCR groups from alignment to plant genomes;

[0054] Step 8: Use the whole genome sequencing data of non-transgenic plants as a background control and remove the SCR groups with the same alignment positions as those in the background plants.

[0055] Step 9: Extract the sequences that are not mapped to the plant genome in each SCR group and arrange them from long to short according to sequence length. If a short sequence appears in a long sequence, the short sequence and the long sequence belong to the same SCR subgroup. Otherwise, the short sequence becomes a new SCR subgroup.

[0056] Step 10: Use ShortRead to search for a perfect match (using default parameters) in the plant genome and vector for the longest unaligned plant genome sequence of each SCR subgroup. If no match is found, relax the matching conditions to allow 1-5 nt mismatches. If no match is found, remove 6-110 nt of sequence at the end of the unaligned sequence and match again. If a match is found in the vector, the full-length sequencing read is searched for a match in the vector. If a match is found, the SCR subgroup is removed. If no match is found, the SCR subgroup is assigned to the TDNA-REF category and the positions aligned to the plant genome and the positions matched to the vector are marked. If a match is found in the plant genome, the SCR subgroup is assigned to the REF-REF category and the positions aligned to the plant genome and the positions matched to the plant genome are marked.

[0057] Step 11: Based on the SCR information in the TDNA-REF category, each SCR subgroup is assigned to the left and right insertion sites of each T-DNA on the genome, and combined with the SCR of REF-REF, the large chromosomal rearrangement of the T-DNA insertion site is located.

[0058] Step 12: Use R language to output the pattern diagram of each T-DNA insertion site, including T-DNA insertion position, T-DNA insertion information, SCR pattern diagram, T-DNA insertion site chimeric sequence, and transgenic plant T-DNA insertion site sequencing coverage.

[0059] Step 13: Output the detailed information of the SCR sequence corresponding to each T-DNA insertion site. Figure 2 Detailed information of the SCR sequence corresponding to one T-DNA insertion site is shown.

[0060] Step 14: Output the fused 2kb sequence of the plant genome and T-DNA spliced ​​at each T-DNA insertion site to provide information for subsequent genotype identification. Figure 3 A T-DNA insertion site in the plant genome and a 2 kb sequence of the T-DNA fusion are shown.

[0061] Step 15: Retrieve the reads generated by sequencing near the T-DNA insertion site and draw a coverage map of the sequencing data at this location to indicate the changes in the plant genome. Figure 4 Shown is the plant genome coverage near a T-DNA insertion site.

[0062] Step 16: Calculate the T-DNA copy number by the coverage of T-DNA between the left and right borders, the plant genome coverage, and the homozygosity of T-DNA insertion in transgenic plants.

[0063] Step 17: Detect the alignment of the paired sequencing reads of the SCR_T_N subgroup from Step 7. If the paired reads are derived from the plant genome or no alignment is found, it indicates that the SCR_T_N subgroup originated from an undetected insertion site in the transgenic plant, and the transgenic plant has an undetected insertion site. If the paired reads are derived from the vector sequence, it indicates that the SCR_T_N subgroup originated from the vector concatenation in the transgenic plant, and the transgenic plant does not have an undetected insertion site. Table 1 summarizes the results of the T-DNA insertion site information analysis of 48 T-DNA transgenic rice Kitaake using the method of the present invention. As can be seen from Table 1, no SCRs of the SCR_T_N subgroup were found in 40 T-DNA transgenic rice Kitaake (indicated by the " / " in the last column of the table), while SCRs of the SCR_T_N subgroup were found in 8 T-DNA transgenic rice Kitaake. The paired sequencing sequences of these SCRs aligned to the transgenic vector, indicating that the present invention can predict all T-DNA insertion sites.

[0064] Example 2 Insertion site PCR detection

[0065] In order to verify the T-DNA insertion site information output by the method of the present invention, three pairs of primers were designed for each of the 23 sites (sequences 17-154 in the sequence listing). Primer pair 1 is used to amplify the T-DNA insertion site on the left; primer pair 2 is used to amplify the T-DNA insertion site on the right; and primer pair 3 is used to amplify a genomic fragment without a T-DNA insertion site. The amplification templates are three plants (S1, S2, and S3) transformed with the same binary vector and Kitaake rice plants that have not been transformed with T-DNA. Since these are T0 generation transgenic rice, the T-DNA is heterozygously inserted. These three pairs of primers can all amplify bands in transgenic plants, while primer pair 1 and primer pair 2 can only amplify bands in the corresponding transgenic plants.

[0066] Figure 5 The results of PCR detection of T-DNA insertion sites are shown. The left image shows the amplified left border fusion sequence, the right image shows the right border fusion sequence, and the WT image shows the amplified genomic sequence of a plant without T-DNA insertion. The templates used in the figure are all from heterozygous transgenic plants with T-DNA insertion. The PCR amplification results of 23 primer pairs are consistent with the T-DNA insertion site predictions using the method of the present invention, demonstrating that the method of the present invention can accurately predict T-DNA insertions.

[0067] Example 3 Detection of multiple T-DNA insertions

[0068] The method of the present invention can detect a variety of T-DNA insertion situations. Figure 6The figure shows various T-DNA insertions in transgenic rice detected by the method of the present invention, wherein 49bTm_s2_TIS1 and 50DEP_s1_TIS1 are T-DNA insertion fragments that do not contain the hygromycin resistance gene.

[0069] Example 4 Detection of Chromosomal Rearrangements at Multiple T-DNA Insertion Sites

[0070] The method of the present invention can detect chromosomal rearrangements produced by T-DNA transgenes. Figure 7 This image shows a T-DNA insertion detected by the present method causing chromosomal rearrangement in transgenic rice. A, Chromosome rearrangement pattern at the T-DNA insertion site; B, SCR pattern; C, Nanopore long sequencing verification.

[0071] Example 5 Detection of genome deletion, duplication, and perfect repair

[0072] The method of the present invention can detect situations such as deletion, duplication and perfect repair of the genome. Figure 8 Showing small genomic fragment duplication at the T-DNA insertion site. A, Schematic diagram of plant genome fragment duplication and sequencing coverage. B, Schematic diagram of SCR.

[0073] Figure 9 This image shows perfect T-DNA repair in plant genomes. A, SCR pattern; B, Distribution of sequencing reads at the T-DNA insertion site. The plant genome exhibits no small deletions or duplications, and its genomic position remains unchanged.

[0074] Comparative Example 1 Comparison between the method of the present invention and TDNAscan

[0075] The results of the T-DNA transgene insert analysis using the present method were compared with those using TDNAscan (https: / / github.com / noble-research-institute / TDNAscan). The rice genome sequence and binary vector sequence described in Example 1, as well as the transgenic rice Kitaake sequencing data, were entered into TDNAscan. Figure 10The figure shows a comparison of the analysis results of the method of the present invention (T-Loc) and TDNAscan. A, T-LOC and TDNAscan have the same T-DNA insertion site, both outputting a complete T-DNA insertion site; B, T-LOC outputs a complete T-DNA insertion site, while TDNAscan outputs a T-DNA insertion site containing half the information; C, T-LOC outputs a T-DNA insertion site, while TDNAscan outputs two incomplete T-DNA insertion sites; D, T-LOC detects a complete T-DNA insertion site, while TDNAscan does not detect an insertion site. It can be seen that TDNAscan cannot correctly assess the T-DNA insertion status, cannot assess multiple T-DNA insertion situations, and cannot assess the changes caused to the transgenic plant genome, while the method of the present invention can detect them.

[0076] Table 1 compares the T-DNA insertion site information analysis results for 48 T-DNA transgenic rice strains, Kitaake, using the method of the present invention and TDNAscan. As can be seen from Table 1, the method of the present invention can predict all T-DNA insertion sites. Furthermore, the method of the present invention can predict complete TIS information, while TDNAscan significantly increases the number of TISs missing half or all of the TIS compared to the method of the present invention.

[0077] Table 1 T-DNA insertion site (TIS) information predicted by the method of the present invention and TDNAscan

[0078]

[0079]

[0080] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention. Sequence Listing <110> Institute of Plant Protection, Chinese Academy of Agricultural Sciences <120> A method for analyzing DNA integration information in transgenic plants and its application <141> 2022-04-12 <160> 139 <170> SIPOSequenceListing 1.0 <210> 1 <211> 18146 <212> DNA <213> Artificial Sequence <400> 1 gtaatcatgt catagctgtt tcctgtgtga aattgttatc cgctcacaat tccacacaac 60 atacgagccg gaagcataaa gtgtaaagcc tggggtgcct aatgagtgag ctaactcaca 120 ttaattgcgt tgcgctcact gcccgctttc cagtcgggaa acctgtcgtg ccagctgcat 180 taatgaatcg gccaacgcgc ggggagaggc ggtttgcgta ttggctagag cagcttgcca 240 acatggtgga gcacgacact ctcgtctact ccaagaatat caaagataca gtctcagaag 300 accaaagggc tattgagact tttcaacaaa gggtaatatc gggaaacctc ctcggattcc 360 attgcccagc tatctgtcac ttcatcaaaa ggacagtaga aaaggaaggt ggcacctaca 420 aatgccatca ttgcgataaa ggaaaggcta tcgttcaaga tgcctctgcc gacagtggtc 480 ccaaagatgg acccccaccc acgaggagca tcgtggaaaa agaagacgtt ccaaccacgt 540 cttcaaagca agtggattga tgtgaacatg gtggagcacg acactctcgt ctactccaag 600 aatatcaaag atacagtctc agaagaccaa agggctattg agacttttca acaaagggta 660 atatcgggaa acctctcgg attccattgc ccagctatct gtcacttcat caaaggaca 720 gtagaaaagg aaggtggcac ctacaaatgc catcattgcg ataaaggaaa ggctatcgtt 780 caagatgcct ctgccgacag tggtcccaaa gatggacccc cacccacgag gagcatcgtg 840 gaaaaaagaag acgttccaac cacgtcttca aagcaagtgg attgatgtga tatctccact 900 gacgtaaggg atgacgcaca atcccactat ccttcgcaag acccttctc tatataagga 960 agttcatttc atttggagag gacacgctga aatcaccagt ctctctctac aaatctatct 1020 ctctcgagct ttcgcagatc cggggggcaa tgagatatga aaaagcctga actcaccgcg 1080 acgtctgtcg agaagtttct gatcgaaaag ttcgacagcg tctccgacct gatgcagctc 1140 tcggagggcg aagaatctcg tgctttcagc ttcgatgtag gagggcgtgg atatgtcctg 1200 cgggtaaata gctgcgccga tggtttctac aaagatcgtt atgtttatcg gcactttgca 1260 tcggccgcgc tcccgattcc ggaagtgctt gacattgggg agtttagcga gagcctgacc 1320 tattgcatct cccgccgttc acagggtgtc acgttgcaag acctgcctga aaccgaactg 1380 cccgctgttc tacaaccggt cgcggaggct atggatgcga tcgctgcggc cgatcttagc 1440 cagacgagcg ggttcggccc attcggaccg caaggaatcg gtcaatacac tacatggcgt 1500 gatttcatat gcgcgattgc tgatccccat gtgtatcact ggcaaactgt gatggacgac 1560 accgtcagtg cgtccgtcgc gcaggctctc gatgagctga tgctttgggc cgaggactgc 1620 cccgaagtcc ggcacctcgt gcacgcggat ttcggctcca acaatgtcct gacggacaat 1680 ggccgcataa cagcggtcat tgactggagc gaggcgatgt tcggggattc ccaatacgag 1740 gtcgccaaca tcttcttctg gaggccgtgg ttggcttgta tggagcagca gacgcgctac 1800 ttcgagcgga ggcatccgga gcttgcagga tcgccacgac tccgggcgta tatgctccgc 1860 attggtcttg accaactcta tcagagcttg gttgacggca atttcgatga tgcagcttgg 1920 gcgcagggtc gatgcgacgc aatcgtccga tccggagccg ggactgtcgg gcgtacacaa 1980 atcgcccgca gaagcgcggc cgtctggacc gatggctgtg tagaagtact cgccgatagt 2040 ggaaaccgac gccccagcac tcgtccgagg gcaaagaaat agagatagag ccgaccggga 2100 tctgtcgatc gacaagctcg agtttctcca taataatgtg tgagtagttc ccagataagg 2160 gaattagggt tcctataggg tttcgctcat gtgttgagca tatagaaac ccttagtatg 2220 tatttgtatt tgtaaaatac ttctatcaat aaaatttcta attcctaaaa ccaaaatcca 2280 gtactaaaaat ccagatcccc cgaattaatt cggcgttaat tcagtacatt aaaaacgtcc 2340 gcaatgtgtt attaagttgt ctaagcgtca atttgtttac accaaatat atcctgccac 2400 cagccagcca acagctcccc gaccggcagc tcggcacaaa atcaccactc gatacaggca 2460 gcccatcagt ccgggacggc gtcagcggga gagccgttgt aaggcggcag actttgctca 2520 tgttaccgat gctattcgga agaacggcaa ctaagctgcc gggtttgaaa cacggatgat 2580 ctcgcggagg gtagcatgtt gattgtaacg atgacagagc gttgctgcct gtgatcaccg 2640 cggtttcaaa atcggctccg tcgatatactat gttatacgcc aactttgaaa acaactttga 2700 aaaagctgtt ttctggtatt taaggtttta gaatgcaagg aacagtgaat tggagttcgt 2760 cttgttataa ttagcttctt ggggtatctt taaatactgt agaaaagagg aaagaataa 2820 taaatggcta aaatgagaat atcaccggaa ttgaaaaaac tgatcgaaaa ataccgctgc 2880 gtaaaagata cggaaggaat gtctcctgct aaggtatata agctggtggg agaaaatgaa 2940 aacctatatt taaaaatgac ggacagccgg taaaaggga ccacctatga tgtggaacgg 3000 gaaaggaca tgatgctatg gctggaagga aagctgcctg ttccaaaggt cctgcacttt 3060 gaacggcatg atgggctggag caatctgctc atgagtgagg ccgatggcgt cctttgctcg 3120 gaagagtatg aagatgaaca aagccctgaa aagattatcg agctgtatgc ggagtgcatc 3180 aggctctttc actccatcga catatcggat tgtccctata cgaatagctt agacagccgc 3240 ttagccgaat tggattactt actgaataac gatctggccg atgtggattg cgaaaactgg 3300 gaagagaca ctccatttaa agatccgcgc gagctgtatg atttttaaa gacggaaaag 3360 cccgaagagg aacttgtctt ttcccacggc gacctgggag acagcaacat ctttgtgaaa 3420 3480 gacattgcct tctgcgtccg gtcgatcagg gaggatatcg gggaagaaca gtatgtcgag 3540 ctattttttg acttactggg gatcaagcct gattggggaga aaataaaata ttatatttta ctggatgaat tgttttagta cctagaatgc atgaccaaaa tcccttaacg ctgagagatc ccctcataat ttccccaaag cgtaaccatg tgtgaataaa ttttgagcta gtagggttgc agccacgagt aagtcttccc ttgttattgt gtagccagaa tgccgcaaaa cttccatgcc 3780. tagcgaact gttgagagta cgtttcgatt tctgactgtg ttagcctgga agtgcttgtc 3840 ccaaccttgt ttctgagcat gaacgcccgc aagccaacat gttagttgaa gcatcagggc gattagcagc atgatatcaa aacgctctga gctgctcgtt cggctatggc gtaggcctag tccgtaggca ggacttttca agtctcgga ggttctctca atctgcattc gcttcgaata gatattaaca agttgtttgg gtgttcgaat ttcaacaggt aagttagttg ctagaatcca tggctccttt gccgacgctg agtagtttt aggtgacggg tggtgacaat gagtccgtgt 4140 cgagcgctga ttttttcggc ctttagagcg agtttatac aatagaattt ggcatgagat tggattgctt ttagtcagcc tcttatagcc taaagtcttt gagtgactag atgacatatc atgtaagttg ctgataggtt tccagttttc cgctcctagg tctgcatatt gtacttttcc 4320 tcttactcga cttaaccagt accaacccag cttctcaacg gatttatacc atggcacttt 4380 aaagccagca tcactgacaa tgagcggtgt ggtgttactc ggtagaatgc tcgcaaggtc 4440 ggctagaaat tggtcatgag ctttctttga acattgctct gaaagcggga acgctttctc 4500 ataaagagta acagaacgac cgtgtagtgc gactgaagct cgcaatacca taagccgttt 4560 ttgctcacgg atatcagacc agtcaacaag tacaatgggc atcgtattgc ccgaacagat 4620 aaagctagca tgccaacggt atacagcgag tcgctctttg tggaggtgac gattacctaa 4680 caatcggtcg attcgtttga tgttatgttt tgttctcgct ttggttggca ggttacggcc 4740 aagttcggta agagtgagag ttttacagtc aagtaaggcg tggcaagcca acgttaagct 4800 gttgagtcgt tttaagtgta attcggggca gaattggtaa agagagtcgt gtaaaatatc 4860 gagttcgcac attttgttgt ctgattattg atttttggcg aaaccatttg atcatatgac 4920 aagatgtgta tctaccttaa cttaatgatt ttgataaaaa tcattagggg attcatcagc 4980 ccttaacgtg agttttcgtt ccactgagcg tcagaccccg tagaaaagat caaaggatct 5040 tcttgagatc ctttttttct gcgcgtaatc tgctgcttgc aaacaaaaaa accaccgcta 5100 ccagcggtgg tttgtttgcc ggatcaagag ctaccaactc tttttccgaa ggtaactggc 5160 ttcagcagag cgcagatacc aaatactgtc cttctagtgt agccgtagtt aggccaccac 5220 ttcaagaact ctgtagcacc gcctacatac ctcgctctgc taatcctgtt accagtggct 5280 gctgccagtg gcgataagtc gtgtcttacc gggttggact caagacgata gttaccggat 5340 aaggcgcagc ggtcgggctg aacggggggt tcgtgcacac agcccagctt ggagcgaacg 5400 acctacaccg aactgagata cctacagcgt gagctatgag aaagcgccac gcttcccgaa 5460 gggagaaagg cggacaggta tccggtaagc ggcagggtcg gaacaggaga gcgcacgagg 5520 gagcttccag ggggaaacgc ctggtatctt tatagtcctg tcgggtttcg ccacctctga 5580 cttgagcgtc gatttttgtg atgctcgtca ggggggcgga gcctatggaa aaacgccagc 5640 aacgcggcct ttttacggtt cctggccttt tgctggcctt ttgctcacat gttctttcct 5700 gcgttatccc ctgattctgt ggataaccgt attaccgcct ttgagtgagc tgataccgct 5760 cgccgcagcc gaacgaccga gcgcagcgag tcagtgagcg aggaagcgga agagcgcctg 5820 atgcggtatt ttctccttac gcatctgtgc ggtatttcac accgcatatg gtgcactctc 5880 agtacaatct gctctgatgc cgcatagtta agccagtata cactccgcta tcgctacgtg 5940 actgggtcat ggctgcgccc cgacacccgc caaccccgc tgacgcgcc tgacgggctt 6000 gtctgctccc ggcatccgct tacagacaag ctgtgaccgt ctccggggagc tgcatgtgtc 6060 agaggttttc accgtcatca ccgaaacgcg cgaggcaggg tgccttgatg tgggcgccgg 6120 cggtcgagtg gcgacggcgc ggcttgtccg cgccctggta gattgcctg ccgtaggcca 6180 gccattttg agcggccagc ggccgcgata ggccgacgcg aagcggcggg gcgtagggag 6240 cgcagcgacc gaagggtagg cgctttttgc agctcttcgg ctgtgcgctg gccagacagt 6300 tatgcacagg ccaggcgggt tttaagagtt ttaataagtt ttaaagagtt ttaggcggaa 6360 aaatcgcctt ttttctcttt tatatcagtc acttacatgt gtgaccggtt cccaatgtac 6420 ggctttgggt tcccaatgta cgggttccgg ttcccaatgt acggctttgg gttcccaatg 6480 tacgtgctat ccacaggaaa cagacctttt cagacctttt tcccctgcta gggcaatttg 6540 ccctagcatc tgctccgtac attaggaacc ggcggatgct tcgccctcga tcaggttgcg 6600 gtagcgcatg actaggatcg ggccagcctg ccccgcctcc tccttcaaat cgtactccgg 6660 caggtcattt gacccgatca gcttgcgcac ggtgaaacag aacttcttga actctccggc 6720 gctgccactg cgttcgtaga tcgtcttgaa caaccatctg gcttctgcct tgcctgcggc 6780 gcggcgtacc aggcggtaga gaaaacggcc gatgccggga tcgatcaaaa agtaatcggg 6840 gtgaaccgtc agcacgtccg ggttcttgcc ttctgtgatc tcgcggtaca tccaatcagc 6900 tagctcgatc tcgatgtact ccggccgccc ggtttcgctc tttacgatct tgtagcggct 6960 aatcaaggct tcaccctcgg ataccgtcac caggcggccg ttcttggcct tcttcgtacg 7020 ctgcatggca acgtgcgtgg tgtttaaccg aatgcaggtt tctaccaggt cgtctttctg 7080 ctttccgcca tcggctcgcc ggcagaactt gagtacgtcc gcaacgtgtg gacggaacac 7140 gcggccgggc ttgtctccct tcccttcccg gtatcggttc atggattcgg ttagatggga 7200 aaccgccatc agtaccaggt cgtaatccca cacactggcc atgccggccg gccctgcgga 7260 aacctctacg tgcccgtctg gaagctcgta gcggatcatc tcgccagctc gtcggtcacg 7320 cttcgacaga cggaaaacgg ccacgtccat gatgctgcga ctatcgcggg tgcccacgtc 7380 atagagcatc ggaacgaaaa aatctggttg ctcgtcgccc ttgggcggct tcctaatcga 7440 cggcgcaccg gctgccggcg gttgccggga ttctttgcgg attcgatcag cggccgcttg 7500 ccacgattca ccggggcgtg cttctgcctc gatgcgttgc cgctgggcgg cctgcgcggc 7560 cttcaacttc tccaccaggt catcacccag cgccgcgccg atttgtaccg ggccggttgg 7620 tttgcgaccg ctcacgccga ttcctcgggc ttgggggttc cagtgccatt gcagggccgg 7680 caggcaaccc agccgcttac gcctggccaa ccgcccgttc ctccacacat ggggcattcc 7740 acggcgtcgg tgcctggttg ttcttgattt tccatgccgc ctcctttagc cgctaaaatt 7800 catctactca tttattcatt tgctcattta ctctggtagc tgcgcgatgt attcagatag 7860 cagctcggta atggtcttgc cttggcgtac cgcgtacatc ttcagcttgg tgtgatcctc 7920 cgccggcaac tgaaagttga cccgcttcat ggctggcgtg tctgtcaggc tggccaacgt 7980 tgcagccttg ctgctgcgtg cgctcggacg gccggcactt agcgtgtttg tgcttttgct 8040 cattttctct ttacctcatt aactcaaata agttttgatt taatttcagc ggccagcgcc 8100 tggacctcgc gggcagcgtc gccctcgggt tctgattcaa gaacggttgt gccggcggcg 8160 gcagtgcctg ggtagctcac gcgctgcgtg atacgggact caagaatggg cagctcgtac 8220 ccggccagcg cctcggcaac ctcaccgccg atgcgcgtgc ctttgatcgc ccgcgacacg 8280 acaaaggccg cttgtagcct tccatccgtg acctcaatgc gctgcttaac cagctccacc 8340 aggtcggcgg tggcccatat gtcgtaaggg cttggctgca ccggaatcag cacgaagtcg 8400 gctgccttga tcgcggacac agccaagtcc gccgcctggg gcgctccgtc gatcactacg 8460 aagtcgcgcc ggccgatggc cttcacgtcg cggtcaatcg tcgggcggtc gatgccgaca 8520 acggttagcg gttgatcttc ccgcacggcc gcccaatcgc gggcactgcc ctggggatcg 8580 gaatcgacta acagaacatc ggccccggcg agttgcaggg cgcgggctag atgggttgcg 8640 atggtcgtct tgcctgaccc gcctttctgg ttaagtacag cgataacctt catgcgttcc 8700 ccttgcgtat ttgtttattt actcatcgca tcatatacgc agcgaccgca tgacgcaagc 8760 tgttttactc aaatacacat caccttttta gacggcggcg ctcggtttct tcagcggcca 8820 agctggccgg ccaggccgcc agcttggcat cagacaaacc ggccaggatt tcatgcagcc 8880 gcacggttga gacgtgcgcg ggcggctcga acacgtaccc ggccgcgatc atctccgcct 8940 cgatctcttc ggtaatgaaa aacggttcgt cctggccgtc ctggtgcggt ttcatgcttg 9000 ttcctcttgg cgttcattct cggcggccgc cagggcgtcg gcctcggtca atgcgtcctc 9060 acggaaggca ccgcgccgcc tggcctcggt gggcgtcact tcctcgctgc gctcaagtgc 9120 gcggtacagg gtcgagcgat gcacgccaag cagtgcagcc gcctctttca cggtgcggcc 9180 ttcctggtcg atcagctcgc gggcgtgcgc gatctgtgcc ggggtgaggg tagggcgggg 9240 gccaaacttc acgcctcggg ccttggcggc ctcgcgcccg ctccgggtgc ggtcgatgat 9300 tagggaacgc tcgaactcgg caatgccggc gaacacggtc aacaccatgc ggccggccgg 9360 cgtggtggtg tcggcccacg gctctgccag gctacgcagg cccgcgccgg cctcctggat 9420 gcgctcggca atgtccagta ggtcgcgggt gctgcgggcc aggcggtcta gcctggtcac 9480 tgtcacaacg tcgccagggc gtaggtggtc aagcatcctg gccagctccg ggcggtcgcg 9540 cctggtgccg gtgatcttct cggaaaatag cttggtgtag ccggccgcgt gcagttcggc 9600 ccgttggttg gtcaagtcct ggtcgtcggt gctgacgcgg ccatagccca ccaggccagc 9660 ggcggcgctc ttgttcatgg cgtaatgtct ccggttctag tcgcaagtat tctactttat 9720 gcgactaaaa cacgcgacaa gaaaacgcca ggaaaagggc agggcggcag cctgtcgcgt 9780 aacttaggac ttgtgcgaca tgtcgttttc agaagacggc tgcactgaac gtcagaagcc 9840 gactgcacta tagcagcgga ggggttggat caaagtactt tgatcccgag gggaaccctg 9900 tggttggcat ccacatacaa atggacgaac ggataagcct tttcacgccc ttttaaatat 9960 ccgattattc taataaacgc tcttttctct taggtttacc cgccaatata tcctgtcaaa 10020 cactgatagt ttaaactgaa ggcgggaaac gacaatctga tccaagctca agctgctcta 10080 gcattcgcca ttcaggctgc gcaactgttg ggaagggcga tcggtgcggg cctcttcgct 10140 attacgccag ctggcgaaag ggggatgtgc tgcaaggcga ttaagttggg taacgccagg 10200 gttttcccag ccacgacgtt gtaaaacgac ggccagtgcc atgctagaga cggggatcac 10260 aagtttgtac aaaaaagcag gctccaccat gggaaccaat tcagtcgact ggatccaagc 10320 ttaagaacga actaagccgg acaaaaaaag gagcacatat acaaaccggt tttattcatg 10380 aatggtcacg atggatgatg gggctcagac ttgagctacg aggccgcagg cgagagaagc 10440 ctagtgtgct ctctgcttgt ttgggccgta acggaggata cggccgacga gcgtgtacta 10500 ccgcgcggga tgccgctggg cgctgcgggg gccgttggat ggggatcggt gggtcgcggg 10560 agcgttgagg ggagacaggt ttagtaccac ctcgcctacc gaacaatgaa gaacccacct 10620 tataaccccg cgcgctgccg cttgtgttgc atgatacatc cctcagaagt tttagagcta 10680 gaaatagcaa gttaaaataa ggctagtccg ttatcaactt gaaaaagtgg caccgagtcg 10740 gtgctttttt ttgagatttc caaccaggtc cctggagccc atagtctagt aacggccgcc 10800 agtgtgctgg aattgccctt ggatcatgaa ccaacggcct ggctgtattt ggtggttgtg 10860 tagggagatg gggagaagaa aagcccgatt ctcttcgctg tgatgggctg gatgcatgcg 10920 ggggagcggg aggcccaagt acgtgcacgg tgagcggccc acagggcgag tgtgagcgcg 10980 agaggcggga ggaacagttt agtaccacat tgcccagcta actcgaacgc gaccaactta 11040 taaacccgcg cgctgtcgct tgtgtggaac atgaacagat tgatagtttt agagctagaa 11100 atagcaagtt aaaataaggc tagtccgtta tcaacttgaa aaagtggcac cgagtcggtg 11160 ctttttttgt cccttcgaag ggcaattctg cagatatcca tcacactggc ggccgctcga 11220 ggtcgacggt atcgataagc ttgatatcga attcgcggcc gcactcgaga tatctagacc 11280 cagctttctt gtacaaagtg gtgaagcttg catgcctgca gtgcagcgtg acccggtcgt 11340 gcccctctct agagataatg agcattgcat gtctaagtta taaaaaatta ccacatattt 11400 ttttgtcac acttgtttga agtgcagtttt atctatcttt atacatatat ttaaacttta 11460 ctctacgaat atataatct atagtactac ataatatca gtgttttaga gatcatata 11520 aatgacagt tagacatggtcaaaggaca attgagtatt ttgacacag gactctacag 11580 ttttattttt ttagtgtgca tgtgttctcc tttttttg caatagctt cacctatata 11640 atactcatc cattttatta gtacatccat ttaggtttta gggttaatgg ttttataga 11700 ctaattttt tagtacatct attttattct ctaattag aaaactaaaa 11760 ctctatttta gtttttttta ttaataattt agataaaa tagaaaa taagtgact 11820 aaaaattaaaaataccct ttaagaattt aaaaaaacta aggaacatt tttcttgttt 11880 cgagtagata atgccagcct gttaaacgcc gtcgacgagt ctaacggaca ccaaccagcg 11940 aaccagcagc gtcgcgtcgg gccaagcgaa gcagacggca cggcatctct gtcgctgcct 12000 ctggacccct ctcgagagtt ccgctccacc gttggacttg ctccgctgtc ggcatccaga 12060 aattgcgtgg cggagcggca gacgtgagcc ggcacggcag gcggctcct cctcctctca 12120 cggcaccggc agctacgggg gattcctttc ccaccgctcc ttcgctttcc cttcctcgcc 12180 cgccgtaata aatagacacc ccctccacac cctctttccc caacctcgtg ttgttcggag 12240 cgcacacaca cacaaccaga tctcccccaa atccacccgt cggcacctcc gcttcaaggt 12300 acgccgctcg tcctcccccc ccccccctct ctaccttctc tagatcggcg ttccggtcca 12360 tggttagggc ccggtagttc tacttctgtt catgtttgtg ttagatccgt gtttgtgtta 12420 gatccgtgct gctagcgttc gtacacggat gcgacctgta cgtcagacac gttctgattg 12480 ctaacttgcc agtgtttctc tttggggaat cctgggatgg ctctagccgt tccgcagacg 12540 ggatcgattt catgattttt tttgtttcgt tgcatagggt ttggtttgcc cttttccttt 12600 atttcaatat atgccgtgca cttgtttgtc gggtcatctt ttcatgcttt tttttgtctt 12660 ggttgtgatg atgtggtgtg gttgggcggt cgttcattcg ttctagatcg gagtagaata 12720 ctgtttcaaa ctacctggtg tatttattaa ttttggaact gtatgtgtgt gtcatacatc 12780 ttcatagtta cgagtttaag atggatggaa atatcgatct aggataggta tacatgttga 12840 tgtgggtttt actgatgcat atacatgatg gcatatgcag catctattca tatgctctaa 12900 ccttgagtac ctatctatta tataaacaa gtatgtttta taattatttt gatcttgata 12960 tacttggatg atggcatatg cagcagctat atgtggattt tttagccct gccttcatac 13020 gctatttatt tgcttggtac tgtctcttt gtcgatgctc accctgttgt ttggtgttac 13080 ttctgcaggt cgactctaga ggatctggaa ttcccgggta ccggatccat gtcagaagtc 13140 gagttctccc atgagtattg gatgaggcac gccctcactc ttgcgaagag ggccagggac 13200 gagaggagg tgccggtcgg tgctgtcctg gtcttgaata acagggtgat aggcgaaggt 13260 tggaacaggg ctattggcct tcatgaccct actgctcatg cggaaatcat ggcacttaga 13320 caggggggcc tcgttatgca aaattaccgc ctgatcgacg ccactcttta tgtcacattt 13380 gaaccatgtg ttatgtgtgc gggcgctatg atccattcac gcataggtcg cgtggtttt 13440 ggagttcgca acagtaaacg tggggctgca ggctctctga tgaacgtttt gaattatccg 13500 ggaatgaacc atagagtcga aatcacagaa gggattttgg cagacgaatg cgcggctctt 13560 ctttgtgatt tttacagaat gccccgccaa gtgtttaatg ctcaaaagaa agcgcagagt 13620 agcatcaact cggggggatc ttctgggggc tcgtctggtt ccgagactcc cggaacttcc 13680 gagtcggcaa cacctgaatc ctccggcggc tcttcgggcg gatctgacaa aaaatactca 13740 attggtctgg ctattgggac aaactctgtg ggctgggcgg taattaccga cgagtacaag 13800 gtgcctagta agaaatttaa agtgctcgga aacactgaca ggcactctat aaagaagaac 13860 ctgatcgggg cactgctttt cgactccgga gagacggcgg aggcgacgcg tctcaagcgt 13920 accgcgcgcc gcaggtacac aagaaggaag aataggatct gctacttgca ggaaatcttc 13980 agtaacgaga tggcgaaggt cgacgatagt ttctttcatc ggttggaaga atcgttcctc 14040 gtagaggagg acaaaaagca cgagcgtcac ccaatattcg ggaatattgt tgacgaggtt 14100 gcctaccatg agaaatatcc tacaatatat cacctccgta agaagcttgt cgattcaact 14160 gataaggctg atctcagact catctatctt gccctcgcac atatgattaa gtttcgtggc 14220 cacttcttga ttgaaggcga cctcaacccg gacaactcag atgttgacaa gctttttata 14280 cagctcgtcc agacatataa ccagctgttt gaagagaatc ccatcaatgc gagtggggtt 14340 gatgctaagg ccattttgtc cgccaggttg tccaaatctc gcagactgga aaacctgatc 14400 gcacagcttc ccggtgaaaa gaaaaacggg ctcttcggca atctcatcgc actgtccctc 14460 ggcctcaccc caaacttcaa gtctaacttc gacctggccg aggatgcgaa gctccagctg 14520 tcaaaagata catacgacga cgatttggac aatctgcttg cgcaaatagg cgaccagtat 14580 gcggacctgt tcctggctgc caaaaatctg tcagatgcaa tcctcctgtc cgatatattg 14640 cgtgtgaaca ccgaaatcac gaaggcaccg cttagcgcat ccatgatcaa gagatacgac 14700 gagcaccatc aggacctcac actcctcaag gcgcttgttc gtcagcagct tcccgagaaa 14760 tataaggaaa tttttttcga tcaaagcaag aatggatatg ctggctatat tgacggtggc 14820 gcttcgcagg aggagttcta taaattcatt aagccgattc tggagaagat ggacggaacg 14880 gaggagctcc tcgtcaagct taaccgggaa gacctgttgc ggaagcagag gactttgat 14940 aacggctcta ttccgcacca aatccatctg ggtgagttgc acgcaatctt gagagacaa 15000 gaggattct acccgttcct taggatac agagaga taggaaaaat actgaccttc 15060 aggataccat actatgtggg cccactggcg cgcggaata gtcgtttcgc atggatgact 15120 agaaagtccg aagaaacgat cacgccatgg aattttgagg aagtggtcga caagggcgcc 15180 tctgcccaga gcttcatcga aaggatgacc aattttgaca aaaatctgcc taacgaaag 15240 gtgcttccga agcacagcct gttgtatgaa tactcacag tttataacga gctcactaag 15300 gtcaagtacg tcacggagggg catgcgtaag cctgctttcc tgtctggtga aaaaaaag 15360 gcgattgtgg acctcctttt cagacgaac cgtaagtta ctgtgaagca actgaagag 15420 gattacttta agaaaattga gtgcttcgac agtgtggaga ttccggtgt cgaggaccgg 15480 tttaacgcca gcctgggtac gtatcatgac ctgcttaaaa tttacagga taagatttc 15540 ctggataatg aagagaacga agatatactg gaggacattg tgttgacttt gaccctcttc 15600 gaggacagag agatgattga ggaaagactg aagacctacg cacacctttt tgatgacaag 15660 gtcatgaaac aactcaagcg ccggcgctat actggctggg gccggctttc tcgcaagctc 15720 atcaatggga ttcgggataa gcaatcaggc aagacaattt tggacttcct caaatccgac 15780 ggattcgcaa ataggaattt tatgcagctg atacatgacg actctttgac attcaaagaa 15840 gacatacaga aggctcaggt ctccggccaa ggagattctt tgcacgagca tatcgctaac 15900 ttggcaggta gccccgccat aaaaaagggc attcttcaaa cggtaaaagt tgttgacgaa 15960 ctcgtgaagg tttgggccg tcataagccg gaaaacattg ttattgaaat ggctagggaa 16020 aatcagacga cccagaaggg acagaaaaat agcagggagc ggatgaagag aattgaagag 16080 ggaattaagg agcttggatc tcagattctt aaggagcacc ctgtggagaa cacccaactt 16140 cagaatgaaa agctctacct ttactacctt caaaacggcc gggatatgta cgtcgatcag 16200 gaacttgaca ttaaccggtt gagcgattat gacgttgacc atattgtgcc ccaatctttc 16260 cttaaagacg actctatcga caataaagtg ctgacgcgca gcgataaaaa tcgcggtaag 16320 tcggataatg tcccgtcgga agaggtggtt aaaaaaatga agaactattg gaggcaactc 16380 ctgaatgcca agctgatcac tcagaggaaa ttcgacaatc tcaccaaggc agaaaggggt 16440 ggacttagcg agctcgacaa ggccggtttt atcaaaagac agctggtgga gacacgccaa 16500 atcaccaaac acgttgccca gatcctggat tcgaggatga acacgaagta tgacgagaac 16560 gacaagttga ttagggaagt caaggtcatc actttgaagt ccaagctggt gagcgacttt 16620 cgcaaagact tccagtttta caaagtcagg gaaattaata actaccacca cgcccacgac 16680 gcctacctta acgccgtggt tggcacagca ctcatcaaga aataccctaa gctcgaatct 16740 gagttcgtct atggcgacta taaggtctac gacgttagaa aaatgatcgc gaaatctgag 16800 caggaaatag gcaaggcaac tgccaagtac ttcttctatt ccaatatcat gaactttttt 16860 aagacggaga ttaccctggc gaatggtgag atccgcaagc gccctttgat tgagacaaac 16920 ggagaaacag gagagatcgt atgggacaaa gggcggact ttgctactgt taggaaggtg 16980 ctctctatgc caaagttaa cattgtcaaa aaaactgaag tgcagacagg tgggtttagc 17040 aaaggaatcta tcctgccgaa gaggaactct gacaagctga tcgcccgcaa gaagattgg 17100 gatccgaaaa agtacggagg attcgactcc cccacagttg cgtactccgt gcttgtcgtg 17160 gccaaagtgg agaagggcaa gtctaagaag ctcaagagcg tcaaagagtt gttggggatc 17220 acgattatgg agcggtcgtc tttcgaaaag aatccgatag attttctcga ggccaagggt 17280 taaaagaag tcaagaagga tcttatcatc aagctcccta agtactccct ctttgagctt 17340 gaaaacggac ggaaaagaat gctggcttca gcgggtgaac ttcagaaggg taatgaactc 17400 gctctgccct caaaatatgt gaatttcctt tacctggcat cacactatga gaagcttaag 17460 gggtctccag agcaacga gcagaagcaa ctgttcgttg aacaacacaa gcactacctt 17520 17580 cttgataagg tccttagcgc ctacaacaag katatagagaca aacccatccg ggagcaggcc 17640 gagaacatta ttcatctctt caccttgacg aatcttgggg ccccggccgc gttcaagtac 17700 ttcgatacta ccatagacag aaagcgctat acatcgacaa aggaagttct tgacgccacg 17760 ctgatccacc aaagtataac aggcctctat gagacacgca tcgacctttc gcagttgggc 17820 ggtgaccgcc ccaaaaagaa gaggaaagtt ggcgggtgaa ctagtgagct cgaatttccc 17880 cgatcgttca aacatttggc aataaagttt cttaagattg aatcctgttg ccggtcttgc 17940 gatgattatc atataatttc tgttgaatta cgttaagcat gtaataatta acatgtaatg 18000 catgacgtta tttatgagat gggtttttat gattagagtc ccgcaattat acatttaata 18060 cgcgatagaa aacaaaatat agcgcgcaaa ctaggataaa ttatcgcgcg cggtgtcatc 18120 tatgttacta gatcgggaat taattc 18146 <210> 17 <211> 19 <212> DNA <213> Artificial Sequence <400> 17 ccaccaacat cgccggtat 19 <210> 18 <211> 20 <212> DNA <213> Artificial Sequence <400> 18 ggcctcgtag ctcaagtctg 20 <210> 19 <211> 20 <212> DNA <213> Artificial Sequence <400> 19 gatagtggaa accgacgccc 20 <210> 20 <211> twenty one <212> DNA <213> Artificial Sequence <400> 20 ggggactacg gggtcaaaat c 21 <210> twenty one <211> 19 <212> DNA <213> Artificial Sequence <400> twenty one ccaccaacat cgccggtat 19 <210> twenty two <211> twenty one <212> DNA <213> Artificial Sequence <400> twenty two ggggactacg gggtcaaaat c 21 <210> twenty three <211> 20 <212> DNA <213> Artificial Sequence <400> twenty three acacttgggc tgatggtagc 20 <210> twenty four <211> 20 <212> DNA <213> Artificial Sequence <400> twenty four gggcgtatat gctccgcatt 20 <210> 25 <211> 20 <212> DNA <213> Artificial Sequence <400> 25 ggcctcgtag ctcaagtctg 20 <210> 26 <211> 20 <212> DNA <213> Artificial Sequence <400> 26 gcagtatcga gcgtggtaca 20 <210> 27 <211> 20 <212> DNA <213> Artificial Sequence <400> 27 acacttgggc tgatggtagc 20 <210> 28 <211> 20 <212> DNA <213> Artificial Sequence <400> 28 gcagtatcga gcgtggtaca 20 <210> 29 <211> 20 <212> DNA <213> Artificial Sequence <400> 29 tctgatctgt agcccgacga 20 <210> 30 <211> 20 <212> DNA <213> Artificial Sequence <400> 30 cgtaatagcg aagaggcccg 20 <210> 31 <211> 20 <212> DNA <213> Artificial Sequence <400> 31 gatagtggaa accgacgccc 20 <210> 32 <211> 20 <212> DNA <213> Artificial Sequence <400> 32 atcaatggct gctcaaggca 20 <210> 33 <211> 20 <212> DNA <213> Artificial Sequence <400> 33 tctgatctgt agcccgacga 20 <210> 34 <211> 20 <212> DNA <213> Artificial Sequence <400> 34 atcaatggct gctcaaggca 20 <210> 35 <211> 20 <212> DNA <213> Artificial Sequence <400> 35 acgaacaaac ggggcagtta 20 <210> 36 <211> 20 <212> DNA <213> Artificial Sequence <400> 36 ctggaccgat ggctgtgtag 20 <210> 37 <211> 20 <212> DNA <213> Artificial Sequence <400> 37 cgtaatagcg aagaggcccg 20 <210> 38 <211> 19 <212> DNA <213> Artificial Sequence <400> 38 ggagggcttg agagccaac 19 <210> 39 <211> 20 <212> DNA <213> Artificial Sequence <400> 39 acgaacaaac ggggcagtta 20 <210> 40 <211> 19 <212> DNA <213> Artificial Sequence <400> 40 ggagggcttg agagccaac 19 <210> 41 <211> twenty one <212> DNA <213> Artificial Sequence <400> 41 gccacttttg cttgcattca g 21 <210> 42 <211> 20 <212> DNA <213> Artificial Sequence <400> 42 tacacaaatc gcccgcagaa 20 <210> 43 <211> 20 <212> DNA <213> Artificial Sequence <400> 43 cgtaatagcg aagaggcccg 20 <210> 44 <211> twenty two <212> DNA <213> Artificial Sequence <400> 44 ggcaatttag gatgcacctt gt 22 <210> 45 <211> twenty one <212> DNA <213> Artificial Sequence <400> 45 gccacttttg cttgcattca g 21 <210> 46 <211> twenty two <212> DNA <213> Artificial Sequence <400> 46 ggcaatttag gatgcacctt gt 22 <210> 47 <211> 20 <212> DNA <213> Artificial Sequence <400> 47 tcgggacagt ctagcatgga 20 <210> 48 <211> 20 <212> DNA <213> Artificial Sequence <400> 48 ggcctcgtag ctcaagtctg 20 <210> 49 <211> 20 <212> DNA <213> Artificial Sequence <400> 49 ctggaccgat ggctgtgtag 20 <210> 50 <211> 20 <212> DNA <213> Artificial Sequence <400> 50 ggaaaggcga cccagaatca 20 <210> 51 <211> 20 <212> DNA <213> Artificial Sequence <400> 51 tcgggacagt ctagcatgga 20 <210> 52 <211> 20 <212> DNA <213> Artificial Sequence <400> 52 ggaaaggcga cccagaatca 20 <210> 53 <211> 20 <212> DNA <213> Artificial Sequence <400> 53 tgtgtttgtg tagccgaacg 20 <210> 54 <211> 20 <212> DNA <213> Artificial Sequence <400> 54 gatagtggaa accgacgccc 20 <210> 55 <211> 20 <212> DNA <213> Artificial Sequence <400> 55 gatagtggaa accgacgccc 20 <210> 56 <211> 20 <212> DNA <213> Artificial Sequence <400> 56 ttgaaggccc agccaactac 20 <210> 57 <211> 20 <212> DNA <213> Artificial Sequence <400> 57 tgtgtttgtg tagccgaacg 20 <210> 58 <211> 20 <212> DNA <213> Artificial Sequence <400> 58 ttgaaggccc agccaactac 20 <210> 59 <211> 20 <212> DNA <213> Artificial Sequence <400> 59 aatgtcactg gtgttccccc 20 <210> 60 <211> 20 <212> DNA <213> Artificial Sequence <400> 60 ctggaccgat ggctgtgtag 20 <210> 61 <211> 20 <212> DNA <213> Artificial Sequence <400> 61 gatagtggaa accgacgccc 20 <210> 62 <211> 20 <212> DNA <213> Artificial Sequence <400> 62 caatgcacgc cctctcagta 20 <210> 63 <211> 20 <212> DNA <213> Artificial Sequence <400> 63 aatgtcactg gtgttccccc 20 <210> 64 <211> 20 <212> DNA <213> Artificial Sequence <400> 64 caatgcacgc cctctcagta 20 <210> 65 <211> twenty one <212> DNA <213> Artificial Sequence <400> 65 tctggaaatc aaggacgagc c 21 <210> 66 <211> 20 <212> DNA <213> Artificial Sequence <400> 66 ctggaccgat ggctgtgtag 20 <210> 67 <211> 20 <212> DNA <213> Artificial Sequence <400> 67 ctggaccgat ggctgtgtag 20 <210> 68 <211> 20 <212> DNA <213> Artificial Sequence <400> 68 cgcttccttt gttttgccgt 20 <210> 69 <211> twenty one <212> DNA <213> Artificial Sequence <400> 69 tctggaaatc aaggacgagc c 21 <210> 70 <211> 20 <212> DNA <213> Artificial Sequence <400> 70 cgcttccttt gttttgccgt 20 <210> 71 <211> twenty one <212> DNA <213> Artificial Sequence <400> 71 cggatgtgga ttcgtgattg c 21 <210> 72 <211> 20 <212> DNA <213> Artificial Sequence <400> 72 ctggaccgat ggctgtgtag 20 <210> 73 <211> 20 <212> DNA <213> Artificial Sequence <400> 73 ctggaccgat ggctgtgtag 20 <210> 74 <211> 20 <212> DNA <213> Artificial Sequence <400> 74 tggcctacgt accgtggaat 20 <210> 75 <211> twenty one <212> DNA <213> Artificial Sequence <400> 75 cggatgtgga ttcgtgattg c 21 <210> 76 <211> 20 <212> DNA <213> Artificial Sequence <400> 76 tggcctacgt accgtggaat 20 <210> 77 <211> 20 <212> DNA <213> Artificial Sequence <400> 77 tcaccaggca agcccataaa 20 <210> 78 <211> 20 <212> DNA <213> Artificial Sequence <400> 78 ctggaccgat ggctgtgtag 20 <210> 79 <211> 20 <212> DNA <213> Artificial Sequence <400> 79 ctggaccgat ggctgtgtag 20 <210> 80 <211> 19 <212> DNA <213> Artificial Sequence <400> 80 tagcaaactc tgcaggggc 19 <210> 81 <211> 20 <212> DNA <213> Artificial Sequence <400> 81 tcaccaggca agcccataaa 20 <210> 82 <211> 19 <212> DNA <213> Artificial Sequence <400> 82 tagcaaactc tgcaggggc 19 <210> 83 <211> twenty two <212> DNA <213> Artificial Sequence <400> 83 tgcattggac tgcaaacata at 22 <210> 84 <211> 20 <212> DNA <213> Artificial Sequence <400> 84 ctggaccgat ggctgtgtag 20 <210> 85 <211> 20 <212> DNA <213> Artificial Sequence <400> 85 ctggaccgat ggctgtgtag 20 <210> 86 <211> 20 <212> DNA <213> Artificial Sequence <400> 86 ggttcggtgt ggctgagtta 20 <210> 87 <211> twenty two <212> DNA <213> Artificial Sequence <400> 87 tgcattggac tgcaaacata at 22 <210> 88 <211> 20 <212> DNA <213> Artificial Sequence <400> 88 ggttcggtgt ggctgagtta 20 <210> 89 <211> 25 <212> DNA <213> Artificial Sequence <400> 89 tcatttagtg gaaaactctt catcc 25 <210> 90 <211> 20 <212> DNA <213> Artificial Sequence <400> 90 ctggaccgat ggctgtgtag 20 <210> 91 <211> 20 <212> DNA <213> Artificial Sequence <400> 91 ctggaccgat ggctgtgtag 20 <210> 92 <211> 20 <212> DNA <213> Artificial Sequence <400> 92 cccagcaaaa acctgacgtg 20 <210> 93 <211> 25 <212> DNA <213> Artificial Sequence <400> 93 tcatttagtg gaaaactctt catcc 25 <210> 94 <211> 20 <212> DNA <213> Artificial Sequence <400> 94 cccagcaaaa acctgacgtg 20 <210> 95 <211> 20 <212> DNA <213> Artificial Sequence <400> 95 ccaaactggg caatccaagc 20 <210> 96 <211> 20 <212> DNA <213> Artificial Sequence <400> 96 ctggaccgat ggctgtgtag 20 <210> 97 <211> 20 <212> DNA <213> Artificial Sequence <400> 97 ctggaccgat ggctgtgtag 20 <210> 98 <211> 20 <212> DNA <213> Artificial Sequence <400> 98 ccgctgataa actggcaagc 20 <210> 99 <211> 20 <212> DNA <213> Artificial Sequence <400> 99 ccaaactggg caatccaagc 20 <210> 100 <211> 20 <212> DNA <213> Artificial Sequence <400> 100 ccgctgataa actggcaagc 20 <210> 101 <211> 20 <212> DNA <213> Artificial Sequence <400> 101 caagcgagac cgttctgtga 20 <210> 102 <211> 20 <212> DNA <213> Artificial Sequence <400> 102 cgcgcggaaa tagtcgtttc 20 <210> 103 <211> 20 <212> DNA <213> Artificial Sequence <400> 103 ggcctcgtag ctcaagtctg 20 <210> 104 <211> 20 <212> DNA <213> Artificial Sequence <400> 104 aaagctctaa ggctgacgcc 20 <210> 105 <211> 20 <212> DNA <213> Artificial Sequence <400> 105 caagcgagac cgttctgtga 20 <210> 106 <211> 20 <212> DNA <213> Artificial Sequence <400> 106 aaagctctaa ggctgacgcc 20 <210> 107 <211> 20 <212> DNA <213> Artificial Sequence <400> 107 ctcctgcctg cttgatctgt 20 <210> 108 <211> 20 <212> DNA <213> Artificial Sequence <400> 108 ggcctcgtag ctcaagtctg 20 <210> 109 <211> 20 <212> DNA <213> Artificial Sequence <400> 109 cctggggtgc ctaatgagtg 20 <210> 110 <211> 20 <212> DNA <213> Artificial Sequence <400> 110 gcgtgtaccg tgtagtcaca 20 <210> 111 <211> 20 <212> DNA <213> Artificial Sequence <400> 111 ctcctgcctg cttgatctgt 20 <210> 112 <211> 20 <212> DNA <213> Artificial Sequence <400> 112 gcgtgtaccg tgtagtcaca 20 <210> 113 <211> 20 <212> DNA <213> Artificial Sequence <400> 113 ccccgtagtg gttttgacga 20 <210> 114 <211> 20 <212> DNA <213> Artificial Sequence <400> 114 ggcctcgtag ctcaagtctg 20 <210> 115 <211> 20 <212> DNA <213> Artificial Sequence <400> 115 tctcctgcac gtatccctca 20 <210> 116 <211> 20 <212> DNA <213> Artificial Sequence <400> 116 gcccggtaca cctacttgtt 20 <210> 117 <211> 20 <212> DNA <213> Artificial Sequence <400> 117 ccccgtagtg gttttgacga 20 <210> 118 <211> 20 <212> DNA <213> Artificial Sequence <400> 118 gcccggtaca cctacttgtt 20 <210> 119 <211> twenty three <212> DNA <213> Artificial Sequence <400> 119 accacgtcat accaataaca cca 23 <210> 120 <211> 20 <212> DNA <213> Artificial Sequence <400> 120 ctggaccgat ggctgtgtag 20 <210> 121 <211> 20 <212> DNA <213> Artificial Sequence <400> 121 aaacacgttg cccagatcct 20 <210> 122 <211> 20 <212> DNA <213> Artificial Sequence <400> 122 gccttagcga gcctcctatc 20 <210> 123 <211> twenty three <212> DNA <213> Artificial Sequence <400> 123 accacgtcat accaataaca cca 23 <210> 124 <211> 20 <212> DNA <213> Artificial Sequence <400> 124 gccttagcga gcctcctatc 20 <210> 125 <211> twenty two <212> DNA <213> Artificial Sequence <400> 125 acagtacaca tggttgctgg tt 22 <210> 126 <211> 20 <212> DNA <213> Artificial Sequence <400> 126 tacacaaatc gcccgcagaa 20 <210> 127 <211> 20 <212> DNA <213> Artificial Sequence <400> 127 tgtcgaggac cggtttaacg 20 <210> 128 <211> 20 <212> DNA <213> Artificial Sequence <400> 128 gttcgttcaa atgccccagc 20 <210> 129 <211> twenty two <212> DNA <213> Artificial Sequence <400> 129 acagtacaca tggttgctgg tt 22 <210> 130 <211> 20 <212> DNA <213> Artificial Sequence <400> 130 gttcgttcaa atgccccagc 20 <210> 131 <211> 20 <212> DNA <213> Artificial Sequence <400> 131 atcgcgatgc agatcgaaca 20 <210> 132 <211> 20 <212> DNA <213> Artificial Sequence <400> 132 ctagccaata cgcaaaccgc 20 <210> 133 <211> 20 <212> DNA <213> Artificial Sequence <400> 133 ctggaccgat ggctgtgtag 20 <210> 134 <211> 20 <212> DNA <213> Artificial Sequence <400> 134 tgcaattcga cgtggagcta 20 <210> 135 <211> 20 <212> DNA <213> Artificial Sequence <400> 135 atcgcgatgc agatcgaaca 20 <210> 136 <211> 20 <212> DNA <213> Artificial Sequence <400> 136 tgcaattcga cgtggagcta 20 <210> 137 <211> 20 <212> DNA <213> Artificial Sequence <400> 137 tcagaaatca ccagcgctca 20 <210> 138 <211> 20 <212> DNA <213> Artificial Sequence <400> 138 ctatccgccc gaagaggaac 20 <210> 139 <211> 20 <212> DNA <213> Artificial Sequence <400> 139 ggcctcgtag ctcaagtctg 20 <210> 140 <211> 20 <212> DNA <213> Artificial Sequence <400> 140 caaatgtggc accacaccac 20 <210> 141 <211> 20 <212> DNA <213> Artificial Sequence <400> 141 tcagaaatca ccagcgctca 20 <210> 142 <211> 20 <212> DNA <213> Artificial Sequence <400> 142 caaatgtggc accacaccac 20 <210> 143 <211> 20 <212> DNA <213> Artificial Sequence <400> 143 atccaatttg ccgagacgga 20 <210> 144 <211> 20 <212> DNA <213> Artificial Sequence <400> 144 aacacggggg actctaggaa 20 <210> 145 <211> 20 <212> DNA <213> Artificial Sequence <400> 145 aaccacggcg ttaaggtagg 20 <210> 146 <211> 20 <212> DNA <213> Artificial Sequence <400> 146 acttggggtg agaacgcaaa 20 <210> 147 <211> 20 <212> DNA <213> Artificial Sequence <400> 147 atccaatttg ccgagacgga 20 <210> 148 <211> 20 <212> DNA <213> Artificial Sequence <400> 148 acttggggtg agaacgcaaa 20 <210> 149 <211> 20 <212> DNA <213> Artificial Sequence <400> 149 caacagcaac cgcatcacat 20 <210> 150 <211> 20 <212> DNA <213> Artificial Sequence <400> 150 gtgctgctag cgttcgtaca 20 <210> 151 <211> 20 <212> DNA <213> Artificial Sequence <400> 151 caggcgagag aagcctagtg 20 <210> 152 <211> 20 <212> DNA <213> Artificial Sequence <400> 152 ggctagtgaa cagcacagca 20 <210> 153 <211> 20 <212> DNA <213> Artificial Sequence <400> 153 caacagcaac cgcatcacat 20 <210> 154 <211> 20 <212> DNA <213> Artificial Sequence <400> 154 ggctagtgaa cagcacagca 20

Claims

1. A method for analyzing DNA integration information in transgenic plants, characterized in that: The steps include: Step 1: aligning the genome sequence of the transgenic plant with the genome sequence of the plant and the transgenic vector sequence, and then screening SCRs aligned to the plant genome and the transgenic vector sequence; Step 2: SCRs are grouped according to the positions of the SCRs aligned to the plant genome and the transgenic vector; the unaligned sequence portions of each SCR in each SCR group are arranged from longest to shortest according to sequence length; if a short sequence appears within a long sequence, the short sequence and the long sequence belong to the same SCR subgroup; otherwise, the short sequence becomes a new SCR subgroup; Step 3: The longest sequence portion that is not aligned to the transgenic vector in each SCR subgroup in the SCR grouping aligned to the transgenic vector sequence is accurately aligned to the plant genome, the longest sequence portion that is not aligned to the plant genome in each SCR subgroup in the SCR grouping aligned to the plant genome is accurately aligned to the transgenic vector sequence, and transgenic insertion site information is obtained based on information on alignment of the SCR subgroups to the transgenic vector and the plant genome; accurately aligning the longest sequence portion not aligned to the plant genome in each SCR subgroup in the SCR group aligned to the plant genome with the plant genome sequence; if the longest sequence portion not aligned to the plant genome in the SCR subgroup is accurately aligned to the plant genome, obtaining information on chromosome rearrangement of the transgenic plant caused by the transgene based on the SCR subgroup information and the transgene insertion site information; Accurately aligning the longest sequence portion not aligned to the transgenic vector in each SCR subgroup in the SCR group aligned to the transgenic vector sequence with the transgenic vector sequence; If the longest sequence portion in the SCR subgroup that is not aligned to the transgenic vector is accurately aligned to the transgenic vector, the transgenic vector fragment tandem insertion information is obtained based on the SCR subgroup information and the insertion site information.

2. The method according to claim 1, characterized in that In step 1, SCR groups containing two or more different fragments are screened, and SCR groups unique to the transgenic plants relative to non-transgenic plants are screened for analysis in step 2.

3. The method according to claim 1, characterized in that In step 3 If the longest sequence portion that is not aligned to the transgenic vector in each SCR subgroup in the SCR grouping that is aligned to the transgenic vector sequence is accurately aligned to the plant genome, the full-length read of the SCR subgroup is aligned to the plant genome sequence, and if the full-length read of the SCR subgroup matches the transgenic vector sequence, the SCR subgroup is deleted; if the longest sequence portion that is not aligned to the plant group in each SCR subgroup in the SCR grouping that is aligned to the genome is accurately aligned to the transgenic vector, the full-length read of the SCR subgroup is aligned to the transgenic vector sequence, and if the full-length read of the SCR subgroup matches the transgenic vector sequence, the SCR subgroup is deleted.

4. The method according to claim 1, wherein In step 3, if the exact alignment does not result in a complete match, the exact alignment parameters allow for a mismatch of 1 to 5 nt; if there is still no match, a 6 to 110 nt sequence is removed from the end of the unaligned sequence and the exact alignment is performed again.

5. The method according to claim 1, wherein In step 1, bwa or bowtie2 software is used for alignment, and in step 3, ShortRead or Perl is used for accurate sequence alignment.

6. The method according to claim 1, characterized in that The transgenic method includes producing transgenic plants through Agrobacterium-mediated transformation; the transgenic vector includes a desired vector transformed through Agrobacterium-mediated transformation; the transgenic insertion site information includes: T-DNA insertion position, SCR pattern, T-DNA insertion site chimeric sequence, transgenic plant T-DNA insertion position sequencing coverage and T-DNA copy number.

7. The method according to claim 1, characterized in that The transgenic plants are plants whose genomes have been sequenced, including rice, Arabidopsis, rapeseed, tomato, soybean, cotton, cucumber, corn or wheat.

8. Use of the method according to any one of claims 1 to 7 in identifying the genotype of transgenic plants.

Citation Information

Patent Citations

  • Method for positioning T-DNA insertion site by flag tag of sam file

    CN110349624A