A primer for detecting huck transposon whole genome site based on second-generation sequencing and a method and application thereof
By designing high-throughput, high-efficiency primer combinations and next-generation sequencing technology, the problems of low detection sites and high cost of Huck transposon sites in existing technologies have been solved, realizing efficient and low-cost whole-genome Huck transposon site detection, which is suitable for agricultural breeding.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-19
- Publication Date
- 2026-04-07
AI Technical Summary
Existing methods for detecting Huck transposons suffer from problems such as low locus count, low marker count, high cost, and low efficiency in agricultural breeding, which limit the development of agricultural breeding.
Using a high-throughput, high-efficiency, and low-cost primer combination, combined with next-generation sequencing technology, maize genomic DNA was fragmented, specific adapters were ligated, and PCR amplification and sequencing were performed to detect Huck transposon sites throughout the genome.
It enables the enrichment of 1/4-1/2 transposon sites at the whole genome level, efficiently isolates individual sample sites, reduces library preparation and sequencing costs, is suitable for high-throughput detection, and is applicable to hybrid sequencing.
Smart Images

Figure CN117887881B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of molecular biology technology, and provides primers, methods and applications for detecting Huck transposon whole genome sites based on next-generation sequencing. Background Technology
[0002] Long terminal repeat (LTR) retrotransposons are a class of mobile DNA sequences that are ubiquitous in eukaryotic genomes. They replicate themselves in the genome through a "copy and paste" mechanism mediated by RNA.
[0003] In higher plants, many active LTR retrotransposons have been extensively studied and applied to molecular marker techniques, gene tags, insertion mutation analysis, and gene function analysis. Currently, more than ten types of active LTR retrotransposons have been identified in various plants.
[0004] Huck is one of the most important LTR transposons discovered in maize, measuring 9439 bp in length. It contains two gag regions, with the coding region structured as gag-gag-pr-int-rt-rh. This transposon exhibits a nearly even distribution across the entire genome, with a copy number of approximately 30,000. The copy number of Huck is generally consistent across different maize varieties, demonstrating good coverage and making it suitable for developing genotyping tools such as molecular markers and liquid-phase microarrays, showing great application potential.
[0005] In agricultural breeding, it is necessary to perform whole-genome loci scanning on parents or offspring, utilize genotypic differences to screen for superior haplotypes, and then develop individual molecular markers for application in breeding. However, existing detection methods, including SSR markers, InDel markers, multiplex PCR, and KASP markers, suffer from low marker counts (generally less than 1000); while using resequencing to determine random sequences leads to low development efficiency and excessive costs. All of these technical methods limit the development of the field of agricultural breeding. Summary of the Invention
[0006] One objective of this invention is to provide a method for capturing more than 10 sites at the whole genome level. 4 Furthermore, the Huck transposon has a fixed insertion site. A second objective of this invention addresses the shortcomings of existing technologies by providing a method for detecting the insertion site of a Huck transposon. To achieve the above objective, this invention adopts the following technical solution:
[0007] In a first aspect, the present invention provides a high-throughput, high-efficiency, and low-cost primer combination for capturing Huck transposon sites, the nucleotide sequence of which is shown in SEQ ID NO.1-3.
[0008] Secondly, the present invention provides a method for detecting whole-genome Huck transposon sites based on next-generation sequencing, comprising: breaking maize genomic DNA into fragments, using the nucleotide sequence shown in SEQ ID NO.4 as adapter 1 and the nucleotide sequence shown in SEQ ID NO.5 as adapter 2, and ligating adapter 1 and adapter 2 to the fragmented DNA fragments to obtain DNA fragments with adapters;
[0009] Using the DNA fragment with the adapter as a template, PCR amplification was performed using the above primer combination. The PCR amplification product was sequenced and compared with the genome sequence to obtain the corresponding Huck site.
[0010] In the method provided by this invention, the mass ratio of adapter 1 and adapter 2 to the fragmented DNA fragment is 10-15 pM: 50-55 ng.
[0011] In the method provided by this invention, the DNA fragments broken are between 250bp and 5kb.
[0012] As a specific embodiment of the present invention, the method provided by the present invention includes:
[0013] (1) Take maize genomic DNA and break it with DNA breaking enzyme NlaIII. The breaking conditions are 37℃ for 1-5 hours to obtain the broken product. The broken DNA fragments are between 250bp and 5kb.
[0014] (2) For the fragmented product obtained in step (1), use DNA ligase to ligate the fragmented product with adapter 1 and adapter 2 to obtain DNA fragments with adapters.
[0015] (3) The DNA fragment with adapter from step (2) is recovered, and the recovered product is then subjected to PCR amplification using the above-mentioned specific primer combination.
[0016] (4) The amplification product from step (3) was recovered, quantified, sequenced using an NGS sequencer, and compared with the genome sequence to obtain the corresponding Huck transposon sites.
[0017] In step (2), the ligase is T4 DNA ligase, and the ligation conditions are 22-25℃ for 1-3 hours; the PCR amplification conditions in step (3) are: 98℃ for 2 minutes of pre-denaturation; 98℃ for 10 seconds, 62℃ for 20 seconds, 72℃ for 1 minute, 15 cycles; extension at 72℃ for 8 minutes.
[0018] As understood by those skilled in the art, the present invention also provides the application of the above primer combinations or the above methods in separating Huck site sequences.
[0019] Furthermore, the application of the aforementioned primer combinations or methods in the identification of Huck transposons at the genome level. And the application of the aforementioned primer combinations or methods in the screening of Huck transposons at the genome level.
[0020] Considering that Huck is the most important LTR transposon in maize, the genome provided in this invention is the whole maize genome.
[0021] The beneficial effects of this invention are as follows:
[0022] The method provided by this invention can enrich 1 / 4 to 1 / 2 of transposon sites on the genome. For example, in the Zheng 58 genome, more than 20,000 Huck sites were detected, demonstrating high efficiency. Furthermore, because the primers contain tag sequences, the method provided by this invention can separate sites from large-scale data after mixed sequencing, making it suitable for high-throughput detection. Moreover, the use of mixed PCR and sequencing in this invention significantly reduces the cost of library preparation and sequencing. Attached Figure Description
[0023] Figure 1 This invention provides a common sequence for the 3' end alignment analysis of different Huck family members in the maize genome.
[0024] Figure 2 This invention provides a common sequence for 5' end alignment analysis of different Huck family members in the maize genome.
[0025] Figure 3 This invention is based on the Illumina sequencing platform and the sequence structure of the sequencing sample, wherein the sequences labeled Huck1-inner and Huck2-inner are the 5' and 3' of the Huck transposon, respectively.
[0026] Figure 4 This is a quality control diagram for the NGS library construction of the Huck transposon of this invention.
[0027] Figure 5 This is the predicted distribution of the Huck transposon in the maize B73 genome according to the present invention.
[0028] Figure 6 This is the distribution of the Huck transposon in the maize DH-1 genome as measured in this invention.
[0029] Figure 7 This is the distribution of the Huck transposon in the maize K17 genome as measured in this invention.
[0030] Figure 8 This is the distribution of the Huck transposon in the maize Zheng58 genome as measured in this invention. Detailed Implementation
[0031] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of this invention, not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0032] The terms “comprising” or “including” in this invention are open-ended descriptions that include the specified ingredients or steps described, as well as other specified ingredients or steps that do not materially affect them.
[0033] The endpoints and any values of the ranges disclosed herein are not limited to the precise ranges or values, and these ranges or values should be understood to include values close to these ranges or values. For numerical ranges, the endpoint values of the various ranges, the endpoint values of the various ranges and individual point values, and individual point values can be combined with each other to obtain one or more new numerical ranges, which should be considered as specifically disclosed herein.
[0034] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "specific implementation," or "some specific implementations," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0035] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of this invention, not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0036] Where specific techniques or conditions are not specified in the examples, they shall be performed in accordance with the techniques or conditions described in the literature in this field, or in accordance with the product instructions. Reagents or instruments whose manufacturers are not specified are all conventional products that can be purchased through legitimate channels.
[0037] Example 1: Design and Testing of Huck-Specific Primers
[0038] Example 1 of this invention provides the design and testing of specific primers for detecting Huck transposons in maize, and the method steps are as follows:
[0039] This embodiment uses the Huck sequences published in the appendix of the 2021 scientific paper "De novo assembly, annotation, and comparative analysis of 26 diverse maize genomes" (https: / / www.science.org / doi / full / 10.1126 / science.abg5289) as references (but not limited to) dozens of Huck transposon family members as preliminary target sites.
[0040] The above sequences were scanned and aligned in the maize B73 genome. The top 20 most abundant Huck transposon family members were selected to ensure a sufficient number of target sequences for detection. The screening results include huck_AC186577_1525#LTR, huck_AC186603_1556#LTR, huck_AC186656_1609#LTR, huck_AC190900_2713#LTR, huck_AC191259_3001#LTR, huck_AC193313_3542#LTR, huck_AC194973_4393#LTR, huck_AC195575_4652#LTR, huck_AC208546_9913#LTR, and huck_AC208842_10038#LTR. LTR, huck_AC210079_10574#LTR, huck_AC210804_10865#LTR, huck_AC212331_11708#LTR, huck_AC213042_12069#LTR , huck_AC213612_12218#LTR, huck_AC213612_12218#LTR, huck_AC214833_12913#LTR, huck_AC216048_13250#LTR and so on.
[0041] To avoid problems such as primer dimer contamination, primer binding to non-specific sites on the genome, and incomplete matching between primers and target Huc k transposon members during PCR, different screening conditions need to be set when designing detection primers. These primers need to be screened and adjusted one by one, and optimized and improved through multiple experiments.
[0042] The primer design process is as follows:
[0043] (1) By performing alignment analysis, the common sequence characteristics at the ends of the Huck sequence were obtained, and the conserved sequences at the 3' and 5' ends were extracted accordingly. This step is significant because it allows for amplification of the most sites with the fewest primers, reducing the possibility of primer dimers. The alignment results are shown below. Figure 1 , Figure 2 .
[0044] (2) When designing primers at the 3' and 5' ends, avoid primers with SNP variations at the 3' end. The significance of this screening step is to increase primer efficiency, avoid a decrease in amplification efficiency due to 3' end mismatch, and make the amplification efficiency of each Huc k type transposon family member as consistent as possible.
[0045] (3) Design 4-6 common sequence primers at both the 3' and 5' ends. The significance of this step is to provide multiple primers for selection, and to screen for efficient primers after experimentation. The 6 primers to be tested at the 5' end are CTGCCCTTGCCCGAGGCTAGGCTC (SEQ ID NO.7), GGCTCGGGCGAGGCGTGATCG(A / T / C)GTC (SEQ ID NO.8), GACTTAATCGCACCCATCAG (SEQ ID NO.9), GACTTAATCGCACCCATCAGG (SEQ ID NO.1), GTCTTGAGGGTACCCCTAATTATGG (SEQ ID NO.10), and CTGATGGGGGTTACC AGCTGAGAATTAGG (SEQ ID NO.11).
[0046] The four primers to be tested at the 3' end are TGGCACACT(C / T)ACTCGTCGG (SEQ ID NO.12), CCCCCCGGTCTCGAAACGCCGAC (SEQ ID NO.13), TCGGCCTCGCGCCGACCCATCT GGG (SEQ ID NO.14), and CGGGACGAACACGAAGGCC (SEQ ID NO.15).
[0047] (4) Introduce degenerate bases that coordinate with different Huck members into the primers. For example, in GGCTCGGGCGAGGCGTGATCG(A / T / C)GTC shown in SEQ ID NO.8, the thickened bases are respectively coordinated with three base types represented by huck_AC186577_1525#LTR, huck_AC208546_9913#LTR, and huck_AC214833_12913#LTR, so as to ensure the amplification efficiency of each family member to the greatest extent.
[0048] (5) Analysis of PCR amplification product yield revealed that the 5' end primers GTCTTGAGGGTACCCCTAATTATGG and GACTTAATCGCACCCATCAGG had high efficiency; the 3' end primer TGGCACACT(C / T)ACTCGTCGG had high efficiency.
[0049] (6) To increase amplification specificity, the 5' end GTCTTGAGGGTACCCCTAATTATGG was set as the inner primer and GACTTAATCGCACCCATCAGG as the outer primer. Nested design can reduce the risk of non-specific primer matching, increase the accuracy of the final library product, and ensure both yield and reduce sequencing costs.
[0050] (7) Based on the library structure characteristics of NGS next-generation sequencing, amplification product sequences with added library adapters were designed, see [reference needed]. Figure 3 .
[0051] This invention will ultimately utilize NGS sequencing technology to obtain tens of thousands of Huck loci at the whole-genome level. Currently, NGS read lengths are generally 150 bp at one end and 300-700 bp for both ends, which are considered short read lengths. To ensure that adjacent genomic sequences longer than 100 bp can be obtained within this read length range, this invention designs the 5' or 3' primers on the Huck locus to be approximately 70 bp from their respective ends, estimating that the obtained genomic sequences will range from 70 to 700 bp.
[0052] In summary, this invention designed 10 pairs of specific primers for the 5' and 3' ends of the Huck transposon. The final selections, Huck1-inner, Huck1-outer, and Huck2-inner, are shown in Table 1. The anchoring sequences and the locations of the transposon terminal specific sequences are marked in Table 1. Furthermore, corresponding adapter sequences, Adaptor1 and Adaptor2, and an adapter PCR primer were designed. The adapter and primer are matched to the anchoring sequences and sequencing primer sequences used in NGS next-generation high-throughput sequencing.
[0053] This invention utilizes NlaIII (NEB, Cat. No. R0125L) to cleave genome fragments and load specific adapters. NlaIII is a highlight of this invention, as this restriction site (4 bp) can be used to obtain a large number of cleavage sites on chromosomes, which is beneficial for capturing Huck sequences.
[0054] This invention simultaneously designs highly specific adapter sequences, adapter1 and adapter2. Adapter1 and adapter2 are decoupled at high temperature and then annealed at 95°C for 2 minutes. After mixing adapter1 and adapter2, the mixture is slowly cooled to room temperature.
[0055] The specific primers, adapters, and adapter primer sequences used for detecting the Huck site in the whole genome are shown in Table 1. The uppercase letters without underscores in Huck1outer, Huck1inner, and Huck2inner correspond to the anchoring sequences, which are used to anneal with the same sequence of the adapter to complete PCR amplification and obtain the product for sequencing; the uppercase letters with underscores are transposon terminal specific sequences.
[0056] Table 1
[0057]
[0058] Example 2: NGS sequencing method to enrich Huck sites from the maize genome
[0059] In this embodiment, K17, Zheng58, and DH-1 maize samples were selected, and the Huck sites of the maize genome were enriched according to the following steps.
[0060] (1) Take 50 ng of maize genomic DNA and break it with NlaIII (NEB, Cat. No R0125L) at 37 degrees for 5 hours to obtain the broken product. The broken DNA fragments are between 250 bp and 5 kb.
[0061] (2) Add 10 pM of adapters (Adaptor1 and Adaptor2) and 1 μl of T4 DNA ligase to the fragmentation product, react at 25°C for 3 hours to ligate, and obtain DNA fragments with adapters; in this step, the ratio of adapters to fragmentation product is 10 pM: 50 ng.
[0062] (3) The PCR product was recovered using a purification kit;
[0063] (4) Add 10pM Huck1outer, 10pM adaptor primer and 1μg of the mixed sample in (3) to the system, and perform PCR. The amplification conditions are 98℃ pre-denaturation for 2 minutes, 3 cycles (98℃ for 10 seconds, 62℃ for 20 seconds, 72℃ for 1 minute), then add Huck1inner and Huck2inner, and continue PCR. The amplification conditions are 3 cycles (98℃ for 10 seconds, 62℃ for 20 seconds, 72℃ for 1 minute) and 72℃ for 8 minutes.
[0064] (5) The product was recovered by equal volume using Beckman's AMPure XP beads;
[0065] (6) NGS novaseq sequencing, data quality control results are shown below. Figure 4 .
[0066] (7) The distribution of Huck in B73 predicted by the reference genome is shown in [reference genome]. Figure 5 According to the method of this invention, using different maize varieties such as Zheng 58, DH-1, and K17 as materials, the actual NGS sequencing analysis results show that the number of effective sequences refers to the genome sequence containing the Huck sequence and the lateral sequences (the number and proportion of effective sequences on each chromosome are shown in Table 2). Figure 6 , Figure 7 , Figure 8 (As shown).
[0067] Table 2 shows the number of Huck loci detected on each chromosome of the three genomes.
[0068]
[0069] Example 3: Other primer combinations for Huck site detection
[0070] In this embodiment, other Huck 5' and 3' end primers designed in Example 1 were used as primer combinations. Specifically, 5' end primers CTGCCCTTGCCCGAGGCTAGGCTC, GGCTCGGGCGAGGCGTGATC G(A / T / C)GTC, GACTTAATCGCACCCATCAG, or CTGATGGGGGTTACCAGCTGAGAATTAG G were used to replace SEQ ID NO. 1. The results showed that SEQ ID NO. 1 amplified the highest amount of product.
[0071] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A primer combination for detecting Huck transposon sites, characterized in that, The nucleotide sequences of the primer combination are shown in SEQ ID NO.1-3.
2. A method for detecting Huck transposon sites in the maize genome based on next-generation sequencing, characterized in that, include: The maize genomic DNA was fragmented, and the nucleotide sequence shown in SEQ ID NO.4 was used as adapter 1 and the nucleotide sequence shown in SEQ ID NO.5 was used as adapter 2. Adapter 1 and adapter 2 were ligated to the fragmented DNA fragment to obtain a DNA fragment with adapters. Using a DNA fragment with a adapter as a template, PCR amplification was performed using the primer combination described in claim 1. The PCR amplification product was sequenced and compared with the genome sequence to obtain the corresponding Huck site.
3. The method according to claim 2, characterized in that, The broken DNA fragments range from 250 bp to 5 kb.
4. The method according to claim 2, characterized in that, include: (1) Take maize genomic DNA and break it with DNA breaking enzyme NlaIII. The breaking conditions are 37℃ for 1-5 hours to obtain the broken product. The broken DNA fragments are between 250 bp and 5 kb. (2) For the fragmented product obtained in step (1), use DNA ligase to ligate the fragmented product with adapter 1 and adapter 2 to obtain DNA fragments with adapters. (3) The DNA fragment with adapter from step (2) is recovered, and the recovered product is then subjected to PCR amplification using the primer combination described in claim 2. (4) The amplification product from step (3) was recovered, quantified, sequenced using an NGS sequencer, and compared with the genome sequence to obtain the corresponding Huck transposon sites.
5. The method according to claim 4, characterized in that, The ligase used in step (2) is T4 DNA ligase, and the ligation conditions are 22-25℃ for 1-3 hours; the PCR amplification conditions in step (3) are: 98℃ for 2 minutes of pre-denaturation; 98℃ for 10 seconds, 62℃ for 20 seconds, 72℃ for 1 minute, 15 cycles; extension at 72℃ for 8 minutes.
6. The application of the primer combination of claim 1 or the method of any one of claims 2-5 in the isolation of maize Huck site sequences.
7. The application of the primer combination of claim 1 or the method of any one of claims 2-5 in the identification of Huck transposons at the maize genome level.
8. The application of the primer combination of claim 1 or the method of any one of claims 2-5 in screening Huck transposons at the maize genome level.
Citation Information
Patent Citations
Method for whole-genome molecular marker detection
CN109295048A
Efficient detection primers and method for AcDs whole genome loci based on NGS sequencing
CN112210620A