Design and application method of plant generic genome exon chip
Through the pan-genome exon chip design technology and the efficient second-generation sequencing platform, the problem of high cost and low efficiency of whole genome sequencing for complex genome crops is solved, and genome sequencing with high coverage, high specificity and low cost is achieved, and genetic research and breeding is supported.
Patent Information
- Application Number
- CN202510319457.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-06
AI Technical Summary
Whole genome sequencing of complex genomic crops such as wheat, oats, and tobacco faces major challenges in cost, data storage and data analysis. There are limitations in existing technologies such as resequencing, SNP chips and exon capture sequencing.
Pan-genome exon chip design technology is adopted to optimize probe design and capture methods, combined with an efficient second-generation sequencing platform, and achieve efficient and low-cost genome sequencing of complex genomic crops. This technology includes multiple rounds of probe design, strict filtration and combination into probes, and combined with a high-throughput sequencing platform for data analysis.
High coverage, high specificity and low cost genome sequencing of complex genomic crops such as wheat, oats, and tobacco has been achieved, providing strong support for genetic research, breeding and genome-wide selection.
Smart Images

Figure FT_1 
Figure FT_2
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biotechnology, and in particular to genome targeted sequencing technology for complex genome crops such as wheat, oats, and tobacco. The present invention specifically relates to a pan-genome exon chip design technology and application, which aims to reduce genome sequencing costs and improve sequencing efficiency. Background Art
[0002] The genomes of plants such as wheat, oats, and tobacco are complex and large, with a large number of repetitive sequences, which makes whole genome sequencing face major challenges in terms of cost, data storage, and data analysis.
[0003] Existing technologies such as resequencing, exon capture sequencing and SNP chips have been widely used in molecular breeding of wheat, but these methods still have certain limitations. Although resequencing can provide more comprehensive genetic information, it is costly and difficult to analyze data. Although SNP chips have a lower cost, they can only detect known polymorphic sites and cannot obtain unknown variation information. Although exon capture technology can solve these problems to a certain extent, it still faces problems such as single probe design and limited capture efficiency. Summary of the invention
[0004] In view of the shortcomings of the prior art described above, the present invention provides a plant pan-genome exon capture sequencing technology, which optimizes probe design and capture methods, combined with an efficient second-generation sequencing platform, to achieve efficient and low-cost genome sequencing of complex genome crops such as wheat, oats, and tobacco. This technology can obtain specific sequences that are not available on the reference genome, while maintaining low cost and high data accuracy and coverage, thereby providing strong support for wheat genetic research, breeding, and whole genome selection.
[0005] Technical Solution
[0006] Step 1: Extract pan-genome exon region sequences using wheat pan-genome sequences and their annotation information. First, obtain the FASTA format sequence of the whole genome of the target plant and the corresponding annotation file from the public database, and use bioinformatics methods to parse the annotation data to accurately determine the location of all exon regions and extract the corresponding sequence from the whole genome. Use a sliding window of a predetermined length (e.g., 120 bp) to segment each exon sequence, calculate the GC content and melting temperature (Tm value) of each window, and strictly filter the fragments containing a large number of uncertain bases (N base ratio exceeds 10%), GC content below 20% or above 80%, and simple repeat sequences exceeding 30% in the window to ensure that the obtained target sequence is of high quality and stability.
[0007] Step 2: Map the exon sequences screened in step 1 to multiple reference genomes, and use an efficient alignment algorithm to count the areas that are not fully covered in each reference genome. For these uncovered areas, use a shorter window (e.g. 80bp) for secondary probe design, calculate the specificity, GC content and Tm value of each window, and screen the sequences according to the filtering criteria set in step 1 to ensure that the secondary designed probe sequences meet the predetermined requirements in terms of coverage and specificity.
[0008] Step 3: For the exon regions that are still not fully covered in step 2, re-align and locate, and use a shorter window (e.g. 60 bp) for the uncovered regions to implement the third round of probe design. The GC content and Tm value of each window are calculated according to a unified algorithm, and fragments containing a large number of uncertain bases, GC content not in the specified range, or too high a proportion of repetitive sequences are eliminated, so as to obtain the final probe sequence with high coverage and high specificity.
[0009] Step 4: Based on the physical and chemical properties of the probe sequences, all the sequences obtained from the first three rounds of design were systematically grouped. Through statistical methods, the probes were classified according to length, GC content and Tm value. It was required that the probes in each group did not differ by more than a predetermined threshold under PCR amplification conditions, and the GC content differences were kept within a small range. Finally, the probes were divided into several homogeneous groups to provide technical support for subsequent unified amplification and application.
[0010] Step 5: Synthesize and amplify the grouped probe sequences. Submit each grouped probe to the synthesis platform, prepare it into an oligonucleotide probe, dissolve it in TE buffer (pH 8.0), and then use a standardized amplification method for micro-amplification. After quantitative and quality testing, the amplified products are mixed at equimolar concentrations to prepare a homogenized capture probe mixture, which is used as the basic material for subsequent targeted capture.
[0011] Step 6: Use the prepared capture probe mixture to perform targeted capture of the genomic DNA of the target sample. First, perform appropriate physical or enzymatic fragmentation on the target plant genomic DNA and construct a sequencing library; then, hybridize the library with the capture probe under specified conditions to allow the probe to specifically bind to the target exon region; after strict elution and amplification steps, obtain a library enriched in the target exon region, thereby improving the efficiency of targeted capture and data quality.
[0012] Step 7: Use a high-throughput sequencing platform to sequence the captured amplified library, and use bioinformatics processing such as data quality control, sequence alignment, variant detection, and functional annotation to comprehensively analyze the sequence information of the target exon region. This method provides high coverage, high specificity, and stable and reliable data support for plant functional genomics research. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 This is a flow chart of the design and application of the plant pan-genome exon chip described in Example 1 of the present invention.
[0014] Figure 2 This is a schematic diagram of the grouping of the probe sequences described in Example 1 of the present invention. DETAILED DESCRIPTION
[0015] The following is a description of the implementation of the present invention by means of specific embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification.
[0016] Embodiment 1:
[0017] A method for designing and applying a plant pan-genome exon chip comprises the following steps:
[0018] Step 1: Extract pan-genome exon region sequences using wheat pan-genome sequences and annotation information. First, download the wheat whole genome FASTA file and the corresponding GFF annotation file from the NCBI or Ensembl database, and use Biopython (version 1.78) and BEDTools (version 2.30.0) configured on a Linux workstation to extract the sequences of all exon regions; then, use the 120 bp sliding window method to analyze each exon sequence window by window, calculate the GC content of each window (the filtering condition is set to GC content less than 20% or greater than 80%) and Tm value (using the Nearest-Neighbor algorithm, predicted by the OligoCalc tool), and strictly exclude fragments with N base ratios exceeding 10% and regions with simple repeat sequences exceeding 30% in the window, so as to ensure that the obtained exon target sequence has high accuracy and stability.
[0019] Step 2: Use BWA (v0.7.17) to align the data extracted and filtered in step 1 on different wheat reference genomes (such as Chinese Spring, Dwarf 58, and Jimai 22), and use SAMtools (v1.10) to calculate the uncovered exon regions in each reference genome; then re-analyze the uncovered regions using an 80 bp sliding window, and use a self-written Python script to calculate the specificity, GC content, and Tm value of each 80 bp window (Tm is calculated according to the SantaLucia method), and strictly screen according to the same filtering criteria as step 1 (N content, GC boundary, repeat sequence ratio) to ensure that the secondary designed probe sequence meets the predetermined requirements in terms of specificity and coverage.
[0020] Step 3. For the exon regions that were still not covered in step 2, align them again on different reference genomes, use Bowtie2 (v2.4.2) to relocate these regions, and design them for the third time with 60 bp as the window after extraction; use the same algorithm to calculate the GC content and Tm value of each 60 bp window (Tm value prediction uses OligoCalc or similar tools), and at the same time eliminate fragments containing a large number of N (ratio > 10%), GC content not in the range of 0.2-0.8, and simple repeat sequences accounting for more than 30%, so as to ensure that the obtained probe sequences have higher coverage and specificity.
[0021] Step 4: All probe sequences obtained from the first three designs were systematically grouped according to their length, GC content and Tm value, and the data were statistically analyzed using R language and data processing packages such as "dplyr". The grouping rules were as follows: the difference in annealing temperature required for PCR amplification within each group did not exceed 5°C, and the difference in GC content did not exceed 0.1. Finally, all probes were divided into 12 groups. The probes in each group maintained a high degree of consistency in physical and chemical properties, providing a guarantee for subsequent amplification under unified conditions.
[0022] Step 5: Submit the probe sequences of the above 12 groups to commercial synthesis manufacturers (such as Integrated DNA Technologies, IDT or Sigma-Aldrich) for independent synthesis. The obtained oligonucleotide probes are dissolved in TE buffer (pH 8.0) and then micro-amplified in the laboratory using a high-fidelity PCR system. The specific reaction system was as follows: the total reaction volume was 25 µL, including 2.5 µL 10× Q5 Reaction Buffer (NEB, catalog number: M0491S), 0.5 µL 10 mM dNTP mixture, 1.0 µL forward and reverse primers with a concentration of 0.5 µM each, 0.5 µL Q5 High-Fidelity DNA Polymerase (2 U / µL, NEB), and about 10 ng of probe template, and the rest was supplemented with RNase-free DNase water to make up to 25 µL; the PCR cycle conditions were as follows: initial denaturation at 95°C for 30 seconds, followed by 30 cycles, each cycle was set to denaturation at 95°C for 10 seconds, annealing temperature at 55–65°C for 30 seconds (the specific temperature was adjusted according to the probe Tm value), extension at 72°C for 1 minute, and finally extension at 72°C for 5 minutes; the amplified product was detected for fragment distribution by Agilent 2100 Bioanalyzer and then analyzed by Qubit 4 Fluorometer was used for quantification, and probes of each group were mixed at equimolar concentrations to prepare a uniform capture probe mixture.
[0023] Step 6. Use the pan-genomic exon capture probe prepared by amplification in step 5 to perform targeted capture of the genomic DNA of the wheat sample. The specific process is: first, use Qiagen DNeasy Plant Mini Kit (Cat. No.: 69104) to extract wheat genomic DNA, and then use Covaris M220 ultrasonic crusher to break the DNA into 200-300 bp fragments; then use NEBNext Ultra II DNA Library Prep Kit (Cat. No.: E7645) to construct a sequencing library, including end repair, A-tailing, adapter connection (using NEBNext Adaptor for Illumina, Cat. No.: E7335S) and PCR amplification; after the library is constructed, mix the library with the capture probe, add 10× Blocking Mix and 10 µg Cot-1 DNA, and hybridize at 65°C for 16 hours according to the standard conditions of the probe supplier; then use Dynabeads MyOneStreptavidin C1 magnetic beads (Cat. No. 65001) capture the probe and target sequence complex, and perform three stringent washes at 65°C for 5 minutes each to remove non-specifically bound sequences; finally, the target sequence is recovered through an elution step and PCR amplified using Q5 High-Fidelity Polymerase to enrich the exon region library.
[0024] Step 7: Use the Illumina NovaSeq 6000 platform to perform high-throughput next-generation sequencing on the library captured and amplified in step 6. After quantification and quality testing by Qubit and Agilent Bioanalyzer, the library was adjusted to a concentration of 10 nM, loaded according to the Illumina standard library sequencing process, and the 150 bp double-end sequencing mode was selected. At the same time, sufficient sequencing depth was set (for example, the coverage of each sample reached at least 30×). After sequencing, preliminary data quality control was performed on the Illumina BaseSpace platform, and the raw data was used for subsequent alignment, variation detection and functional annotation analysis, so as to comprehensively analyze the sequence information of the wheat exon region.
Claims
1. A method for designing and applying a plant pan-genome exon chip, characterized in that: The following steps are involved: Step 1: Design efficient liquid capture probes through iterative algorithms to ensure that the probes can cover more than 99% of the exon regions of the target genome; Step 2: using an algorithm to calculate and group the original probe templates for the synthesized exon chip; Step 3: Prepare liquid capture probes using ultra-microamplification method; Step 4: Perform experimental hybridization by group gradient hybridization method to ensure efficient capture of the target area; Step 5: Use the second-generation sequencing platform to perform high-throughput sequencing on the captured samples to obtain high-quality exon region data.
2. The method for designing and applying a plant pan-genome exon chip according to claim 1, characterized in that: Multiple assembled and annotated reference genomes were used for sequence extraction and analysis of exon regions, and an iterative algorithm was used for pan-genome design to ensure coverage of exon regions of each genome while reducing probe redundancy.
3. The design and application method of the plant pan-genome exon chip according to claim 1, characterized in that: According to the characteristics of the probe sequences, an algorithm is used to score and group them to ensure that the characteristic information of each group of probe sequences is consistent and the amplification and hybridization conditions are consistent.
4. The method for designing and applying a plant pan-genome exon chip according to claim 1, characterized in that: The probes grouped according to the algorithm are subjected to ultramicro amplification using corresponding amplification conditions.
5. The method for designing and applying a plant pan-genome exon chip according to claim 1, characterized in that: The probes grouped according to the algorithm are subjected to hybridization capture experiments using corresponding hybridization reaction conditions.
6. The method for designing and applying a plant pan-genome exon chip according to claim 1, characterized in that: The DNA products after group hybridization are recovered and mixed, and then sequenced using a high-throughput sequencing platform.