A wheat D-genome pan-gene exon capture sequencing chip and its design method
By designing a pan-gene exon capture sequencing chip for wheat-like D genomes, integrating the genetic information of multiple wheat-like and wheat-like D subgenomes, the problem that existing chips cannot effectively cover the mutation sites of wheat-like, and achieving efficient and low-cost genome analysis.
Patent Information
- Application Number
- CN202211147663.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-19
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-09-19
AI Technical Summary
Existing wheat genomic chips cannot effectively cover the genetic mutation sites of wheat populations, and too many repeat sequences lead to high cost and low efficiency of variant detection.
A wheat-type D genome pan-gene exon capture sequencing chip was designed. By integrating the gene information of multiple wheat-type genomes and wheat-type D subgenomes, the gene coding region sequences were extracted by iterative method, and the Kmer classification count group amplification was used to obtain the wheat-type pan-genome capture probe.
High coverage of more than 98% of the gene coding regions of the wheat D genome is achieved, reducing sequencing and analysis costs and improving analysis efficiency.
Smart Images

Figure CN115323046B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of crop molecular breeding and functional research, and particularly relates to a Triticeae D genome pan-gene exon capture sequencing chip and a design method thereof. Background Art
[0002] Wheat is one of the important food crops in the world, and China is the country with the largest planting area and the highest total output of wheat in the world. Common wheat is an allohexaploid composed of three sub-genomes A, B, and D. Due to the long-term use of backbone parents in the breeding process and the evolutionary bottleneck during the formation of common wheat, its genetic basis has become increasingly narrow, restricting the improvement of wheat varieties. The wheat genome is complex and huge (16GB), which is the main factor restricting the large-scale study of genetic diversity of its germplasm resources. Obtaining variation information through whole-genome or reduced-representation genome sequencing methods is too costly, with serious locus deletions, relatively large sequencing data volume, and a large amount of redundant data requiring a large amount of analysis cost and data storage cost, and having high requirements for the software and hardware of data storage and analysis calculation.
[0003] Aegilops tauschii is the donor species of the D sub-genome of wheat. However, perhaps only a very small number of Aegilops tauschii participated in the formation of hexaploid wheat. Its wild population contains rich gene resources related to resistance, yield, and quality, and is a secondary germplasm resource bank for wheat genetic improvement. How to achieve the analysis of genetic variation in the Aegilops tauschii population and the exploration and utilization of germplasm resources based on the whole-genome map information is the key. Currently, a series of derived new synthetic species or introgression line materials have been created using Aegilops tauschii. These valuable germplasm resources have important application values for studying wheat domestication, artificial selection, and genetic improvement using excellent genes of Aegilops tauschii. Therefore, there is an urgent need for a low-cost and high-precision method for whole-genome genetic analysis and identification of Aegilops tauschii-derived germplasm materials.
[0004] High-density single nucleotide polymorphism (SNP) markers have been widely used in molecular marker-assisted selection, backcross breeding and its background selection, multi-gene pyramiding breeding, genome-wide association analysis, QTL mapping, genome-wide selection, species evolution analysis, germplasm resource identification, etc. Currently, the acquisition of variant genotypes across the whole genome is mainly through re-sequencing and custom gene chips. However, SNP chips are more flexible in the preparation of samples and detection sites, and have higher detection rates and stability compared with sequencing. Currently, several chips have been developed in the field of wheat molecular breeding to characterize the variant information of test materials, including the Wheat Illumina Wheat 90K iSelect SNP genotyping array (90K), Wheat 660K SNP array (660K), Wheat 55K SNP array (55K), HD Wheat genotyping (820K) array (820K), Wheat 50K Triticum Trait Breed array (50K) and other chips have made important research progress in the field of wheat molecular breeding. However, they also have the following disadvantages:
[0005] First, the preparation of the above chips requires the use of specialized supporting detection equipment and specific analysis software to analyze the genotype results. The technology is mostly monopolized by foreign countries, with many usage restrictions and inconvenience. Second, the probes of the current chips are unevenly distributed on the chromosomes, mostly located in non-gene coding regions such as chromosome repeat regions with more variation information, while for chromosomal regions with more genes, the marker distribution is less. Third, the current chips have a low coverage of functional genes and are mostly located in gene non-coding intervals. Summary of the Invention
[0006] The present invention provides a Triticeae D genome pan-gene exon capture sequencing chip, which is used to solve the problem that existing chips based solely on the wheat genome cannot cover the effective variation sites of Aegilops tauschii populations, and the redundant genomic data with a repeat sequence of more than 85% increases the cost of variation detection and reduces the analysis efficiency. The present invention adopts the design of pan-genomic probes, which can be better applied to the detection of different D sub-genome-derived materials and can cover the genetic variations of different Aegilops tauschii. The exon capture sequencing chip of the present invention can not only achieve a high coverage of the whole-genome exon region derived from Aegilops tauschii, but also accurately identify the exogenous fragments of the introgression lines or synthetic materials derived from Aegilops tauschii, and greatly reduces the sequencing and analysis costs.
[0007] The present invention integrates multiple published Aegilops tauschii genomes and the wheat D sub-genome, extracts the gene coding region sequences by the iterative method, and uses Kmer classification counting to group and amplify to obtain the Aegilops tauschii pan-genome capture probes.
[0008] The Aegilops tauschii genomes include the Aegilops tauschii T093 genome, the Aegilops tauschii AY61 genome, the Aegilops tauschii AL878 genome, the Aegilops tauschii AY17 genome, and the Aegilops tauschii XJ02 genome. In the embodiments of the present invention, the gene coding region sequences of 5 Aegilops tauschii materials and the gene coding region sequences of 1 wheat D sub-genome are integrated in total.
[0009] The specific design process of the Triticeae D genome pan-gene exon capture sequencing chip of the present invention is as follows:
[0010] Step 1: Use the gene coding region sequence of Aegilops tauschii material A as the starting template P1 of the probe.
[0011] Step 2: Use the starting template P1 to align and cover the gene coding region sequence of Aegilops tauschii B material. Specifically, use Blat, BLAST, and Minimap2 software for homologous alignment. The alignment parameters are: the alignment identity is not less than 98%, and the alignment length is not less than 100 bp. Add the regions on the gene coding region sequence of Aegilops tauschii B material that cannot be covered by the starting template P1 to the starting template P1 to obtain a more comprehensive template P2.
[0012] Step 3: Use the template P2 to align and cover the gene coding region sequence of Aegilops tauschii C material. Specifically, use Blat, BLAST, and Minimap2 software for homologous alignment. The alignment parameters are: the alignment identity is not less than 98%, and the alignment length is not less than 100 bp. Add the regions on the gene coding region sequence of Aegilops tauschii C material that cannot be covered to the template P2 to obtain a more comprehensive template P3.
[0013] Step 4: Repeat Step 3, and use the iterative method to complete the alignment and coverage of the gene coding region sequences of N Aegilops tauschii materials and the gene coding region sequence of the wheat D subgenome. Add the regions on the gene coding region sequences of the Aegilops tauschii materials and the gene coding region sequence of the wheat D subgenome that cannot be covered in this step to the template to obtain the final pan-genome template. Here, N is a positive integer.
[0014] Step 5: Use K-mer to perform statistical classification of the template sequence and determine the probe copy number. The specific method is to use jellyfish to separately count the K-mer sequences with lengths of 60 bp, 80 bp, 100 bp, and 120 bp in the pan-genome template in Step 4, delete the fragments containing N and simple repeat sequences in the template, determine the synthesis concentration ratio of different probe sequences, and the synthesis concentration ratio is the ratio of the statistical value of each Kmer fragment divided by the GC value of the fragment.
[0015] Step 6: Synthesize and amplify the template probes in groups according to the synthesis concentration ratio of different probe sequences in Step 5 to obtain the capture probes.
[0016] The gene coding region sequences of the above Aegilops tauschii materials are from the Aegilops tauschii T093 genome, Aegilops tauschii AY61 genome, Aegilops tauschii AL878 genome, Aegilops tauschii AY17 genome, and Aegilops tauschii XJ02 genome (Zhou, Y., Bai, S., Li, H., et al. Introgressing the Aegilops tauschii genome into wheat as a basis for cereal improvement. Nat Plants. 7, 774–786 (2021); Luo, M.-C., et al. Genome sequence of the progenitor of the wheat D genome Aegilops tauschii. Nature 551, 498–502 (2017).).
[0017] The D sub-genome of wheat uses the D sub-genome data in the Aikang 58 genome (Jia J, X.Y., Cheng J, et al. Homology-mediated inter-chromosomal interactions in hexaploid wheat lead to specific sub-genome territories following polyploidization and introgression. Genome Biol. 22(1):26 (2021); Zhengang Ru, Angela Juhasz, Danping Li, Pingchuan Deng, Jing Zhao, et al. 1RS.1BL molecular resolution provides novel contributions to wheat improvement. BioRxiv. (2020).).
[0018] The beneficial effects of the present invention are as follows:
[0019] The capture sequencing chip of the present invention uses gene information of multiple Aegilops tauschii genomes and the wheat D genome, and can cover more than 98% of the gene coding regions of the wheat D genome.
[0020] The present invention uses an iterative method to extract gene coding region sequences of different genomes, which can avoid the generation of redundant sequences. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 It is the chip design process in Example 1 of the present invention.
[0022] Figure 2 Sequencing depth diagrams of Aegilops tauschii AL878, AY61, T093, XJ02, AY17 and wheat Aikang 58 (AK58).
[0023] Figure 3 Venn diagram for comparison of the deletion regions of T093. In the figure, 1-1 and 1-2 represent the first and second technical replicates of the first biological sample; 2-1 and 2-2 represent the first and second technical replicates of the second biological sample; 3-1 and 3-2 represent the first and second technical replicates of the second biological sample. Other numbers in the figure represent the number of gene coding regions. Specific implementation manners
[0024] The present invention will be described in more detail below through specific implementation manners, so as to facilitate the understanding of the technical solution of the present invention, but it is not used to limit the protection scope of the present invention.
[0025] Example 1
[0026] Step 1: Use the gene coding region sequence of Aegilops tauschii material A as the starting template P1 of the probe. In this example, Aegilops tauschii T093 material is used as Aegilops tauschii material A, and the high-confidence gene HC gene coding region sequence of Aegilops tauschii T093 material is used as the starting template P1.
[0027] Step 2: Use the starting template P1 to align and cover the high-confidence gene HC gene coding region sequence of Aegilops tauschii material B, and add the regions on the high-confidence gene HC gene coding region sequence of Aegilops tauschii material B that cannot be covered by the starting template P1 to the starting template P1 to obtain a more comprehensive template P2. In this example, Aegilops tauschii AY61 is used as Aegilops tauschii material B, and Blat and Blast software are used to perform coverage statistics on the template P1, and the core parameters are that the alignment sequence length > 100bp and the alignment identity >= 98%.
[0028] Step 3: Use the template P2 to align and cover the high-confidence gene HC gene coding region sequence of Aegilops tauschii material C, and add the regions on the high-confidence gene HC gene coding region sequence of Aegilops tauschii material C that cannot be covered by the template P2 to the template P2 to obtain a more comprehensive template P3. In this example, Aegilops tauschii AL878 is used as Aegilops tauschii material C, and Blat and Blast software are used to perform coverage statistics on the template P2, and the core parameters are that the alignment sequence length > 100bp and the alignment identity >= 98%.
[0029] Step 4. Repeat Step 3, and use the iterative method to complete the alignment coverage detection of the high-confidence gene HC gene coding region sequences of Aegilops tauschii AY17, the high-confidence gene HC gene coding region sequences of Aegilops tauschii XJ02, and the high-confidence gene HC gene coding region sequences of wheat Aikang 58. Add the regions not covered by the high-confidence gene HC gene coding region sequences of Aegilops tauschii AY17, the high-confidence gene HC gene coding region sequences of Aegilops tauschii XJ02, and the high-confidence gene HC gene coding region sequences of wheat Aikang 58 to the template to obtain the final pan-genome template. Also use Blat and Blast software for coverage statistics, with the core parameters being that the alignment sequence length > 100 bp and the alignment identity >= 98%.
[0030] Step 5. Use jellyfish to respectively count the K-mer sequences with lengths of 60 bp, 80 bp, 100 bp, and 120 bp in the pan-genome template in Step 4. Delete the fragments containing N in the template and simple repetitive sequencing, and determine the synthesis concentration ratio of different probe sequences. The synthesis concentration ratio is the ratio of the statistical value of Kmer fragments with different lengths divided by the GC value of the fragment.
[0031] Step 6. According to the synthesis concentration ratio of different probe sequences, perform the synthesis and amplification of grouped template probes to obtain the capture sequences. Add biotin modification to the amplified sequences to complete the probe preparation, and thus obtain the wheat D genome pan-gene exon capture sequencing chip, where the biotin is hydrophilic biotin.
[0032] Example 2
[0033] Evaluate the sequencing cost and coverage in Aegilops tauschii and wheat.
[0034] Use the wheat D genome pan-gene exon capture sequencing chip designed in Example 1 to perform capture sequencing on five Aegilops tauschii (Aegilops tauschii T093, AY61, AL878, AY17, XJ02) and one wheat material (common wheat Aikang 58), and evaluate the probe quality in terms of sequencing depth and coverage.
[0035] The specific steps are as follows:
[0036] Step 1: Use the BWA software to align the sequences of different samples (Aegilops tauschii T093, AY61, AL878, AY17, XJ02, and common wheat Aikang 58) to their respective reference genomes (Zhou, Y., Bai, S., Li, H., et al. Introgressing the Aegilops tauschii genome into wheat as a basis for cereal improvement. Nat Plants. 7, 774–786 (2021); Luo, M.-C., et al. Genome sequence of the progenitor of the wheat D genome Aegilops tauschii. Nature 551, 498–502 (2017).; Jia J, X.Y., Cheng J., et al. Homology-mediated inter-chromosomal interactions in hexaploid wheat lead to specific subgenome territories following polyploidization and introgression. Genome Biol. 22(1):26 (2021); Zhengang Ru, Angela Juhasz, Danping Li, Pingchuan Deng, Jing Zhao, et al. 1RS.1BL molecular resolution provides novel contributions to wheat improvement. BioRxiv. (2020).), and use the SAMTOOLS software to sort the alignment results.
[0037] Step 2: Use the Picard software to remove duplicate sequences, re-align the Indel regions, and perform baseline correction. The parameters used in this step are all default parameters.
[0038] Step 3: Use the BEDCOV function of the SAMTOOLS software to calculate the sequencing depth of different samples in the protein-coding regions of the genome in the deduplicated alignment results. Except for chromosome 1B of common wheat Aikang 58, there are no significant differences between different homologous groups of the same subgenome in other materials; the insufficient depth of chromosome 1B of wheat Aikang 58 is caused by the 1BS-1RS translocation. The sequencing depths of different samples are as Figure 2 shown.
[0039] Step 4: Use the COVERAGE function of BEDTOOLS software to calculate the coverage ratio of the protein-coding regions in the alignment results. If the segment coverage ratio is less than 50%, the segment is considered not covered and marked as missing. The differences in the missing segments among different samples of the same material and different technical replicates of the same sample are counted, and the results are shown in Table 1. Taking the T093 material as an example, there are 2,193 segments that are commonly missing in 6 samples, with a segment length of 1.50788 Mb, accounting for 3.1% of the coding region of T093 (48.519 M); the total length of all deletions is 1.7329 Mb, accounting for 3.57%. The evaluation results show that the capture rate of the present invention in the Aegilops tauschii gene coding region is above 96%. The Venn diagram of the comparison of the T093 deletion segments is as Figure 3 shown.
[0040] Table 1 Deletion ratios of five Aegilops tauschii
[0041]
[0042]
[0043] The above evaluation results show that the capture rate of the chip of the present invention for the exons of Aegilops tauschii is above 96%, and the capture depth is relatively high in different materials. The insufficient depth of chromosome 1B of wheat dwarf-resistant 58 is caused by the 1BS-1RS translocation. When using the present invention for capture sequencing, 13-15 G of sequencing data can achieve an average sequencing depth of 80× for Aegilops tauschii and an average sequencing depth of 30× for the D subgenome of common wheat. However, if the whole-genome resequencing method is used to reach the same depth, more than 320 G of Aegilops tauschii sequencing data and more than 420 G of wheat sequencing data are required.
[0044] The above embodiments are only preferred embodiments of the present invention and do not limit the scope of implementation of the present invention. Therefore, any equivalent changes or modifications made according to the structure, characteristics, and principles described in the scope of the present invention should be included in the scope of the present invention for patent application.
Claims
1. A wheat D genome pan - gene exon capture sequencing chip, Characterized in that, It is designed and obtained by the following method: integrating the Aegilops tauschii genome and the wheat D sub - genome, extracting gene coding region sequences by the iterative method, and obtaining Aegilops tauschii pan - genome capture probes through Kmer classification and counting - based group amplification, which is the wheat D genome pan - gene exon capture sequencing chip; Among them, integrating the Aegilops tauschii genome and the wheat D sub - genome and extracting gene coding region sequences by the iterative method includes: Step 1: Using the gene coding region sequence of Aegilops tauschii material A as the starting template P1 of the probe; Step 2: Using the starting template P1 to align and cover the gene coding region sequence of Aegilops tauschii material B, and adding the regions on the gene coding region sequence of Aegilops tauschii material B that cannot be covered by the starting template P1 into the starting template P1 to obtain template P2; Step 3: Using template P2 to align and cover the gene coding region sequence of Aegilops tauschii material C, and adding the regions on the gene coding region sequence of Aegilops tauschii material C that cannot be covered into template P2 to obtain template P3; Step 4: Repeat step 3, using the iterative method to complete the alignment and coverage of the gene coding region sequences of N Aegilops tauschii materials and the gene coding region sequences of the wheat D sub - genome, and adding the regions on the gene coding region sequences of the Aegilops tauschii materials and the wheat D sub - genome that cannot be covered in this step into the template to obtain the final pan - genome template; N is a positive integer; In steps 2, 3, and 4, the alignment and coverage specifically are: using Blat, BLAST, and Minimap2 software for homologous alignment, and the alignment parameters are: the alignment identity is not less than 98%, and the alignment length is not less than 100 bp.
2. The wheat D genome pan - gene exon capture sequencing chip according to claim 1, Characterized in that, The Aegilops tauschii genome includes the Aegilops tauschii T093 genome, the Aegilops tauschii AY61 genome, the Aegilops tauschii AL878 genome, the Aegilops tauschii AY17 genome, and the Aegilops tauschii XJ02 genome.
3. The wheat D genome pan - gene exon capture sequencing chip according to claim 1, Characterized in that, The wheat D sub - genome uses the D sub - genome data in the Aikang 58 genome.
4. The design method of the wheat D genome pan - gene exon capture sequencing chip according to claim 1, Characterized in that, It includes the following steps: Step 1: Using the gene coding region sequence of Aegilops tauschii material A as the starting template P1 of the probe; Step 2: Using the starting template P1 to align and cover the gene coding region sequence of Aegilops tauschii material B, and adding the regions on the gene coding region sequence of Aegilops tauschii material B that cannot be covered by the starting template P1 into the starting template P1 to obtain template P2; Step 3: Using template P2 to align and cover the gene coding region sequence of Aegilops tauschii material C, and adding the regions on the gene coding region sequence of Aegilops tauschii material C that cannot be covered into template P2 to obtain template P3; Step 4: Repeat Step 3. Using the iterative method, complete the alignment coverage of the gene coding region sequences of N Aegilops tauschii materials and the alignment coverage of the gene coding region sequences of the D subgenome of wheat. Add the regions not covered by the gene coding region sequences of the Aegilops tauschii materials and the gene coding region sequences of the D subgenome of wheat detected in this step to the template to obtain the final pan-genome template; wherein N is a positive integer; Step 5: Use K-mer to perform statistical classification of the template sequences to determine the probe copy number; Step 6: Perform template probe synthesis and amplification to obtain capture probes, which are the exon capture sequencing chips for the pan-genes of the wheat D genome.
5. According to the design method described in claim 4, it is characterized in that the gene coding region sequence is the gene coding region sequence of a high-confidence gene.
6. According to the design method described in claim 4, it is characterized in that in Steps 2, 3, and 4, the alignment coverage specifically is: perform homologous alignment using Blat, BLAST, and Minimap2 software, and the alignment parameters are: the alignment identity is not less than 98%, and the alignment length is not less than 100 bp.
7. According to the design method described in claim 4, it is characterized in that Step 5 specifically is: use jellyfish to respectively count the K-mer sequences of 60 bp, 80 bp, 100 bp, and 120 bp for the pan-genome template in Step 4, delete the fragments containing N in the template and simple repetitive sequencing, and determine the synthesis concentration ratio of different probe sequences.
8. According to the design method described in claim 7, it is characterized in that the synthesis concentration ratio is the ratio of the statistical value of Kmer fragments of different lengths divided by the GC value of the fragment.
9. According to the design method described in claim 7, it is characterized in that in Step 6, according to the synthesis concentration ratio of different probe sequences in Step 5, perform grouped synthesis and amplification of the template probes.
Citation Information
Patent Citations
Probe design method and positioning method for wheat exon sequencing gene positioning
CN112837746A