Method for Identifying Aegilops tauschii - Derived Germplasm Fragments by Using Exon Capture Sequencing Chip of D Genome of Wheat - Related Crops
By using the wheat-like D genome pan-gene exon capture sequencing chip and second-generation sequencing technology, combined with the differential site tracking method, high-precision identification and whole-genome genetic analysis of wheat-derived germplasm materials are achieved, solving the problem of difficult to achieve high-precision identification in the existing technology and reducing sequencing costs.
Patent Information
- Application Number
- CN202211137466.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-19
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2042-09-19
AI Technical Summary
The prior art is difficult to achieve genome-wide genetic analysis and identification of high-precision and low-cost genome-derived germplasm materials, especially when sequencing using a wheat-like D-genome pan-genome capture probe, it is impossible to accurately identify infiltration fragments from jejuni to wheat genome.
Exon region capture and sequencing was performed using the wheat-like D genome pan-gene exon capture sequencing chip. Combined with second-generation sequencing and comparative genome analysis procedures, exogenous fragments of wheat-derived infiltration systems or artificial synthetic materials were identified through differential site tracking methods.
The high coverage exon region sequencing and accurate identification of exogenous fragments of the introspective materials of the quadrilateral-derived introspective materials was achieved, which solved the bottleneck that the prior art was unable to identify the quadrilateral-derived population with high accuracy, and reduced the sequencing cost.
Smart Images

Figure CN115449545B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of crop molecular breeding and functional research, and particularly relates to a method for identifying and analyzing germplasm materials derived from the hybridization of Aegilops tauschii and common wheat based on a wheat D-genome pan-gene exon capture sequencing chip. Background Art
[0002] Common wheat varieties are allopolyploid hexaploids composed of three sub-genomes, A, B, and D. Due to the long-term use of backbone parents in the breeding process and the evolutionary bottleneck during the formation of common wheat, its genetic basis has become increasingly narrow, restricting wheat variety improvement. The wheat genome is complex and huge (16 GB), which is the main factor limiting large-scale whole-genome genetic diversity research of its germplasm resources. Obtaining variant information through whole-genome or reduced-representation genome sequencing methods is too costly, with serious locus deletions, relatively large sequencing data volumes, and a large amount of redundant data requiring a large amount of analysis costs and data storage costs, and having high software and hardware requirements for data storage and analysis calculations.
[0003] Aegilops tauschii is the donor species of the wheat D sub-genome. However, perhaps only a very small number of Aegilops tauschii participated in the formation of hexaploid wheat. Its wild population contains rich gene resources related to resistance, yield, and quality, and is a secondary germplasm resource bank for wheat genetic improvement. How to achieve the analysis of genetic variation in the Aegilops tauschii population and the exploration and utilization of germplasm resources based on the whole-genome map information is the key. Currently, a series of derived new synthetic species or introgression line materials have been created using Aegilops tauschii. These valuable germplasm resources have important application values for studying wheat domestication, artificial selection, and genetic improvement using excellent genes of Aegilops tauschii (Zhou, Y., Bai, S., Li, H., et al. Introgressing the Aegilops tauschii genome into wheat as a basis for cereal improvement. Nat Plants. 7, 774–786 (2021)). Therefore, there is an urgent need for a low-cost and high-precision method for whole-genome genetic analysis and identification of Aegilops tauschii-derived germplasm materials. The applicant previously extracted gene coding region sequences using an iterative method based on multiple published Aegilops tauschii genomes and the wheat D genome, and obtained wheat D-genome pan-genome capture probes by Kmer classification counting and group amplification.
[0004] Therefore, there is an urgent need and application value for how to use wheat D-genome pan-genome capture probes for sequencing and develop an accurate method for analyzing introgression fragments from Aegilops tauschii to the wheat genome based on the sequencing data. Summary of the Invention
[0005] The present invention provides a method for exon region capture sequencing using a wheat D genome pan-gene exon capture sequencing chip, and an analysis process based on comparative genomics, which can accurately identify exogenous fragments in introgression lines or synthetic materials derived from Aegilops tauschii, so as to solve the bottleneck that existing wheat genome-based chips cannot be used for high-precision identification of Aegilops tauschii-derived populations.
[0006] The present invention provides an application of a wheat D genome pan-gene exon capture sequencing chip. Specifically, the present invention uses a wheat D genome pan-gene exon capture sequencing chip, combines second-generation sequencing for molecular breeding and functional genomics research, and adopts a differential locus tracking method to identify exogenous fragments in introgression lines derived from Aegilops tauschii. The method includes the following steps:
[0007] Step 1: Use a wheat D genome pan-gene exon capture sequencing chip to perform capture sequencing on the parents of the introgression line material and the introgression line material. The parents of the introgression line material are wheat parent P1 and Aegilops tauschii parent P2;
[0008] Step 2: Align the capture sequencing data of wheat parent P1 and Aegilops tauschii parent P2 of the introgression line material to the wheat reference genome, and retain the sites where the genotypes of the two parent samples are homozygous and the genotypes of the two parents are inconsistent, to obtain a parental differential locus file. The wheat reference genome uses the wheat parent of the introgression line material as the reference genome.
[0009] Step 3: Map and align the capture sequencing data of the introgression line material to the wheat reference genome to obtain a differential locus file of the introgression line material. Filter the differential locus file of the introgression line material with allele frequency and Genotype. The filtering parameter allele frequency is 0.5, and Genotype is 0 / 1.
[0010] Specifically, for the alignment results in Steps 2 and 3, use the Picard software to remove duplicate sequences, re-align the Indel regions, and perform baseline correction, all with default parameters. Then use the GATK software to perform Variants Calling on individual sample materials of the introgression line to obtain a parental differential locus file; at the same time, perform population Variants Calling on the parents to obtain a differential locus file of the introgression line material. Parameters: MQ≥30, BQ≥20, and other parameters are all default parameters.
[0011] Step 4: Use the parental differential locus file to filter the differential locus file of the introgression line material. Use the BCFTOOLS software to retain the intersection of the two differential locus files, and use the intersection file as the set of differential loci caused by exogenous introgression; the differential loci that are not in the intersection are not considered.
[0012] Step 5: Using the intersection file and the parental differential locus file, calculate the exogenous introgression ratio of the derived introgression line materials within the window. Use the BEDTOOLS software to obtain a sliding window of 100k windows with a step size of 20k for the genome, and calculate the number of loci in the intersection file and the parental differential locus file within the window respectively. Within the same window, use the number of parental differential loci as the denominator and the number of intersection differential loci as the numerator to obtain the introgression ratio of the derived introgression line materials.
[0013] Step 6: Use the introgression ratio file to draw a dot plot, and define the region with continuously high introgression ratio as the exogenous fragment. In the present invention, the introgression ratio with ratio > 0.2 is listed as the high introgression ratio.
[0014] The above-mentioned wheat D genome pan-gene exon capture sequencing chip is prepared by the following method: Integrate multiple published Aegilops tauschii genomes and the wheat D sub-genome, extract the gene coding region sequences by the iterative method, and use Kmer classification counting and grouped amplification to obtain the Aegilops tauschii pan-genome capture probes.
[0015] The Aegilops tauschii genomes include the Aegilops tauschii T093 genome, the Aegilops tauschii AY61 genome, the Aegilops tauschii AL878 genome, the Aegilops tauschii AY17 genome, and the Aegilops tauschii XJ02 genome. In the examples of the present invention, the gene coding region sequences of 5 Aegilops tauschii materials and the gene coding region sequences of 1 wheat D sub-genome are integrated in total.
[0016] The specific design process of the wheat D genome pan-gene exon capture sequencing chip of the present invention is as follows:
[0017] Step 1: Use the gene coding region sequence of Aegilops tauschii material A as the starting template P1 of the probe.
[0018] Step 2: Use the starting template P1 to align and cover the gene coding region sequence of Aegilops tauschii material B. Specifically, use software such as Blat, BLAST, and Minimap2 for homologous alignment, and the alignment parameters are: the alignment identity is not less than 98%, and the alignment length is not less than 100bp. Add the regions on the gene coding region sequence of Aegilops tauschii material B that cannot be covered by the starting template P1 to the starting template P1 to obtain a more comprehensive template P2.
[0019] Step 3: Use the template P2 to align and cover the gene coding region sequence of Aegilops tauschii material C. Specifically, use software such as Blat, BLAST, and Minimap2 for homologous alignment, and the alignment parameters are: the alignment identity is not less than 98%, and the alignment length is not less than 100bp. Add the regions on the gene coding region sequence of Aegilops tauschii material C that cannot be covered to the template P2 to obtain a more comprehensive template P3.
[0020] Step 4: Repeat Step 3. Using the iterative method, complete the alignment coverage of the gene coding region sequences of N Aegilops tauschii materials and the alignment coverage of the gene coding region sequences of the D sub-genome of wheat. Add the regions that are not covered by the gene coding region sequences of the Aegilops tauschii materials and the gene coding region sequences of the D sub-genome of wheat detected in this step to the template to obtain the final pan-genome template. Here, N is a positive integer.
[0021] Step 5: Use K-mer to perform statistical classification of the template sequences and determine the probe copy number. The specific method is to use jellyfish to respectively count the K-mer sequences with lengths of 60bp, 80bp, 100bp, and 120bp in the pan-genome template in Step 4, delete the fragments containing N and simple repeat sequences in the template, determine the synthesis concentration ratio of different probe sequences, and the synthesis concentration ratio is the ratio of the statistical value of each Kmer fragment divided by the GC value of the fragment.
[0022] Step 6: Synthesize and amplify the template probes in groups according to the synthesis concentration ratio of different probe sequences in Step 5 to obtain the capture probes.
[0023] The gene coding region sequences of the above Aegilops tauschii materials are from the Aegilops tauschii T093 genome, the Aegilops tauschii AY61 genome, the Aegilops tauschii AL878 genome, the Aegilops tauschii AY17 genome, and the Aegilops tauschii XJ02 genome (Zhou, Y., Bai, S., Li, H. et al. Introgressing the Aegilops tauschii genome into wheat as a basis for cereal improvement. Nat Plants. 7, 774–786 (2021); Luo, M.-C. et al. Genome sequence of the progenitor of the wheat D genome Aegilops tauschii. Nature 551, 498–502 (2017).).
[0024] The D sub-genome of wheat uses the D sub-genome data in the genome of Aikang 58 (Jia J, X.Y., Cheng J. et al. Homology-mediated inter-chromosomal interactions in hexaploid wheat lead to specific sub-genome territories following polyploidization and introgression. Genome Biol. 22(1):26(2021).; Zhengang Ru, Angela Juhasz, Danping Li, Pingchuan Deng, Jing Zhao, et al. 1RS.1BL molecular resolution provides novel contributions to wheat improvement. BioRxiv. (2020).).
[0025] The beneficial effects of the present invention are as follows:
[0026] The method of the present invention can not only achieve high coverage of the whole-genome exon region derived from Aegilops tauschii using an exon capture sequencing chip, but also be accurately used for the identification and genetic analysis of exogenous fragments derived from Aegilops tauschii.
[0027] The present invention uses precise tracking of variant sites to trace exogenous fragments in derivative materials and eliminates background effects in a proportional manner. Description of the Drawings
[0028] Figure 1 It is the chip design process in Example 1 of the present invention.
[0029] Figure 2 It is the sequencing depth map of Aegilops tauschii AL878, AY61, T093, XJ02, AY17 and wheat Aikang 58 (AK58).
[0030] Figure 3 It is a Venn diagram for comparing the deletion segments of T093. In the figure, 1-1 and 1-2 represent the first and second technical replicates of the first biological sample; 2-1 and 2-2 represent the first and second technical replicates of the second biological sample; 3-1 and 3-2 represent the first and second technical replicates of the second biological sample. Other numbers in the figure represent the number of gene coding regions.
[0031] Figure 4 It is the analysis process of exon capture sequencing data.
[0032] Figure 5Identification result of exogenous fragment for a single introgression line.
[0033] Figure 6 Tracking exogenous fragments in families. Detailed implementation manners
[0034] The present invention will be described in more detail below through specific implementation manners, so as to facilitate the understanding of the technical solution of the present invention, but it is not used to limit the protection scope of the present invention.
[0035] Example 1
[0036] Step 1: Use the gene coding region sequence of Aegilops tauschii material as the starting template P1 of the probe. In this example, Aegilops tauschii T093 material is used as the Aegilops tauschii A material, and the high-confidence gene HC gene coding region sequence of Aegilops tauschii T093 material is used as the starting template P1.
[0037] Step 2: Use the starting template P1 to align and cover the high-confidence gene HC gene coding region sequence of Aegilops tauschii B material, and add the regions on the high-confidence gene HC gene coding region sequence of Aegilops tauschii B material that cannot be covered by the starting template P1 to the starting template P1 to obtain a more comprehensive template P2. In this example, Aegilops tauschii AY61 is used as the Aegilops tauschii B material, and Blat and Blast software are used to perform coverage statistics on the template P1. The core parameters are that the alignment sequence length > 100bp and the alignment identity >= 98%.
[0038] Step 3: Use the template P2 to align and cover the high-confidence gene HC gene coding region sequence of Aegilops tauschii C material, and add the regions on the high-confidence gene HC gene coding region sequence of Aegilops tauschii C material that cannot be covered by the template P2 to the template P2 to obtain a more comprehensive template P3. In this example, Aegilops tauschii AL878 is used as the Aegilops tauschii C material, and Blat and Blast software are used to perform coverage statistics on the template P2. The core parameters are that the alignment sequence length > 100bp and the alignment identity >= 98%.
[0039] Step 4: Repeat Step 3, and use the iterative method to complete the alignment and coverage detection of the high-confidence gene HC gene coding region sequences of Aegilops tauschii AY17, Aegilops tauschii XJ02, and the high-confidence gene HC gene coding region sequence of wheat Aikang 58. Add the regions on the high-confidence gene HC gene coding region sequences of Aegilops tauschii AY17, Aegilops tauschii XJ02, and the high-confidence gene HC gene coding region sequence of wheat Aikang 58 that cannot be covered to the template to obtain the final pan-genome template. Similarly, Blat and Blast software are used to perform coverage statistics. The core parameters are that the alignment sequence length > 100bp and the alignment identity >= 98%.
[0040] Step 5: Use jellyfish to separately count the K-mer sequences with lengths of 60bp, 80bp, 100bp, and 120bp in the pan-genome template in Step 4, delete the fragments containing N in the template and simple repetitive sequencing, determine the synthesis concentration ratio of different probe sequences, and the synthesis concentration ratio is the statistical value of Kmer fragments of different lengths divided by the GC value of the fragment.
[0041] Step 6: According to the synthesis concentration ratio of different probe sequences, group and synthesize and amplify the template probes to obtain the capture sequences, add biotin modification to the amplified sequences to complete the probe preparation, and obtain the wheat D-genome pan-gene exon capture sequencing chip, where the biotin is hydrophilic biotin.
[0042] Example 2
[0043] Evaluate the sequencing cost and coverage in Aegilops tauschii and wheat.
[0044] Use the wheat D-genome pan-gene exon capture sequencing chip designed in Example 1 to perform capture sequencing on five Aegilops tauschii (Aegilops tauschii T093, AY61, AL878, AY17, XJ02) and one wheat material (common wheat Aikang 58), and evaluate the probe quality in terms of sequencing depth and coverage.
[0045] The specific steps are as follows:
[0046] Step 1. Use the BWA software to align the sequences of different samples (Aegilops tauschii T093, AY61, AL878, AY17, XJ02, and common wheat Aikang 58) to their respective reference genomes (Zhou, Y., Bai, S., Li, H., et al. Introgressing the Aegilops tauschii genome into wheat as a basis for cereal improvement. Nat Plants. 7, 774–786 (2021); Luo, M.-C., et al. Genome sequence of the progenitor of the wheat D genome Aegilops tauschii. Nature 551, 498–502 (2017).; Jia J, X.Y., Cheng J., et al. Homology-mediated inter-chromosomal interactions in hexaploid wheat lead to specific subgenome territories following polyploidization and introgression. Genome Biol. 22(1):26 (2021); Zhengang Ru, Angela Juhasz, Danping Li, Pingchuan Deng, Jing Zhao, et al. 1RS.1BL molecular resolution provides novel contributions to wheat improvement. BioRxiv. (2020).), and use the SAMTOOLS software to sort the alignment results.
[0047] Step 2. Use the Picard software to remove duplicate sequences, re-align the Indel regions, and perform baseline correction. The parameters used in this step are all default parameters.
[0048] Step 3. Use the BEDCOV function of the SAMTOOLS software to count the sequencing depths of different samples in the protein-coding regions of the genome in the deduplicated alignment results. Except for chromosome 1B of common wheat Aikang 58, there are no significant differences between different homologous groups of the same subgenome in the other materials; the insufficient depth of chromosome 1B of wheat Aikang 58 is caused by the 1BS-1RS translocation. The sequencing depths of different samples are as Figure 2 shown.
[0049] Step 4: Use the COVERAGE function of BEDTOOLS software to calculate the coverage ratio of the protein-coding regions in the alignment results. If the segment coverage ratio is lower than 50%, it is considered that this segment is not covered and is recorded as missing. Calculate the differences in the missing segments among different samples of the same material and different technical replicates of the same sample. The results are shown in Table 1. Taking the T093 material as an example, there are 2,193 segments that are commonly missing in 6 samples, with a segment length of 1.50788 Mb, accounting for 3.1% of the coding region of T093 (48.519 M); the total length of all deletions is 1.7329 Mb, accounting for 3.57%. The evaluation results show that the capture rate of the present invention in the Aegilops tauschii gene coding region is above 96%. The Venn diagram of the comparison of the T093 deletion segments is as Figure 3 shown.
[0050] Table 1 Deletion ratios of five Aegilops tauschii
[0051]
[0052] The above evaluation results show that the capture rate of the chip of the present invention for the exons of Aegilops tauschii is above 96%, and the capture depth is relatively high in different materials. The insufficient depth of chromosome 1B of wheat Aikang 58 is caused by the 1BS-1RS translocation. When using the present invention for capture sequencing, 13-15 G of sequencing data can reach an average sequencing depth of 80× for Aegilops tauschii and 30× for the D subgenome of common wheat. However, if the whole-genome resequencing method is used to reach the same depth, more than 320 G of Aegilops tauschii sequencing data and more than 420 G of wheat sequencing data are required.
[0053] Example 3
[0054] Use the exon capture sequencing chip of Example 1 to perform capture sequencing on a batch of materials containing Aegilops tauschii genomic introgression (T093-AK58 introgression line), and accurately identify the exogenous fragments by comparing the variant sites between the parents and the offspring. As Figure 4 shown, the specific steps are as follows:
[0055] Step 1: Use the exon capture sequencing chip of Example 1 to perform capture sequencing on a batch of materials containing Aegilops tauschii genomic introgression (T093-AK58 introgression line) and the parents of the T093-AK58 introgression line, common wheat Aikang 58 and T093. Use the FASTP software to perform quality control on the sequencing data of all T093-AK58 introgression line materials, removing low-quality sequences and adapter sequences, using the default parameters.
[0056] Step 2: Using the MEM program of BWA software, align the capture sequencing data of the two parents of the T093-AK58 introgression line obtained in Step 1, namely common wheat Aikang 58, T093, and the T093-AK58 introgression line material, to the reference genome of common wheat Aikang 58 (Jia J, X.Y., Cheng J.et al. Homology-mediated inter-chromosomal interactions in hexaploid wheat lead to specific subgenome territories following polyploidization and introgression. Genome Biol. 22(1):26(2021).; Zhengang Ru, Angela Juhasz, Danping Li, Pingchuan Deng, Jing Zhao, et al. 1RS.1BL molecular resolution provides novel contributions to wheat improvement. BioRxiv.(2020).), and use SAMTOOLS software to sort the alignment results.
[0057] Alignment and sorting commands: bwamem -t 38 -M -Y -R '@RG\tID:${sample}\tPL:ILLUMINA\tSM:${sample}' ${reference}.facleanfq / ${sample}.R1.fq.gz cleanfq / ${sample}.R2.fq.gz|samtools view -F 4 -q 30 -Sb >${sample}.bam && samtools sort -@10 ${sample}.bam -o ${sample}.sort.bam.
[0058] Step 3: For the alignment results in Step 2, use Picard software to remove duplicate sequences, re-align the Indel regions, and perform baseline correction and other processes. Use default parameters.
[0059] Step 4: Use GATK software to perform Variants Calling on the individual sample materials of the derived introgression line to obtain the parental differential site file, and at the same time perform Variants Calling on the parents for the population to obtain the differential site file of the derived introgression line material. Parameters: MQ≥30, BQ≥20, and other parameters are all default parameters.
[0060] Step 5: Use a PYTHON file to filter the parental differential locus file and the introgression line material differential locus file. Specifically, for the parental differential locus file, retain the loci where the genotypes of the two parents are homozygous and the genotypes of the two parents are inconsistent. Each derived introgression line material differential locus file retains the loci with a heterozygous genotype (0 / 1) and the same depth of 0 and 1 in the heterozygous genotype.
[0061] Step 6: Use isec of the BCFTOOLS software to take the intersection of the derived introgression line material differential locus file and the parental differential locus file.
[0062] Step 7: Use the BEDTOOLS software to count the number of differential loci in the window. Specifically, use the BEDTOOLS software to obtain a sliding window of 100k windows with a step size of 20k for the genome, and count the number of loci in the window for the intersection file and the parental differential locus file, namely the number of intersection differential loci and the number of parental differential loci. Parameters: -w 100000, -s 20000.
[0063] Step 8: In the same window, use the number of parental differential loci as the denominator and the number of intersection differential loci as the numerator to calculate the introgression ratio of the derived introgression line material individuals. The regions with continuous high introgression ratios are defined as exogenous fragments. In this embodiment, the introgression ratio with ratio > 0.2 is listed as a high introgression ratio. In the non-exogenous fragment region, the points gather into a line at y = 0; in the exogenous introgression region, the points are dispersed between 0 and 1, as Figure 5 shown.
[0064] Step 9: Select the offspring of the same sample in different generations to form a family and conduct pedigree tracking of the exogenous fragments. The results of pedigree tracking of the exogenous fragments are as Figure 6 shown. In the figure, the red dots represent the exogenous fragments.
[0065] The above evaluation results show that the present invention can obtain the specific positions where exogenous fragments are integrated into wheat chromosomes by using the exon capture sequencing chip of the D genome of Triticeae crops, which is convenient for tracking exogenous fragments by pedigree and beneficial to the integration and utilization of gene resources of the entire D genome of the Triticeae tribe.
[0066] The above-described embodiments are only preferred embodiments of the present invention and do not limit the scope of implementation of the present invention. Therefore, any equivalent changes or modifications made according to the structure, characteristics, and principles described in the scope of the present invention patent should be included in the scope of the present invention's patent application.
Claims
1. Method for identifying Aegilops tauschii-derived germplasm fragments using exon capture sequencing chip of wheat D genome, characterized in that, it comprises the following steps: Step 1: Use the wheat D genome pan-gene exon capture sequencing chip to perform capture sequencing on the parents of the derived introgression line material and the derived introgression line material; The parents of the derived introgression line material are wheat parent P1 and Aegilops tauschii parent P2; The wheat D genome pan-gene exon capture sequencing chip is prepared by the following method: Integrate multiple published Aegilops tauschii genomes and the wheat D sub-genome, use the iterative method to extract gene coding region sequences, and use Kmer classification counting and grouping amplification to obtain Aegilops tauschii pan-genome capture probes; Step 2: Align the capture sequencing data of wheat parent P1 of the derived introgression line material and the capture sequencing data of Aegilops tauschii parent P2 to the wheat reference genome to obtain the parental differential site file; Step 3: Align the capture sequencing data of the derived introgression line material to the wheat reference genome to obtain the derived introgression line material differential site file; Step 4: Filter the parental differential site file and the introgression line material differential site file, retain the intersection of the two differential site files, and use the intersection file as the differential site set file caused by exogenous introgression; Step 5: Use the intersection file and the parental differential site file to statistically calculate the exogenous introgression ratio of the derived introgression line material within the window; Use the number of parental differential sites as the denominator and the number of intersection differential sites as the numerator to obtain the introgression ratio of the derived introgression line material; Step 6: Use the introgression ratio file to draw a dot plot, and define the region with continuously high introgression ratio as the exogenous fragment; among them, the introgression ratio with ratio > 0.2 is listed as the high introgression ratio.
2. The method according to claim 1, characterized in that, in Step 2, the wheat reference genome uses the wheat parent of the derived introgression line material as the reference genome.
3. The method according to claim 1, characterized in that, After aligning the derived introgression line material, wheat parent P1 of the derived introgression line material, and Aegilops tauschii parent P2 to the wheat reference genome respectively, for the alignment results, use Picard software to remove duplicate sequences, re-align the Indel regions, and perform baseline correction; then use GATK software to perform Variants Calling on the single sample material of the derived introgression line to obtain the derived introgression line material differential site file; perform Variants Calling on the parents for the population; Obtain the parental differential site file.
4. The method according to claim 1, characterized in that, when filtering in Step 4, for the parental differential site file, retain the sites where the genotypes of the two parents are homozygous and the genotypes of the two parents are inconsistent; for the derived introgression line material differential site file, retain the sites where the genotype is heterozygous and the depths of 0 and 1 in the heterozygous genotype are the same.
5. The method according to claim 1, characterized in that, In Step 5, the BEDTOOLS software is used to obtain a 100k window of the genome with a 20k-step sliding window. The number of sites in the window is separately counted for the intersection file and the parental differential site file. In the same window, the infiltration ratio of the derived introgression line material is obtained by using the number of parental differential sites as the denominator and the number of intersection differential sites as the numerator.