InDel molecular marker primer pair combinations for constructing DNA fingerprint of fresh soybean
By constructing the DNA fingerprint map of fresh soybean ‘Tongsu No. 1’, using whole genome resequencing to screen InDel markers, the problem of morphological markers being susceptible to environmental impact was solved, and efficient identification and utilization of soybean germplasm resources was achieved, and breeding efficiency was improved.
Patent Information
- Application Number
- CN202411609772.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-12
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2044-11-12
AI Technical Summary
In the prior art, fresh soybean breeding mainly relies on morphological markers, is susceptible to environmental influences, lacks effective molecular marking methods, and is difficult to accurately identify and utilize germplasm resources.
The InDel molecular marker combination was developed for the DNA fingerprint of fresh soybean ‘Tongsu No. 1’, and 31 InDel markers with ≥30 bp were screened through whole genome resequencing, and primer pair combinations were designed for genetic diversity analysis and germplasm resource utilization.
It has achieved efficient and accurate identification of fresh soybeans and the utilization of germplasm resources, improved breeding efficiency, shortened the variety approval cycle, and provided a basis for molecular marker assisted breeding.
Smart Images

Figure CN119710057B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of DNA fingerprints, and in particular relates to a primer pair combination of an InDel molecular marker combination for constructing a DNA fingerprint of fresh soybean "Tongsu No. 1" and an application thereof. Background Art
[0002] Currently, cultivated soybeans grown on a large scale are derived from wild soybeans through long-term breeding and systematic improvement. They are divided into fresh-eating soybeans and grain-producing soybeans, depending on their harvest time. Fresh-eating soybeans are harvested between R6 and R7 of the growth period and eaten as a fresh vegetable. They have a soft texture, excellent flavor, and are rich in nutrients. They are high in vitamins, sucrose, and starch, while low in indigestible oligosaccharides and anti-nutrients.
[0003] Traditional breeding methods primarily rely on morphological markers for selection. However, morphological traits are the result of a combination of genetic and environmental influences and are susceptible to environmental changes, resulting in significant limitations. However, with the advancement of molecular biology, molecular marker technology has become an important tool for crop genetic improvement. The combined use of morphological and molecular markers will play a significant role in crop genetic improvement and the utilization of germplasm resources. Therefore, genomic-level genetic variation analysis in soybeans will be of great significance for population genetics research, marker-assisted breeding, and gene identification.
[0004] In recent years, a number of fresh soybean varieties have been developed through physiological breeding combined with molecular-assisted screening. Developing molecular markers to analyze the genetic diversity of fresh soybeans not only facilitates the full utilization of high-quality germplasm resources but also provides important guidance for broadening the genetic background and selecting superior varieties (lines) for use in fresh soybean breeding and production. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to address the deficiencies of the above-mentioned prior art and provide a primer pair combination of an InDel molecular marker combination for constructing a DNA fingerprint of the fresh soybean "Tongsu No. 1" and its application. The primer pair combination of the InDel molecular marker combination is used for genetic diversity analysis, identification and germplasm resource utilization.
[0006] To solve the above technical problems, the technical solution adopted by the present invention is: a primer pair combination for constructing an InDel molecular marker combination of the DNA fingerprint of the fresh soybean "Tongsu No. 1", wherein the DNA fingerprint of the fresh soybean "Tongsu No. 1" is composed of 31 InDel molecular markers of ≥30bp distributed on 15 chromosomes;
[0007] The nucleotide sequence of the first primer pair used to amplify the Gm02_687775_687827 molecular marker is shown in SEQ ID NOs. 1-2;
[0008] The nucleotide sequence of the second primer pair used to amplify the Gm02_38671726_38671769 molecular marker is shown in SEQ ID NOs. 3-4;
[0009] The nucleotide sequence of the third primer pair used to amplify the Gm02_46235870_46235916 molecular marker is shown in SEQ ID NOs. 5-6;
[0010] The nucleotide sequences of the fourth primer pair used to amplify the Gm02_48831944_48831999 molecular marker are shown in SEQ ID NOs. 7 to 8;
[0011] The nucleotide sequences of the fifth primer pair used to amplify the Gm02_50061608_50061609 molecular marker are shown in SEQ ID NOs. 9 to 10;
[0012] The nucleotide sequence of the sixth primer pair used to amplify the Gm03_3866902_3866939 molecular marker is shown in SEQ ID NOs. 11-12;
[0013] The nucleotide sequence of the seventh primer pair used to amplify the Gm03_40000161_40000198 molecular marker is shown in SEQ ID NOs. 13-14;
[0014] The nucleotide sequence of the eighth primer pair used to amplify the Gm04_1014450_1014481 molecular marker is shown in SEQ ID NOs. 15-16;
[0015] The nucleotide sequence of the ninth primer pair used to amplify the Gm04_47260900_47260937 molecular marker is shown in SEQ ID NOs. 17-18;
[0016] The nucleotide sequences of the tenth primer pair used to amplify the Gm07_20149006_20149007 molecular marker are shown in SEQ ID NOs. 19-20;
[0017] The nucleotide sequence of the 11th primer pair used to amplify the Gm08_44385355_44385399 molecular marker is shown in SEQ ID NOs. 21-22;
[0018] The nucleotide sequence of the 12th primer pair used to amplify the Gm09_12267638_12267680 molecular marker is shown in SEQ ID NOs. 23-24;
[0019] The nucleotide sequence of the 13th primer pair used to amplify the Gm09_33743573_33743612 molecular marker is shown in SEQ ID NOs. 25-26;
[0020] The nucleotide sequence of the 14th primer pair used to amplify the Gm09_37146920_37146952 molecular marker is shown in SEQ ID NOs. 27-28;
[0021] The nucleotide sequence of the 15th primer pair used to amplify the Gm10_21414835_21414878 molecular marker is shown in SEQ ID NOs. 29-30;
[0022] The nucleotide sequence of the 16th primer pair used to amplify the Gm11_9027110_9027111 molecular marker is shown in SEQ ID NOs. 31-32;
[0023] The nucleotide sequence of the 17th primer pair used to amplify the Gm13_28033266_28033267 molecular marker is shown in SEQ ID NOs. 33-34;
[0024] The nucleotide sequence of the 18th primer pair used to amplify the Gm13_32370867_32370901 molecular marker is shown in SEQ ID NOs. 35-36;
[0025] The nucleotide sequence of the 19th primer pair used to amplify the Gm13_34120611_34120657 molecular marker is shown in SEQ ID NOs. 37-38;
[0026] The nucleotide sequences of the 20th primer pair used to amplify the Gm13_42180045_42180091 molecular marker are shown in SEQ ID NOs. 39-40;
[0027] The nucleotide sequence of the 21st primer pair used to amplify the Gm14_7027282_7027313 molecular marker is shown in SEQ ID NOs. 41-42;
[0028] The nucleotide sequence of the 22nd primer pair used to amplify the Gm16_3283288_3283319 molecular marker is shown in SEQ ID NOs. 43-44;
[0029] The nucleotide sequence of the 23rd primer pair used to amplify the Gm16_32569599_32569630 molecular marker is shown in SEQ ID NOs. 45-46;
[0030] The nucleotide sequence of the 24th primer pair used to amplify the Gm16_35889588_35889638 molecular marker is shown in SEQ ID NOs. 47-48;
[0031] The nucleotide sequences of the 25th primer pair used to amplify the Gm16_37504818_37504854 molecular marker are shown in SEQ ID NOs. 49-50;
[0032] The nucleotide sequence of the 26th primer pair used to amplify the Gm17_36063299_36063330 molecular marker is shown in SEQ ID NOs. 51-52;
[0033] The nucleotide sequence of the 27th primer pair used to amplify the Gm17_39292061_39292062 molecular marker is shown in SEQ ID NOs. 53-54;
[0034] The nucleotide sequence of the 28th primer pair used to amplify the Gm17_39474344_39474381 molecular marker is shown in SEQ ID NOs. 55-56;
[0035] The nucleotide sequence of the 29th primer pair used to amplify the Gm18_2089656_2089688 molecular marker is shown in SEQ ID NOs. 57-58;
[0036] The nucleotide sequences of the 30th primer pair used to amplify the Gm19_47148869_47148912 molecular marker are shown in SEQ ID NOs. 59-60;
[0037] The nucleotide sequence of the 31st primer pair used to amplify the Gm20_47615536_47615570 molecular marker is shown in SEQ ID NOs. 61-62.
[0038] The present invention also provides the use of the primer pair combination of the above-mentioned InDel molecular marker combination for constructing the DNA fingerprint of the fresh soybean "Tongsu No. 1", and the primer pair combination of the InDel molecular marker combination is used to identify the fresh soybean "Tongsu No. 1".
[0039] The present invention also provides an application of the primer pair combination of the above-mentioned InDel molecular marker combination for constructing a DNA fingerprint of the fresh soybean "Tongsu No. 1", characterized in that the primer pair combination of the InDel molecular marker combination is used to identify true and false hybrids of the F1 hybrid offspring whose parents are the fresh soybean "Tongsu No. 1" and "Williams 82".
[0040] Preferably, if the PCR amplification band of the F1 hybrid offspring is not exactly the same as that of the female parent of the two parents "Williams 82" and "Tongsu No. 1", it is a true F1 hybrid;
[0041] If the PCR amplification band of the F1 hybrid offspring is completely identical only to the female parent of the two parents "Williams 82" and "Tongsu No. 1", it is a pseudo-hybrid.
[0042] The present invention also provides the application of the primer pair combination of the above-mentioned InDel molecular marker combination for constructing the DNA fingerprint of the fresh soybean "Tongsu No. 1", and the primer pair combination of the InDel molecular marker combination is used for analyzing the genetic diversity of soybean germplasm resources and the evolutionary tree of soybean variety resources.
[0043] The fingerprint of the fresh soybean "Tongsu No. 1" based on whole genome resequencing of the present invention can be used for genetic diversity analysis, identification and germplasm resource utilization.
[0044] The primer pair combination based on 31 InDel molecular markers of the present invention is not only suitable for the accurate identification of the authenticity of soybean varieties, but can also be widely used in related fields such as molecular marker-assisted breeding of soybeans.
[0045] The specific method is:
[0046] DNA extraction, library construction and whole genome resequencing of fresh soybean "Tongsu No. 1" samples were performed:
[0047] S1. Use a DNA extraction kit (Hi-DNAsecure Plant Kit, Tiangen DP350) to extract DNA from fresh young soybean leaves.
[0048] S2. The total amount of DNA was detected using a fluorescent dye (Quant-iT PicoGreen dsDNA Assay Kit); the quality of DNA extraction was detected by 0.8% agarose gel electrophoresis, and the DNA was quantified using an ultraviolet spectrophotometer.
[0049] S3. Prepare the sequencing library using the standard library construction process of Illumina's TruSeq DNA PCR-free prep kit.
[0050] S4. First, perform quality control on the library on Agilent Bioanalyzer. After gradient dilution of qualified sequencing libraries (Index sequences cannot be repeated), perform high-throughput sequencing on the machine.
[0051] Analyze the data from the resequencing of "Tongsu No. 1" and construct a fingerprint map:
[0052] S1. Further process the data to remove 3' end adapter contamination, use the sliding window method for quality filtering, and generate high-quality sequences.
[0053] S2. The resequencing data were aligned with the Williams whole genome sequence using the bwa (0.7.17-r1188) mem program. Duplicates were removed using the "MarkDuplicates" program in the Picard software package. SNP and InDel variant sites were screened and identified using the GATK program. CNV variants were identified and detected using CNVnator 0.2.7. SV variants in the genome were detected using BreakDancer 1.1 software.
[0054] S3. Use ANNOVAR software to annotate SNP and InDel sites.
[0055] S4. Based on the analysis results of SNPs, InDels, CNVs, and SVs, use the reference sequence as a ruler and use R circlize software to display them on a circular plot to visually compare the variation analysis results.
[0056] Application of fingerprint mapping based on resequencing data of Tongsu No. 1:
[0057] S1. Screen the entire genome of "Tongsu No. 1" for 31 homozygous indel variants ≥30 bp distributed across 15 chromosomes. Design primers for these indel markers and perform e-PCR amplification of the genome sequences of "Williams 82" and "Tongsu No. 1" using the Primer Check plug-in. If the PCR-amplified bands for the 31 indel markers in the F1 hybrid are identical only to the maternal parent, the hybrid is considered a pseudohybrid. If they are not identical to the maternal parent, the hybrid is considered a true F1 hybrid.
[0058] S2. Use 14 highly specific InDel molecular markers to identify the genetic diversity of 29 soybean germplasm resources, including Jidou 17, through PCR amplification or e-PCR amplification based on whole genome sequences, and construct a phylogenetic tree for evolutionary analysis.
[0059] Compared with the prior art, the present invention has the following advantages:
[0060] 1. The present invention realizes the identification and screening of SNP, InDel, CNV, and SV markers across the entire genome based on resequencing data analysis methods, with higher density and larger quantity, and more options for developing specific molecular markers, laying the foundation for the classification, identification, and utilization of soybean germplasm resources.
[0061] 2. The present invention integrates data quality control, InDel screening, and primer design to develop a specific InDel molecular marker for the fresh soybean "Tongsu No. 1". Different molecular marker combinations can be used to identify the authenticity of the hybrid F1 generation and multiple soybean germplasm resource materials. This method has good reproducibility, is fast and easy to operate, and can identify more soybean germplasm materials as needed.
[0062] 3. The primer set constructed in this invention exhibits high amplification efficiency. The products amplified by PCR or e-PCR produce clear bands with high size resolution, demonstrating significant polymorphism among different soybean varieties. This not only improves the efficiency of variety identification but also provides a strong molecular foundation for soybean genetic research and breeding.
[0063] 4. The soybean DNA fingerprint and specific molecular markers constructed by the present invention can be used for identification at the soybean seedling stage, which is of great significance for early and rapid identification. It not only improves work efficiency, but also helps to shorten the variety approval cycle and accelerate the promotion and application of new varieties. The present invention is further described in detail below with reference to the accompanying drawings and examples. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 This is the original resequencing data statistics and quality control chart of the "Tongsu No. 1" sample in Example 3 of the present invention.
[0065] Figure 2 This is the sequencing depth and coverage diagram of the "Tongsu No. 1" sample of Example 3 of the present invention.
[0066] Figure 3 This is a density distribution diagram of SNPs on chromosomes of the "Tongsu No. 1" sample of Example 3 of the present invention.
[0067] Figure 4 It is the number of different types of SNPs in the coding region of the "Tongsu No. 1" sample of Example 3 of the present invention and the number of SNPs in different regions of the genome.
[0068] Figure 5 This is the mutated base preference statistics of the "Tongsu No. 1" sample of Example 3 of the present invention.
[0069] Figure 6This is the SNP mutation spectrum analysis of the "Tongsu No. 1" sample of Example 3 of the present invention.
[0070] Figure 7 This is the InDel length distribution diagram of the "Tongsu No. 1" sample of Example 3 of the present invention.
[0071] Figure 8 It is the number of different types of Indels in the coding region of the "Tongsu No. 1" sample of Example 3 of the present invention and the number of Indels in different regions of the genome.
[0072] Figure 9 This is the SV length distribution of the "Tongsu No. 1" sample of Example 3 of the present invention.
[0073] Figure 10 This is a circos diagram of the SNP, InDel, CNV and SV variation results of the "Tongsu No. 1" sample of Example 3 of the present invention.
[0074] Figure 11 This is the e-PCR amplification band of "Williams 82" in Example 4 of the present invention based on 31 InDel molecular markers.
[0075] Figure 12 This is the e-PCR amplification band of "Tongsu No. 1" in Example 4 of the present invention based on 31 InDel molecular markers.
[0076] Figure 13 This is the genetic diversity analysis of soybean germplasm resources such as 29 soybean variety resources based on 14 InDel molecular markers in Example 4 of the present invention.
[0077] Figure 14 This is the phylogenetic tree analysis of 29 soybean variety resources based on 14 InDel molecular markers in Example 4 of the present invention. DETAILED DESCRIPTION
[0078] The fresh soybean "Tongsu No. 1" in the present invention is commercially available, and "Williams 82" is provided by the germplasm bank of the Institute of Crop Germplasm Resources, Shandong Academy of Agricultural Sciences.
[0079] Example 1
[0080] This example describes the DNA extraction, library construction, and sequencing of fresh soybean "Tongsu No. 1" leaf samples.
[0081] 1. Sample DNA extraction, purification and quality inspection
[0082] (1) DNA extraction from samples: DNA was extracted from fresh young soybean leaves using a DNA extraction kit (Hi-DNAsecure Plant Kit, Tiangen DP350). The main extraction process was as follows: about 100 mg of fresh plant tissue was taken and ground thoroughly with liquid nitrogen; GP1 was preheated at 65°C in advance and then mercaptoethanol was added to a final concentration of 0.1%. The ground tissue powder was quickly transferred to 700 μl of GP1 preheated buffer, mixed by inversion, and placed in a 65°C water bath for 20 min. The centrifuge tube was inverted 2 to 3 times; 700 μl of chloroform was added, mixed by inversion, and centrifuged at 12,000 rpm for 5 min; the upper aqueous phase was transferred to a new centrifuge tube, 700 μl of buffer GP2 was added, mixed thoroughly, and then transferred to the adsorption column CB3, centrifuged at 12,000 rpm for 30 sec, and the waste liquid was discarded; 500 μl of buffer G was added to the adsorption column CB3. D, centrifuge at 12,000 rpm for 30 seconds, discard the waste liquid, and place the adsorption column CB3 in a collection tube; add 600 μl of rinse buffer PW to the adsorption column CB3, centrifuge at 12,000 rpm for 30 seconds, discard the waste liquid, and place the adsorption column CB3 in a collection tube; return the adsorption column CB3 to the collection tube, centrifuge at 12,000 rpm for 2 minutes, and discard the waste liquid; let the adsorption column CB3 stand at room temperature for several minutes to completely dry any residual rinse liquid from the adsorption material. Note: Before use, add anhydrous ethanol to buffer GD and rinse buffer PW.
[0083] (2) Sample DNA quality inspection: The total amount of DNA was detected using a fluorescent dye (Quant-iT PicoGreen dsDNA Assay Kit); the quality of DNA extraction was detected by 0.8% agarose gel electrophoresis, and the DNA was quantified using an ultraviolet spectrophotometer; the DNA dilution working solution was stored at 4°C, and the storage solution was stored at -20°C.
[0084] 2. Second-generation library preparation
[0085] The sequencing library was prepared using the standard library construction process of Illumina's TruSeq DNA PCR-free prep kit. The main experimental process is as follows:
[0086] (1) First, the extracted DNA was randomly sheared by ultrasound and the sequence ends were repaired. The protruding bases at the 5' end of the DNA sequence were removed by EndRepair Mix2 in the kit, and a phosphate group was added to fill the missing bases at the 3' end.
[0087] (2) Adding an A base to the 3' end of the DNA sequence to prevent the DNA fragment from self-ligating and to ensure that the target sequence can be connected to the sequencing adapter (there is a protruding T base at the 3' end of the sequencing adapter).
[0088] (3) A sequencing adapter containing a library-specific tag (i.e., index sequence) is added to the 5' end of the sequence so that the DNA molecule can be fixed on the flow cell.
[0089] (4) BECKMAN AMPure XP Beads were used to remove the self-ligated fragments by magnetic bead screening and purify the library system after adding the adapter.
[0090] (5) PCR amplification of the DNA fragments connected to the adapter was performed to enrich the sequencing library template, and the library enrichment product was purified again using BECKMAN AMPure XP Beads.
[0091] (6) The library was subjected to final fragment selection and purification by 2% agarose gel electrophoresis. The insert fragment of the library was ∼450 bp.
[0092] 3. Perform high-throughput sequencing on the machine
[0093] (1) Before sequencing, the library was quality checked on an Agilent Bioanalyzer using the Agilent High Sensitivity DNA Kit. Qualified libraries had a single peak and no adapter dimers. The library was quantified using the Quant-iT PicoGreen dsDNA Assay Kit on a Promega QuantiFluor fluorescence quantification system. Qualified libraries had a concentration of 2 nM or higher.
[0094] (2) After gradient dilution of the qualified sequencing library (Index sequence cannot be repeated), mix them in the corresponding proportion according to the required sequencing amount, use NaOH to denature the double-stranded DNA library into single-stranded DNA, and use Nova Seq sequencer to perform 2×150bp double-end sequencing.
[0095] Example 2
[0096] This example analyzes the data from the resequencing of the fresh soybean "Tongsu No. 1".
[0097] 1. Obtaining high-quality resequencing data
[0098] In order to ensure the quality of subsequent information analysis, it is necessary to further filter the offline data, that is, filter the original offline data to generate high-quality sequences. The data filtering criteria mainly include:
[0099] (1) Use AdapterRemoval (version 2) to remove 3' end adapter contamination.
[0100] (2) A sliding window method was used for quality filtering, with a window size of 5 bp and a step size of 1 bp. Each time a base was moved forward, the average Q value of the window was calculated for 5 bases. If the Q value of the last base was ≤ 2, only the bases before that position were retained. If the average Q value of the window was ≤ 20, only the bases before and after the second-to-last base in the window were retained.
[0101] (3) If the length of any read in the paired-end sequence is ≤50 bp, the paired-end read is removed.
[0102] 2. Compare the resequencing data with the whole genome sequence
[0103] (1) The high-quality data obtained after filtering were aligned to the reference genome using the bwamem program (0.7.17-r1188). The alignment parameters were all based on the default parameters of bwamem.
[0104] (2) Picard 1.107 software was used to sequence the sam files and convert them into bam files. The "FixMateInformation" command was used to ensure consistency between all paired-end read information. The generation of sequencing data involves library amplification and cluster formation, and these two steps are prone to generate some duplicates.
[0105] (3) Duplicates were removed using the "MarkDuplicates" function in the Picard software package. That is, if multiple paired reads had the same chromosome coordinates after alignment, only the paired reads with the highest score were retained. Reads near InDels are most prone to mapping errors. To minimize SNPs caused by mapping errors, reads near InDels need to be re-aligned to improve the accuracy of SNP calling. The IndelRealigner command in the GATK program was used to re-align all reads near InDels to improve the accuracy of SNP prediction.
[0106] 3. SNP detection and annotation
[0107] GATK software is used to detect SNPs. The specific steps are as follows:
[0108] (1) Realignment using known InDel information. This is done in two steps: first, use the RealignerTargetCreator command in the GAT K software package to output a file containing all possible InDels; second, use the IndelRealigner command to realign all reads near InDels to improve the accuracy of SNP prediction.
[0109] (2) The UnifiedGenotyper program was used to obtain the SNP sites of the samples, with stand_call_conf set to 30 and stand_emit_conf set to 10. In order to ensure the reliability of the SNP sites, the obtained SNP sites were further filtered, and the filtering criteria were: Fisher test of strand bias (FS) ≤ 60; HaplotypeScore ≤ 13.0; MappingQuality (MQ) ≥ 40; Quality Dept h (QD) ≥ 2; ReadPosRankSum ≥ -8.0; MQRankSum > -12.5; DP ≥ 4 reads.
[0110] (3) ANNOVAR software was used to annotate SNP sites.
[0111] 4. InDel detection and annotation
[0112] (1) The UnifiedGenotyper program in the GenomeAnalysisTK v3.8 software package was used to obtain all mutation points of the sample. The stand_call_conf was set to 30, the stand_emit_conf was set to 10, and then Select Variants was used to extract InDel mutation information.
[0113] (2) To ensure the reliability of the InDel results, the InDel sites were further filtered according to the following filtering criteria: Fisher Test of Strand Bias (FS) ≤ 200; Read depth (DP) > 4; Quality Depth (QD) ≥ 2; ReadPosRankSum ≥ -20.
[0114] (3) ANNOVAR software was used to annotate INDEL sites.
[0115] 5. CNV detection and annotation
[0116] CNVnator 0.2.7 was used to detect copy number variations (CNVs) in the whole genome, with the bin_size set to 100. The results obtained by CNVnator were screened to remove the multi-copy regions in the genome. The Read Depth value was set (the Read Depth of the normal chromosome region was defined as 1), and regions with RD values ≤ 0.05 or ≥ 1.8 were retained.
[0117] 6.SV detection and annotation
[0118] BreakDancer 1.1 software was used to detect chromosome structural variations in the genome. The principle of this software for detecting chromosome structural variations is based on pair-end mapping, which uses abnormal length of pair-end read inserts for detection.
[0119] Example 3
[0120] This example constructs a fingerprint based on the resequencing data of the fresh soybean "Tongsu No. 1".
[0121] 1. Quality assessment results of resequencing raw data
[0122] Statistics and quality control of the original resequencing data of the "Tongsu No. 1" sample ( Figure 1 ), the GC content was 39.28%, the percentage of ambiguous bases (N) was 0.01%, the percentage of bases with a base call accuracy of 99% or higher (Q20) was 98.33%, and the percentage of bases with a base call accuracy of 99.9% or higher (Q30) was 95.78%. The average quality of the single bases in the R1 and R2 sequencing reads ranged from 35 to 40, with an average estimated error rate of less than 0.016%. The sequencing data showed no AT / GC splitting, with a nearly consistent AT and GC count. These results indicate good sequencing quality.
[0123] 2. Filtering and preprocessing results of resequencing raw data
[0124] The resequencing raw data were further filtered to generate high-quality sequences. The number of high-quality reads was 172,796,170, the percentage of high-quality reads to the original reads was 96.57%, the number of high-quality read bases was 2,580,886,9698, the percentage of high-quality read bases to the total number of original bases was 95.52%, and the sequencing depth was 27.08×.
[0125] 3. Comparison of resequencing data with whole genome sequences
[0126] (1) Overview of comparison results
[0127] From the sequence alignment results, we can see that the total number of reads is 172796170, the number of reads aligned to the reference genome is 172015034, and the percentage of reads aligned to the reference genome is 99.55% of the total number of reads. The number of repeat sequence reads is 36610129, and the percentage of repeat sequence reads to the total number of reads is 21.19%.
[0128] From the sequencing depth and coverage statistics, we can see that ( Figure 2 ), with an average coverage depth of 27.8×, and the percentages of sites covered by at least 1, 4, 10, and 20 sequences in the reference genome were 93.04%, 85.35%, 71.06%, and 38.70% of the genes, respectively.
[0129] (2) SNP detection, analysis, and annotation
[0130] SNP statistics using a sliding window revealed a total of 1,744,120 SNPs, 354,610 heterozygous genotypes, and 1,389,510 homozygous genotypes. The number of transitions between purines or pyrimidines, Ts, was 1,150,576, the number of transversions between purines and pyrimidines, Tv, was 593,544, and the transition / transversion ratio, Ts / Tv, was 1.93.
[0131] Statistical analysis of the distribution of SNPs on each chromosome ( Figure 3 ), it can be seen that there are a total of 1720673 SNP sites distributed in 20 nuclear chromosomes and chloroplast and mitochondrial genomes, among which 92186, 93083, 105015, 72168, 52741, 100165, 68067, 52508, 74465, 85153, 42687, 43660, 90144, 59929, 149728, 114377, 62241, 172572, 120913, and 68781 SNP sites are found in chromosomes 1-20, respectively. The distribution densities were 628.43, 541.46, 447.10, 709.50, 801.55, 508.62, 660.37, 899.43, 679.15, 606.42, 928.71, 951.24, 501.70, 832.54, 359.01, 333.21, 670.63, 337.75, 424.05, and 695.63 / bp, respectively. There were 45 SNP sites in both the chloroplast and chromosome genomes, with distribution densities of 8945.73 and 3382.62 / bp, respectively.
[0132] According to the SNP annotation results ( Figure 4 ) showed that a total of 71,135 SNP sites were detected in the exon region, of which 30,793 were synonymous mutations, 38,288 were non-synonymous mutations, 835 SNPs resulted in the acquisition of stop codons, 155 SNPs resulted in the loss of stop codons, and 1,064 sites of unknown function due to errors in the gene structure annotation database used for annotation. There are 391, 3605, 1873, 6, 2, 1724, 188028, 1284767, 17844, 18951, 43, 78470, 75944, and 4942 SNP sites located in the 2bp region of splicing junction, non-coding ncRNA region, exon region of non-coding ncRNA region, 2bp region of splicing junction of non-coding ncRNA, intron region of non-coding ncRNA region, intron region, intergenic region, 5'UTR region, 3'UTR region, region between 5'UTR and 3'UTR, region 1kb upstream of transcription start site, region 1kb downstream of transcription stop site, and regions upstream and downstream of transcription start site, respectively.
[0133] According to the statistical results of SNP mutation base preference ( Figure 5 ) It can be seen that compared with the base sequences before and after the SNP mutation site, after the SNP mutation occurs, the A and T bases decrease, and the G and C bases increase.
[0134] The results of SNP mutation spectrum analysis ( Figure 6 ) shows that genomic SNP mutations are divided into 6 categories: T:A>C:G, C:G>T:A, T:A>G:C, T:A>A:T, C:G>G:C, C:G>A:T, among which Ts transition mutations T:A>C:G and C:G>T:A are the main SNP mutation types, accounting for 65.97%, and the other four Tv transversion mutations account for 34.03%.
[0135] (3) InDel detection, analysis, and annotation
[0136] After filtering, a total of 232,213 InDel sites were obtained, including 28,149 heterozygous sites and 204,064 homozygous sites; 113,091 Insertion sites and 119,122 Deletion sites. Figure 7 It can be seen that the ratios of Insertion sites and Deletion sites of different lengths are basically the same.
[0137] Annotate the sites ( Figure 8), there are 2986 InDel sites distributed in the exon region, 652 frameshift deletions lead to changes in the reading frame of protein-coding genes (non-integer multiples of length 3), 521 frameshift insertions lead to changes in the reading frame of protein-coding genes (non-integer multiples of length 3), 900 non-frameshift deletions do not change the reading frame of protein-coding genes (integer multiples of length 3), 810 non-frameshift insertions do not change the reading frame of protein-coding genes (integer multiples of length 3), 39 InDels lead to the acquisition of stop codons, 5 InDels lead to the loss of stop codons, 143 are located in the splicing junction 2bp region, 565 are located in the non-coding RNA region, 276 are located in the exon region of the non-coding region, and 1 is located in the splicing junction of the non-coding RNA. 2 bp regions, 288 were located in intronic regions of noncoding regions, 36,503 were located in intronic regions, 149,916 were located in intergenic regions, 5,835 were located in 5'UTR regions, 4,857 were located in 3'UTR regions, 4 were located in the region between 5'UTR and 3'UTR, 15,025 were located in the region 1 kb upstream of the transcription start site, 15,277 were located in the region 1 kb downstream of the transcription stop site, and 1,102 were located in the regions upstream and downstream of the transcription start site.
[0138] (4) CNV and SV detection and analysis
[0139] After identification, it was found that the total number of CNV variations was 8732, of which 739 were copy number deletions and 7993 were copy number gains.
[0140] After detection, it was found that the total number of SV mutations was 5749, including 1926 deletions (DEL), 178 inversions (INV), 2255 intra-chromosomal translocations (ITX), and 1390 inter-chromosomal translocations (CTX). Figure 9 ), 201-300bp and >1000bp were the largest, both greater than 1000.
[0141] (5) Draw genomic SNP, InDel, CNV, and SV site variation maps
[0142] According to the analysis results of SNP, InDel, CNV and SV, the reference sequence was used as the ruler and the R circlize software was used to display them on a circular diagram ( Figure 10 ), the distribution of variation analysis results can be compared intuitively.
[0143] Example 4
[0144] This embodiment is an application of a fingerprint map constructed from resequencing data.
[0145] 1. Identify the authenticity of soybean F1 hybrid seeds
[0146] The whole genomes of Williams 82 and Tongsu 1 were screened for 31 homozygous InDel variants ≥30 bp distributed in 15 chromosomes (Table 1), and primers were designed for InDel molecular markers using the Batch Target Region Primer Design plug-in of TBtools software (Table 2).
[0147] Based on 31 InDel molecular markers, the genome sequences of Williams 82 and Tongsu No. 1 were amplified by e-PCR using the Primer Check plug-in ( Figure 11 , Figure 12 ), and obtain fingerprints of "Williams 82" and "Tongsu No. 1" based on 31 InDel molecular markers, respectively. Amplify the fingerprints of the F1 hybrid offspring to be identified based on the 31 InDel molecular markers. If the 31 InDel molecular marker PCR amplification bands of the F1 hybrid offspring are completely identical only to those of the maternal parent of "Williams 82" and "Tongsu No. 1", it is a pseudo-hybrid. If the 31 InDel molecular marker PCR amplification bands of the F1 hybrid offspring are not completely identical to those of the maternal parent of "Williams 82" and "Tongsu No. 1", it can be determined to be a true F1 hybrid.
[0148] Table 1 Information of 31 InDel molecular markers
[0149]
[0150]
[0151] Table 231 InDel molecular marker PCR amplification primer information
[0152]
[0153]
[0154] 2. Conduct genetic diversity analysis of soybean germplasm resources
[0155] Use Gm02_48831944_48831999, Gm02_50061608_50061609, Gm03_3866902_3866939, Gm04_1014450_1014481, Gm07_20149006_20149007, Gm08_44385355_44385399, Gm09_12267638_12267680, Gm10_21 Fourteen InDel molecular markers, including 414835_21414878, Gm13_28033266_28033267, Gm13_32370867_32370901, Gm13_42180045_42180091, Gm16_35889588_35889638, Gm17_39474344_39474381, and Gm18_2089656_2089688, were detected by PCR amplification or e-PCR amplification based on whole genome (SoyOmics, https: / / ngdc.cncb.ac.cn / soyomics / index; LIS, https: / / www.legumeinfo.org / collections / glycine / ) sequence was used to identify Jidou 17, Keshan 1, Qihuang 34, 58-161, Amsoy, Fengdihuang, Handou 5, The genetic diversity of 29 soybean germplasm resources including Heihe 43, Jindou 23, Juxuan 23, PI_398296, PI_548362, PI_562565, PI_578357, PI483463, Zhonghuang 13, Shisheng Changye, Zihua 4, Yudou 22, Mancangjin, Huaxia 3, Tiefeng 18, Wandou 28, Zhonghuang 35, Hwangkeum, Hefeng 25, Fiskeby III, Tongshan Tianedan, and Wenfeng 7 was analyzed. Figure 13 ), the evolutionary analysis of germplasm resources was performed using the R language ggtree package, and an evolutionary tree was constructed ( Figure 14 ).
[0156] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any way. Any simple modification, change and equivalent variation made to the above embodiment based on the essence of the invention technology shall still fall within the scope of protection of the technical solution of the present invention.
Claims
1. A primer pair combination of an InDel molecular marker combination for constructing a DNA fingerprint of fresh soybean "Tongsu No. 1", characterized in that: The primer pair combination consists of primers shown in SEQ ID NOs. 1 to 62.
2. A primer pair combination for constructing the DNA fingerprint of fresh soybean "Tongsu No. 1" as claimed in claim 1, characterized in that: The application is to identify the fresh soybean "Tongsu No. 1".
3. A use of a primer pair combination of an InDel molecular marker combination for constructing a DNA fingerprint of fresh soybean "Tongsu No. 1" as claimed in claim 1, characterized in that: The application is to identify true and false hybrids of F1 hybrid offspring whose parents are fresh soybeans "Tongsu No. 1" and "Williams 82".
4. The use according to claim 3, characterized in that If the PCR amplified bands of the F1 hybrid offspring are not identical to those of the female parent of the two parents "Williams 82" and "Tongsu No. 1", it is a true F1 hybrid; If the PCR amplification band of the F1 hybrid offspring is completely identical only to the female parent of the two parents "Williams 82" and "Tongsu No. 1", it is a pseudo-hybrid.
5. A use of a primer pair combination of an InDel molecular marker combination for constructing a DNA fingerprint of fresh soybean "Tongsu No. 1" as claimed in claim 1, characterized in that: The application is the genetic diversity of soybean germplasm resources and the evolutionary tree analysis of soybean variety resources.