A primer pair combination for constructing a SNP molecular marker combination of peony DNA fingerprint
Patent Information
- Application Number
- CN202511388441.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-26
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2045-09-26
AI Technical Summary
现有技术中尚未见报道能够专门用于构建牡丹品种DNA身份认证体系,并可在苗期实现早期、快速、精准鉴定,同时兼顾群体遗传结构解析与真假杂种高效鉴定的SNP分子标记解决方案
1、本发明基于牡丹全基因组重测序数据,实现了在全基因组范围内高效鉴定与筛选SNP标记。该方法可获得密度更高、数量更丰富的SNP位点,显著增加了特异性分子标记的开发可选性,为牡丹种质资源的精准分类、鉴定及高效利用奠定了坚实的技术基础。
Smart Images

Figure CN121450825B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of DNA fingerprinting technology, specifically relating to a primer pair combination for constructing a DNA fingerprint of peony and its application. Background Technology
[0002] Peony (Paeonia suffruticosa) is a famous ornamental and medicinal plant unique to my country, possessing extremely high economic and cultural value. With the vigorous development of breeding work, the number of new peony varieties has surged, making variety identification and intellectual property protection key issues restricting the healthy development of the industry. Traditional identification methods rely heavily on morphological characteristics such as flower shape, color, and flowering period. However, these traits are easily affected by environmental conditions, cultivation practices, and plant development stages, resulting in poor stability and long identification cycles, which cannot meet the urgent needs of modern seed industry for variety rights protection, market supervision, and breeding efficiency.
[0003] Molecular marker technology provides a stable and reliable new approach for variety identification. Currently, studies have utilized molecular marker technologies such as SSR (simple sequence repeat), ISSR (inter-simple sequence repeat), and SRAP (sequence-related amplified polymorphism) to construct DNA fingerprints for peonies and applied them to genetic diversity analysis and hybrid progeny identification. However, these marker types generally suffer from limitations such as low polymorphism, limited detection throughput, poor reproducibility of experimental results, or high development costs, making it difficult to achieve high-throughput, standardized, and cross-platform accurate identification. Especially for superior varieties with highly similar genetic backgrounds, the resolution of these markers is often insufficient, failing to provide enough information to accurately reveal the genetic structure of the population in genetic diversity assessment, and also making it difficult to efficiently and accurately distinguish between true hybrids and false hybrids or self-pollinated seedlings in hybrid identification.
[0004] Single nucleotide polymorphisms (SNPs), as third-generation molecular markers, possess inherent advantages such as abundant quantity, wide distribution, high genetic stability, and ease of high-throughput automated detection, providing a new technical approach for high-precision genetic diversity analysis and early identification of the authenticity of hybrid offspring. Although SNP markers are widely used in major crops, a fully validated, high-resolution, and widely applicable core SNP marker combination and corresponding specific primer pairs are still lacking in most peony germplasm resources. Currently, there are no reported SNP molecular marker solutions specifically designed for constructing a DNA identification system for peony varieties, enabling early, rapid, and accurate identification at the seedling stage, while also facilitating population genetic structure analysis and efficient identification of true and false hybrids.
[0005] Therefore, there is an urgent need in this field to develop a set of specific SNP molecular markers and their primer pairs for peonies that are highly polymorphic and stable, in order to overcome the shortcomings of existing technologies. This would not only be used to construct a standardized peony DNA fingerprint database to support variety approval, rights protection and market traceability, but also serve the diversity assessment of peony genetic resources, the construction of core germplasm and the identification of the authenticity of offspring in hybridization breeding, thus comprehensively promoting the process of peony molecular breeding and intellectual property protection. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a primer pair combination for constructing a DNA fingerprint of peony by means of SNP molecular marker combination and its application, which is used for the analysis, identification and utilization of genetic diversity of peony varieties and germplasm resources.
[0007] This invention aims to identify peony varieties “Huafen Xinzhuang” and / or “Hezhi Chuyan”; to identify true and false hybrids in the pairwise hybrids of “Fengdan”, “Huafen Xinzhuang”, and “Hezhi Chuyan”; and to analyze the genetic diversity of peony germplasm resources and the phylogenetic tree of peony variety resources. The primer pairs of the SNP molecular marker combinations in this invention are used for genetic diversity analysis, identification, and germplasm resource utilization; and to identify true and false hybrids in the hybrid offspring of peony parents with different genotypes at multiple SNP loci.
[0008] To solve the above technical problems, the technical solution adopted by the present invention is: a primer pair combination for constructing a peony DNA fingerprint pattern using SNP molecular marker combinations, wherein the primer pair combination consists of primers shown in SEQ ID NO.1 to 122; The DNA fingerprint of peony consists of 61 molecular markers developed from loss-of-function SNP mutations located in exons across 5 chromosomes, all of which cause premature codon termination or termination gain. The nucleotide sequences of the first primer pair used to amplify the Chr01_13974972_G / A molecular marker are shown in SEQ ID NO. 1-2; The nucleotide sequences of the second primer pair used to amplify the Chr01_52764705_C / T molecular marker are shown in SEQ ID NO. 3-4; The nucleotide sequence of the third primer pair used to amplify the Chr01_301779146_T / G molecular marker is shown in SEQ ID NO. 5-6; The nucleotide sequence of the fourth primer pair used to amplify the Chr01_582511719_C / T molecular marker is shown in SEQ ID NO. 7-8; The nucleotide sequence of the fifth primer pair used to amplify the Chr01_604023402_C / T molecular marker is shown in SEQ ID NO. 9-10; The nucleotide sequence of the sixth primer pair used to amplify the Chr01_1806171054_T / C molecular marker is shown in SEQ ID NO. 11-12; The nucleotide sequence of the 7th primer pair used to amplify the Chr01_1918223993_C / T molecular marker is shown in SEQ ID NO. 13-14; The nucleotide sequence of the 8th primer pair used to amplify the Chr01_1959229511_C / T molecular marker is shown in SEQ ID NO. 15-16; The nucleotide sequence of the 9th primer pair used to amplify the Chr01_2120097634_T / A molecular marker is shown in SEQ ID NO. 17-18; The nucleotide sequence of the 10th primer pair used to amplify the Chr01_2345730173_A / T molecular marker is shown in SEQ ID NO.19-20; The nucleotide sequence of the 11th primer pair used to amplify the Chr02_340870593_G / A molecular marker is shown in SEQ ID NO. 21-22; The nucleotide sequence of the 12th primer pair used to amplify the Chr02_407187313_T / G molecular marker is shown in SEQ ID NO. 23-24; The nucleotide sequence of the 13th primer pair used to amplify the Chr02_540706395_A / G molecular marker is shown in SEQ ID NO. 25-26; The nucleotide sequence of the 14th primer pair used to amplify the Chr02_650741859_C / T molecular marker is shown in SEQ ID NO. 27-28; The nucleotide sequence of the 15th primer pair used to amplify the Chr02_702603442_A / G molecular marker is shown in SEQ ID NO. 29-30; The nucleotide sequence of the 16th primer pair used to amplify the Chr02_1710187145_C / A molecular marker is shown in SEQ ID NO.31-32; The nucleotide sequence of the 17th primer pair used to amplify the Chr02_2097523064_T / C molecular marker is shown in SEQ ID NO.33-34; The nucleotide sequence of the 18th primer pair used to amplify the Chr03_33351259_C / T molecular marker is shown in SEQ ID NO. 35-36; The nucleotide sequence of the 19th primer pair used to amplify the Chr03_87443773_T / G molecular marker is shown in SEQ ID NO. 37-38; The nucleotide sequence of the 20th primer pair used to amplify the Chr03_132275045_C / T molecular marker is shown in SEQ ID NO. 39-40; The nucleotide sequence of the 21st primer pair used to amplify the Chr03_280143517_A / G molecular marker is shown in SEQ ID NO. 41-42; The nucleotide sequence of the 22nd primer pair used to amplify the Chr03_500972943_T / C molecular marker is shown in SEQ ID NO. 43-44; The nucleotide sequence of the 23rd primer pair used to amplify the Chr03_699473250_C / T molecular marker is shown in SEQ ID NO. 45-46; The nucleotide sequence of the 24th primer pair used to amplify the Chr03_851012098_C / T molecular marker is shown in SEQ ID NO. 47-48; The nucleotide sequence of the 25th primer pair used to amplify the Chr03_855211440_C / A molecular marker is shown in SEQ ID NO. 49-50; The nucleotide sequence of the 26th primer pair used to amplify the Chr03_936166517_C / A molecular marker is shown in SEQ ID NO. 51-52; The nucleotide sequence of the 27th primer pair used to amplify the Chr03_936166860_A / G molecular marker is shown in SEQ ID NO. 53-54; The nucleotide sequence of the 28th primer pair used to amplify the Chr03_1055906550_G / A molecular marker is shown in SEQ ID NO. 55-56; The nucleotide sequence of the 29th primer pair used to amplify the Chr03_1173465922_A / C molecular marker is shown in SEQ ID NO. 57-58; The nucleotide sequence of the 30th primer pair used to amplify the Chr03_1173465930_T / A molecular marker is shown in SEQ ID NO.59-60; The nucleotide sequence of the 31st primer pair used to amplify the Chr03_1216736027_T / A molecular marker is shown in SEQ ID NO. 61-62; The nucleotide sequence of the 32nd primer pair used to amplify the Chr03_1234693042_G / T molecular marker is shown in SEQ ID NO. 63-64; The nucleotide sequence of the 33rd primer pair used to amplify the Chr03_1246856404_C / T molecular marker is shown in SEQ ID NO. 65-66; The nucleotide sequence of the 34th primer pair used to amplify the Chr03_1250250494_C / T molecular marker is shown in SEQ ID NO. 67-68; The nucleotide sequence of the 35th primer pair used to amplify the Chr03_1414051447_C / T molecular marker is shown in SEQ ID NO. 69-70; The nucleotide sequence of the 36th primer pair used to amplify the Chr03_1622577802_C / T molecular marker is shown in SEQ ID NO.71-72; The nucleotide sequence of the 37th primer pair used to amplify the Chr03_1704789394_T / C molecular marker is shown in SEQ ID NO.73-74; The nucleotide sequence of the 38th primer pair used to amplify the Chr04_108838226_T / A molecular marker is shown in SEQ ID NO. 75-76; The nucleotide sequence of the 39th primer pair used to amplify the Chr04_151772272_G / A molecular marker is shown in SEQ ID NO. 77-78; The nucleotide sequence of the 40th primer pair used to amplify the Chr04_520822950_T / C molecular marker is shown in SEQ ID NO. 79-80; The nucleotide sequence of the 41st primer pair used to amplify the Chr04_680727619_A / T molecular marker is shown in SEQ ID NO. 81-82; The nucleotide sequence of the 42nd primer pair used to amplify the Chr04_1030650704_G / T molecular marker is shown in SEQ ID NO. 83-84; The nucleotide sequence of the 43rd primer pair used to amplify the Chr04_1071214861_G / A molecular marker is shown in SEQ ID NO. 85-86; The nucleotide sequence of the 44th primer pair used to amplify the Chr04_1109376597_C / A molecular marker is shown in SEQ ID NO. 87-88.
[0009] The nucleotide sequence of the 45th primer pair used to amplify the Chr04_1109376647_G / T molecular marker is shown in SEQ ID NO. 89-90; The nucleotide sequence of the 46th primer pair used to amplify the Chr04_1140925370_G / A molecular marker is shown in SEQ ID NO. 91-92; The nucleotide sequence of the 47th primer pair used to amplify the Chr04_1553032924_A / G molecular marker is shown in SEQ ID NO. 93-94; The nucleotide sequence of the 48th primer pair used to amplify the Chr04_1787787355_T / A molecular marker is shown in SEQ ID NO. 95-96; The nucleotide sequence of the 49th primer pair used to amplify the Chr04_2017263232_C / A molecular marker is shown in SEQ ID NO. 97-98; The nucleotide sequence of the 50th primer pair used to amplify the Chr04_2230098669_A / G molecular marker is shown in SEQ ID NO. 99-100; The nucleotide sequence of the 51st primer pair used to amplify the Chr04_2285760072_A / G molecular marker is shown in SEQ ID NO. 101-102; The nucleotide sequence of the 52nd primer pair used to amplify the Chr04_2330967850_G / T molecular marker is shown in SEQ ID NO. 103-104; The nucleotide sequence of the 53rd primer pair used to amplify the Chr04_2397495861_A / G molecular marker is shown in SEQ ID NO. 105-106; The nucleotide sequence of the 54th primer pair used to amplify the Chr05_123298388_A / C molecular marker is shown in SEQ ID NO. 107-108; The nucleotide sequence of the 55th primer pair used to amplify the Chr05_248171257_G / C molecular marker is shown in SEQ ID NO. 109-110; The nucleotide sequence of the 56th primer pair used to amplify the Chr05_337081376_G / T molecular marker is shown in SEQ ID NO. 111-112; The nucleotide sequence of the 57th primer pair used to amplify the Chr05_807239075_G / A molecular marker is shown in SEQ ID NO. 113-114; The nucleotide sequence of the 58th primer pair used to amplify the Chr05_1529394053_A / T molecular marker is shown in SEQ ID NO. 115-116; The nucleotide sequence of the 59th primer pair used to amplify the Chr05_1535053858_G / A molecular marker is shown in SEQ ID NO. 117-118; The nucleotide sequence of the 60th primer pair used to amplify the Chr05_1741394421_G / A molecular marker is shown in SEQ ID NO. 119-120; The nucleotide sequence of the 61st primer pair used to amplify the Chr05_2326548240_T / C molecular marker is shown in SEQ ID NO. 121-122.
[0010] This invention also provides the application of the primer pair combination of the above-mentioned SNP molecular marker combination for constructing the DNA fingerprint of peony, characterized in that the application is for identifying the peony varieties “Huafen Xinzhuang” and / or “Hezhi Chuyan”.
[0011] This invention also provides the application of the primer pair combination of the above-mentioned SNP molecular marker combination for constructing the DNA fingerprint of peony, characterized in that the application is to identify the true and false hybrids of the offspring of pairwise hybridization in "Fengdan", "Huafen Xinzhuang" and "Hezhi Chuyan".
[0012] Preferably, the hybrid offspring exhibit a heterozygous genotype at all loci where the genotypes of the parents differ. If the hybrid offspring exhibit only a homozygous genotype that is completely identical to the maternal parent, then it is a false hybrid.
[0013] This invention also provides the application of the primer pair combination of the above-mentioned SNP molecular marker combination for constructing the DNA fingerprint of peony, the application being the analysis of genetic diversity of peony germplasm resources and phylogenetic tree of peony variety resources.
[0014] This invention utilizes primer pairs based on 61 SNP molecular markers. It is applicable not only to genetic diversity analysis of different peony varieties but also to phylogenetic assessment, core germplasm construction, molecular breeding-assisted selection, and hybrid authenticity identification. It boasts advantages such as high throughput, good reproducibility, and strong resolution. This set of markers comprises 61 polymorphic SNPs, covering all five chromosomes of the peony genome. The rich polymorphism significantly improves the accuracy and reliability of identification, effectively avoiding misjudgments caused by individual marker failure or mutation. It provides stable and reliable technical support for peony germplasm resource management, new variety protection, and genetic research. The specific method is as follows: Whole-genome resequencing was performed using DNA extracted from two peony varieties, “Huafen Xinzhuang” and “Hezhi Chuyan”, followed by library construction and sequencing. S1. Extract DNA from fresh, tender peony leaves.
[0015] S2. Detect the quality of DNA extraction and quantify the DNA.
[0016] S3. Prepare sequencing libraries using the standard library preparation procedure of the reagent kit.
[0017] S4. Perform quality control on the library, then dilute the qualified sequencing library in a gradient and sequence it.
[0018] The resequencing data of two peony varieties, “Huafen Xinzhuang” and “Hezhi Chuyan”, were analyzed to construct fingerprint profiles. S1. Remove contaminants from the 3' end connector using a sliding window method for quality filtration.
[0019] S2. First, the resequencing data was compared with the reference genome of *Paeonia ostii* using bwa software. Then, the "MarkDuplicates" module in the Picard toolkit was used to mark and remove PCR repetitive sequences to reduce interference from repetitive fragments in variant detection. Finally, the GATK workflow was used to screen and identify SNP variant sites. Based on the above high-quality SNP dataset, a DNA fingerprint was further constructed.
[0020] S3. SNP sites were annotated using ANNOVAR software.
[0021] Application of fingerprinting based on resequencing data of two peony varieties, “Huafen Xinzhuang” and “Hezhi Chuyan”: S1. Two peony varieties, “Huafen Xinzhuang” and “Hezhi Chuyan”, were selected to identify homozygous loss-of-function SNPs that were distributed across five chromosomes throughout the entire genome, located in exons, and caused premature codon termination or termination gain. Primers were designed for the SNP molecular markers, and the genetic diversity of different peony germplasm resources was identified by PCR amplification and Sanger sequencing. A phylogenetic tree was constructed for evolutionary analysis.
[0022] S2. Using 61 specific SNP molecular markers, the genotypes of the hybrid offspring at different variation sites were determined by PCR amplification and Sanger sequencing. If the hybrid offspring exhibited a heterozygous genotype at all sites where the genotypes of the parents differed, it was a true hybrid; if it only exhibited a homozygous genotype that was completely identical to the maternal parent, it was a false hybrid (maternal self-fertilization or parthenogenesis).
[0023] Specifically, the method for determining a true hybrid is as follows: First, identify all loci where the parents' genotypes differ. Then, test the genotypes of the hybrid offspring at these specific loci. If the offspring exhibit 100% heterozygosity at all such loci, then the offspring can be determined to be a true hybrid.
[0024] Compared with the prior art, the present invention has the following advantages: 1. This invention, based on whole-genome resequencing data of peony, achieves efficient identification and screening of SNP markers across the entire genome. This method yields SNP loci with higher density and greater abundance, significantly increasing the availability of specific molecular markers and laying a solid technical foundation for the accurate classification, identification, and efficient utilization of peony germplasm resources.
[0025] 2. This invention successfully developed a specific SNP molecular marker system for the peony varieties "Huafen Xinzhuang" and "Hezhi Chuyan" by integrating data quality control, SNP screening, and primer design. Using different combinations of molecular markers, the authenticity of hybrid offspring can be rapidly identified, and the method is applicable to the identification of multiple peony germplasm resources. This method exhibits good reproducibility, is simple and efficient to operate, and has good scalability, enabling it to be widely applied to the identification of more peony germplasm materials as needed.
[0026] 3. The marker set constructed in this invention consists of 61 polymorphic SNPs, covering all five chromosomes of the peony genome. This rich polymorphism significantly improves the accuracy and reliability of identification, effectively avoiding misjudgments caused by individual marker failure or mutation. Furthermore, these 61 SNPs were developed from loss-of-function mutant SNPs, which directly lead to gene function interruption. Their advantage lies in directly linking diversity analysis with phenotypic variation, efficiently identifying key germplasm resources related to important agronomic traits (such as flower color, resistance, and flowering time) at the genetic level. This not only enhances the efficiency and accuracy of marker-assisted breeding but also provides crucial predictive power for designing hybrid combinations and introducing new alleles.
[0027] 4. The peony DNA fingerprint and specific molecular markers constructed in this invention can achieve accurate identification in the peony seedling stage, and have important value for early and rapid identification. This technology not only significantly improves the efficiency of variety identification, but also helps to shorten the variety approval cycle and accelerate the promotion and application of new varieties.
[0028] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. Attached Figure Description
[0029] Figure 1 These are the single base mass distribution map, read average error rate distribution map, and read sequencing base content distribution map of the peony variety "Huafen Xinzhuang" sample in Example 2 of this invention.
[0030] Figure 2 These are the single base mass distribution map, read average error rate distribution map, and read sequencing base content distribution map of the peony variety "Hezhi Chuyan" sample in Example 2 of this invention.
[0031] Figure 3 This is a sequencing depth and coverage map of the peony variety "Huafen Xinzhuang" sample from Example 3 of this invention. The left figure shows the sequencing depth distribution of the sample, with the horizontal axis representing sequencing depth and the vertical axis representing the fraction of basses at that sequencing depth. The right figure shows the cumulative base ratio at different sequencing depths, with the horizontal axis representing sequencing depth and the vertical axis representing the fraction of basses exceeding that sequencing depth.
[0032] Figure 4 This is a sequencing depth and coverage map of the peony variety "Hezhi Chuyan" from Example 3 of this invention. The left figure shows the sequencing depth distribution of the sample, with the horizontal axis representing sequencing depth and the vertical axis representing the fraction of basses at that sequencing depth. The right figure shows the cumulative sequencing depth at different sequencing depths, with the horizontal axis representing sequencing depth and the vertical axis representing the fraction of basses exceeding that sequencing depth.
[0033] Figure 5 This is a density distribution diagram of SNPs on chromosomes for two peony varieties, “Hua Fen Xin Zhuang” and “He Zhi Chu Yan”, from Example 3 of this invention.
[0034] Figure 6 This refers to the number of different types of SNPs in the coding regions and the number of SNPs in different regions of the genome of two peony varieties, “Hua Fen Xin Zhuang” and “He Zhi Chu Yan”, in Example 3 of this invention.
[0035] Figure 7 This is a statistical analysis of the mutation base preference of two peony varieties, “Hua Fen Xin Zhuang” and “He Zhi Chu Yan”, in Example 3 of this invention.
[0036] Figure 8 This is an SNP mutation spectrum analysis of two peony varieties, “Hua Fen Xin Zhuang” and “He Zhi Chu Yan”, from Example 3 of this invention.
[0037] Figure 9This is the distribution of the 61 SNP molecular markers on the chromosome in Example 4 of the present invention.
[0038] Figure 10 This is the sequence alignment result of three peony cultivar resources, “Danfeng”, “Huafen Xinzhuang”, and “Hezhi Chuyan”, based on 61 SNP molecular markers, according to Example 4 of the present invention.
[0039] Figure 11 This is an phylogenetic analysis of three peony varieties, “Danfeng”, “Huafen Xinzhuang”, and “Hezhi Chuyan”, based on 61 SNP molecular markers, according to Example 4 of the present invention. Detailed Implementation
[0040] In this invention, the two peony varieties "Hua Fen Xin Zhuang" and "He Zhi Chu Yan" were both bred by the Heze Academy of Agricultural Sciences through hybridization of "Feng Dan Bai" and "Jitsugetsunishiki".
[0041] 'Hua Fen Xin Zhuang' is a hybrid of 'Feng Dan Bai' and 'Jitsugetsunishiki'. It first bloomed in 2018 and began propagation in 2020. The plant can grow up to 1 meter tall and blooms in the early to mid-stages. The flowers are rose pink with the edges of the petals gradually turning light pink. There are pinkish-purple spots at the base of the petals, and the flowers can reach a diameter of 17 cm.
[0042] "Hezhi Chuyan": A hybrid of "Fengdanbai" and "Jitsugetsunishiki", it first bloomed in 2018 and began to propagate in 2020. It has an upright plant shape, strong growth, and the stem can grow up to 85cm tall. The flowers are pink with a rose-red blush at the base and the flower diameter is 15cm.
[0043] Example 1 This example describes the extraction, library construction, and high-throughput sequencing of total genomic DNA from leaf samples of two peony varieties, “Huafen Xinzhuang” and “Hezhi Chuyan”.
[0044] 1. Total genomic DNA extraction Young leaves of the peony varieties “Huafen Xinzhuang” and “Hezhi Chuyan” were taken respectively, and total genomic DNA was extracted from them. The DNA extraction quality was detected by 0.8% agarose gel electrophoresis, and the DNA was quantified by ultraviolet spectrophotometer. The DNA was stored at -80℃ for sequencing.
[0045] 2. Sequencing library preparation Libraries with 400 inserts were constructed, and paired-end (PE) sequencing was performed on these libraries using next-generation sequencing (NGS) on the BGI T7 sequencing platform. One library was constructed for each sample.
[0046] (1) First, the extracted DNA was randomly broken by ultrasound and then the sequence ends were repaired by end repair reaction solution.
[0047] (2) Add MGIEasy PF Adapters to DNA for adapter ligation.
[0048] (3) Magnetic bead purification was performed using DNA Clean Beads from the MGIEasy DNA Purification Magnetic Bead Kit.
[0049] (4) Perform PCR amplification on the above DNA fragments to enrich the sequencing library template, and then purify the library enrichment product again with En-Beads.
[0050] (5) The library was finally selected and purified by 2% agarose gel electrophoresis. The inserted fragments in the library were ~400bp.
[0051] (6) The above library was subjected to hybridization capture of exon regions using the Agilent SureSelect Human All Exon V6 kit.
[0052] 3. High-throughput sequencing (1) Before sequencing, the library must be strictly inspected. First, the Agilent High Sensitivity DNA Kit is used to detect the library on the Agilent Bioanalyzer. A qualified library should show a single peak and have no adapter dimers.
[0053] (2) Subsequently, fluorescence quantification was performed on the Promega QuantiFluor system using the Quant-iT PicoGreen dsDNA Assay Kit to ensure that the library concentration was not less than 2 nM.
[0054] (3) After the qualified libraries are serially diluted, they are mixed in a specific ratio according to the sequencing volume requirements and denatured into single strands by NaOH.
[0055] (4) Finally, 2×150 bp paired-end sequencing was performed on the DNBSEQ-T7 sequencer.
[0056] The reference genome sequence was downloaded from the genome database of the National Gene Bank Life Big Data Platform (CNGBdb) (https: / / db.cngb.org / search / project / CNP0003098 / ), and is the cultivated peony Paeonia ostii “Fengdan”.
[0057] Example 2 This example is to conduct a preliminary evaluation and filtering analysis on the original off-machine data of high-throughput sequencing of two peony varieties, "Huafen Xinzhuang" and "Hezhi Chuyan", to obtain high-quality data.
[0058] 1. Quality assessment of raw re-sequencing data Statistical analysis was performed on the original re-sequencing data of the samples ( Figure 1 , Figure 2 , Table 1). It can be seen that the proportion of ambiguous bases in both "Huafen Xinzhuang" and "Hezhi Chuyan" is 0, the GC contents are 36.35% and 36.02% respectively, the percentages of bases with base recognition accuracy above 99% are 99.57% and 99.53% respectively, and the percentages of bases with base recognition accuracy above 99.9% are 98.23% and 98.11% respectively. The above results indicate that the sequencing results have good quality.
[0059] Table 1 Statistical analysis of sequencing data 2. Further filtering of the raw data to obtain high-quality data To ensure the quality of subsequent analysis, it is necessary to filter the off-machine raw data to obtain high-quality sequences. The filtering criteria are as follows: (1) Use AdapterRemoval (version 2) to remove 3'-end adapter contamination.
[0060] (2) Perform quality filtering using the sliding window method (window size 5 bp, step size 1 bp): If the average Q value of the window ≤ 20, retain the bases up to the second-to-last base of the window; if the Q value of the terminal base ≤ 2, retain all sequences before that base.
[0061] (3) Remove sequences with any one length ≤ 50 bp in paired-end reads.
[0062] (4) After filtering the raw data, high-quality sequences were obtained (Table 2). The numbers of high-quality reads of "Huafen Xinzhuang" and "Hezhi Chuyan" are 993,886,138 and 1,169,884,600 respectively. The percentages of high-quality reads in the original reads are 99.62% and 99.57% respectively. The numbers of bases of high-quality reads are 148,540,017,935 and 174,856,234,261 respectively. The percentages of the numbers of bases of high-quality reads in the total number of original bases are 99.26% and 99.22% respectively. The sequencing depths are 12.1× and 14.24× respectively.
[0063] Table 2 Statistical analysis of high-quality data Example 3 This embodiment constructs fingerprint maps based on high-throughput resequencing data of two peony varieties, "Huafen Xinzhuang" and "Hezhi Chuyan".
[0064] 1. Align the resequencing data with the reference genome. (1) Use the bwa(0.7.17-r1188)mem program with default parameters to align the filtered high-quality data to the reference genome.
[0065] (2) Use picard 1.107 to sort the generated sam file and convert it to bam format, and execute the "FixMateInformation" command to ensure that the reads information on both ends is consistent.
[0066] (3) To reduce the possible introduction of repetitive sequences during the amplification process, the “MarkDuplicates” module in Picard is used to remove redundant paired reads with the same alignment position, and only the best alignment result is retained.
[0067] The sequence alignment results (Table 3) show that the total number of reads for the two peony varieties "Huafen Xinzhuang" and "Hezhi Chuyan" were 993,886,138 and 1,169,884,600, respectively. The number of reads aligned to the reference genome were 991,894,352 and 1,167,388,515, respectively, representing 99.80% and 99.79% of the total reads. The number of repetitive sequence reads were 174,985,005 and 218,056,912, respectively, representing 17.61% and 18.64% of the total reads.
[0068] Table 3. Statistical analysis of sequence alignment results The average coverage depth and coverage of the statistical samples can be used to determine ( Figure 3 , Figure 4 (Table 4) The percentages of sites covered by at least 1, 4, 10, and 20 sequences in the reference genomes of the two peony varieties “Huafen Xinzhuang” and “Hezhi Chuyan” were 86.52% and 82.14%, 72.59% and 67.36%, 42.42% and 11.70%, and 33.03% and 8.83%, respectively.
[0069] Table 4. Statistics on alignment and sequencing depth and coverage results 2. SNP detection and mapping of genomic SNP variant sites The specific steps for SNP detection using GATK software are as follows: (1) Alignments near InDels are usually unreliable and need to be realigned using known InDel information. This is done in two steps: First, the RealignerTargetCreator command in the GATK package is used to output a file containing all possible InDels; second, the IndelRealigner command is used to re-align all reads near InDels to improve the accuracy of SNP prediction.
[0070] (2) The SNP sites of the sample were obtained using the UnifiedGenotyper program, and stand_call_conf was set to 30.
[0071] To ensure the reliability of SNP sites, the obtained SNP sites were further filtered using the following criteria: (1) Fisher test of strand bias (FS) ≤ 60; (2) HaplotypeScore ≤ 13.0; (3)Mapping Quality (MQ)≥ 40; (4) Quality Depth(QD) ≥ 2; (5) ReadPosRankSum ≥ -8.0; (6) MQRankSum > -12.5.
[0072] Statistical analysis of the detected SNPs (Table 5) revealed that the genomic variations of the two peony samples, “Huafen Xinzhuang” and “Hezhi Chuyan”, showed a high degree of consistency, but some differences also existed.
[0073] A total of 224,884,780 SNP loci were detected in the "Huafen Xinzhuang" sample (homozygous genotype consistent with the reference genome + heterozygous genotype + unknown genotype + inconsistent homozygous genotype), with a transition to transversion ratio (Ts / Tv) of 1.8024. The "Hezhi Chuyan" sample, on the other hand, had 225,904,820 SNP loci detected, with a Ts / Tv ratio of 1.8346. The Ts / Tv ratios of both samples are close to 2:1, consistent with the general rule that SNP transitions are usually more frequent than transversions in eukaryotic genomes, indicating that the genotyping results are accurate and reliable.
[0074] Comparing the two samples, the number of homozygous genotypes consistent with the reference genome in "Hezhi Chuyan" (73,818,019) was significantly higher than that in "Huafen Xinzhuang" (50,860,339), indicating that it may be more closely related to the reference genome. Meanwhile, the number of heterozygous genotypes in "Huafen Xinzhuang" (139,515,279) and the number of homozygous genotypes inconsistent with the reference genome (34,509,547) were both higher than those in "Hezhi Chuyan" (121,751,808 and 29,731,917 respectively), suggesting that "Huafen Xinzhuang" may have higher genetic heterozygosity and richer genetic variation. These differences provide an important data foundation for subsequent analysis of the genetic characteristics and phenotypic differences between the two peony varieties.
[0075] Table 5. Statistics of SNP results in each sample of the population Based on the SNP identification and analysis results, SNP fingerprint maps of two peony samples, "Huafen Xinzhuang" and "Hezhi Chuyan," were plotted to visually observe the distribution of variations. Figure 5 It is known that the distribution of SNPs on chromosomes is not uniform, but rather exhibits a clear regional clustering characteristic. This uneven distribution may be related to the functional structural domains of chromosomes and the differences in selection pressure in different regions.
[0076] The distribution of SNPs on each nuclear genome chromosome (excluding SNPs on unknown chromosomes) was statistically analyzed (Table 6). It can be seen that the overall number of SNP variations (approximately 165 million) and density (14.34 / Kb) of "Huafen Xinzhuang" are higher than those of "Hezhi Chuyan" (approximately 143 million, 12.48 / Kb). Furthermore, the SNP density of "Huafen Xinzhuang" is consistently higher than that of "Hezhi Chuyan" on all five chromosomes. This indicates that the overall genetic diversity of "Huafen Xinzhuang" is significantly higher than that of "Hezhi Chuyan".
[0077] The two varieties maintained the highest and lowest SNP densities on Chr04 and Chr03, respectively. This consistency suggests that inherent characteristics of certain chromosomes (such as recombination rate, gene density, and selection pressure) may have contributed to this shared variation pattern. While the overall SNP distribution trends of the two peony varieties were similar, differences existed in localized areas, indicating that they shared a common genetic background during genomic evolution while also accumulating their own unique variations. These specific regions of variation may be an important genetic basis for the phenotypic differences (such as flower color, flower shape, and resistance) between the two peony varieties.
[0078] Table 6. Statistics of Chromosomal SNP Variants 3. SNP site annotation Functional annotation of SNP sites in two peony samples, “Huafen Xinzhuang” and “Hezhi Chuyan”, was performed using ANNOVAR software.
[0079] Based on SNP annotation results ( Figure 6 As can be seen, most SNPs are located in intergenic regions, while intronic regions, synonymous and nonsynonymous variants are also distributed in gene-related regions. The potential impact of SNPs in different regions on gene function varies. SNPs account for a very high proportion in intergenic regions of the genome, reaching 94.39%, indicating that most SNPs are located between genes, and their direct impact on gene coding and other functions may be relatively indirect. SNPs account for 3.65% in intronic regions, indicating that some SNPs exist in intronic regions of genes. Although introns do not directly encode proteins, they may play a role in gene transcription regulation, mRNA processing, etc., and SNPs in these regions may affect gene expression. SNPs account for 0.14% in synonymous single nucleotide variants (SNVs). Synonymous variants refer to SNPs that cause codon changes but do not change the encoded amino acid; they generally have no direct impact on the amino acid sequence and function of proteins. SNPs account for a significant portion of nonsynonymous single nucleotide variants (SNVs). SNVs account for 0.28% of all mutations. These variations alter the amino acids encoded by codons, potentially affecting protein structure and function, and consequently influencing the traits of organisms. SNPs account for 0.49% upstream and 0.48% downstream. SNPs in these regions may participate in gene expression regulation, such as affecting the function of regulatory elements like promoters and enhancers, thus impacting transcription initiation. The UTR (untranslated region) is divided into the 5'UTR and 3'UTR. SNPs account for 0.03% in the 5'UTR (UTR5) and 0.05% in the 3'UTR (UTR3). The UTR plays a role in mRNA stability and translation initiation, and SNPs in these regions may affect mRNA function and thus gene expression. Other SNPs, such as those affecting frameshift and splicing, account for very low percentages, at 0.02% and 0.01% respectively. However, these variations often have a significant impact on gene function, potentially leading to premature termination of protein synthesis or abnormal mRNA splicing.
[0080] Based on the statistical results of SNP mutation base preference ( Figure 7 As can be seen, at the mutation site (position 0), there is a significant difference in the number of different bases, which indicates that there is a clear base preference in mutations, and some bases may be more likely to be mutated, or some mutation types may be more common.
[0081] Based on the SNP mutation spectrum analysis results ( Figure 8 The results show that SNP mutations in the genome are divided into two main categories: Ts transitions (T:A>C:G, C:G>T:A) and Tv transversions (T:A>G:C, T:A>A:T, C:G>G:C, C:G>A:T). Ts transitions (T:A>C:G and C:G>T:A) are the predominant SNP mutations, accounting for 64.24%, while Tv transversions account for 35.76%. The higher prevalence of transitions (Ts:Tv ≈ 1.8:1) is consistent with the general pattern of plant genomes, indicating that the sequencing and mutation detection results meet general biological expectations and the data quality is reliable. Among these, the high-frequency specific transversion types may be mutation hotspots in the peony genome influenced by specific biological or environmental factors, and can serve as key candidate regions for developing specific molecular markers for variety identification and genetic map construction.
[0082] Example 4 This example demonstrates the application of fingerprint maps constructed from resequencing data of two peony samples, "Huafen Xinzhuang" and "Hezhi Chuyan".
[0083] 1. Screening and primers for 61 SNP functional gene molecular markers By using parameter thresholds of DP (Depth, coverage of the locus) > 12 and GQ (Genotype Quality, genotype quality value) > 20, 61 homozygous high-quality SNP variants at exons were screened (Table 7). All of these variants were of the stopgain or stoploss type, indicating that all 61 SNPs were loss-of-function mutations and were all homozygous. This means that both alleles on the positive and negative strands of DNA have undergone the same functional disruption, which is very likely to lead to complete loss of gene function and thus have a significant impact on individual phenotype.
[0084] The 1500bp sequences upstream and downstream of each SNP variant site in the reference genome were extracted using Samtools software, and primers were designed using Primer Premier 6 software (Table 8). DNA samples from different peony varieties were extracted and amplified by PCR, followed by Sanger sequencing. The sequencing results were compared and analyzed using DNAMAN software. The distribution of the 61 SNP variant sites on the chromosome is shown below. Figure 9 As shown. Taking "Fengdan", "Huafen Xinzhuang", and "Hezhi Chuyan" as examples, as per the present invention, Figure 10 As shown, the genotypes of different peony varieties at 61 SNP loci can be obtained through sequence alignment analysis.
[0085] Table 7. Homozygous SNPs in 61 functional gene exons Table 8 Primers for homozygous SNP variants in 61 functional gene exons 2. Conduct genetic diversity analysis of peony germplasm resources. Using the aforementioned 61 SNP molecular marker combinations and primers, the genetic diversity of different peony germplasm resources was identified by PCR amplification and Sanger sequencing. Germplasm resource evolution analysis was performed using MEGA software, and a phylogenetic tree of kinship relationships was constructed. Taking the varieties "Fengdan," "Huafen Xinzhuang," and "Hezhi Chuyan" involved in this invention as examples, such as... Figure 11 As shown, since the two peony varieties "Huafen Xinzhuang" and "Hezhi Chuyan" share the same parent, they are more closely related.
[0086] 3. Conduct authenticity verification of peony hybrids. DNA was extracted from fresh, healthy young leaves of the hybrid offspring seedlings and the two parent plants (“Fengdan”, “Huafen Xinzhuang”, and “Hezhi Chuyan”). Using the aforementioned 61 SNP molecular marker combinations and primers, PCR amplification and Sanger sequencing were performed to identify the genotypes of the hybrid offspring and the two parent plants at multiple SNP loci. If the hybrid offspring exhibited heterozygous genotypes at all loci where the parental genotypes differed, it was a true hybrid; if the hybrid offspring only exhibited homozygous genotypes completely identical to the maternal parent, it was a false hybrid (maternal self-fertilization or parthenogenesis).
[0087] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention in any way. Any simple modifications, alterations, and equivalent changes made to the above embodiments based on the inventive essence shall still fall within the protection scope of the present invention.
Claims
1. A primer pair combination for constructing a DNA fingerprint of peony using SNP molecular marker combinations, characterized in that, The primer pair combination is the first primer pair used to amplify the Chr01_13974972_G / A molecular marker, and the nucleotide sequence is shown in SEQ ID NO. 1-2; The second primer pair used to amplify the Chr01_52764705_C / T molecular marker has the nucleotide sequence shown in SEQ ID NO.3-4; The third primer pair used to amplify the Chr01_301779146_T / G molecular marker has the nucleotide sequence shown in SEQ ID NO.5-6; The fourth primer pair used to amplify the Chr01_582511719_C / T molecular marker has the nucleotide sequence shown in SEQ ID NO.7-8; The fifth primer pair used to amplify the Chr01_604023402_C / T molecular marker has the nucleotide sequence shown in SEQ ID NO.9-10; The sixth primer pair used to amplify the Chr01_1806171054_T / C molecular marker has the nucleotide sequence shown in SEQ ID NO. 11-12; The 7th primer pair used to amplify the Chr01_1918223993_C / T molecular marker has the nucleotide sequence shown in SEQ ID NO. 13-14; The 8th primer pair used to amplify the Chr01_1959229511_C / T molecular marker has the nucleotide sequence shown in SEQ ID NO. 15-16; The 9th primer pair used to amplify the Chr01_2120097634_T / A molecular marker has the nucleotide sequence shown in SEQ ID NO. 17-18; The 10th primer pair used to amplify the Chr01_2345730173_A / T molecular marker has the nucleotide sequence shown in SEQ ID NO. 19-20; The 11th primer pair used to amplify the Chr02_340870593_G / A molecular marker has the nucleotide sequence shown in SEQ ID NO. 21-22; The 12th primer pair used to amplify the Chr02_407187313_T / G molecular marker has the nucleotide sequence shown in SEQ ID NO. 23-24; The 13th primer pair used to amplify the Chr02_540706395_A / G molecular marker has the nucleotide sequence shown in SEQ ID NO. 25-26; The 14th primer pair used to amplify the Chr02_650741859_C / T molecular marker has the nucleotide sequence shown in SEQ ID NO. 27-28; The 15th primer pair used to amplify the Chr02_702603442_A / G molecular marker has the nucleotide sequence shown in SEQ ID NO. 29-30; The 16th primer pair used to amplify the Chr02_1710187145_C / A molecular marker has the nucleotide sequence shown in SEQ ID NO. 31-32; The 17th primer pair used to amplify the Chr02_2097523064_T / C molecular marker has the nucleotide sequence shown in SEQ ID NO. 33-34; The 18th primer pair used to amplify the Chr03_33351259_C / T molecular marker has the nucleotide sequence shown in SEQ ID NO.35-36; The 19th primer pair used to amplify the Chr03_87443773_T / G molecular marker has the nucleotide sequence shown in SEQ ID NO.37-38; The 20th primer pair used to amplify the Chr03_132275045_C / T molecular marker has the nucleotide sequence shown in SEQ ID NO. 39-40; The 21st primer pair used to amplify the Chr03_280143517_A / G molecular marker has the nucleotide sequence shown in SEQ ID NO. 41-42. The 22nd primer pair used to amplify the Chr03_500972943_T / C molecular marker has the nucleotide sequence shown in SEQ ID NO. 43-44; The 23rd primer pair used to amplify the Chr03_699473250_C / T molecular marker has the nucleotide sequence shown in SEQ ID NO. 45-46; The 24th primer pair used to amplify the Chr03_851012098_C / T molecular marker has the nucleotide sequence shown in SEQ ID NO. 47-48; The 25th primer pair used to amplify the Chr03_855211440_C / A molecular marker has the nucleotide sequence shown in SEQ ID NO. 49-50; The 26th primer pair used to amplify the Chr03_936166517_C / A molecular marker has the nucleotide sequence shown in SEQ ID NO. 51-52; The 27th primer pair used to amplify the Chr03_936166860_A / G molecular marker has the nucleotide sequence shown in SEQ ID NO. 53-54; The 28th primer pair used to amplify the Chr03_1055906550_G / A molecular marker has the nucleotide sequence shown in SEQ ID NO. 55-56; The 29th primer pair used to amplify the Chr03_1173465922_A / C molecular marker has the nucleotide sequence shown in SEQ ID NO. 57-58; The 30th primer pair used to amplify the Chr03_1173465930_T / A molecular marker has the nucleotide sequence shown in SEQ ID NO. 59-60; The 31st primer pair used to amplify the Chr03_1216736027_T / A molecular marker has the nucleotide sequence shown in SEQ ID NO. 61-62; The 32nd primer pair used to amplify the Chr03_1234693042_G / T molecular marker has the nucleotide sequence shown in SEQ ID NO. 63-64; The 33rd primer pair used to amplify the Chr03_1246856404_C / T molecular marker has the nucleotide sequence shown in SEQ ID NO. 65-66; The 34th primer pair used to amplify the Chr03_1250250494_C / T molecular marker has the nucleotide sequence shown in SEQ ID NO. 67-68; The 35th primer pair used to amplify the Chr03_1414051447_C / T molecular marker has the nucleotide sequence shown in SEQ ID NO. 69-70; The 36th primer pair used to amplify the Chr03_1622577802_C / T molecular marker has the nucleotide sequence shown in SEQ ID NO. 71-72; The 37th primer pair used to amplify the Chr03_1704789394_T / C molecular marker has the nucleotide sequence shown in SEQ ID NO. 73-74; The 38th primer pair used to amplify the Chr04_108838226_T / A molecular marker has the nucleotide sequence shown in SEQ ID NO. 75-76; The 39th primer pair used to amplify the Chr04_151772272_G / A molecular marker has the nucleotide sequence shown in SEQ ID NO. 77-78; The 40th primer pair used to amplify the Chr04_520822950_T / C molecular marker has the nucleotide sequence shown in SEQ ID NO. 79-80; The 41st primer pair used to amplify the Chr04_680727619_A / T molecular marker has the nucleotide sequence shown in SEQ ID NO. 81-82; The 42nd primer pair used to amplify the Chr04_1030650704_G / T molecular marker has the nucleotide sequence shown in SEQ ID NO. 83-84; The 43rd primer pair used to amplify the Chr04_1071214861_G / A molecular marker has the nucleotide sequence shown in SEQ ID NO. 85-86; The 44th primer pair used to amplify the Chr04_1109376597_C / A molecular marker has the nucleotide sequence shown in SEQ ID NO. 87-88; The 45th primer pair used to amplify the Chr04_1109376647_G / T molecular marker has the nucleotide sequence shown in SEQ ID NO. 89-90; The 46th primer pair used to amplify the Chr04_1140925370_G / A molecular marker has the nucleotide sequence shown in SEQ ID NO. 91-92; The 47th primer pair used to amplify the Chr04_1553032924_A / G molecular marker has the nucleotide sequence shown in SEQ ID NO. 93-94; The 48th primer pair used to amplify the Chr04_1787787355_T / A molecular marker has the nucleotide sequence shown in SEQ ID NO. 95-96; The 49th primer pair used to amplify the Chr04_2017263232_C / A molecular marker has the nucleotide sequence shown in SEQ ID NO. 97-98; The 50th primer pair used to amplify the Chr04_2230098669_A / G molecular marker has the nucleotide sequence shown in SEQ ID NO. 99-100; The 51st primer pair used to amplify the Chr04_2285760072_A / G molecular marker has the nucleotide sequence shown in SEQ ID NO. 101-102; The 52nd primer pair used to amplify the Chr04_2330967850_G / T molecular marker has the nucleotide sequence shown in SEQ ID NO. 103-104; The 53rd primer pair used to amplify the Chr04_2397495861_A / G molecular marker has the nucleotide sequence shown in SEQ ID NO. 105-106; The 54th primer pair used to amplify the Chr05_123298388_A / C molecular marker has the nucleotide sequence shown in SEQ ID NO. 107-108; The 55th primer pair used to amplify the Chr05_248171257_G / C molecular marker has the nucleotide sequence shown in SEQ ID NO. 109-110; The 56th primer pair used to amplify the Chr05_337081376_G / T molecular marker has the nucleotide sequence shown in SEQ ID NO. 111-112; The 57th primer pair used to amplify the Chr05_807239075_G / A molecular marker has the nucleotide sequence shown in SEQ ID NO. 113-114; The 58th primer pair used to amplify the Chr05_1529394053_A / T molecular marker has the nucleotide sequence shown in SEQ ID NO. 115-116; The 59th primer pair used to amplify the Chr05_1535053858_G / A molecular marker has the nucleotide sequence shown in SEQ ID NO. 117-118; The 60th primer pair used to amplify the Chr05_1741394421_G / A molecular marker has the nucleotide sequence shown in SEQ ID NO. 119-120; The 61st primer pair used to amplify the Chr05_2326548240_T / C molecular marker has the nucleotide sequence shown in SEQ ID NO. 121-122; The reference genome is CNGBdb:CNP0003098.
2. The application of a primer pair combination for constructing a peony DNA fingerprint as described in claim 1, characterized in that, The application is for identifying the peony varieties "Huafen Xinzhuang" and / or "Hezhi Chuyan".
3. The application of a primer pair combination for constructing a peony DNA fingerprint as described in claim 1, characterized in that, The application is used to identify true and false hybrids among the offspring of pairwise hybridization of "Fengdan", "Huafen Xinzhuang" and "Hezhi Chuyan".
4. The application according to claim 3, characterized in that, If the hybrid offspring exhibit a heterozygous genotype at all loci where the genotypes of the parents differ, they are a true hybrid. If the hybrid offspring exhibit only a homozygous genotype that is completely identical to the maternal parent, then it is a false hybrid.
5. The application of a primer pair combination for constructing a peony DNA fingerprint as described in claim 1, characterized in that, The application is for the genetic diversity of peony germplasm resources and the phylogenetic tree analysis of peony variety resources.
Citation Information
Patent Citations
KR20230100949A