Snps molecular marker related to identification of hybrid between corylus avellana and c. heterophylla, and identification method and application thereof
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- LIAONING INST OF ECONOMIC FORESTRY
- Filing Date
- 2026-06-02
- Publication Date
- 2026-08-07
AI Technical Summary
然而,对于平欧杂种榛这一我国特色种间杂交栽培类型,目前仍缺乏基于大规模全基因组重测序数据筛选、同时经过实验分型验证、能够直接用于产业化品种鉴定的核心SNP标记组合
本发明技术方案是针对平欧杂种榛品种数量持续增加、苗木市场品种混杂风险上升以及现有SSR体系在标准化和产业化应用方面存在不足的问题,建立的一种以全基因组变异数据为基础、以少量高判别力SNP位点为核心、以KASP为检测平台的品种分子鉴定技术体系。基于266份平欧杂种榛种质资源的全基因组重测序数据,通过严格质控、核心种质构建、PIC筛选、物理距离过滤和算法优化筛选候选位点,并进一步经KASP实验验证确定稳定可用的核心SNP组合,3个核心判别SNP位点(SNP03、SNP05、SNP11)和1个辅助参考SNP位点(SNP09)。
Smart Images

Figure CN122521889A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of molecular biology detection technology, and in particular to SNP molecular markers related to the identification of European hybrid hazelnut varieties, as well as their identification methods and applications. Background Technology
[0002] European hybrid hazelnut ( Corylus heterophylla × C. avellana ) is my country's hazelnut ( Corylus heterophylla Fisch. ) and European hazelnut ( Corylus avellana L. This species is an important woody oilseed and nut-producing economic tree species developed through interspecific hybridization and selection of parent lines. This type of germplasm combines the cold hardiness and strong adaptability of Hazelnut (Corylus heterophylla) with the advantages of large, high-quality, and commercially viable European hazelnuts, making it one of the main cultivated types for the large-scale development of the hazelnut industry in northern my country. my country has abundant Hazelnut resources, and since the 1980s, hybridization, regional trials, and industrial promotion of Hazelnut (Corylus heterophylla) have been continuously carried out. The cultivation area has gradually expanded from Northeast and North China to multiple suitable ecological zones, and the commercial cultivation area continues to increase. With the registration of new varieties, the promotion of superior lines, and the expansion of the seedling market, the demand for variety authenticity identification, parent line confirmation, standardized management of germplasm resource banks, and traceability in seedling production and sales is becoming increasingly prominent.
[0003] Currently, the identification of hybrid hazelnut varieties still mainly relies on morphological traits, pedigree records, and limited molecular marker information. Morphological identification is usually based on phenotypic indicators such as tree posture, leaves, nut size, fruit shape, shell thickness, bract morphology, and maturity period. However, these traits are easily affected by tree age, cultivation management, environmental conditions, and year effects, and most key traits can only be observed stably during the fruiting period or specific phenological stages. Therefore, traditional phenotypic identification is difficult to meet the requirements of seedling identification, rapid detection of asexually propagated materials, comparison of multiple batches of samples from different locations, and the determination of variety authenticity at the judicial / administrative level.
[0004] Existing studies have used codominant molecular markers such as SSR and EST-SSR for genetic diversity analysis, phylogenetic evaluation, and varietal fingerprinting of European hazelnut and European-hybrid hazelnut. SSR markers have advantages such as high polymorphism, codominant inheritance, and relatively mature experimental systems, which can improve the objectivity of varietal identification to a certain extent. For example, domestic studies have used SSR or EST-SSR markers from European hazelnuts to analyze the genetic differences of main European-hybrid hazelnut cultivation materials, and related patents have proposed using several EST-SSR primers to identify European-hybrid hazelnut varieties. These works provide a foundation for hazelnut germplasm resource evaluation and varietal identification. However, the SSR identification system still has significant limitations in practical application: First, detection usually relies on polyacrylamide gel electrophoresis, capillary electrophoresis, or fragment analysis platforms, making the experimental process relatively complex; second, different laboratories are prone to differences in instrument models, internal standard systems, fragment interpretation, and allele nomenclature, leading to difficulties in data sharing and cross-platform comparison; third, the number of SSR loci is limited, and without unified standard samples and databases, problems such as difficulty in distinguishing homonyms, homonyms, and closely related strains can easily arise; fourth, for large-scale seedling testing and industrialized variety traceability, the SSR system still cannot meet the needs of rapid detection in terms of throughput, cost, automation, and result standardization.
[0005] With the development of high-throughput sequencing and reference genome resources, single nucleotide polymorphisms (SNPs) have become important marker types for plant variety identification, genetic diversity analysis, molecular fingerprinting, and seed purity detection due to their wide distribution in the genome, genetic stability, ease of automated detection, and cross-laboratory standardization. Studies on European hazelnuts based on simplified genome sequencing or high-density SNP data for genetic diversity assessment and DNA typing have demonstrated the feasibility and application potential of SNP markers in hazelnut germplasm identification. However, for the hybrid European hazelnut, a unique interspecific hybrid cultivation type in my country, there is currently a lack of core SNP marker combinations that are screened based on large-scale whole-genome resequencing data, validated by experimental typing, and directly applicable to industrial variety identification.
[0006] Kompetitive allele-specific PCR (KASP) is a fluorescent endpoint genotyping technique suitable for SNP locus detection. This technique amplifies the target locus using two allele-specific primers and one universal reverse primer, and determines the genotype based on FAM / HEX fluorescence signal clustering. Compared to sequencing genotyping or SSR fragment analysis, KASP offers advantages such as lower cost per locus detection, stable reaction system, flexible throughput, objective data interpretation, relatively low equipment requirements, and easy result standardization. It has been widely used in variety identification, seed purity testing, parental authenticity verification, and assisted selection breeding for various agricultural and horticultural crops.
[0007] Therefore, in response to the continuous increase in the number of hybrid hazelnut varieties, the rising risk of mixed varieties in the seedling market, and the shortcomings of the existing SSR system in terms of standardization and industrial application, it is necessary to establish a molecular identification technology system for varieties based on whole genome variation data, with a small number of high-discrimination SNP sites as the core, and KASP as the detection platform. Summary of the Invention
[0008] The purpose of this invention is to provide a combination of SNP molecular markers and KASP primers for amplifying the corresponding SNP molecular markers related to the identification of European hybrid hazelnut varieties. By utilizing the molecular marker sites and KASP primers, rapid, accurate and low-cost identification of European hybrid hazelnut varieties can be achieved with fewer detection sites.
[0009] To achieve the above-mentioned objectives, the present invention provides the following technical solution: This invention provides an application of a reagent for detecting SNP molecular markers in the preparation of products related to the identification of Hazel hybrid hazel varieties, wherein the SNP molecular markers include HazelID_SNP03, HazelID_SNP05, HazelID_SNP09, and HazelID_SNP11.
[0010] Preferably, the HazelID_SNP03 is located at the 25,552,398th base on chromosome 1 of the Jefferson HAP1_v4.0 reference genome of the European hazelnut, with a mutated base of G or C; the HazelID_SNP05 is located at the 15,913,337th base on chromosome 4 of the Jefferson HAP1_v4.0 reference genome of the European hazelnut, with a mutated base of T or C; the HazelID_SNP09 is located at the 22,599,658th base on chromosome 6 of the Jefferson HAP1_v4.0 reference genome of the European hazelnut, with a mutated base of G or A; and the HazelID_SNP11 is located at the 17,014,214th base on chromosome 9 of the Jefferson HAP1_v4.0 reference genome of the European hazelnut, with a mutated base of A or G.
[0011] The present invention also provides KASP primer compositions for amplifying the SNP molecular markers, including HazelID_SNP03 primer composition, HazelID_SNP05 primer composition, HazelID_SNP09 primer composition and HazelID_SNP11 primer composition.
[0012] Preferably, the HazelID_SNP03 primer composition comprises a forward primer sequence as shown in SEQ ID NO:7, a reverse primer sequence as shown in SEQ ID NO:8, and a universal reverse primer sequence as shown in SEQ ID NO:9; the HazelID_SNP05 primer composition comprises a forward primer sequence as shown in SEQ ID NO:13, a reverse primer sequence as shown in SEQ ID NO:14, and a universal reverse sequence as shown in SEQ ID NO:15; the HazelID_SNP09 primer composition comprises a forward primer sequence as shown in SEQ ID NO:25, a reverse primer sequence as shown in SEQ ID NO:26, and a universal reverse sequence as shown in SEQ ID NO:27; and the HazelID_SNP11 primer composition comprises a forward primer sequence as shown in SEQ ID NO:31, a reverse primer sequence as shown in SEQ ID NO:32, and a universal reverse sequence as shown in SEQ ID NO:33.
[0013] This invention also provides the application of the KASP primer composition in the preparation of products related to the identification of European hybrid hazelnut varieties.
[0014] The present invention also provides a kit for identifying European hybrid hazelnut varieties, comprising a reagent for detecting the SNP molecular marker or the KASP primer composition.
[0015] The present invention also provides a method for identifying European hybrid hazelnut varieties using the KASP primer composition, comprising the following steps: (1) Extract genomic DNA from the young leaves of the test subject; (2) Using the young leaf genomic DNA obtained in step (1) as a template, amplification was performed using the KASP primer composition to obtain the amplification product; (3) Perform genotyping analysis on the amplification products obtained in step (2), and then compare the genotypes of the SNP molecular marker sites to determine the different varieties of hazelnuts.
[0016] Preferably, the amplification reaction system in step (2) consists of 10 μL: PARMS 2×Master Mix 5 μL, primer mixture 0.14 μL, DNA template 1 μL, and ddH2O added to 10 μL; the primer mixture consists of forward primer sequence, reverse primer sequence, and universal reverse primer sequence with a concentration ratio of (8-16) μM:(8-16) μM:(20-40) μM.
[0017] Preferably, the amplification reaction program in step (2) is: 94℃ pre-denaturation for 15 min; 94℃ denaturation for 20 s, 65℃ annealing / extension for 60 s, with a decrease of 0.8℃ per cycle, for 10 cycles; 94℃ denaturation for 20 s, 57℃ annealing / extension for 60 s, for 26 cycles.
[0018] By adopting the above technical solution, the present invention has the following beneficial effects: This invention addresses the challenges posed by the increasing number of hybrid hazelnut varieties, the rising risk of mixed varieties in the seedling market, and the shortcomings of existing SSR systems in standardization and industrial application. It establishes a molecular identification technology system for hazelnut varieties based on whole-genome variation data, with a small number of high-discrimination SNP loci at its core, and KASP as the detection platform. Based on whole-genome resequencing data from 266 hybrid hazelnut germplasm resources, candidate loci were screened through rigorous quality control, core germplasm construction, PIC screening, physical distance filtering, and algorithm optimization. Furthermore, stable and usable core SNP combinations were determined through KASP experiments, including three core discriminant SNP loci (SNP03, SNP05, and SNP11) and one auxiliary reference SNP locus (SNP09).
[0019] This invention utilizes a combination of screened core SNPs and KASP primers to identify European hybrid hazelnut varieties. This identification system features clear locus origins, high discrimination efficiency, simple detection process, low cost, strong reproducibility, and ease of establishing a standard genotype database. It provides technical support for the authenticity identification of European hybrid hazelnut varieties, management of breeding materials, supervision of seedling quality, and digital protection of germplasm resources.
[0020] The core SNP combinations screened by this invention have the following technical advantages: (1) High discrimination efficiency: the discrimination rate reached 98.23% in 66 core germplasms and 95.26% in 20 verification samples; (2) Stable performance as verified by experiments: the KASP genotyping clusters of the three core discriminative SNP sites are clear, SNP09 can provide auxiliary reference information but is not used as an independent discrimination basis, and the first genotyping success rate is 98.75%; (3) Low cost and high efficiency: only 4 KASP reactions are required for a single variety identification, and the detection cycle is 2-3 hours; (4) Good genome representativeness: the four markers are distributed on four different chromosomes, avoiding linkage interference.
[0021] Compared to the closest existing technology CN 120519617 A (which uses 50 core SNP loci to construct a DNA fingerprint of hazelnut plants), this invention, for the identification of European hybrid hazelnut varieties, uses 3 core discriminant SNP loci (SNP03, SNP05, SNP11) and 1 auxiliary reference SNP locus (SNP09) to achieve a discrimination rate of 98.23% in 66 core germplasm samples and 95.26% in 20 KASP validation samples. It maintains high discrimination efficiency while significantly reducing the number of markers and detection reactions, and has the technical effects of low marker count, high discrimination, and suitability for rapid KASP detection and industrial promotion.
[0022] Compared to traditional SSR marker identification methods, which require capillary electrophoresis for fragment size analysis and suffer from issues such as subjective interpretation, data incompatibility between laboratories, and low throughput, the KASP-SNP method of this invention uses fluorescence endpoint reading, providing objective two-dimensional clustering interpretation. This eliminates human interpretation errors, allows for genotyping in a single reaction, has a short detection cycle, low cost, and is fully compatible with data from different laboratories. Attached Figure Description
[0023] Figure 1 Scatter plot of KASP genotyping for HazelID_SNP01 locus ( Figure 1 The black dots in the image represent samples that were not successfully genotyped, and the cyan squares represent NTC negative controls. Figure 2 Scatter plot of KASP genotyping for HazelID_SNP08 locus ( Figure 2 The black dots in the image represent samples that were not successfully genotyped, and the cyan squares represent NTC negative controls. Figure 3 Scatter plot of KASP genotyping for HazelID_SNP02 locus ( Figure 3 The black dots in the image represent non-specific amplification samples, and the cyan squares represent NTC negative controls. Figure 4 Scatter plot of KASP genotyping for HazelID_SNP03 locus ( Figure 4 In the diagram, red dots represent homozygous allele 1, green dots represent heterozygous genotypes, and cyan squares represent NTC negative controls. Figure 5 Scatter plot of KASP genotyping for HazelID_SNP05 locus ( Figure 5 The red dots represent homozygous allele 1, the blue dots represent homozygous allele 2, the green dots represent heterozygous genotypes, and the cyan squares represent NTC negative controls. Figure 6 Scatter plot of KASP genotyping for HazelID_SNP09 locus ( Figure 6 In the diagram, the blue dots represent homozygous allele 2, the green dots represent heterozygous genotypes, and the cyan squares represent NTC negative controls. Figure 7 Scatter plot of KASP genotyping for HazelID_SNP11 locus ( Figure 7 The red dots represent homozygous allele 1, the blue dots represent homozygous allele 2, the green dots represent heterozygous genotypes, and the cyan squares represent NTC negative controls. Detailed Implementation
[0024] This invention provides an application of a reagent for detecting SNP molecular markers in the preparation of products related to the identification of Hazel hybrid hazel varieties, wherein the SNP molecular markers include HazelID_SNP03, HazelID_SNP05, HazelID_SNP09, and HazelID_SNP11.
[0025] In this invention, HazelID_SNP03, HazelID_SNP05, and HazelID_SNP11 are the core discriminant SNP sites, and HazelID_SNP09 is the auxiliary reference SNP site. Overall, genetic linkage interference between markers is avoided, and the genome has good representativeness.
[0026] In this invention, the HazelID_SNP03 is located at the 25,552,398th base on chromosome 1 of the Jefferson HAP1_v4.0 reference genome of the European hazelnut, and the mutated base is G or C.
[0027] In this invention, the HazelID_SNP05 is located at the 15,913,337th base on chromosome 4 of the Jefferson HAP1_v4.0 reference genome of the European hazelnut, and the mutated base is T or C.
[0028] In this invention, the HazelID_SNP09 is located at the 22,599,658th base on chromosome 6 of the Jefferson HAP1_v4.0 reference genome of the European hazelnut, and the mutated base is G or A.
[0029] In this invention, the HazelID_SNP11 is located at the 17014214th base on chromosome 9 of the Jefferson HAP1_v4.0 reference genome of the European hazelnut, and the mutated base is A or G.
[0030] In this invention, the reference genotype combination table of the 20 validation samples is as follows:
[0031] The present invention also provides KASP primer compositions for amplifying the SNP molecular markers, including HazelID_SNP03 primer composition, HazelID_SNP05 primer composition, HazelID_SNP09 primer composition and HazelID_SNP11 primer composition.
[0032] In this invention, the HazelID_SNP03 primer composition includes a forward primer sequence, a reverse primer sequence, and a universal reverse primer sequence. The forward primer sequence is shown in SEQ ID NO:7, specifically 5'-GAAGGTGACCAAGTTCATGCTGGAATGAGTAACCATGGAATGGGTAG-3'; the reverse primer sequence is shown in SEQ ID NO:8, specifically 5'-GAAGGTCGGAGTCAACGGATTGGAATGAGTAACCATGGAATGGGTAC-3'; and the universal reverse primer sequence is shown in SEQ ID NO:9, specifically 5'-GCTCACATGTCCCTCCAAGCT-3'. The forward primer sequence of this invention contains a FAM marker at its 5' end, and the reverse primer sequence contains a HEX marker at its 5' end.
[0033] In this invention, the HazelID_SNP05 primer composition includes a forward primer sequence, a reverse primer sequence, and a universal reverse primer sequence. The forward primer sequence is shown in SEQ ID NO:13, specifically 5'-GAAGGTGACCAAGTTCATGCTTTTCGTCAGCCTATGGACATTTGGTT-3'; the reverse primer sequence is shown in SEQ ID NO:14, specifically 5'-GAAGGTCGGAGTCAACGGATTTTTCGTCAGCCTATGGACATTTGGTC-3'; and the universal reverse sequence is shown in SEQ ID NO:15, specifically 5'-CATGACCACCCATGGGAACAG-3'. The forward primer sequence of this invention contains a FAM marker at its 5' end, and the reverse primer sequence contains a HEX marker at its 5' end.
[0034] In this invention, the HazelID_SNP09 primer composition includes a forward primer sequence, a reverse primer sequence, and a universal reverse primer sequence. The forward primer sequence is shown in SEQ ID NO:25, specifically 5'-GAAGGTGACCAAGTTCATGCTCGATCCCAGACCAGCCAATAAATAAG-3'; the reverse primer sequence is shown in SEQ ID NO:26, specifically 5'-GAAGGTCGGAGTCAACGGATTCGATCCCAGACCAGCCAATAAATAAA-3'; and the universal reverse sequence is shown in SEQ ID NO:27, specifically 5'-CCTCTCAATGACTTCGAAACGATCT-3'. The forward primer sequence of this invention contains a FAM marker at its 5' end, and the reverse primer sequence contains a HEX marker at its 5' end.
[0035] In this invention, the HazelID_SNP11 primer composition includes a forward primer sequence, a reverse primer sequence, and a universal reverse primer sequence. The forward primer sequence is shown in SEQ ID NO:31, specifically 5'-GAAGGTGACCAAGTTCATGCTCAATTCGAGTTGGCGTATTCGGTTTA-3'; the reverse primer sequence is shown in SEQ ID NO:32, specifically 5'-GAAGGTCGGAGTCAACGGATTCAATTCGAGTTGGCGTATTCGGTTTG-3'; and the universal reverse sequence is shown in SEQ ID NO:33, specifically 5'-CCTTCAAAGGCCGATGAATTGTCC-3'. The forward primer sequence of this invention contains a FAM marker at its 5' end, and the reverse primer sequence contains a HEX marker at its 5' end.
[0036] This invention also provides the application of the KASP primer composition in the preparation of products related to the identification of European hybrid hazelnut varieties.
[0037] The present invention also provides a kit for identifying European hybrid hazelnut varieties, comprising a reagent for detecting the SNP molecular marker or the KASP primer composition.
[0038] The present invention also provides a method for identifying European hybrid hazelnut varieties using the KASP primer composition, comprising the following steps: (1) Extract genomic DNA from the young leaves of the test subject; (2) Using the young leaf genomic DNA obtained in step (1) as a template, amplification was performed using the KASP primer composition to obtain the amplification product; (3) Perform genotyping analysis on the amplification products obtained in step (2), and then compare the genotypes of the SNP molecular marker sites to determine the different varieties of hazelnuts.
[0039] In this invention, the amplification reaction system in step (2) consists of 10 μL: PARMS 2×Master Mix 5 μL, primer mixture 0.14 μL, DNA template 1 μL, and ddH2O added to 10 μL.
[0040] In this invention, the primer mixture is a premix of a forward primer sequence, a reverse primer sequence, and a universal reverse primer sequence. The preferred concentration ratio of the forward primer sequence, the reverse primer sequence, and the universal reverse primer sequence is (8-16) μM:(8-16) μM:(20-40) μM, more preferably (10-14) μM:(10-14) μM:(25-35) μM, and even more preferably 12 μM:12 μM:30 μM.
[0041] In this invention, the final concentration of the forward primer sequence in the reaction system is 0.168 μM, the final concentration of the reverse primer sequence is 0.168 μM, and the final concentration of the universal reverse primer sequence is 0.42 μM.
[0042] In this invention, the amount of DNA template used is preferably 10-50 ng, more preferably 20-40 ng, and even more preferably 30 ng.
[0043] In this invention, the amplification reaction program in step (2) is as follows: pre-denaturation at 94°C for 15 min; denaturation at 94°C for 20 s, annealing / extension at 65°C for 60 s, with a temperature decrease of 0.8°C per cycle, for 10 cycles; denaturation at 94°C for 20 s, annealing / extension at 57°C for 60 s, for 26 cycles. Each cycle in this invention includes two steps: denaturation and annealing / extension.
[0044] In this invention, the amplification product obtained in step (2) is subjected to genotyping analysis, and then the genotypes of the SNP molecular marker sites are compared to determine the molecular fingerprint type of the test subject, so as to assist in determining its variety identity.
[0045] The technical solutions provided by the present invention will be described in detail below with reference to the embodiments, but they should not be construed as limiting the scope of protection of the present invention.
[0046] Example 1: Determination of Core SNP Marker Sites
[0047] 1. Materials and Sequencing
[0048] Using 266 accessions of European-Ping hybrid hazelnut germplasm preserved in the National Hazelnut Germplasm Resource Nursery of the Liaoning Provincial Institute of Economic Forestry as experimental materials, this study covered approved varieties, superior lines, and offspring of different hybrid combinations. Genomic DNA was extracted from fresh young leaves using a modified CTAB method and whole-genome resequencing (PE150 strategy) was performed on the Illumina NovaSeq 6000 platform, generating a total of 2.98 Tb of clean data with an average sequencing depth of 11.2×.
[0049] Table 1. List of 266 hybrid hazelnut germplasm resources from Europe
[0050] Note: The bolded parts in Table 1 represent the 66 core germplasm accessions constructed.
[0051] 2. SNP Detection and Filtering
[0052] Using the Jefferson HAP1 v4.0 genome (genotype PRJNA1107824) of the European hazelnut variety as a reference, sequence alignment was performed using BWA-MEM v0.7.17, PCR duplicates were removed using Picard v2.27.5, and SNPs were detected using GATK v4.2.3 HaplotypeCaller. After hard filtering, removal of SNPs within 10 bp of the InDel region, and retention of biallelic loci with MAF > 0.05 and a deletion rate < 10%, a total of 10,218,717 high-quality biallelic SNP loci were obtained. The genome-wide Ts / Tv ratio was 2.48.
[0053] 3. Core Germplasm Construction
[0054] CoreHunter 3 software was used to construct core germplasm using a weighted combination strategy of genotypic distance and phenotypic distance (G:P=0.7:0.3) with a 25% sampling ratio. The optimized core germplasm comprised 66 accessions (24.8% of all materials), retaining 99.5% of the expected heterozygosity and 95.47% of the alleles.
[0055] 4. Three-level candidate SNP filtering
[0056] Using 66 core germplasm accessions as subjects, three levels of filtering were performed sequentially: 8,750,267 SNPs were retained after filtering with a deletion rate of 0; 1,797,187 SNPs were retained after filtering with a PIC ≥ 0.35; and 658 candidate SNPs were retained after filtering with an adjacent marker spacing ≥ 500 kb.
[0057] 5. Greedy Algorithm Selection
[0058] A greedy forward selection algorithm was used to screen 658 candidate SNPs to obtain candidate loci. Using the genotype matrix of 66 core germplasms as input, and aiming to maximize the pairwise discrimination rate (C(66,2)=2145 pairs), the algorithm selected the SNP with the highest marginal discriminative power from the remaining candidate SNPs in each round and added it to the panel. When discriminative power was the same, the algorithm prioritized the SNP with the higher PIC. The iteration terminated when the discrimination rate reached 100%. The results showed that a minimum of 5 SNPs were needed to achieve the theoretical discrimination rate of 100%. Subsequently, one locus with PIC=0.375 was added to each of the 8 previously uncovered chromosomes, ultimately obtaining 13 candidate SNP loci. The results are shown in Table 2.
[0059] Table 2 13 SNP candidate sites
[0060] Example 2: KASP Primer Design
[0061] Based on the Jefferson HAP1 v4.0 reference genome (gene number PRJNA1107824), KASP competitive allele-specific PCR primers were designed for the 13 candidate SNP sites identified in Example 1. Each site contains: a forward primer sequence (with a FAM marker in the 5' segment), a reverse primer sequence (with a HEX marker in the 5' segment), and a universal reverse primer (Common_R). Specific primer sequences are shown in Table 3.
[0062] Table 3. KASP primer sequences for different SNP molecular markers.
[0063] Example 3: KASP typing experiment verification
[0064] KASP primers were designed for all 13 sites and two rounds of genotyping experiments were conducted for verification. (one)
[0066] Genomic DNA was extracted from fresh young hazelnut leaves using a modified CTAB method and used as a DNA template for PCR amplification. The quality requirements for the genomic DNA were an A260 / A280 ratio of 1.8–2.0 and a concentration of ≥20 ng / μL.
[0067] Amplification reaction system (total volume 10 μL): PARMS 2×Master Mix 5 μL, primer mixture 0.14 μL (FAM forward primer 12 μM, HEX reverse primer 12 μM, universal reverse primer 30 μM premixed), DNA template 1 μL (10~50 ng), ddH2O 3.86 μL.
[0068] PCR program: 94℃ pre-denaturation for 15 min; 10 Touchdown cycles (94℃ 20s → annealing / extension 60s, annealing temperature decreased by 0.8℃ per cycle from 65℃); 26 isothermal cycles (94℃ 20s → 57℃ 60s).
[0069] After amplification, FAM and HEX fluorescence values were read using a fluorescence plate reader, and a two-dimensional scatter plot was plotted. Genotypes were determined based on cluster positions. Results showed that the three core discriminant loci (SNP03, SNP05, and SNP11) produced clear fluorescence clustering signals. SNP09 provided auxiliary reference information but was not used as an independent discriminant. The remaining nine loci failed to achieve reliable genotyping under the same conditions. The results are shown in Table 4.
[0070] Table 4. Experimental validation results of KASP genotyping for 13 candidate SNP loci.
[0071] The KASP genotyping results of the above 13 sites showed that the three core discriminant sites (SNP03, SNP05, SNP11) could produce clear fluorescent clustering signals, SNP09 could provide auxiliary reference information but not be used as an independent genotyping basis, and the KASP primers of the other 9 sites failed to achieve reliable genotyping under the same reaction system and amplification procedure conditions. A systematic analysis of the scatter plots of the failed sites showed the following two main characteristics: (1) Amplification failure type (SNP01, SNP04, SNP06, SNP07, SNP08, SNP10): all sample points were concentrated near the origin of the scatter plot (FAM<1.0, HEX<0.3), and the signals of the two fluorescent channels FAM and HEX were extremely weak, indicating that the PCR reaction failed to start effectively and the primers failed to anneal and bind successfully with the template DNA. A representative comparison of the failed scatter plots is shown in Figure 1 (SNP01 KASP genotyping scatter plot; all sample points are clustered near the origin, FAM and HEX fluorescence signals are extremely weak, no genotype clustering is formed, belonging to the typical amplification failure type; the primer GC content of this site is about 24%, Tm value is 43~49℃, far below the annealing temperature requirement) and Figure 2 (SNP08 KASP genotyping scatter plot; the sample points are also concentrated in the origin region, and there is no effective amplification signal in both channels; the primer GC content of this site is only about 14%, and the Tm value is as low as 36.8℃, which is the site with the highest AT content and the lowest Tm value among all candidate sites). (2) Non-specific amplification types (SNP02, SNP12, SNP13): Although the sample spots show some fluorescence signal, they are scattered and cannot form clear genotype clusters, indicating that the primers have amplified non-target regions or that primer dimers have interfered. A scatter plot of representative comparison failures is shown below. Figure 3 (SNP02 KASP genotyping scatter plot; the sample points have some fluorescence signal but are scattered randomly along the VIC axis, and cannot form clear genotype clusters, which belongs to non-specific amplification failure).
[0072] Further analysis of the primer sequence characteristics of the above-mentioned failed sites revealed that the nine failed or non-specific amplification sites shared the following common characteristics: (a) The AT base content in the flanking regions is generally high (65%~85%), with the AT content of primers for SNP06 and SNP08 exceeding 80%, resulting in extremely low GC content in the primers (14%~24%), which is far below the recommended value for KASP primer design (40%~60%). (b) The primer Tm values are generally low. The primer Tm values for SNP06 (38.7~41.8℃), SNP08 (36.8~43.7℃), and SNP10 (42.6~44.6℃) are much lower than the minimum annealing temperature (57℃) of the KASP standard Touchdown PCR program, which means that the primers cannot form stable base pairs with the template DNA at this temperature window. (c) AT-rich sequences are prone to forming hairpin structures within primers and dimers between primers, which further reduces amplification efficiency and allele specificity; (d) Although some sites can produce fluorescent signals (such as SNP02 and SNP12), the non-specific amplification results in insufficient separation between different genotype clusters, making it difficult to form standardized typing results that can be stably used across batches and laboratories.
[0073] All nine loci were true polymorphic loci in the whole-genome resequencing data (PIC=0.375, with balanced distribution of the three genotypes), and the loci themselves have the ability to distinguish varieties. The fundamental reason for the failure of KASP typing is that the nucleotide composition of the flanking regions of the loci is not suitable for KASP primer design, rather than the loss of locus polymorphism. (two)
[0075] KASP samples were used to validate the four SNP loci (SNP03, SNP05, SNP09, and SNP11) identified above.
[0076] 1. Verification materials
[0077] Twenty representative materials were selected from 66 core germplasm accessions, covering all comprehensive evaluation levels (excellent, good, qualified, and pending qualification), and including major approved varieties.
[0078] Table 5 Verification Materials
[0079] 2. DNA extraction and KASP reaction
[0080] Genomic DNA was extracted from fresh young hazelnut leaves using a modified CTAB method and used as a DNA template for PCR amplification. The quality requirements for the genomic DNA were an A260 / A280 ratio of 1.8–2.0 and a concentration of ≥20 ng / μL.
[0081] Amplification reaction system (total volume 10 μL): PARMS 2× Master Mix 5 μL; primer mixture 0.14 μL (12 μM forward primer, 12 μM reverse primer, 30 μM universal reverse primer); DNA template 1 μL; ddH2O to 10 μL. One NTC negative control was included for each site. PCR program: 94℃ pre-denaturation for 15 min; 94℃ denaturation for 20 s, 65℃ annealing / extension for 60 s (decreasing by 0.8℃ per cycle), 10 cycles; 94℃ denaturation for 20 s, 57℃ annealing / extension for 60 s, 26 cycles.
[0082] 3. KASP Fractal Scatter Plot
[0083] After amplification, FAM and HEX fluorescence values were read using a fluorescence plate reader, and a two-dimensional scatter plot was plotted. Genotypes were determined based on cluster positions. The genotype combinations of the test samples at the four SNP loci were compared with a standard genotype database to determine the variety identity.
[0084] The KASP genotyping scatter plots for the three core discriminant SNP sites and one auxiliary reference SNP site are shown below. Figures 4-7 As shown in the scatter plot, red dots represent homozygous allele 1, blue dots represent homozygous allele 2, green dots represent heterozygous genotypes, and cyan squares represent NTC negative controls. Figure 4 Among them, 20 samples showed two clear clusters, with 10 samples each of C / C homozygous and C / G heterozygous. Figure 5 Among them, 20 samples showed clear separation of three genotype clusters: 5 T / T, 10 C / T, and 5 C / C. Figure 6 Of the 20 samples, 19 were homozygous for G / G and 1 was heterozygous for A / G. The FAM channel signal was normal, while the HEX channel signal was weak. Figure 7 Among the 20 samples, three genotype clusters were clearly separated: 8 A / A, 7 G / A, and 5 G / G.
[0085] Genotyping results are shown in Table 6. Of the 80 reactions (4 loci per sample) from the initial testing of 20 samples, 79 yielded valid genotypes, resulting in an initial genotyping success rate of 98.75%. The SNP11 locus of SID_60 was retested and yielded a valid genotype. All NTC negative controls showed no signal.
[0086] Table 6. Genotyping Results
[0087] 5. Validation of consistency between KASP genotyping and resequencing genotypes
[0088] The KASP genotypes of the 20 samples were compared one by one with the corresponding locus genotypes obtained by whole-genome resequencing (WGS). The results of the locus analysis are shown in Table 7.
[0089] Table 7. Comparison results of KASP genotyping and whole-genome resequencing genotypes
[0090] For the three core discriminant markers SNP03, SNP05, and SNP11, the concordance rate between KASP genotyping and resequencing genotypes was 86.67% (52 out of 60 data points were consistent). The eight inconsistencies all showed a one-way bias where resequencing indicated heterozygosity (0 / 1) while KASP interpreted it as homozygous. This is a systematic phenomenon where differences in KASP fluorescence signal intensity cause heterozygous samples to lean towards a homozygous cluster, and it does not affect the effective differentiation between varieties.
[0091] For the auxiliary marker SNP09, the low GC content (approximately 35%) in the flanking region of this locus leads to insufficient amplification efficiency of the HEX channel (for detecting the A allele), resulting in underestimation of the signal in some samples containing the ALT allele. In this invention, this marker is positioned as an auxiliary reference marker and is not used as an independent criterion; its genotyping results are interpreted in conjunction with the other three core markers.
[0092] Three core discriminant SNP sites that could be stably amplified and made major discriminative contributions were finally identified (HazelID_SNP03, HazelID_SNP05, HazelID_SNP11), and one auxiliary reference site (HazelID_SNP09) that could provide supplementary discriminative information was retained.
[0093] Table 8. SNP molecular marker sites
[0094] The four SNP markers mentioned above are distributed on four different chromosomes (Chr01, Chr04, Chr06, and Chr09). Among them, SNP03, SNP05, and SNP11 serve as core discriminant markers, while SNP09 serves as an auxiliary reference marker. This approach avoids genetic linkage interference between markers and demonstrates good genomic representativeness. The four SNP markers balance discriminative efficiency and experimental reliability, making them more suitable for practical applications.
[0095] Example 4: Evaluation of the variety differentiation ability of core SNP combinations
[0096] 1. Discrimination rate of all 66 core germplasm accessions
[0097] Based on whole-genome resequencing data, the genotypes of all 66 core germplasms were extracted at 3 core discriminant SNP sites and 1 auxiliary reference SNP site, and the pairwise discrimination rate of each variety was calculated.
[0098] Total number of pairs: 2145 (C(66,2) represents the total number of combinations of randomly selecting 2 from 66 core germplasms and pairing them together, calculated as C(66,2)=66! / (2!×64!)=66×65 / 2=2145)
[0099] Distinguishing pair count: 2107 pairs
[0100] Discrimination rate = 2107 / 2145 × 100% = 98.23%
[0101] The 66 core accessions formed 40 unique genotype combinations at 4 SNP loci. The genotype frequencies at the 4 loci were evenly distributed, and the three genotypes (homozygous REF, heterozygous, and homozygous ALT) were fully representative at each locus.
[0102] 2.20 KASP validation samples' discriminant rate
[0103] Based on the 4-SNP KASP genotype combinations of the 20 core germplasms in Example 3 (each sample has one genotype at each of the four SNP loci; stringing these four genotypes together gives the "genotype combination" of the sample, equivalent to its molecular identification number), the pairwise variety discrimination rate was 95.26% (out of 190 pairwise combinations C(20,2) = 190, 181 genotype combinations were different and could be effectively distinguished). The discrimination rate using only the three core discriminant markers, SNP03, SNP05, and SNP11, was 94.21% (179 out of 190 pairs were distinguishable), indicating that the core discriminant marker combination has sufficient variety discrimination capability.
[0104] In summary, the use of the SNP molecular marker and KASP primer combination described in this invention for the identification of European hybrid hazelnut varieties can achieve rapid, accurate, and low-cost identification of European hybrid hazelnut varieties with fewer detection sites. It has high discrimination efficiency, simple detection process, and strong repeatability, and has clear industrial application value. This also provides technical support for the authenticity identification of European hybrid hazelnut varieties, breeding material management, seedling quality supervision, and digital protection of germplasm resources.
[0105] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. The application of a reagent for detecting SNP molecular markers in the preparation of products related to the identification of European hybrid hazelnut varieties, characterized in that, The SNP molecular markers include HazelID_SNP03, HazelID_SNP05, HazelID_SNP09, and HazelID_SNP11; The HazelID_SNP03 is located at the 25,552,398th base on chromosome 1 of the Jefferson HAP1_v4.0 reference genome of the European hazelnut, with the mutated base being G or C; The HazelID_SNP05 is located at the 15,913,337th base on chromosome 4 of the Jefferson HAP1_v4.0 reference genome of the European hazelnut, with the mutated base being either T or C; The HazelID_SNP09 is located at the 22,599,658th base on chromosome 6 of the Jefferson HAP1_v4.0 reference genome of the European hazelnut, with the mutated base being G or A; The HazelID_SNP11 is located at the 17014214th base on chromosome 9 of the Jefferson HAP1_v4.0 reference genome of the European hazelnut, and the mutated base is either A or G.
2. The KASP primer composition for amplifying the SNP molecular marker as described in claim 1, characterized in that, This includes primer compositions for HazelID_SNP03, HazelID_SNP05, HazelID_SNP09, and HazelID_SNP11; The HazelID_SNP03 primer composition includes a forward primer sequence as shown in SEQ ID NO:7, a reverse primer sequence as shown in SEQ ID NO:8, and a universal reverse primer sequence as shown in SEQ ID NO:9; The HazelID_SNP05 primer composition includes a forward primer sequence as shown in SEQ ID NO:13, a reverse primer sequence as shown in SEQ ID NO:14, and a universal reverse sequence as shown in SEQ ID NO:15; The HazelID_SNP09 primer composition includes a forward primer sequence as shown in SEQ ID NO:25, a reverse primer sequence as shown in SEQ ID NO:26, and a universal reverse sequence as shown in SEQ ID NO:27; The HazelID_SNP11 primer composition includes a forward primer sequence as shown in SEQ ID NO:31, a reverse primer sequence as shown in SEQ ID NO:32, and a universal reverse sequence as shown in SEQ ID NO:
33.
3. The application of the KASP primer composition of claim 2 in the preparation of products related to the identification of European hybrid hazelnut varieties.
4. A kit for identifying hybrid hazelnut varieties, characterized in that, This includes reagents for detecting the SNP molecular markers of claim 1, or the KASP primer composition of claim 2.
5. A method for identifying European hybrid hazelnut varieties using the KASP primer composition of claim 2, characterized in that, Includes the following steps: (1) Extract genomic DNA from the young leaves of the test subject; (2) Using the young leaf genomic DNA obtained in step (1) as a template, amplification is performed using the KASP primer composition described in claim 2 to obtain the amplification product; (3) Perform genotyping analysis on the amplification products obtained in step (2), and then compare the genotypes of the SNP molecular marker sites described in claim 1 to determine different varieties of hazelnuts.
6. The method according to claim 5, characterized in that, The amplification reaction system described in step (2) consists of 10 μL: PARMS 2×Master Mix 5 μL, primer mixture 0.14 μL, DNA template 1 μL, and ddH2O added to 10 μL. The primer mixture consists of a forward primer sequence, a reverse primer sequence, and a universal reverse primer sequence with a concentration ratio of (8-16) μM:(8-16) μM:(20-40) μM.
7. The method according to claim 5, characterized in that, The amplification reaction program in step (2) is as follows: pre-denaturation at 94℃ for 15 min; denaturation at 94℃ for 20 s, annealing / extension at 65℃ for 60 s, with a decrease of 0.8℃ per cycle, for 10 cycles; denaturation at 94℃ for 20 s, annealing / extension at 57℃ for 60 s, for 26 cycles.
Citation Information
Patent Citations
SNP (Single Nucleotide Polymorphism) molecular marker combination for constructing hazel DNA (Deoxyribonucleic Acid) fingerprint spectrum and application thereof
CN120519617A