A kasp technology-based pumpkin snp marker combination and application thereof in germplasm identification
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NORTHWEST A & F UNIV
- Filing Date
- 2026-07-02
- Publication Date
- 2026-08-04
AI Technical Summary
[0005]本发明要解决的技术问题是:针对目前印度南瓜缺乏高效、准确、低成本的分子鉴定标记的现状,以及现有SSR标记操作繁琐、通量低、主观性强,SNP芯片成本高、物种特异性强无法跨物种通用等技术缺陷,提供一种基于KASP技术的印度南瓜SNP标记组合及其在种质鉴定中的应用
[0044] 1. This invention is the first to develop a dedicated KASP-SNP marker combination for Indian pumpkin, solving the technical problem that markers in existing technologies cannot be universally applied across species. Currently available SNP chips for Chinese pumpkin have an amplification success rate of less than 50% in Indian pumpkin due to the only 70%-80% homology between the genome sequences of Indian and Chinese pumpkins, making them unusable directly. Furthermore, KASP markers for other crops, such as tobacco, are completely different from those for pumpkin and have no universality. This invention clearly defines 20 SNP sites and their primer sequences specifically for Indian pumpkin, filling the technical gap of lacking efficient KASP molecular identification markers for Indian pumpkin.
Smart Images

Figure CN122503544A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of molecular marker-related technology, and in particular to a KASP-based SNP marker combination for Indian pumpkin and its application in germplasm identification. Background Technology
[0002] Pumpkin, an annual vine belonging to the Cucurbita genus of the Cucurbitaceae family, has a long history of cultivation and is grown worldwide. Highly nutritious, pumpkin offers various health benefits, including lowering blood pressure, blood sugar, and blood lipids, and protecting eyesight, making it a popular vegetable in my country. However, in recent years, with the development and utilization of pumpkin germplasm resources, problems such as varietal confusion and unclear genetic backgrounds have become increasingly prominent. The market is rife with counterfeit products and inferior goods being sold as superior ones, which not only seriously infringes upon the legitimate rights and interests of breeders and consumers but also hinders the promotion of superior varieties, impedes pumpkin breeding efforts, and ultimately negatively impacts the healthy development of the pumpkin industry. Therefore, clarifying the genetic background of pumpkin germplasm resources and breeding materials, and establishing efficient and accurate methods for identifying seed purity and authenticity, is crucial.
[0003] Currently, the construction of fingerprint profiles and variety identification for pumpkin germplasm resources and varieties mainly employs technologies such as SSR markers and SNP chips. For example, patent application CN107805672A discloses a method for identifying the authenticity of rootstock varieties of Indian pumpkin × Chinese pumpkin hybrids using five pairs of SSR primers. However, SSR markers require gel electrophoresis detection, which is cumbersome, has low throughput, and results are highly subjective. Furthermore, this technology is developed specifically for hybrid rootstocks and cannot be directly applied to the broad-spectrum identification of pure Indian pumpkin germplasm. Patent application CN118497397A discloses a Chinese pumpkin SNP chip containing 20,213 SNPs, which can be used for population structure analysis, genome-wide association analysis, etc. However, this chip is designed for Chinese pumpkin (Cucurbita amoschata), while Indian pumpkin (Cucurbita maxima) is a different species with approximately 70%-80% genome sequence homology. The amplification success rate of Chinese pumpkin SNP markers in Indian pumpkin is usually less than 50%, making it unsuitable for cross-species application. Furthermore, high-throughput chip detection is costly, involves a large number of markers, and suffers from information redundancy, making it unsuitable for routine variety identification scenarios. In addition, patent application CN114507747A discloses tobacco SNP markers developed based on KASP technology, which can be used for tobacco variety fingerprinting and purity detection. However, this technology is developed specifically for tobacco and is completely different from pumpkin species, resulting in a lack of marker universality.
[0004] In summary, current molecular marker research on pumpkins faces several challenges. SSR markers are cumbersome and inefficient, failing to meet the demands of high-throughput detection. SNP chip technology is costly, species-specific, and chips developed for Chinese pumpkins cannot be used for Indian pumpkins. While KASP technology has been applied in other crops, it remains a blank area in Indian pumpkins. Currently, there is a lack of a core SNP marker combination specifically for Indian pumpkins (Cucurbita maxima) based on KASP technology that balances detection efficiency and cost, making it difficult to meet the practical needs of germplasm resource identification, varietal authenticity testing, and seed purity assessment. Therefore, developing a KASP-SNP core marker combination suitable for Indian pumpkins has significant research value and practical implications. Summary of the Invention
[0005] The technical problem to be solved by this invention is: to address the current lack of efficient, accurate and low-cost molecular identification markers for Indian pumpkin, as well as the technical shortcomings of existing SSR markers such as cumbersome operation, low throughput and strong subjectivity, and SNP chips with high cost and strong species specificity that cannot be used across species, this invention provides an Indian pumpkin SNP marker combination based on KASP technology and its application in germplasm identification.
[0006] To solve the above-mentioned technical problems, the present invention adopts the following technical solution: a KASP-based SNP marker combinatorial for Cucurbita maxima, comprising 20 SNP loci, and its detection primers consisting of 20 independent KASP primer sets, each primer set being used to detect a genome-specific SNP locus of Cucurbita maxima. Each primer set includes two competitive forward primers F1 and F2 and one reverse universal primer R, wherein the 5' end of F1 carries the tail sequence corresponding to the first fluorescent reporter group, and the 5' end of F2 carries the tail sequence corresponding to the second fluorescent reporter group. The 3' terminal bases of F1 and F2 correspond to the two alleles of the SNP locus, respectively, to achieve dual-color fluorescent competitive allele-specific PCR typing.
[0007] Furthermore, the 20 SNP loci are evenly distributed across the 20 chromosomes of *Swiss bipinnatus*, with one SNP locus per chromosome, achieving uniform coverage across the entire genome. The polymorphism information content (PIC) value for each locus is ≥0.35, and the minimum allele frequency (MAF) value is ≥0.3, demonstrating high polymorphism and resolution.
[0008] Furthermore, the 5' tail sequence of F1 is GAAGGTGACCAAGTTCATGCT, and the 5' tail sequence of F2 is GAAGGTCGGAGTCAACGGATT. The 3' core region length of F1 and F2, excluding the tail, is 21 to 27 nt, and F1 and F2 differ by only one base at the SNP identification site.
[0009] Specifically, the nucleotide sequences of F1, F2, and R in the 20 primer sets are shown in the following SEQ ID NO, or their complementary sequences:
[0010] Cma1_2: F1 as SEQ ID NO:1, F2 as SEQ ID NO:2, R as SEQ ID NO:3;
[0011] Cma2_2: F1 as SEQ ID NO:4, F2 as SEQ ID NO:5, R as SEQ ID NO:6;
[0012] Cma3_2: F1 as SEQ ID NO:7, F2 as SEQ ID NO:8, R as SEQ ID NO:9;
[0013] Cma4_2: F1 as SEQ ID NO:10, F2 as SEQ ID NO:11, R as SEQ ID NO:12;
[0014] Cma5_1: F1 as in SEQ ID NO:13, F2 as in SEQ ID NO:14, R as in SEQ ID NO:15;
[0015] Cma6_1: F1 as SEQ ID NO:16, F2 as SEQ ID NO:17, R as SEQ ID NO:18;
[0016] Cma7_2: F1 as SEQ ID NO:19, F2 as SEQ ID NO:20, R as SEQ ID NO:21;
[0017] Cma8_2: F1 as SEQ ID NO:22, F2 as SEQ ID NO:23, R as SEQ ID NO:24;
[0018] Cma9_1: F1 as SEQ ID NO:25, F2 as SEQ ID NO:26, R as SEQ ID NO:27;
[0019] Cma10_2: F1 as SEQ ID NO:28, F2 as SEQ ID NO:29, R as SEQ ID NO:30;
[0020] Cma11_2: F1 as SEQ ID NO:31, F2 as SEQ ID NO:32, R as SEQ ID NO:33;
[0021] Cma12_2: F1 as SEQ ID NO:34, F2 as SEQ ID NO:35, R as SEQ ID NO:36;
[0022] Cma13_2: F1 as SEQ ID NO:37, F2 as SEQ ID NO:38, R as SEQ ID NO:39;
[0023] Cma14_2: F1 as SEQ ID NO:40, F2 as SEQ ID NO:41, R as SEQ ID NO:42;
[0024] Cma15_1: F1 as SEQ ID NO:43, F2 as SEQ ID NO:44, R as SEQ ID NO:45;
[0025] Cma16_2: F1 as SEQ ID NO:46, F2 as SEQ ID NO:47, R as SEQ ID NO:48;
[0026] Cma17_1: F1 as SEQ ID NO:49, F2 as SEQ ID NO:50, R as SEQ ID NO:51;
[0027] Cma18_1: F1 as SEQ ID NO:52, F2 as SEQ ID NO:53, R as SEQ ID NO:54;
[0028] Cma19_4: F1 as SEQ ID NO:55, F2 as SEQ ID NO:56, R as SEQ ID NO:57;
[0029] Cma20_2: F1 as SEQ ID NO:58, F2 as SEQ ID NO:59, R as SEQ ID NO:60.
[0030] Furthermore, the 20 SNP loci are located on chromosomes 1 to 20 of the Indian squash genome, with the following specific physical locations: Cma1_2 is located at position 9975272 on chromosome 1, Cma2_2 is located at position 973733 on chromosome 2, Cma3_2 is located at position 8907238 on chromosome 3, Cma4_2 is located at position 980502 on chromosome 4, Cma5_1 is located at position 1000157 on chromosome 5, Cma6_1 is located at position 1023295 on chromosome 6, Cma7_2 is located at position 856994 on chromosome 7, Cma8_2 is located at position 925481 on chromosome 8, Cma9_1 is located at position 1030918 on chromosome 9, and Cma10_2 is located on chromosome 10. Cma11_2 is located at position 999475 on chromosome 11, Cma12_2 is located at position 9717178 on chromosome 12, Cma13_2 is located at position 852718 on chromosome 13, Cma14_2 is located at position 9917851 on chromosome 14, Cma15_1 is located at position 1058773 on chromosome 15, Cma16_2 is located at position 980007 on chromosome 16, Cma17_1 is located at position 1126728 on chromosome 17, Cma18_1 is located at position 10017687 on chromosome 18, Cma19_4 is located at position 1336123 on chromosome 19, and Cma20_2 is located at position 1384468 on chromosome 20.
[0031] A method for identifying the authenticity / purity of Indian squash germplasm, varieties, or seeds using the above-mentioned SNP marker combinations includes the following steps:
[0032] (1) Extract genomic DNA from the sample of the Indian pumpkin to be tested;
[0033] (2) Using the above-mentioned marker combination, competitive allele-specific PCR amplification was performed on the genomic DNA;
[0034] The reaction system is 10 μL, containing: 1 μL DNA template, 0.14 μL primer mixture, 5 μL 2×KASP Mix, and 3.86 μL distilled water; the volume ratio of Forward1, Forward2, and reverse primers in the primer mixture is 12:12:30.
[0035] The amplification program was as follows: 95℃ pre-denaturation for 10 min; (95℃ denaturation for 15 s, 61℃→55℃ touchdown annealing extension for 1 min, decreasing by 0.6℃ per cycle) × 10 cycles; (95℃ denaturation for 15 s, 55℃ annealing for 40 s) × 35 cycles;
[0036] (3) Read the FAM / HEX fluorescence signal, determine the genotype of each SNP site, and obtain the genotyping results of the sample to be tested at the 20 SNP sites;
[0037] The classification rules are as follows: if the signal falls into the FAM cluster, it is recorded as homozygous isotype 1; if it falls into the HEX cluster, it is recorded as homozygous isotype 2; if it falls in between, it is recorded as heterozygous; if it cannot form a cluster, it is recorded as NN. The reaction well is then re-read after a cycle.
[0038] (4) Arrange the genotyping results of the 20 SNP loci in a fixed order of chromosomes 1 to 20 into a string of length 20, with each digit taken from {A,G,T,C,NN}, to generate the SNP fingerprint code of the sample to be tested; compare the SNP fingerprint code with the SNP fingerprint database of standard varieties / germplasm, and when the matching degree reaches the set threshold, it is determined to be the same variety / germplasm or the seed is genuine and qualified; otherwise, it is determined to be different varieties / germplasm or the authenticity / purity is unqualified.
[0039] Furthermore, step (4) also includes converting the SNP fingerprint code into a QR code or barcode to realize the digital identification and management of germplasm resources.
[0040] Furthermore, the matching threshold is set as follows: when identifying the authenticity of a variety, a matching degree of ≥95% indicates that the variety is the same; when identifying the purity of a seed, a hybrid rate of ≤5% indicates that the variety is qualified.
[0041] A kit comprising the above-described marker combination and a KASP PCR reaction mixture.
[0042] The use of a KASP-based SNP marker combinatorial system for Indian squash in the preparation of kits for identifying the authenticity of Indian squash varieties, determining seed purity, constructing SNP fingerprints for germplasm resources, or analyzing genetic diversity.
[0043] Compared with the prior art, the beneficial effects of the present invention are:
[0044] 1. This invention is the first to develop a dedicated KASP-SNP marker combination for Indian pumpkin, solving the technical problem that markers in existing technologies cannot be universally applied across species. Currently available SNP chips for Chinese pumpkin have an amplification success rate of less than 50% in Indian pumpkin due to the only 70%-80% homology between the genome sequences of Indian and Chinese pumpkins, making them unusable directly. Furthermore, KASP markers for other crops, such as tobacco, are completely different from those for pumpkin and have no universality. This invention clearly defines 20 SNP sites and their primer sequences specifically for Indian pumpkin, filling the technical gap of lacking efficient KASP molecular identification markers for Indian pumpkin.
[0045] 2. The 20 SNP markers of this invention have undergone rigorous screening, with each locus having a polymorphism information content (PIC) value ≥ 0.35 and a minimum allele frequency (MAF) value ≥ 0.3. They are evenly distributed across the 20 chromosomes of *Cucurbita indicus*, with one locus per chromosome, achieving uniform coverage across the entire genome. Compared to traditional SSR markers with only a few pairs, the markers of this invention are more numerous, more evenly distributed, and exhibit higher polymorphism. They can distinguish 100% of the 28 tested *Cucurbita indicus* germplasms, including closely related varieties, demonstrating significantly superior discrimination ability compared to traditional SSR markers and comprehensively and accurately reflecting the genetic background of the germplasm.
[0046] 3. This invention utilizes the KASP technology platform, combined with an optimized 10μL reaction system, a 12:12:30 primer ratio, and a touchdown amplification program, achieving efficient and low-cost detection. Compared to the drawbacks of traditional SSR labeling, which requires gel electrophoresis, is cumbersome, and takes 2-3 days per batch, this invention requires only 4-6 hours per batch, reducing detection time by more than 80%. The cost per sample is approximately 15-20 yuan, about 50% lower than SSR labeling. High-throughput detection can be performed on 96-well or 384-well plates, with dozens to hundreds of samples per batch, increasing throughput by 2-4 times, meeting the needs of large-scale routine testing in modern seed industry.
[0047] 4. This invention, based on automated genotyping of fluorescence signals, avoids the subjectivity issues of manual band interpretation in traditional SSR labeling. This invention clearly defines specific genotyping rules: signals falling into the FAM cluster are denoted as homozygous allele 1, those falling into the HEX cluster are denoted as homozygous allele 2, those falling in between are denoted as heterozygous, and those not forming clusters are denoted as NN with additional cyclic readout. This standardized genotyping method ensures a genotyping accuracy of ≥99.5%, and the optimized reaction system and amplification procedure guarantee the specificity and stability of amplification. Results from different laboratories and different operators are highly comparable, facilitating mutual recognition and standardization of detection results.
[0048] 5. This invention can be used not only for variety identification but also for the digital management of germplasm resources. This invention further specifies the technical solutions for converting SNP fingerprint codes into QR codes or barcodes, and for setting matching thresholds (authenticity ≥95%, purity hybridization rate ≤5%). Compared to traditional SSR markers that can only perform simple variety identification and high-throughput chips used only for scientific research, this invention can be widely applied to multiple scenarios such as variety authenticity identification, seed purity detection, germplasm resource genetic diversity analysis, molecular marker-assisted breeding, and digital management of germplasm resources, demonstrating significant practical value and promising prospects for widespread application. Attached Figure Description
[0049] Figure 1 This is a SNP clustering analysis diagram of 28 pumpkin germplasm accessions in this invention;
[0050] Figure 2 This is a schematic diagram of the fingerprint spectrum QR code of 28 pumpkin germplasms of the present invention. Detailed Implementation
[0051] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0052] Example 1: Design and Synthesis of KASP Tags
[0053] 1.1 SNP tag selection process
[0054] The 20 SNP markers in this invention were not simply randomly selected from the genome, but underwent rigorous multi-round screening and experimental verification to ensure high quality, high polymorphism, and high stability. The specific screening process is as follows:
[0055] Round 1: Mining of raw SNP sites
[0056] Based on the whole-genome resequencing data of 28 pumpkin germplasm resources provided by the Watermelon and Melon Research Group of the College of Horticulture, Northwest A&F University, the sequencing sequences were aligned to the pumpkin reference genome ("Rimu", V1) using BWA software. Variation detection and filtering were performed using the GATK tool to screen for high-quality SNP loci, resulting in approximately 1.25 million original SNP loci.
[0057] Second round: Initial screening using bioinformatics
[0058] The original SNP loci were initially screened according to the following strict criteria:
[0059] (1) Flanking sequence specificity: The sequences within 100 bp upstream and downstream of the SNP site must be unique and free of homologous repeat sequences to ensure specific amplification of the primers;
[0060] (2) Appropriate GC content: The GC content of the flanking sequences should be between 40% and 60% to avoid the formation of secondary structures and ensure PCR amplification efficiency;
[0061] (3) High polymorphism: Minimum allele frequency (MAF) ≥ 0.3, ensuring that the locus has high polymorphism in the tested germplasm;
[0062] (4) High polymorphic information content: Polymorphic information content (PIC) ≥ 0.35, ensuring the resolving power of the marker;
[0063] (5) Prioritize functional relevance: Prioritize SNP sites located in gene coding regions or regulatory regions, especially non-synonymous mutation sites, to increase the functional significance of the marker;
[0064] (6) Avoid InDel interference: There are no insertion / deletion (InDel) variations within 50bp upstream and downstream of the SNP site, thus avoiding affecting primer binding and amplification efficiency.
[0065] After the above screening, approximately 3,200 candidate SNP sites were selected from 1.25 million original sites.
[0066] Third round: Chromosome distribution screening
[0067] To achieve uniform coverage across the entire genome, screening was conducted according to the principle of "selecting the optimal site for each chromosome":
[0068] (1) The 3200 candidate loci were classified according to their chromosomal location;
[0069] (2) A comprehensive score is given to the candidate sites on each chromosome. The scoring indicators include: PIC value, MAF value, flanking sequence specificity, GC content, distance from the gene, etc.
[0070] (3) Select the 3-5 sites with the highest scores for each chromosome as candidates, and obtain a total of about 80 candidate sites.
[0071] Fourth round: Experimental verification and screening
[0072] KASP primers were designed for the above 80 candidate sites, and experimental verification was performed using 28 pumpkin germplasm accessions. The selection was based on the following criteria:
[0073] (1) High amplification efficiency: high fluorescence signal intensity, signal-to-noise ratio ≥5:1;
[0074] (2) Clear classification: good clustering effect, homozygous and heterozygous types can be clearly distinguished, and there are no obvious intermediate or diffuse types;
[0075] (3) No template control (NTC) no amplification: The negative control showed no fluorescence signal and no nonspecific amplification;
[0076] (4) Good stability: The consistency rate of typing results in repeated experiments is ≥99%;
[0077] (5) High detection rate: The detection success rate in 28 tested germplasm samples is ≥95%.
[0078] Experimental results showed that 42 out of 80 candidate sites met all of the above criteria.
[0079] Round 5: Final Core Marker Determined
[0080] From the 42 successfully validated loci, 20 core SNP markers were ultimately determined according to the following principles:
[0081] (1) Uniform distribution throughout the genome: ensure that each chromosome has at least one marker, achieving full coverage of all 20 chromosomes;
[0082] (2) Optimal polymorphism: Prioritize sites with the highest PIC and MAF values;
[0083] (3) Optimal typing effect: Prioritize the selection of sites with the clearest clustering and the strongest signal;
[0084] (4) The combination has the strongest distinguishing power: Through simulation analysis, the combination of 20 sites can distinguish 28 tested germplasms with 100% accuracy, and also has a good distinguishing ability for closely related varieties.
[0085] The 20 core SNP markers obtained in the final screening are characterized by uniform distribution, high polymorphism, high stability, and high resolution, which can meet the needs of Indian pumpkin variety identification and fingerprint map construction.
[0086] 1.2 KASP Primer Design and Synthesis
[0087] KASP primers were designed using the SNP_Primer_Pipeline2-master process.
[0088] Each set of KASP primers contains two forward competing primers (F1, F2) and one reverse universal primer (R). The 5' end of F1 is fitted with a specific tail sequence (GAAGGTGACCAAGTTCATGCT) that binds to FAM fluorescence, while the 5' end of F2 is fitted with a specific tail sequence (GAAGGTCGGAGTCAACGGATT) that binds to HEX fluorescence. The primers were synthesized by a commercial company; their specific nucleotide sequences are detailed in the sequence listing.
[0089] Example 2: Cultivation of test materials and DNA extraction
[0090] 2.1 Test Materials
[0091] The germplasm information of the 28 Indian squash accessions tested is shown in Table 1. All germplasm was provided by the Watermelon and Melon Research Group of the College of Horticulture, Northwest A&F University.
[0092] Table 1. Pumpkin germplasm resources tested
[0093]
[0094] 2.2 Material Cultivation
[0095] All pumpkin germplasm was soaked in warm water (55℃) for 8 hours, rinsed 2-3 times with sterile water, and then placed in a 30℃ constant temperature and light incubator to germinate. Once the seeds showed signs of germination, they were sown individually in 50-cell seedling trays in a greenhouse, with 5 seeds sown for each germplasm type. The greenhouse temperature was maintained at 25-30℃. After sowing, normal water and fertilizer management was carried out to ensure uniform seedling emergence.
[0096] 2.3 DNA Extraction
[0097] When pumpkin seedlings reach the two-leaf-one-heart stage, fresh tender leaves are collected as DNA extraction material. Total DNA is extracted from the leaves using the CTAB method: approximately 1g of tender leaf tissue is weighed and placed in a cryogenic grinding tube pre-added with 2% CTAB extraction solution; the grinding tube is placed in a 65℃ constant temperature oven for 30 minutes to soften the cell walls and accelerate DNA release; after incubation, 800μL of a 24:1 chloroform-isoamyl alcohol mixture is added to the tube, mixed well, and centrifuged at 12000r / min for 15 minutes; the supernatant is transferred to a new centrifuge tube, and 2 / 3 volume of isopropanol is added. The mixture is gently inverted to mix well; after centrifugation again, the supernatant is discarded, and the precipitate in the tube is the DNA precipitate; the precipitate is washed twice with 75% ethanol to remove residual impurities; after drying, it is dissolved in 200μL ddH2O and stored at -20℃ for later use. The DNA concentration is measured using a UV spectrophotometer to ensure that the DNA quality meets the requirements for PCR amplification (OD200). 260 / 280 The ratio is between 1.8 and 2.0.
[0098] Example 3: KASP-PCR reaction and genotyping
[0099] 3.1 Primer preparation
[0100] First, dilute the KASP primers to a working concentration of 100 μmol·L⁻¹ and prepare the primer mixture as follows: 12 μL each of Forward1 (F1) and Forward2 (F2) forward primers, and 30 μL of reverse primer (R). Then, bring the volume to 100 μL with distilled water. After preparation, store the primer mixture at 4℃ for later use. If long-term storage is required, store it at -20℃.
[0101] 3.2 PCR reaction system and procedure
[0102] PCR reactions were performed in 96-well plates with the following reaction mixture (10 μL): 1 μL DNA template, 0.14 μL primer mixture, 5 μL 2×KASP Mix (provided by Beijing Jiacheng Biotechnology Co., Ltd.), and 3.86 μL distilled water. Three template-free negative controls (NTCs) were included in each DNA sample. After sample loading, the reaction plate was sealed with a membrane.
[0103] The PCR amplification reaction was performed using a real-time quantitative PCR instrument, and the program settings were as follows:
[0104] Stage 1: Pre-denaturation at 95℃ for 10 min;
[0105] Stage 2 (Touchdown): Denaturation at 95℃ for 15 seconds, annealing and extension at 61~55℃ for 1 minute (decreasing by 0.6℃ per cycle), for a total of 10 cycles;
[0106] Stage 3: denaturation at 95℃ for 15 seconds, annealing at 55℃ for 40 seconds, for a total of 35 cycles.
[0107] After the PCR reaction is complete, read the fluorescence typing data. If the typing effect is not good (weak signal or unclear clustering), continue the amplification. The amplification program is 95℃ denaturation for 20s, 55℃ annealing for 40s. Check the typing status every 4 cycles until a clear genotyping result is obtained.
[0108] 3.3 Genotyping
[0109] After quantitative real-time PCR amplification, genotyping and analysis were performed using a quantitative real-time PCR instrument (QuantStudio3, Thermo Fisher Scientific) to obtain the genotype of each pumpkin germplasm at the designed locus. The genotyping results are presented in two-base format, such as "AA" representing homozygotes and "AT" representing heterozygotes. Genotypes that were not successfully detected are marked with "NN".
[0110] The classification rules are as follows: if the signal falls into the FAM cluster, it is recorded as homozygous isotype 1; if it falls into the HEX cluster, it is recorded as homozygous isotype 2; if it falls in between, it is recorded as heterozygous; if it cannot form a cluster, it is recorded as NN. The reaction well is then re-read after a cycle.
[0111] Example 4: SNP marker details and typing results
[0112] 4.1 Detailed information on the 20 KASP-SNP markers
[0113] The 20 SNP markers selected in this invention are located on different chromosomes of pumpkin (Swiss squash), and their specific physical locations and flanking sequence information are as follows (where the SNP sites and their two alleles are in parentheses):
[0114] The primer numbers, chromosome locations, and sequence information for the 20 successfully genotyped primers are as follows:
[0115] Cma1_2: Located at position 9975272 on chromosome 1, its nucleotide is A or G. The sequence extending 50 bp before and after this position is: TGATATAGTAAAGAAAAGAATAGATTACCATGGCAATGTGATAGATTCAA[G / A]ACCAGATGGGATTGGATCATCTAAGGTAATTACATATTCTTTTGAAGGCT
[0116] Cma2_2: Located at position 973733 on chromosome 2, its nucleotides are T or G. The sequence extending 50 bp before and after this position is GTGAACTATCTTTCCTTTCTGTTTCAATTTCATTACAATGAGAACGAGGC[T / G]AGAGGGGCAATTCAGCTTAGAGATAAATGGGGACGTTTCGTTCGATAACT
[0117] Cma3_2: Located at position 8907238 on chromosome 3, its nucleotide is A or T. The sequence extending 50 bp before and after this position is TTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTT[T / A]AATGCCATTTAGGTGTGTGCTTTTGTGAAATGTTTCCCATATTGTGGTAT
[0118] Cma4_2: Located at position 980502 on chromosome 4, its nucleotides are T or C. The sequence extending 50 bp before and after this position is CCGCTAATAAATATTATTTGTTTTGACTCGTTACGTATCGTCGTTAGTCT[C / T]ACTATTTTAAAACGTATTTACTAAGGAGAGATTTTTATCTTCTTGTAAAA
[0119] Cma5_1: Located at position 1000157 on chromosome 5, its nucleotides are T or C. The sequence extending 50 bp before and after this position is: AGAGAATGGCAGAAAGTTAACTACTTAGTTTAGATTCACCCCGACTTAC[C / T]GGTATGTTTTTCTCAAAAGGTGAGGAGTAGAAATGCTGGTCAATGCTCGA
[0120] Cma6_1: Located at position 1023295 on chromosome 6, its nucleotides are A or G. The sequence extending 50 bp before and after this site is AAATTGGGCGGCGGTAAATGCTCCGTTTCAATTAAAAATGCCCTTCCGAT[A / G]AACCCTTCGCCGTACTCCGCCGGCACTGCCAGAGTAACATCGTGACCAGG
[0121] Cma7_2: Located at position 856994 on chromosome 7, its nucleotides are G or C. The sequence extending 50 bp before and after this position is TAGATGCATTTTAAAATCTGGAGGGGAAACCCTTGAAAGGAATACCCAAA[G / C]AGGACAATATCAGCTCGCGGTCGGTCGGTTTAAGTAGTTACATGAAAACT
[0122] Cma8_2: Located at position 925481 on chromosome 8, its nucleotides are A or G. The sequence extending 50 bp before and after this position is TCGAGATTTTGGGGAATTTTAAACTGTTTCATAGATGCTGGAATCTGAGA[A / G]TCTGGGCTACACTTTGCTTGGAAAAGCGGAGCATCTAATACAGAGAAAAT
[0123] Cma9_1: Located at position 1030918 on chromosome 9, its nucleotide is A or G. The sequence extending 50 bp before and after this site is CTCTAAATTCATAGGTATGGAAGTTTATATGTAAAAGGTATACTTTGCAC[G / A]ACCCAAACATACGCTTAACTCATATATACACTTACTCTTATGGTAATTAT
[0124] Cma10_2: Located at position 999475 on chromosome 10, its nucleotides are T or C. The sequence extending 50 bp before and after this position is: ATTTGTCATATGATTGATTTTAAAATACCAGTCCTACCCCTAATAATGAT[C / T]AGTCTGTGCCTTTATTTATTCCACTACTTCTATTAAAATAAAATAAATTT
[0125] Cma11_2: Located at position 12503969 on chromosome 11, its nucleotides are A or G. The sequence extending 50 bp before and after this site is: CTTTCTATGCTGAAGTTGCTTCTGCAGTAAGAGATGTGCAAGAACCGTTG[A / G]CAAACATGGGTCGGTCGAGCAAAGTGTTCGCTATGTACTTGAACTCTTCA
[0126] Cma12_2: Located at position 9717178 on chromosome 12, its nucleotides are A or G. The sequence extending 50 bp before and after this position is GGAGGAGAGTACATTGGTGGTGGTGGTGGAGAGTTGGGCGGTGGCGGTGA[G / A]TTGTAATAGTAAACAGGTGGGGGTGGAGAGAATAATGGAGGAGGAGGTGG
[0127] Cma13_2: Located at position 852718 on chromosome 13, its nucleotides are T or G. The sequence extending 50 bp before and after this site is GATTCTTTTTTAGAGCCAACAGACATCTTGAACCTAGAAAATCACATAAG[T / G]TTAGTCTAGGTAACTGGTAAGCCACAATAATGGTAGTCAAATACTGTGAG
[0128] Cma14_2: Located at position 9917851 on chromosome 14, its nucleotides are A or G. The sequence extending 50 bp before and after this position is: AGTACGAAAGAGAGAGTTTAAAGGACAATATCTACTAGTGGTGGACTTAT[G / A]GACTTAGACGGTTACAATTCTACCACACTTCTTCCAATAGAGTCAAAGAA
[0129] Cma15_1: Located at position 1058773 on chromosome 15, its nucleotides are T or C. The sequence extending 50 bp before and after this position is TGTCCAGGTGCACTCCTAGGATGTGATTTCTCGATGGCAGACACGAACTT[T / C]TTGAGATGTCTTATGTACTGTGGGTATAGCTCTAAGTCACAGTGGTTCCC
[0130] Cma16_2: Located at position 980007 on chromosome 16, its nucleotides are A or G. The sequence extending 50 bp before and after this position is: AACCGCAAGAGTCATCTGGTTCAATAGGTGGCTTGGAGCCTGTGAACTTC[A / G]TCACTGTTGCCACCCCGCACCTTGGTTCAAGGGGTAATAAGCAGGTATCT
[0131] Cma17_1: Located at position 1126728 on chromosome 17, its nucleotides are A or G. The sequence extending 50 bp before and after this site is TCGAAGCATCTGCAATGGAGGCACAATCAGTTATAGAATCCCCACTCACC[G / A]CATTGATTCCATCCTTGGAGTTGAAAGAAAACCAAGCAGGAATTTCAATG
[0132] Cma18_1: Located at position 10017687 on chromosome 18, its nucleotide is A or G. The sequence extending 50 bp before and after this site is CCAAAATCAACGTAAAAGTAGTCTATTTGTAACAACCAAGGCCACCCTTA[G / A]CAGATATTGTCCTCTTTGGGCTTTCTCTTTTGGGTTTCCCCTCAAGGTTT
[0133] Cma19_4: Located at position 1336123 on chromosome 19, its nucleotides are T or C. The sequence extending 50 bp before and after this site is TGTTTGGCCTACCTAAAAAAACAAGTACACGAATACTAATAAAATAAACT[C / T]GAGAAATTCACACCATCCTTAGGGAAATTGGCAAATGCAACCAATGAGAG
[0134] Cma20_2: Located at position 1384468 on chromosome 20, its nucleotide is A or G. The sequence extending 50 bp before and after this position is GATATTAGAGTCATCCTAATCGAATCCTCAAGTGTCGAATAAAGAAGTTG[G / A]TAGCCTCAAAGGTGTAGTCAAAAGTAAATCAAGTGTCGAACAAATGGTGT.
[0135] 4.2 Typing results of 28 germplasm accessions
[0136] Table 2 shows the KASP genotyping results of 28 Indian pumpkin germplasms at 20 SNP loci.
[0137] Table 2. KASP genotyping results at 20 SNP loci in 28 Indian squash germplasm accessions.
[0138]
[0139] Analysis of these loci reveals genetic diversity among different materials. For example, six materials—CMA07, CMA08, CMA15, CMA19, CMA20, and CMA26—showed homozygous genotypes at all 20 loci; after excluding undetected data, four materials—CMA13, CMA14, CMA25, and CMA27—also exhibited homozygous genotypes. This genetic diversity is of significant value for pumpkin breeding and genetic improvement.
[0140] Example 5: Cluster Analysis and Fingerprint Map Construction
[0141] 5.1 Cluster Analysis
[0142] Based on the KASP-SNP genotyping results, a phylogenetic tree was constructed using MEGA software. The results showed that the 28 *Cucurbita indica* accessions formed two main subgroups, comprising 18 and 10 accessions respectively. For example, CMA07 and CMA27 clustered together, indicating high genotypic similarity at 20 KASP-SNP loci and close phylogenetic relationship.
[0143] 5.2 Fingerprint Map Construction and QR Code Generation
[0144] A SNP fingerprint map of pumpkin germplasm was constructed based on KASP genotyping results. The genotypic information of 20 loci in each material was converted into unique string codes according to a fixed order on chromosomes 1 to 20, and then further generated into corresponding QR codes. The genetic characteristics of the germplasm can be quickly obtained by scanning the codes. This fingerprint map not only reveals the genetic diversity of the tested materials but also provides a molecular basis for the selection of superior germplasm, enabling rapid identification and sharing of pumpkin germplasm genetic information and improving the utilization efficiency of germplasm resources.
[0145] Specifically, the genotyping results of 20 SNP loci are arranged in a fixed order on chromosomes 1 to 20 into a string of length 20, with each digit taken from {A, G, T, C, NN}, generating an SNP fingerprint code for the sample to be tested. This SNP fingerprint code is compared with the SNP fingerprint database of standard varieties / germplasm. When the matching degree reaches a set threshold, it is determined to be the same variety / germplasm or the seed is genuine and qualified; otherwise, it is determined to be a different variety / germplasm or the authenticity / purity is unqualified. The matching degree threshold is set as follows: for variety authenticity identification, a matching degree ≥ 95% is considered the same variety; for seed purity identification, a heterosis rate ≤ 5% is considered qualified.
[0146] Furthermore, the SNP fingerprint code can be converted into a QR code or barcode to realize the digital identification and management of germplasm resources.
[0147] Comparative Example 1: Comparison with SSR Markup Technology
[0148] To verify the technical advantages of the KASP-SNP markers of this invention, 20 KASP-SNP markers of this invention and 20 pairs of traditional SSR markers were used to detect the same 28 pumpkin germplasm accessions. The results were compared in terms of detection time, detection cost, throughput, accuracy, and resolution, and are as follows:
[0149] Single batch testing time: The KASP-SNP of this invention takes 4-6 hours, while the traditional SSR labeling takes 48-72 hours (2-3 days), with a relative advantage of shortening the time by more than 80%.
[0150] Single-sample detection cost: The cost of the KASP-SNP of this invention is about 15-20 yuan, while that of traditional SSR markers is about 30-40 yuan, representing a relative advantage of about 50% reduction;
[0151] Single batch throughput (96-well plate): The KASP-SNP of this invention requires 94 samples (2 controls), while the traditional SSR label requires 24-48 samples, representing a relative advantage of 2-4 times.
[0152] Result interpretation method: The KASP-SNP of this invention is automatically interpreted by the instrument, which is objective and accurate. The traditional SSR marking is manually interpreted, which is highly subjective. Its relative advantage is that it eliminates human error.
[0153] Genome identification accuracy: The KASP-SNP accuracy of this invention is ≥99.5%, while that of traditional SSR markers is approximately 90-95%, representing a relative advantage of 4.5-9.5 percentage points.
[0154] Experimental repeatability: The KASP-SNP of this invention has good performance and consistent results between different laboratories. Traditional SSR markers are generally less effective and are greatly affected by experimental conditions. The relative advantage of this invention is a significant improvement.
[0155] Distinguishing ability of 28 accessions: The KASP-SNP of this invention has a 100% distinguishing ability (28 / 28), while the traditional SSR marker has a distinguishing ability of about 75% (21 / 28), and the relative advantage is a significant improvement in resolution;
[0156] Automation level: The KASP-SNP of this invention is high, which can realize high-throughput automated detection. Traditional SSR marking is low and relies on manual operation. Its relative advantage is that it is easy to apply on a large scale.
[0157] Experimental results show that the KASP-SNP marker of the present invention is significantly superior to the traditional SSR marker in terms of detection efficiency, detection cost, throughput, accuracy, and resolution, and can better meet the detection needs of modern seed industry for high throughput, high efficiency, and low cost.
[0158] Comparative Example 2: Verification of universality with Chinese pumpkin SNP chips
[0159] To verify the applicability of Chinese pumpkin SNP markers in Indian pumpkin, 50 publicly available Chinese pumpkin SNP loci were selected and amplified in 28 Indian pumpkin germplasm accessions. The results are as follows:
[0160] (1) Low amplification success rate: Of the 50 SNP sites in Chinese pumpkin, only 22 sites could be amplified in Indian pumpkin, with an amplification success rate of 44%; the remaining 28 sites could not be amplified at all due to large differences in the primer binding region sequence.
[0161] (2) Low typing accuracy: Among the 22 successfully amplified sites, only 11 sites could obtain clear typing results, with a typing accuracy of 50% (accounting for 22% of the total number of sites); the remaining 11 sites had serious problems of non-specific amplification or unclear typing.
[0162] (3) Poor polymorphism: Among the 11 successfully genotyped loci, 4 loci were monomorphic (no polymorphism) in 28 Indian squash germplasm accessions, only 7 loci were polymorphic, and the PIC values were generally low (average about 0.25).
[0163] The above results clearly demonstrate that SNP markers from Chinese pumpkin cannot be directly applied to Indian pumpkin; specific SNP markers must be developed for the Indian pumpkin genome. This invention fills this technological gap by developing 20 KASP-SNP markers for Indian pumpkin.
[0164] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A combination of SNP markers for Indian pumpkin based on KASP technology, characterized in that, The marker combination contains 20 SNP sites, and its detection primers consist of the following 20 sets of KASP primers, each set used to detect one SNP site in Cucurbita maxima. Each primer set includes two competitive forward primers F1 and F2 and one reverse primer R. The 5′ end of F1 carries the tail sequence corresponding to the first fluorescent reporter group, and the 5′ end of F2 carries the tail sequence corresponding to the second fluorescent reporter group. The 3′ ends of F1 and F2 correspond to the two alleles of the SNP site, respectively, to achieve two-color fluorescent competitive allele-specific PCR typing. Furthermore, the 20 SNP loci are evenly distributed on the 20 chromosomes of the Indian squash, with one SNP locus corresponding to each chromosome. The polymorphism information content (PIC) value of each locus is ≥0.35, and the minimum allele frequency (MAF) is ≥0.
3. The nucleotide sequences of F1, F2, and R in each primer set are shown in the following SEQ ID NO: Cma1_2: F1 as SEQ ID NO:1, F2 as SEQ ID NO:2, R as SEQ ID NO:3; Cma2_2: F1 as SEQ ID NO:4, F2 as SEQ ID NO:5, R as SEQ ID NO:6; Cma3_2: F1 as SEQ ID NO:7, F2 as SEQ ID NO:8, R as SEQ ID NO:9; Cma4_2: F1 as SEQ ID NO:10, F2 as SEQ ID NO:11, R as SEQ ID NO:12; Cma5_1: F1 as in SEQ ID NO:13, F2 as in SEQ ID NO:14, R as in SEQ ID NO:15; Cma6_1: F1 as SEQ ID NO:16, F2 as SEQ ID NO:17, R as SEQ ID NO:18; Cma7_2: F1 as SEQ ID NO:19, F2 as SEQ ID NO:20, R as SEQ ID NO:21; Cma8_2: F1 as SEQ ID NO:22, F2 as SEQ ID NO:23, R as SEQ ID NO:24; Cma9_1: F1 as SEQ ID NO:25, F2 as SEQ ID NO:26, R as SEQ ID NO:27; Cma10_2: F1 as SEQ ID NO:28, F2 as SEQ ID NO:29, R as SEQ ID NO:30; Cma11_2: F1 as SEQ ID NO:31, F2 as SEQ ID NO:32, R as SEQ ID NO:33; Cma12_2: F1 as SEQ ID NO:34, F2 as SEQ ID NO:35, R as SEQ ID NO:36; Cma13_2: F1 as SEQ ID NO:37, F2 as SEQ ID NO:38, R as SEQ ID NO:39; Cma14_2: F1 as SEQ ID NO:40, F2 as SEQ ID NO:41, R as SEQ ID NO:42; Cma15_1: F1 as SEQ ID NO:43, F2 as SEQ ID NO:44, R as SEQ ID NO:45; Cma16_2: F1 as SEQ ID NO:46, F2 as SEQ ID NO:47, R as SEQ ID NO:48; Cma17_1: F1 as SEQ ID NO:49, F2 as SEQ ID NO:50, R as SEQ ID NO:51; Cma18_1: F1 as SEQ ID NO:52, F2 as SEQ ID NO:53, R as SEQ ID NO:54; Cma19_4: F1 as SEQ ID NO:55, F2 as SEQ ID NO:56, R as SEQ ID NO:57; Cma20_2: F1 as SEQ ID NO:58, F2 as SEQ ID NO:59, R as SEQ ID NO:60; or its complementary sequence.
2. The SNP marker combination for Indian pumpkin based on KASP technology according to claim 1, characterized in that, The 5′ tail sequence of F1 is GAAGGTGACCAAGTTCATGCT, and the 5′ tail sequence of F2 is GAAGGTCGGAGTCAACGGATT. The 3′ core region length of F1 and F2, excluding the tail, is 21nt to 27nt, and F1 and F2 differ by only one base at the SNP identification site.
3. The SNP marker combination for Indian pumpkin based on KASP technology according to claim 1, characterized in that, The 20 SNP loci are located on chromosomes 1 through 20 of the Indian squash strain, with the following specific physical locations: Cma1_2 is located at position 9975272 on chromosome 1, Cma2_2 is located at position 973733 on chromosome 2, Cma3_2 is located at position 8907238 on chromosome 3, Cma4_2 is located at position 980502 on chromosome 4, Cma5_1 is located at position 1000157 on chromosome 5, Cma6_1 is located at position 1023295 on chromosome 6, Cma7_2 is located at position 856994 on chromosome 7, Cma8_2 is located at position 925481 on chromosome 8, Cma9_1 is located at position 1030918 on chromosome 9, Cma10_2 is located at position 999475 on chromosome 10, and Cma1_2_3_4_4_5_6_6_7_8_9_9_9_9_9_9_9_1 ...9_1_9_9_9_9_9_1_9_9_9_9_9_1_9_9_9_9_9_1_9_9_9_9_9_9_1_9_9_9_9_9_9_1_9_9_9_9_9_9_ Cma12_2 is located at position 12503969 on chromosome 11, Cma12_2 is located at position 9717178 on chromosome 12, Cma13_2 is located at position 852718 on chromosome 13, Cma14_2 is located at position 9917851 on chromosome 14, Cma15_1 is located at position 1058773 on chromosome 15, Cma16_2 is located at position 980007 on chromosome 16, Cma17_1 is located at position 1126728 on chromosome 17, Cma18_1 is located at position 10017687 on chromosome 18, Cma19_4 is located at position 1336123 on chromosome 19, and Cma20_2 is located at position 1384468 on chromosome 20. These marker combinations are used to construct the SNP fingerprint of Pumpkin germplasm.
4. A method for identifying the authenticity / purity of Indian pumpkin germplasm, varieties, or seeds, characterized in that, Includes the following steps: (1) Extract genomic DNA from the Indian pumpkin sample to be tested; (2) Using the SNP marker combination of Indian pumpkin based on KASP technology as described in any one of claims 1 to 3, competitive allele-specific PCR amplification is performed on the genomic DNA; The reaction system is 10 μL, containing: 1 μL DNA template, 0.14 μL primer mixture, 5 μL 2×KASP Mix, and 3.86 μL distilled water. The volume ratio of Forward1, Forward2, and reverse primers in the primer mixture is 12:12:
30. The amplification program was as follows: 95℃ pre-denaturation for 10 min; (95℃ denaturation for 15 s, 61℃→55℃ touchdown annealing extension for 1 min, decreasing by 0.6℃ per cycle) × 10 cycles; (95℃ denaturation for 15 s, 55℃ annealing for 40 s) × 35 cycles; (3) Read the FAM / HEX fluorescence signal, determine the type of each SNP site, and obtain the typing results of the sample to be tested at the 20 SNP sites; The classification rules are as follows: if the signal falls into the FAM cluster, it is recorded as homozygous isotype 1; if it falls into the HEX cluster, it is recorded as homozygous isotype 2; if it falls in between, it is recorded as heterozygous; if it cannot form a cluster, it is recorded as NN. The reaction well is then re-read after a cycle. (4) Arrange the genotyping results of the 20 SNP loci in a fixed order of chromosomes 1 to 20 into a string of length 20, with each digit taken from {A,G,T,C,NN}, to generate the SNP fingerprint code of the sample to be tested; compare the SNP fingerprint code with the SNP fingerprint database of standard varieties / germplasm, and when the matching degree reaches the set threshold, determine that it is the same variety / germplasm or the seed is genuine and qualified; Otherwise, it will be judged as a different variety / germination or as unqualified in authenticity / purity.
5. The method for identifying the authenticity / purity of Indian pumpkin germplasm, variety, or seeds according to claim 4, characterized in that, Step (4) also includes converting the SNP fingerprint code into a QR code or barcode to realize the digital identification and management of germplasm resources.
6. The method for identifying the authenticity / purity of Indian pumpkin germplasm, variety, or seeds according to claim 4, characterized in that, The matching threshold is set as follows: when identifying the authenticity of a variety, a matching degree of ≥95% indicates that the variety is the same; when identifying the purity of a seed, a hybrid rate of ≤5% indicates that the seed is qualified.
7. A reagent kit, characterized in that, It comprises a combination of SNP markers for Indian pumpkin based on KASP technology as described in any one of claims 1 to 3, and a KASP PCR reaction mixture.
8. The use of the KASP-based SNP marker combination of Indian pumpkin as described in any one of claims 1 to 3 in the preparation of a kit for identifying the authenticity of Indian pumpkin varieties, identifying seed purity, constructing SNP fingerprints for germplasm resources, or analyzing genetic diversity.