A SNP molecular marker combination for identifying aloe vera germplasm resource
Patent Information
- Application Number
- CN202611045280.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-14
- Publication Date
- 2026-09-25
AI Technical Summary
朱顶红属的DUS测试指南农业行业标准(NY/T 3508-2019)中虽然对繁殖材料的要求、测试方法、测试的时间和地点、观测方法以及特异性、一致性、稳定性判定的标准等都有明确的规定,然而朱顶红的DUS测试主要依赖于大田栽培,测试周期较长,且许多目测性状在实际测试中容易受到气候和环境等影响,也具有人为的偶然因素,客观上增加了测试的难度,容易造成测试误差和误判
1)通过采用不同品种82个标记信息的区别统计,实现了辅助朱顶红品种区分和辅助DUS测试的有益效果。
Smart Images

Figure CN122811403A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of molecular biology, and more specifically, to a combination of SNP molecular markers for identifying amaryllis germplasm resources. Background Technology
[0002] Amaryllis (Hippeastrum Herb.) is a perennial herbaceous plant belonging to the genus Hippeastrum in the family Amaryllidaceae. It is native to tropical and subtropical regions, mainly including northern Argentina to Mexico and the Caribbean Sea. It has a spherical underground bulb and large, full flowers with red corollas or white or purple stripes. Due to its bright flower colors, it is suitable for planting in the ground and can form a community landscape, adding to the scenery of the garden.
[0003] Since the beginning of the 21st century, with the introduction of hybrid varieties of amaryllis, my country has achieved rapid development in the breeding and cultivation techniques of amaryllis. Currently, domestic and international research on amaryllis mainly focuses on cultivation techniques, breeding, propagation techniques, and pest and disease control. The breeding of new amaryllis varieties is an important part of the construction of my country's germplasm resource bank, and the protection of varieties after breeding is crucial for the development of the amaryllis genus.
[0004] The Distinctness, Uniformity, and Stability (DUS) test is an analytical test conducted on the distinctiveness, uniformity, and stability of a proposed new plant variety. It is the core basis for determining whether an applied variety is a "new variety." The DUS test for Amaryllis covers quantitative characteristics such as leaf length and width, flower diameter, and the length and width of the outer perianth segments, as well as qualitative characteristics such as anthocyanin coloration, flower color, and flower posture. It also includes four optional test traits: leaf height and dominant color of the flower throat. While the agricultural industry standard (NY / T 3508-2019) for the DUS test of the Amaryllis genus clearly stipulates the requirements for propagation materials, test methods, test time and location, observation methods, and standards for determining distinctiveness, uniformity, and stability, the DUS test for Amaryllis mainly relies on field cultivation. The testing cycle is long, and many visually apparent traits are easily affected by climate and environment in actual testing, and there are also accidental human factors, objectively increasing the difficulty of the test and making it prone to testing errors and misjudgments. Summary of the Invention
[0005] To address the aforementioned technical problems, the present invention aims to provide an SNP molecular marker combination for identifying amaryllis germplasm resources, thereby using molecular markers to assist in the variety identification of amaryllis and solving the shortcomings of current DUS testing for new amaryllis varieties.
[0006] The objective of this invention is achieved through the following technical solution: In a first aspect, the present invention provides an SNP molecular marker combination for identifying amaryllis germplasm resources, the SNP molecular marker combination comprising 82 SNP molecular markers, the flanking sequences and variant types of the 82 SNP molecular markers being as follows:
[0007] In the table above, the flanking sequences are the flanking nucleotide sequences of the SNP molecular markers; "[ ]" marks the site positions of the SNP molecular markers, and the bases in the marks represent the polymorphisms at that site.
[0008] Secondly, the present invention provides a primer pair combination for amplifying the SNP molecular marker combination as described above, comprising:
[0009] Thirdly, the present invention provides a kit for detecting the above-described SNP molecular marker combinations, the kit comprising the primer pair combinations described above.
[0010] Fourthly, the present invention provides a chip for detecting SNP molecular marker combinations as described above, the chip comprising primer pair combinations as described above.
[0011] Fifthly, the present invention provides the application of the SNP molecular marker combination, primer pair combination, kit or chip as described above in constructing an amaryllis DNA fingerprint.
[0012] Sixthly, the present invention provides a method for constructing an amaryllis DNA fingerprint, comprising: S1. Extract genomic DNA from amaryllis and perform PCR amplification using the 82 primer pairs mentioned above. Mix the products amplified by different primer pair combinations and perform high-throughput sequencing on the mixed products. S2. Identify the genotypes of the sites where the SNP molecular marker combinations described above are located in the amaryllis genome; construct the DNA fingerprint map of amaryllis using the identified SNP genotype combinations of different amaryllis varieties.
[0013] In a seventh aspect, the present invention provides an amaryllis DNA fingerprint, which is constructed using the construction method described above.
[0014] Eighthly, the present invention provides an application of the SNP molecular marker combination, primer pair combination, kit, chip or fingerprint spectrum as described above in the identification of amaryllis germplasm resources.
[0015] As some specific embodiments of the present invention, the amaryllis germplasm is selected from 163 amaryllis varieties, specifically from Sunong Hongdie, Apple, Meng Tong, Xiu Xiu, Minerva, Xiyan, Peacock flower, Seeking the truth, Jing Qi, Blossom Peacock, Xia Bing, Spring Search, Happy Lamb, Choir, Yi Yi, Whirlpool, Double Dragon, Yujie, Spicy, Xin Tong, LeZhen, Red Daruma, Picotee, Merry Christmas, Exposure, Star language, Meng Yuan, Xiao Ge, Ru Yi, Yun Xi, Spring Apricot, Xia Tong, Yunhan, An Ran, and Chan Xi. Xī, Lia, Elegantand graceful, Feifei, Rookie, Yu'e, Jiaxin, Shan Shan, Wanqing, Little Apple, Afu lei, Spring glow, Wuhuatianbao, Ice Summer, Masonry red, Xin Yu, Carimero, Pink Surprise, Bloodline, Wedding dance music, Sunong RedSilk, Red lips, Double Happiness, Graceful dancing, Goodluck, Xiang Mei, Red Peafowl, Beautiful Lady, Sea star, Green dress, Smiley, Magic, Pink delicate, Sunong Hongwu, Child, Childhood, Fish Sister, Fairy Taletale, Marilyn, RuiLi, Daphne, Qixia, Pipi, Flame Peacock, Jiari, Springsnow, Hongguang, Sunset glow, Red Sun, Jingxiu, Ruoxi, Provocative, Yiting, Mengqi, Xiuyan, Spring Dependence, RedDawn, Feiyi, Xia Lan, Meilin, Xiutong, Yitong, Snow Bath, Xia Xue, Ruoqing, Ruoxi, Ziyan, Xingxiu, Witch, Sunong Red Wind, Radiant glow, Colored clouds, Red glow, Xiao Ai, Elf, Little Dream, Grow Chu yun, Emerald Green, Picasso, Berry Core, Ruby, Ambiance, Sweet Nivea, Jingya, Pink Dress, Volcanic Eruption, Yiting, Lan Yan, Sweet Fairy, Lux, Jia Zhi, Dance Queen, Lucky Angel, Black Swan, Phantasm, Cupid, Tika Gloria, Tika Fantasy, Tika Kam, Tika Star, Tika Trisco, Tika Lutz, Tika Yum Pairot, Tika McGeeck, Tika Quin, Tika Sarris, Tika Sa Sha, Tika Auror, Tika Biyouti, Tika Ram, Tika Farrey, Sir Tika, Tika Flamingo, Tika Pu Sais, Tika RangRoot, Yihan, Beautiful, Anhui, Yirui, Chen Yan, Yashi, Red Sleeve, ShiYan, The Paper, Dawn glow, Yunxuan, Tika c Ruo, Sunshining, Princess Xiang Lei.
[0016] As some specific embodiments of the present invention, the method for identifying amaryllis germplasm resources includes the following steps: A1. Extract genomic DNA from the germplasm of *Amaryllis* to be identified (DNA concentration ≥ 50 ng / μl; OD260 / 280 between 1.7 and 2.0). A2. SNP primer amplification and purification: The genomic DNA of the amaryllis variety extracted using the 82 SNP primer pairs as described above was amplified by PCR. The amplification products were mixed in equal volume proportions and then purified. A3. The purified mixed product was subjected to high-throughput sequencing. The sequencing results were compared with the flanking sequences of each SNP site to obtain the actual site information of the amaryllis variety. The results were then compared with the fingerprint map to identify the amaryllis germplasm.
[0017] Compared with the prior art, the present invention has the following beneficial effects: 1) By using statistical analysis of 82 different markers from different varieties, the study achieved the beneficial effects of assisting in the differentiation of amaryllis varieties and assisting in DUS testing.
[0018] 2) By using mixed PCR product purification and high-throughput sequencing to compare flanking sequences, the efficiency of variety identification was improved and the cost of molecular detection was reduced.
[0019] 3) By constructing fingerprint maps and analyzing genetic diversity through SNP marker information, the genetic background of amaryllis was successfully analyzed. Some of the 82 SNP markers were found to have potential association analysis with agronomic traits, providing a useful reference for the future development of amaryllis germplasm resources. Attached Figure Description
[0020] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 A distribution map of SNP site mutation types; Figure 2The graphs show the genetic diversity analysis of 82 SNP markers, where A is the observed heterozygosity (Ho) result, B is the Shannon diversity index (I) result, C is the polymorphism information content (PIC) result, and D is the effective allele count (Ae) result. Figure 3 A heatmap of genetic distances among 163 amaryllis varieties based on SNP markers; Figure 4 A phylogenetic tree of 163 amaryllis varieties based on 82 SNP markers; Figure 5 These are the principal component analysis (PCA) plots, where A is the ungrouped principal component plot and B is the principal component plot grouped by phylogenetic tree. Figure 6 Line graphs showing DeltaK and LnP(D) values for different K values, where A is the K-DeltaK line graph and B is the K-LnP(D) line graph; Figure 7 This is a population structure diagram for K=2-8; Figure 8 Fingerprint profiles of 163 amaryllis varieties based on 82 core SNP markers; Figure 9 SNP fingerprint QR codes for 6 amaryllis varieties. Detailed Implementation
[0021] The present invention will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present invention, but do not limit the invention in any way. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention. These all fall within the scope of protection of the present invention.
[0022] Example 1: SNP marker screening and primer design 1. Test materials This invention uses the 163 amaryllis varieties listed in Table 1 below as a basis for screening SNP markers: Table 1. List of Amaryllis Varieties
[0023] In the table above, FT represents the petal type, S represents single petals, and D represents double petals.
[0024] Based on the results of the phenotypic genetic diversity analysis, 50 amaryllis varieties were selected from the aforementioned 163 varieties according to criteria such as smaller genetic distance, larger phenotypic differences, and wider coverage of phenotypic grading codes. These 50 varieties were then divided into 5 groups using principal component analysis with 8 major traits as variables, as shown in Table 2 below. Table 250 Groups of Amaryllis Varieties
[0025] 2. Extraction of amaryllis genomic DNA For the 50 amaryllis varieties shown in Table 2 above, 10 individuals from each amaryllis material were randomly selected for leaf extraction. The leaves were then placed in self-sealing bags and transported to the laboratory using liquid nitrogen and stored at -80°C for subsequent genomic DNA extraction.
[0026] The genomic DNA extraction procedure is as follows: Take 200 mg–300 mg of leaf or young flower bud and place it in a 2.0 mL centrifuge tube. Grind thoroughly with liquid nitrogen and transfer the ground mixture to another 2.0 mL centrifuge tube. Add 500–700 μL of 2×CTAB extraction buffer preheated to 65 °C to each tube, mix thoroughly, and incubate at 65 °C for 45 min, gently inverting the tube several times during incubation. After the water bath, add an equal volume of chloroform and isoamyl alcohol mixture (24:1 volume ratio) to each tube, vortex vigorously to mix, let stand for 10 min, and centrifuge at 12,000 r / min for 10 min. After centrifugation, transfer the supernatant to a new 2.0 mL centrifuge tube, add 1.0 mL of pre-chilled anhydrous ethanol, gently invert to mix, incubate at -20 °C for 40 min, and centrifuge at 12,000 r / min for 10 min. Discard the supernatant, wash twice with 70% ethanol solution, discard the ethanol solution, and place in a sterile operating table to air dry. After the ethanol has fully evaporated, add 200 µL of double-distilled water or TE buffer to dissolve completely. Detect the DNA concentration and quality (the ratio of OD260 to OD280 of the DNA solution should be between 1.7 and 2.0). Store at -20℃ for later use.
[0027] 3. Simplified genome sequencing The extracted amaryllis genomic DNA sequencing material was sent to Shanghai Paiseno Biotechnology Co., Ltd., and simplified genome sequencing was performed according to the following steps: (1) First, the extracted DNA is digested with restriction endonucleases. Based on the selected endonuclease, the most suitable reaction conditions are chosen, and the digestion temperature and time are set, generally 37℃ for 5h. After digestion, 5μl of each sample is taken for gel testing. The DNA digestion should be complete, that is, there should be no main band after digestion, and there should be a clearly diffuse band between 250-1000bp.
[0028] (2) Use magnetic bead centrifugation (VAHTS™ DNA Clean Beads) to remove fragments that are too large or too small after enzyme digestion.
[0029] (3) Connect the P1 adapter and the P2 adapter. The P1 adapter contains the amplification primer, sample barcode, and enzyme 1 restriction site; the P2 adapter contains the amplification primer and enzyme 2 restriction site.
[0030] (4) The DNA fragments with ligated adapters were amplified by PCR, the sequencing library template was enriched, and sequencing adapters were ligated. Then, the PCR products were purified using magnetic beads (VAHTS™ DNA Clean Beads); (5) The library was finally selected and purified by 2% agarose gel electrophoresis. The inserted fragments in the library are generally 220~450bp.
[0031] (6) Before sequencing, the library needs to be quality checked on the Agilent Bioanalyzer using the Agilent High Sensitivity DNA Kit. A qualified library has only a single peak and no adapter dimers.
[0032] (7) After that, the library was quantified using the Quant-iT PicoGreen dsDNA Assay Kit on the Promega QuantiFluor fluorescence quantitative system. The concentration of the qualified library should be above 2nM.
[0033] (8) After gradient dilution of each qualified sequencing library (index sequence is not reproducible), mix them according to the required sequencing volume in the appropriate proportion, and denature them into single strands with NaOH for sequencing. (9) Perform 2×150bp paired-end sequencing using the NovaSeq sequencer.
[0034] This simplified genome sequencing yielded 2,324,855,618 raw reads, ranging from 31,595,920 to 142,383,514 reads across different samples. The average GC content was 37.36%, and under a normal distribution, the GC content ranged from 34.97% to 40.84%. The average Q20 of the raw data was 98.88%, and the average Q30 was 96.70%, indicating high accuracy in this simplified genome sequencing, ensuring a low base error rate, and meeting the requirements for subsequent analysis.
[0035] 4. Data filtering to generate high-quality sequences To ensure the quality of subsequent information analysis, the raw data is further filtered to generate high-quality sequences. FastP (v0.20.0) is used with a sliding window method to filter the raw data, and the main standards include the following: (1) Remove contaminants from the connector, specifically the 3' end. (2) Quality filtering: The sliding window method is used for quality filtering, and the window size is set to 5 bp. The window is slid from the 3' end to the 5' end, and the average Q value of the bases in the window is calculated. If the Q value is <20, the bases in the window are deleted; if the Q value is ≥20, the sliding is stopped.
[0036] (3) Length filtering: If the length of any one of the reads in the double-ended terminal is ≤50 bp, then the double-ended read is removed.
[0037] (4) Blurred base N filtering: If the number of N bases in the double-ended reads is ≥5, then the double-ended reads are removed.
[0038] After quality control filtering of the raw data, high-quality sequences were generated, resulting in 2,268,422,316 high-quality reads. The number of high-quality reads ranged from 30,904,026 to 138,220,870 across different samples. The average percentage of high-quality reads was 97.62%, ranging from 95.75% to 98.36% across different samples.
[0039] 5. SNP site generation (1) SNP sites were obtained using Stacks software: First, cstacks are used to merge the loci of all samples to obtain the catalog consensus sequence of each loci. The main parameter settings are as follows: -n: a maximum of 4 mismatches are allowed between different sample loci; Then, sstacks were used to align the loci sequence of each sample with the catalog consensus sequence.
[0040] Finally, populations were used to filter and output the SNP loci of all samples. The main parameters were set as follows: -p: a locus must appear in at least half of the population; -m: minimum sample depth, set to 3; -r: minimum proportion of the locus in all individuals of a single population, set to 1; --min_maf: minimum minor allele frequency, set to 0.05.
[0041] Finally, a group of exploitable SNPs was obtained.
[0042] (2) Principal component analysis was performed using GCTA (v4.6.0.0) software with SNP data (SNPs with MAF less than 0.05 were removed).
[0043] (3) The phylogenetic tree was constructed using the Maximum likelihood algorithm in fastTree (v2.1.11). After the tree was constructed, the reliability of the phylogenetic tree branches was verified (bootstrap, 1000 replications).
[0044] Through the above steps, a total of 247,899 SNP sites were obtained from 50 amaryllis samples. The distribution of mutation types at all SNP sites is as follows: Figure 1 As shown. By Figure 1 It is evident that T:A>C:G and C:G>T:A are the main SNP mutant types, accounting for 68.54%, while the proportion of other mutant types is significantly lower than these two types. Furthermore, T:A>C:G and C:G>T:A both belong to the Ts mutant type, while the other types belong to the Tv type. In this study, the Ts / Tv ratio was approximately 2.05, indicating high sequencing quality and a high degree of similarity between the sequencing results and the actual results.
[0045] 6. Core SNP tagging and filtering Based on existing SNP loci, core SNP markers were screened according to the following parameters: (1) allele frequency greater than 0.2; (2) marker deletion site less than 10%; (3) marker is biallelic; (4) marker PIC value greater than 0.05; (5) no other site mutations within 100 bp before and after the marker; (6) the marker is not located at the edge of the corresponding fragment within 30 bp. A total of 186 SNP markers were obtained.
[0046] Table 3 below shows the Ho, Ae, PIC, and I index information for 186 SNP tags: Table 3. Ho, Ae, PIC, and I index information for 3186 SNP loci
[0047] As shown in Table 3 above, the average observed heterozygosity of the 186 markers was 0.49, the PIC value ranged from 0.27 to 0.38, the average number of effective alleles was 1.75, and the Shannon diversity index was greater than 0.5. It can be seen that the polymorphism and variety differentiation ability of the 186 selected markers are quite remarkable.
[0048] To ensure the feasibility of subsequent primer design, the GC content of the flanking sequences of 186 markers was detected. It was found that 73 sites had abnormal GC content (greater than 70% and less than 30%), hairpin structures (such as 5'-GATC...CTAG-3'), etc. After removing these sites, 113 SNP markers remained.
[0049] 7. SNP primer design Based on the location of the selected SNP markers in the sequencing fragment, the corresponding fragment sequence for each site was extracted as a reference sequence for primer design. SNP primers were designed using NCBI's Primer-BLAST online tool, with the upstream and downstream primer reference sequences set 30 bp away from the SNP marker. Primers with similar product length and flanking sequence length were selected as SNP primers for that site.
[0050] The parameters for SNP primer design were as follows: TM value between 55 and 62°C, with the difference between the upstream and downstream primer TM values not exceeding 2°C; GC content between 30% and 62%; primer length between 18 and 25 bp; and PCR product size between 70 and 150 bp. The designed primers were synthesized by Beijing Qingke Biotechnology Co., Ltd.
[0051] Primers were designed for 113 SNP markers using NCBI's Primer-BLAST online tool. Primer pairs were successfully designed for 83 SNP markers. Since the upstream and downstream primers need to be on both sides of the SNP site and at a certain distance, some markers could not be designed with primer pairs that could achieve the target TM value and GC content without affecting the SNP site.
[0052] Table 483 SNP markers and primer information
[0053] Note: Example of site location: 80317 (this number is the sequence number of the digestion fragment)_103 (this number is the sequence position of the target marker in the corresponding digestion fragment). This number indicates that the SNP marker is at the 103rd base of fragment 80317 in this sequencing result. In the table above, REF indicates SNP alleles consistent with the reference genome; ALT indicates SNP alleles inconsistent with the reference genome.
[0054] The flanking sequence information of the above 83 SNP markers is shown in Table 5 below: Table 583 SNP marker flanking sequence information
[0055] In the flanking sequences in the table above, the parentheses indicate the position of the SNP site within that flanking sequence. Taking SNP01 as an example, [C / A] indicates that the nucleotide at this SNP site can be either cytosine or adenine, while the nucleotide at this site in the reference genome is cytosine. The sequence listing only includes flanking sequences consistent with the reference genome. For example, "SEQ ID NO.167" represents the flanking sequence of SNP01. When the allele at this SNP site is replaced by A instead of C, only the nucleotide at that site changes, while the rest of the flanking sequences remain unchanged.
[0056] 8. SNP primer verification PCR amplification experiments were performed using the designed SNP primers and all the amaryllis materials to be validated. The specific steps are as follows: (1) PCR amplification The genomic DNA of each amaryllis variety was amplified by PCR using 83 pairs of primers on a separate 96-well plate.
[0057] Prepare the PCR amplification reaction system as shown in Table 6 below: Table 6. PCR Amplification Reaction System
[0058] In Table 6 above, the genomic DNA of Amaryllis was extracted from 163 Amaryllis varieties using the Plant Genomic DNA Extraction Kit (catalog number D101-01) from Beijing Kangrun Chengye Biotechnology Co., Ltd., yielding DNA stock solutions with concentrations of 200–2000 ng / μL. These stock solutions were then diluted with enzyme-free double-distilled water to a concentration of 50–100 ng / μL and used as templates for PCR amplification.
[0059] The 2×PCR MIX was a 2×SanTaqPCR Mix premix purchased from Sangon Biotech Co., Ltd., with the product number CAT.NO.:B532061. Its main components include DNA polymerase, magnesium chloride, potassium chloride, dNTPs, and nuclease-free water.
[0060] The PCR reaction procedure is as follows: Pre-denaturation at 94℃ for 10 min; The process involved denaturation at 94℃ for 30 seconds, annealing at 58℃ for 30 seconds, and extension at 72℃ for 30 seconds, for a total of 30 cycles. Final extension at 72℃ for 10 min, then store at 4℃.
[0061] (2) Purification of PCR products PCR products were purified using a magnetic bead purification method. The specific steps are as follows: Take 25 μL of the PCR amplified product from a 96-well plate and add 50 μL of magnetic bead suspension. Gently pipette 5-8 times to mix thoroughly. Let stand at room temperature for 5 minutes to allow the magnetic beads to fully bind to the DNA in the PCR product. Then place the 96-well plate on a magnetic rack and let it stand for 2 minutes until the magnetic beads are completely adsorbed to one side of the well wall. Carefully aspirate the supernatant. Next, add 150 μL of freshly prepared 70% ethanol to each well, let it stand for 30 seconds, and then aspirate the ethanol. Repeat this washing step once to ensure the removal of residual salt ions and PCR reaction residues. After washing, keep the 96-well plate on the magnetic rack and let it stand at room temperature for 5-10 minutes with the lid open until the magnetic beads are completely dry to avoid ethanol residue affecting subsequent experiments. Finally, add 30 μL of nuclease-free water to each well, gently pipette the magnetic beads to suspend them completely, let stand at room temperature for 2 minutes, then place the 96-well plate back on the magnetic rack. After the magnetic beads are completely adsorbed, transfer the supernatant (i.e., the purified PCR product) to a new 96-well plate for subsequent high-throughput sequencing experiments. Use this method to purify the PCR amplification products of all amaryllis materials and each primer pair. The purification rate of the PCR amplification products should be ≥70%. The purified products are stored at -20°C for subsequent labeling detection.
[0062] (3) Agarose gel electrophoresis detection After purification of the PCR products, agarose gel electrophoresis was performed. The results showed that among the 83 primer pairs, primer SNP53 had a darker band, while all other primers had clear and clean bands, and the band sizes were all within the range of the target product size.
[0063] (4) High-throughput sequencing The purified products amplified by 83 primer pairs for each variety were mixed in equal volumes, resulting in 163 tubes of mixed products. These tubes were sent to Sangon Biotech (Shanghai) Co., Ltd. for SNP site detection based on high-throughput sequencing. The sequencing results were compared with the flanking sequences of the SNP marker sites in Table 5. The results showed that the amplification product of primer SNP53 did not detect the target SNP site in any of the 163 Amaryllis varieties, which was consistent with the agarose gel electrophoresis results. The amplification products of the remaining 82 primer pairs all detected the target SNP site. A very small number of sites showed missing results in some varieties, but all were within the normal range.
[0064] 9. Differential locus analysis Among all 163 amaryllis varieties, pairwise differential loci analysis revealed an average of 30 differential loci, a minimum of 2, and a maximum of 56. The varieties with the highest number of differential loci were ZDH111 and ZDH003, whose genetic distance of 0.178, while not the maximum, was still higher than the genetic distance between most varieties. Varieties with the minimum number of differential loci included ZDH060 Xiangmei and ZDH104 Nüwu, ZDH070 Tongnian and ZDH123 Tonghua, ZDH124 Tianmi Xiannv and ZDH118 Tianmi Nifu, ZDH015 Yiyi and ZDH042 Shanshan, ZDH007 Kongquehua and ZDH010 Hua Kongque, and ZDH154 Anhui and ZDH102 Ziyan, with a minimum of 2 differential loci. These varieties exhibit phenotypic specificity and can be distinguished as different varieties.
[0065] Example 2: Genetic diversity analysis and fingerprinting based on SNP markers 1. Test materials This experiment used 163 amaryllis varieties listed in Table 1 as test materials, all of which were tested by the Plant Variety Testing Center (Shanghai) of the Ministry of Agriculture and Rural Affairs.
[0066] 2. Genetic diversity analysis Genetic diversity statistics were performed on the detection results of 163 amaryllis varieties using 82 SNP markers (excluding SNP053) screened in Example 1. PolyGene V1.7 software was used to calculate the AF, PIC, heterozygosity, etc. of the SNP markers.
[0067] The observed heterozygosity (Ho) results are as follows Figure 2 As shown in Figure A: 5 out of 82 markers had observed heterozygosity below 0.3, accounting for 6%, 22 markers were between 0.3 and 0.5, and the markers with heterozygosity greater than 0.5 were the most numerous, accounting for 67%.
[0068] The results of the Shannon Diversity Index (I) are as follows: Figure 2 As shown in Figure B, the average Shannon diversity index of the 82 markers is 0.6, indicating a high overall genetic diversity. Among them, 77 markers are above 0.4, indicating that most alleles are nearly evenly distributed and the population has a high degree of genetic variation.
[0069] Polymorphic information content (PIC) is an important indicator for measuring the richness of genetic information provided by SNP markers. The calculation results are as follows: Figure 2 As shown in Figure C, the average PIC of the 82 markers in this invention is 0.324, with 74% of the markers having a PIC greater than 0.3 and only 4 markers having a PIC less than 0.2, accounting for 5%. This result indicates that the vast majority of markers can provide relatively rich genetic information and can provide a reasonable basis for subsequent fingerprinting construction.
[0070] Effective allele count (Ae) is a key indicator for measuring the genetic contribution of alleles. The analysis results are as follows: Figure 2 As shown in Figure D, the number of effective alleles among the 82 markers is mostly above 1.4, accounting for 91%, of which 58% are above 1.8. This indicates that the allele distribution among the 82 markers is relatively uniform and the genetic diversity is good.
[0071] 3. Group Structure Analysis (1) Calculate the genetic distance of varieties Genetic distances among 163 amaryllis varieties were calculated using PolyGene. The genetic distances of these varieties ranged from 0.003 to 0.218. A distance heatmap was then plotted based on these genetic distances. Figure 3As shown, the 163 amaryllis varieties are arranged in the same order from top to bottom and from left to right on both the horizontal and vertical axes. The order is as follows: ZDH111, ZDH136, ZDH028, ZDH161, ZDH075, ZDH157, ZDH133, ZDH130, ZDH143, ZDH141, ZDH007, ZDH135, ZDH142, ZDH145, ZDH052, ZDH013, ZDH043, ZDH004, ZDH038, ZDH039, ZDH060, ZDH104, ZDH025, ZDH044, ZDH099, ZDH110, ZDH109, ZDH121, ZDH085, ZDH108 , ZDH128, ZDH159, ZDH027, ZDH103, ZDH008, ZDH035, ZDH107, ZDH156, ZDH066, ZDH024, ZDH076, ZDH017, ZDH150, ZDH054, ZDH046, ZDH049, ZDH036, ZDH08 6. ZDH032, ZDH050, ZDH071, ZDH061, ZDH114, ZDH047, ZDH056, ZDH089, ZDH092, ZDH119, ZDH018, ZDH062, ZDH003, ZDH106, ZDH014, ZDH087, ZDH009, ZDH0 26. ZDH078, ZDH132, ZDH146, ZDH033, ZDH162, ZDH057, ZDH077, ZDH097, ZDH163, ZDH010, ZDH069, ZDH151, ZDH112, ZDH152, ZDH020, ZDH048, ZDH100, ZDH 031, ZDH093, ZDH153, ZDH030, ZDH160, ZDH015, ZDH042, ZDH084, ZDH058, ZDH068, ZDH079, ZDH016, ZDH022, ZDH045, ZDH059, ZDH155, ZDH064, ZDH140, ZD H117, ZDH126, ZDH041, ZDH005, ZDH083, ZDH164, ZDH081, ZDH021, ZDH122, ZDH011, ZDH080, ZDH098, ZDH055, ZDH139, ZDH125, ZDH127, ZDH037, ZDH006, Z DH096, ZDH063, ZDH116, ZDH073, ZDH134, ZDH053, ZDH113, ZDH051, ZDH001, ZDH065, ZDH105, ZDH102, ZDH154, ZDH158, ZDH129, ZDH091, ZDH095, ZDH002,ZDH019, ZDH094, ZDH118, ZDH124, ZDH067, ZDH120, ZDH082, ZDH029, ZDH088, ZDH034, ZDH144, ZDH090, ZDH074, ZDH149, ZDH137, ZDH138, ZDH147, ZDH148, ZDH131, ZDH012, ZDH115, ZDH070, ZDH123, ZDH072, ZDH023, ZDH040. The color intensity of each cell represents the genetic distance between the corresponding two varieties—the smaller the value, the closer the genetic relationship, and the color tends to be brown (light brown to dark brown); the larger the value, the more distant the genetic relationship, and the color tends to be blue (light blue to dark blue). Simultaneously, the heatmap's row- and column-side data were supplemented with a phylogenetic clustering tree. By analyzing the clustering patterns of varieties along the branches, multiple groups could be clearly identified. The average genetic distance between varieties within a group was significantly smaller than the average genetic distance between groups. The results showed that among the 163 amaryllis varieties, 44 pairs had a genetic distance of less than 0.05, and 6 pairs had relatively close genetic distances, as shown in Table 7 below. Table 76 Genetic Distance Table for Variety Combinations
[0072] The above results indicate that these six pairs of varieties are genetically close, consistent with the analysis of the number of differential loci.
[0073] (2) Constructing a phylogenetic tree Genetic distances between varieties were calculated using the Neighbor Joining (NJ) algorithm, and a clustering graph was constructed using MEGA 11.0 software. The phylogenetic tree was then beautified using Chiplot (v2.6.1) software.
[0074] The results are as follows Figure 4 As shown: The phylogenetic tree divides 163 amaryllis varieties into six categories, with the variety ZDH111, which has the greatest genetic distance, being classified into a separate category. This aligns with the genetic distance heatmap for this variety. Figure 3The results were consistent with those observed in the previous data. The six categories included Cluster 1 (pink), encompassing 25 varieties; Cluster 2 (orange), encompassing 55 varieties; Cluster 3 (light purple), encompassing 20 varieties; Cluster 4 (yellow), encompassing 29 varieties; Cluster 5 (dark purple), encompassing 33 varieties; and Cluster 6 (green), encompassing 1 variety. Differentiation among the six groups based on key phenotypic traits of *Amaryllis* revealed significant differences in six traits: inflorescence number; corolla shape; flower color distribution; flower diameter; style color; and anthocyanin coloration at the base of the peduncle. For example, Clusters 1 and 4 showed a small number of anthocyanin-colored peduncle traits, while almost all other groups showed anthocyanin coloration at the base of the peduncle. Cluster 2 primarily exhibited round corolla shapes, Clusters 1, 4, and 5 primarily exhibited triangular shapes, and Clusters 3 and 6 primarily exhibited star-shaped shapes. The results indicate that there are some associations between the series of SNPs and phenotypic traits, but there are also a few similar varieties among the six clusters, indicating that the 163 amaryllis varieties have complex parental origins and rich genetic diversity. This indirectly proves the reliability and applicability of the 82 SNP markers for the analysis of genetic diversity and kinship of amaryllis germplasm resources.
[0075] (3) PCA distribution map Based on the SNP information obtained from the above analysis, eigenvectors and eigenvalues were calculated using PolyGene V1.7 software, and the PCA distribution was plotted using Chiplot. The plotted PCA principal component analysis scatter plot is shown below. Figure 5 As shown in Figure A: The cumulative contribution rates of the first three principal components in the principal component analysis are 6.06% (PC1), 5.46% (PC2), and 4.47% (PC3), respectively. These three principal components mainly explain 16.00% of the genetic variation.
[0076] Import the phylogenetic tree classification results into the PCA diagram, such as... Figure 5 As shown in Figure B: Except for clusters 5, 3, and 6, which are relatively compact, the other clusters are more dispersed and not clearly distinguishable. This may be because the genetic differentiation signals provided by SNPs are affected by the number of markers, or because some loci are derived from shared ancestral loci. The dispersed distribution of clusters 1, 2, and 4 indicates less gene flow within these three clusters and more distant kinship among them. However, the complex mixing of the three clusters suggests a close association between the three groups, possibly indicating significant gene flow and similar genetic backgrounds.
[0077] (4) Analyze population structure Based on the SNP marker information obtained from the above analysis, the population structure was analyzed using Structure (v.2.3.4) and StructureSelect software.
[0078] The LnP(D) likelihood and deltaK values of 163 amaryllis varieties with different K values (assuming ancestral population size) were calculated using Sructure software and plotted as line graphs (e.g.) Figure 6 As shown), the population structure diagram for K=2-8 is plotted using the Q value. Figure 7 ).like Figure 6 As shown in Figure A, DeltaK is maximized when K=6. DeltaK and LnP(D) are negatively correlated. The LnP(D) value corresponding to K=6 is as follows: Figure 6 As shown in Figure B, this indicates that the 163 amaryllis varieties are most likely divided into six groups, consistent with the results of phylogenetic analysis. (In the population structure diagram...) Figure 7 The phylogenetic trees and varietal distributions among different groups show some variety movement, particularly in cluster 2 corresponding to the blue variety, which exhibits a significant mixed ancestral origin with other groups, indicating substantial inter-group interaction and a trend towards similar genetic backgrounds. When K=6, combining the population structure diagram and the Q-value of the population genetic component ratio reveals that the vast majority of varieties possess multiple genetic backgrounds, reflecting the directional integration characteristics of artificial breeding in ornamental horticulture. A few standard varieties possess a single or limited genetic background, primarily diploid primitive varieties resulting from natural selection.
[0079] 4. Fingerprint mapping Among 163 amaryllis varieties, the 82 SNP markers had an average of 30 differentially expressed loci, a maximum of 56 differentially expressed loci, a minimum of 2 differentially expressed loci, and an average PIC greater than 0.3. The varietal differentiation of the SNP markers was analyzed and verified using SNPT software, and these 82 markers were found to be suitable as core markers for fingerprinting.
[0080] Use Excel to draw fingerprint patterns, such as Figure 8 As shown. Figure 8 In the graph, the vertical axis lists 163 amaryllis varieties (ZDH001~ZDH164, excluding ZDH101) in numerical order, and the horizontal axis lists 82 valid SNP molecular markers (SNP1~SNP83, excluding SNP53) in sequential order, with each marker occupying one column. The graph is plotted according to the three types of each SNP marker, where green represents wild-type homozygous (0 / 0), pink represents mutant heterozygous (0 / 1), yellow represents mutant homozygous (1 / 1), and missing data are represented in gray.
[0081] The fingerprint data was converted into a QR code to construct the amaryllis fingerprint QR code, such as... Figure 9 The image shows the SNP fingerprint QR codes of six amaryllis varieties representing different phylogenetic groups.
[0082] Taking ZDH020 'Xintong' as an example, the QR code label shows ZDH020 as the variety number of the amaryllis in this experiment, and 'Xintong' as the variety name. The data in the QR code is: ZDH020 (Amaryllis variety number) XinTong (Amaryllis variety English name) CATT##CCTAGATTGTCCAGCCCTCCCTTCGATTTTGAAGGTTCAGGACGGGTTAATCCTAACCGGTGCTAATTGACACGAGATTCGC##ATGGTTTGCCCTCTTTAAAAGG##ATGTCTCATTCCCTGAGACCCCCCTAACATTCAGACAAGTAACTTAAGCA (82 SNP marker allele combinations, starting from the first base, every two bases represent a specific allele of a SNP marker for this variety, and so on, forming the SNP fingerprint information. Among them, "##" represents the missing allele).
[0083] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art can make various modifications or variations within the scope of the claims, which do not affect the essence of the present invention.
Claims
1. A combination of SNP molecular markers for identifying amaryllis germplasm resources, characterized in that, The SNP molecular marker combination includes 82 SNP molecular markers, and the flanking sequences and variant types of the 82 SNP molecular markers are shown below: In the table above, the flanking sequences are the flanking nucleotide sequences of the SNP molecular markers; "[ ]" marks the site positions of the SNP molecular markers, and the bases in the marks represent the polymorphisms at that site.
2. A primer pair combination for amplifying the SNP molecular marker combination as described in claim 1, characterized in that, include: 。 3. A kit for detecting the SNP molecular marker combination as described in claim 1, characterized in that, The kit includes the primer pair combination as described in claim 2.
4. A chip for detecting SNP molecular marker combinations as described in claim 1, characterized in that, The chip comprises the primer pair combination as described in claim 2.
5. The application of the SNP molecular marker combination as described in claim 1, the primer pair combination as described in claim 2, the kit as described in claim 3, or the chip as described in claim 4 in constructing an amaryllis DNA fingerprint.
6. A method for constructing a DNA fingerprint of amaryllis, characterized in that, include: S1. Extract genomic DNA from amaryllis and perform PCR amplification using the 82 primer pairs as described in claim 2. Mix the products amplified by different primer pair combinations and perform high-throughput sequencing on the mixed products. S2. Identify the genotype of the site where the SNP molecular marker combination as described in claim 1 is located in the amaryllis genome; construct the DNA fingerprint map of amaryllis using the identified SNP genotype combinations of different amaryllis varieties.
7. A DNA fingerprint of amaryllis, characterized in that, It is constructed using the construction method described in claim 6.
8. The application of the SNP molecular marker combination as described in claim 1, the primer pair combination as described in claim 2, the kit as described in claim 3, the chip as described in claim 4, or the fingerprint spectrum as described in claim 7 in the identification of amaryllis germplasm resources.
9. The application according to claim 7, characterized in that, The amaryllis germplasm was selected from 163 amaryllis varieties: Sunong Hongdie, Apple, Meng Tong, Xiu Xiu, Minerva, Xiyan, Peacock flower, Seeking the truth, Jing Qi, Blossom Peacock, Xia Bing, Spring Search, Happy Lamb, Choir, Yi Yi, Whirlpool, Double Dragon, Yu Jie, Spicy, Xin Tong, Le Zhen, Red Daruma, Picotee, Merry Christmas, Exposure, Star language, Meng Yuan, Xiao Ge, Ru Yi, Yun Xi, Spring Apricot, Xia Tong, Yun Han, An Ran, Chan Xi Xī, Lia, Elegantand graceful, Feifei, Rookie, Yu'e, Jiaxin, Shan Shan, Wanqing, Little Apple, Afu lei, Spring glow, Wuhuatianbao, Ice Summer, Masonry red, Xin Yu, Carimero, Pink Surprise, Bloodline, Wedding dance music, Sunong RedSilk, Red lips, Double Happiness, Graceful dancing, Goodluck, Xiang Mei, Red Peafowl, Beautiful Lady, Sea star, Green dress, Smiley, Magic, Pink delicate, Sunong Hongwu, Child, Childhood, Fish Sister, Fairy Tale, Marilyn, RuiLi, Daphne, Qixia, Pipi, Flame PeacockPeacock, Jiari, Springsnow, Hongguang, Sunset glow, Red Sun, Jingxiu, Ruoxi, Provocative, Yiting, Mengqi, Xiuyan, Spring Dependence, RedDawn, Feiyi, Xia Lan, Meilin, Xiutong, Yitong, Snow Bath, Xia Xue, Ruoqing, Ruoxi, Ziyan, Xingxiu, Witch, Sunong Red Wind, Radiant glow, Colored clouds, Red glow, Xiao Ai, Elf, Little Dream, Grow Chuyun, Emerald Green, Picasso, Berry core, Ruby, Ambiance, Sweet Nivea, Jingya, Pink dress, Volcanic eruption, Yiting, Lan Yan, Sweet Fairy, Lux, Jia Zhi, Dance Queen, Lucky Angel, Black Swan, Phantasm, Cupid, Tika Gloria, Tika Fantasy, Tika kam, Tika star, Tika Trisco, Tika lutz, Tika Yum Pairot, Tika McGeeck, Tika Quin, Tika Sarris, Tika sa Sha, Tika Auror, Tika Biyouti, Tika ram, Tika Farrey, Sir Tika, Tika Flamingo, Tika pu Sais, Tika rang Root, Yihan, Beautiful, Anhui, Yirui, Chen Yan, Yashi, RedAny one of the following: Sleeve, ShiYan, The Paper, Dawn glow, Yunxuan, Tika c Ruo, Sunshining, Princess Xiang Lei.
10. The application according to claim 8, characterized in that, The method for identifying amaryllis germplasm resources includes the following steps: A1. Extract genomic DNA from the germplasm of Amaryllis to be identified; A2. SNP primer amplification and purification: The genomic DNA of the amaryllis variety extracted using the 82 SNP primer pairs as described in claim 2 was amplified by PCR. The amplification products were mixed in equal volume proportions and then purified. A3. The purified mixed product is subjected to high-throughput sequencing. The sequencing results are compared with the flanking sequences of each SNP site to obtain the actual site information of the amaryllis variety. The results are then compared with the fingerprint pattern described in claim 7 to identify the amaryllis germplasm.