A SNP molecular marker combination for identifying taro germplasm resources and application thereof

By developing 27 SNP marker combinations for taro germplasm resources and PARMS-SNP marker technology, the problem of inaccurate identification of taro germplasm resources in traditional methods has been solved, enabling rapid and accurate resource identification and evaluation, and establishing a core taro germplasm resource bank.

CN120776052BActive Publication Date: 2026-02-24HUNAN VEGETABLE RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511207780.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2026-02-24
Estimated Expiration
2045-08-27

AI Technical Summary

Technical Problem

Existing technologies are insufficient for efficiently and accurately identifying and evaluating taro germplasm resources. Traditional methods are greatly affected by environmental and human factors, and the application of SNP molecular marker technology in taro germplasm resources is inadequate.

Method used

A molecular marker combinatorial system comprising 27 SNP markers was developed for genotyping and DNA fingerprinting of taro germplasm resources, and combined with PARMS-SNP marker technology for genetic diversity analysis and identification.

Benefits of technology

Rapid and accurate identification and evaluation of taro germplasm resources has been achieved, a core taro germplasm resource bank has been established, and identification efficiency and accuracy have been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120776052B_ABST
    Figure CN120776052B_ABST
Patent Text Reader

Abstract

The application discloses a SNP molecular marker combination for identifying taro germplasm resources and application, and belongs to the technical field of molecular biology. 252 SNP sites are screened for the development of PARMS-SNP markers, and 27 original PARMS-SNP markers are successfully obtained. 72 taro resources are subjected to genetic typing by using the 27 SNP markers, and population structure analysis, PCA principal component analysis and genetic evolution analysis are further carried out. The taro germplasm resources can be better distinguished, and reference is provided for the evaluation and identification of taro germplasm resources and the breeding of new germplasm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of molecular biology, specifically to a combination of SNP molecular markers for identifying taro germplasm resources and its application. Background Technology

[0002] Taro (Colocasia esculenta) belongs to the Araceae family and is a perennial plant, but is usually grown as an annual. It contains more starch than potatoes, sweet potatoes, and other crops. According to the Food and Agriculture Organization of the United Nations, the global taro planting area in 2020 was 39,000 hectares, making it considered the fifth most important root crop in the world and a vital food crop in tropical and subtropical countries. Taro has a cultivation history of over 2,300 years in China. Over this long period, various local varieties have developed due to different regional environmental factors and human selection.

[0003] Germplasm resources are a prerequisite for breeding, and accurate assessment of germplasm resources can greatly improve the utilization rate of materials by breeders or researchers. Traditional field phenotypic evaluation and identification can only visually investigate resources and are greatly affected by subjective personal judgment and objective environmental factors. However, currently popular DNA molecular marker technologies distinguish individual genotype differences within and between species based on crop genetic information. They have short cycles, are not affected by environmental and biological developmental cycles, have no tissue specificity, and offer high-throughput detection. They have been widely used for genetic diversity analysis and fingerprinting of germplasm resources. SNPs, as a third-generation molecular marker technology, have the advantages of large quantity, stable heritability, simple and rapid operation, and wide application.

[0004] Research on the genetic diversity analysis of taro resources is limited. Wang et al. analyzed the genetic diversity of 234 taro germplasm resources from 16 provinces in China using genome-wide SNP markers, concluding that the experimental materials could be divided into eight populations. Pan et al. developed a high-density InDel SSR locus covering the entire genome based on resequencing data of core taro germplasm and divided 121 resources into three groups, but they failed to screen out core markers for subsequent effective utilization. Guo et al. screened 22 pairs from 100 pairs of SSR molecular markers and completed the genetic diversity analysis of 109 taro resources, providing useful markers for the identification and evaluation of taro resources. However, SSR markers rely on PCR amplification and electrophoresis analysis, which, compared with SNPs, have lower throughput, longer processing time, and lower genetic stability. Summary of the Invention

[0005] The purpose of this invention is to provide a combination of SNP molecular markers for identifying taro germplasm resources and its application. Using the SNP molecular marker combination of this invention, genetic diversity analysis and DNA fingerprinting of taro resources can be performed, and an accurate, simple and rapid method for identifying taro germplasm resources or varieties can be established, thereby enabling more efficient identification of taro germplasm resources and helping to establish a core taro germplasm resource bank.

[0006] To achieve the above objectives, the present invention provides the following technical solution:

[0007] In a first aspect, the present invention provides a combination of SNP molecular markers for identifying taro germplasm resources, the combination of SNP molecular markers comprising 27 SNP markers, as follows:

[0008] The first locus is located at position 142740049 on chromosome LG01, and the variant type is C / T.

[0009] The second site is located at position 208737408 on chromosome LG01, and the variant type is C / T.

[0010] The third site is located at position 64737033 on chromosome LG02, and the variant type is T / C.

[0011] The fourth locus is located at position 65417421 on chromosome LG02, and the variant type is T / C.

[0012] The fifth site is located at position 8979131 on chromosome LG08, and the variant type is C / G.

[0013] The 6th locus is located at position 22020255 on chromosome LG01, and the variant type is G / T.

[0014] The 7th locus is located at position 48214697 on chromosome LG01, and the variant type is A / T.

[0015] The 8th position is located at position 99501267 on chromosome LG12, and the variant type is G / A.

[0016] The 9th position is located at 109034243 on chromosome LG10, and the variant type is T / G.

[0017] The 10th position is located at position 165161699 on chromosome LG05, and the variant type is A / C.

[0018] The 11th position is located at position 144128461 on chromosome LG04, and the variant type is T / C.

[0019] The 12th position is located at position 867311 on chromosome LG14, and the variant type is A / G.

[0020] The 13th position is located at 148250105 on chromosome LG04, and the variant type is T / G.

[0021] The 14th position is located at position 153062049 on chromosome LG05, and the variant type is C / T.

[0022] The 15th position is located at position 89325049 on chromosome LG11, and the variant type is T / C.

[0023] The 16th position is located at position 170119348 on chromosome LG08, and the variant type is C / A.

[0024] The 17th position is located at position 7248248 on chromosome LG10, and the variant type is A / G.

[0025] The 18th position is located at position 130247652 on chromosome LG06, and the variant type is G / T.

[0026] The 19th position is located at position 68668940 on chromosome LG08, and the variant type is G / A.

[0027] The 20th position is located at 17307221 on chromosome LG08, and the variant type is T / C.

[0028] The 21st position is located at position 102781792 on chromosome LG14, and the variant type is T / C.

[0029] The 22nd position is located at position 51666285 on chromosome LG13, and the variant type is T / A.

[0030] The 23rd position is located at position 206486904 on chromosome LG01, and the variant type is T / A.

[0031] The 24th position is located at position 64252159 on chromosome LG04, and the variant type is T / C.

[0032] The 25th position is located at 185181757 on chromosome LG02, and the variant type is G / A.

[0033] The 26th position is located at position 83510795 on chromosome LG12, and the variant type is A / G.

[0034] The 27th position is located at 179908017 on chromosome LG01, and the variant type is A / T.

[0035] Secondly, the present invention provides a taro DNA fingerprint, which is constructed using the aforementioned SNP molecular marker combination.

[0036] Thirdly, the present invention provides the application of the aforementioned SNP molecular marker combination in the genotyping of taro germplasm resources.

[0037] Fourthly, the present invention provides the application of the aforementioned SNP molecular marker combination in constructing a taro DNA fingerprint.

[0038] Fifthly, the present invention provides the application of the aforementioned SNP molecular marker combination in the identification of taro resources.

[0039] Sixthly, the present invention provides a method for identifying taro resources, comprising:

[0040] Total DNA was extracted from the sample to be identified, and the genotype of the site where the SNP molecular marker combination was located on the genome of the sample to be identified was determined.

[0041] The genotype of the individual sample to be identified is compared with the fingerprint spectrum; the variety of the sample to be identified is determined based on the comparison results.

[0042] Based on the above technical solution, the embodiments of the present invention can produce at least the following technical effects:

[0043] This invention is based on simplified genome resequencing of 29 taro germplasm samples collected from Hunan and other provinces during the Third National Crop Resources Survey. A total of 86,596,792 single nucleotide polymorphisms (SNPs) were identified. Further SNP screening identified 252 SNP sites for PARMS-SNP marker development, successfully obtaining 27 original PARMS-SNP markers. These 27 SNP markers were used to perform genetic typing on 72 taro resources, followed by population structure analysis, principal component analysis (PCA), and evolutionary genetic analysis. Combining field phenotypes and clustering results, significant gene mixing was found between diploid and triploid taro. A taro germplasm fingerprint was constructed, providing a reference for the evaluation and identification of taro germplasm resources and the breeding of new germplasm. Attached Figure Description

[0044] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.

[0045] Figure 1This invention uses 27 pairs of primers to perform genotyping on 72 taro resources;

[0046] Figure 2 In the diagram, A is the cross-validation error rate plot; B is the clustering plot based on the number of subgroups (K value).

[0047] Figure 3 A is the PAC diagram obtained based on SNP data; B is the phylogenetic tree constructed based on SNP data; C is the phylogenetic tree of the taro sample.

[0048] Figure 4 It is a fingerprint spectrum of taro. Detailed Implementation

[0049] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. In addition, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such combination of technical solutions does not exist and is not within the scope of protection claimed by the present invention.

[0050] The objective of this invention is achieved through the following technical solution:

[0051] 1. Materials and Methods

[0052] 1.1 Materials

[0053] The materials used in this experiment were taro germplasm resources provided by the Vegetable Research Institute of Hunan Academy of Agricultural Sciences (Table 1). Based on previous ploidy testing, the 29 resources were divided into 3 groups, where A was diploid, and B and C were triploid. All materials were grown at the Gaoqiao Research Base of Hunan Academy of Agricultural Sciences in Gaoqiao Town, Changsha County, Hunan Province.

[0054] Table 1. Basic information on 29 taro resources

[0055]

[0056]

[0057] 1.2 Genome Sequencing

[0058] Fresh leaves were collected and rapidly frozen in liquid nitrogen. Genomic DNA was extracted using a kit, and DNA quality was assessed by 0.8% agarose gel electrophoresis. After quantification using a UV spectrophotometer, a library with an insert size of 400 was constructed according to the standard library construction protocol of the Illumina TruSeq DNA PCR-freeprep kit. The library was then sequenced using Next-Generation Sequencing on the Illumina NovaSeq sequencing platform.

[0059] 1.3 SNP Analysis

[0060] To ensure the quality of subsequent analyses, the raw data was filtered using the fastp (v0.20.0) sliding window method to remove contaminating adapters. Reads with an average Q value <20 and a length <50 bp within the window were deleted to generate high-quality sequences. The filtered high-quality data was then aligned to the taro reference genome using the bwa (0.7.12-r1039) mem program (https: / / ftp.cngb.org / pub / CNSA / data2 / CNP0001082 / CNS0231432 / CNA0014218 / ). SNP detection was performed using GATK software. To ensure the reliability of SNP loci, the obtained SNP loci were further screened according to the following criteria: (1) Fisher test of strand bias (FS) ≤ 60, (2) HaplotypeScore ≤ 13.0, (3) Mapping Quality (MQ) ≥ 40, (4) Quality Depth (QD) ≥ 2, (5) ReadPosRankSum ≥ -8.0, (6) MQRankSum > -12.5. The ANNOVAR software was used to annotate the filtered SNP loci.

[0061] 1.4 SNP screening and primer design

[0062] Further preliminary screening of SNP sites: (1) Minimum allele frequency (MAF) > 0.4; (2) False negative rate < 0.25; (3) Heterozygosity < 0.6; (4) No other mutations 50 bp flanking the SNP site. 50 bp sequences upstream and downstream of the SNP were extracted. The sequences were aligned to a reference genome using TBtools software, and primers were designed using a single aligned sequence selected on Primer5.

[0063] 1.5 Development of PARMS-SNP markers

[0064] The PARMS-SNP genotyping reagents were obtained from Wuhan Jingtai Biotechnology Co., Ltd. Twenty-two samples were randomly selected for preliminary primer screening, with water used as a negative control. Primers with good genotyping results were selected for genetic analysis of 72 samples. Powermarker software was used to calculate PIC, MAF (minimum allele frequency), GD (genetic diversity), and HE (heterozygosity).

[0065] 1.6 Genetic diversity analysis

[0066] Genetic diversity analysis was performed on the genotyping results of 72 materials. PCA analysis was conducted using the ade4R software package. Phylogenetic trees were constructed using Powermarker. The K-value was calculated using Structure Harvester software, and the optimal K-value was obtained. Simultaneously, Admixture software (http: / / dalexander.github.io / admixture) was used, incorporating SNP information, to set the K=2-10 model (assuming 2-7 ancestral populations) as a mixture model, and the genetic structure of the population was analyzed using the software's default settings. Cluster analysis of phenotypic traits of taro resources was performed using DPS software, utilizing Euclidean distance and variance.

[0067] 1.7 Construction of 72 Taro Fingerprint Maps

[0068] Based on the results, fingerprint maps were constructed using Excel. Table 2 shows the relevant sequences of the 27 core primers. A heatmap of the genotypes for the optimal combination of markers was plotted for fingerprint identification, with each row representing a SNP locus and each column representing a sample. Genotypes were color-coded as follows: CT = yellow, TT = light green, CC = blue, GG = purple, GC = black, GA = dark red, TG = green, AT = gray, AA = red, CA = white; and no calling genotypes were specified as NN = brown.

[0069] Table 2 Information on 27 core PARMS-SNP markers

[0070]

[0071]

[0072]

[0073] Note: Chr: Chromosome sequence name; Position: Marker location; Ref: Base of reference genome; Alt: Variant base; Primer sequence X adds FAM fluorescent matching adapter: GAAGGTGACCAAGTTCATGCT; Primer sequence Y adds HEX fluorescent matching adapter: GAAGGTCGGAGTCAACGGATT; Z is the reverse complementary strand.

[0074] 2 Results

[0075] 2.1 Genome sequencing and reference genome alignment

[0076] A total of 10Gb of clean data was obtained, with a Q30 of over 95%. After filtering the raw data using the fastp (v0.20.0) sliding window method, over 55,978,374 high-quality reads were obtained, with the lowest high-quality read alignment rate at 91.79%.

[0077] Alignment of high-quality reads with the reference genome showed low alignment rates of approximately 60% for reads 60, 34, and 45. Other samples achieved alignment rates of at least 99.3%, with duplicate alignment rates not exceeding 18%. The sequencing depth of the samples was above 3×, indicating that the data can be used for the development and screening of SNP molecular marker sites.

[0078] 2.3 Development of PARMS-SNP markers

[0079] Based on the simplified genome sequencing results, 252 SNP sites were screened from 86,596,792 SNPs, and these 252 SNP sites were further screened to obtain PARMS-SNP markers. The 252 primer sets were validated using 22 random samples, resulting in 36 usable PARMS-SNP markers. Genotyping of these 36 primer sets was performed on 72 taro samples, and the results of each genotyping were recorded. Nine primer sets with a loss rate >15% were removed, resulting in 27 usable primer sets.

[0080] Genotyping of 72 taro resources was performed using these 27 primer pairs, and the results are as follows: Figure 1 As shown, the MAF values ​​of these 27 primer sets ranged from 0.51 to 0.92, with an average MAF value of 0.63; the PIC values ​​ranged from 0.14 to 0.37, with an average PIC value of 0.34; the HE values ​​ranged from 0.07 to 0.90, with an average of 0.49; and the GD values ​​ranged from 0.15 to 0.50, with an average of 0.45. The results indicate that these primers possess good genetic diversity and can be used for genetic diversity analysis of taro.

[0081] 2.4 Genetic diversity analysis of 72 taro resources

[0082] like Figure 2 As shown, when analyzing the population structure and selecting the K=2-7 model as the mixed model, it was found that when K=6, the CV error value is closer to 0, indicating that it is more appropriate to divide these 72 taro resources into 6 groups.

[0083] Principal component analysis was performed using GCTA software and SNP data (excluding SNPs with a MAF less than 0.05). Figure 3 As shown, 72 resources were found to be divided into multiple groups, with diploids and triploids not clustered together. This conclusion was further confirmed using a phylogenetic tree. Cluster analysis was performed based on the qualitative trait results from the field phenotypic survey of the 72 resources. The study found that the constructed phylogenetic tree mostly matched the phenotypic clusters, but some resource differential classifications showed significant differences between the two.

[0084] 2.5 Construction of fingerprint maps of 72 taro resources

[0085] like Figure 4 As shown, fingerprint profiles of 72 local resources were constructed using 27 primer pairs. The results indicate that these 27 markers can effectively distinguish between resources. Among these resources, samples 40 and 121, 97 and 17, and 59 and 37 showed differences at only one locus, indicating potential similarities between varieties. These varieties were also very similar in phenotypic traits.

[0086] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A combination of SNP primers for identifying taro germplasm resources, characterized in that, It includes 27 sets of primers, and their sequence information is as follows: SNP1, the nucleotide sequences of the forward primer are shown in SEQ ID No. 1 and SEQ ID No. 2, and the nucleotide sequence of the reverse primer is shown in SEQ ID No. 3; SNP2, the nucleotide sequences of the forward primer are shown in SEQ ID No. 4 and SEQ ID No. 5, and the nucleotide sequence of the reverse primer is shown in SEQ ID No. 6; SNP3, the nucleotide sequences of the forward primer are shown in SEQ ID No. 7 and SEQ ID No. 8, and the nucleotide sequence of the reverse primer is shown in SEQ ID No. 9; SNP4, the nucleotide sequences of the forward primer are shown in SEQ ID No. 10 and SEQ ID No. 11, and the nucleotide sequence of the reverse primer is shown in SEQ ID No. 12; SNP5, the nucleotide sequences of the forward primer are shown in SEQ ID No. 13 and SEQ ID No. 14, and the nucleotide sequence of the reverse primer is shown in SEQ ID No. 15; SNP6, the nucleotide sequences of the forward primer are shown in SEQ ID No. 16 and SEQ ID No. 17, and the nucleotide sequence of the reverse primer is shown in SEQ ID No. 18; SNP7, the nucleotide sequences of the forward primer are shown in SEQ ID No. 19 and SEQ ID No. 20, and the nucleotide sequence of the reverse primer is shown in SEQ ID No. 21; SNP8, the nucleotide sequences of the forward primer are shown in SEQ ID No. 22 and SEQ ID No. 23, and the nucleotide sequence of the reverse primer is shown in SEQ ID No. 24; SNP9, the nucleotide sequences of the forward primer are shown in SEQ ID No. 25 and SEQ ID No. 26, and the nucleotide sequence of the reverse primer is shown in SEQ ID No. 27; SNP10, the nucleotide sequences of the forward primer are shown in SEQ ID No. 28 and SEQ ID No. 29, and the nucleotide sequence of the reverse primer is shown in SEQ ID No. 30; SNP11, the nucleotide sequences of the forward primer are shown in SEQ ID No. 31 and SEQ ID No. 32, and the nucleotide sequence of the reverse primer is shown in SEQ ID No. 33; SNP12, the nucleotide sequences of the forward primer are shown in SEQ ID No. 34 and SEQ ID No. 35, and the nucleotide sequence of the reverse primer is shown in SEQ ID No. 36; SNP13, the nucleotide sequences of the forward primer are shown in SEQ ID No. 37 and SEQ ID No. 38, and the nucleotide sequence of the reverse primer is shown in SEQ ID No. 39; SNP14, the nucleotide sequences of the forward primer are shown in SEQ ID No. 40 and SEQ ID No. 41, and the nucleotide sequence of the reverse primer is shown in SEQ ID No. 42; SNP15, the nucleotide sequences of the forward primer are shown in SEQ ID No. 43 and SEQ ID No. 44, and the nucleotide sequence of the reverse primer is shown in SEQ ID No. 45; SNP16, the nucleotide sequences of the forward primer are shown in SEQ ID No. 46 and SEQ ID No. 47, and the nucleotide sequence of the reverse primer is shown in SEQ ID No. 48; SNP17, the nucleotide sequences of the forward primer are shown in SEQ ID No. 49 and SEQ ID No. 50, and the nucleotide sequence of the reverse primer is shown in SEQ ID No. 51; SNP18, the nucleotide sequences of the forward primer are shown in SEQ ID No. 52 and SEQ ID No. 53, and the nucleotide sequence of the reverse primer is shown in SEQ ID No. 54; SNP19, the nucleotide sequences of the forward primer are shown in SEQ ID No. 55 and SEQ ID No. 56, and the nucleotide sequence of the reverse primer is shown in SEQ ID No. 57; SNP20, the nucleotide sequences of the forward primer are shown in SEQ ID No. 58 and SEQ ID No. 59, and the nucleotide sequence of the reverse primer is shown in SEQ ID No. 60; SNP21, the nucleotide sequences of the forward primer are shown in SEQ ID No. 61 and SEQ ID No. 62, and the nucleotide sequence of the reverse primer is shown in SEQ ID No. 63; SNP22, the nucleotide sequences of the forward primer are shown in SEQ ID No. 64 and SEQ ID No. 65, and the nucleotide sequence of the reverse primer is shown in SEQ ID No. 66; SNP23, the nucleotide sequences of the forward primer are shown in SEQ ID No. 67 and SEQ ID No. 68, and the nucleotide sequence of the reverse primer is shown in SEQ ID No. 69; SNP24, the nucleotide sequences of the forward primer are shown in SEQ ID No. 70 and SEQ ID No. 71, and the nucleotide sequence of the reverse primer is shown in SEQ ID No. 72; SNP25, the nucleotide sequences of the forward primer are shown in SEQ ID No. 73 and SEQ ID No. 74, and the nucleotide sequence of the reverse primer is shown in SEQ ID No. 75; SNP26, the nucleotide sequences of the forward primer are shown in SEQ ID No. 76 and SEQ ID No. 77, and the nucleotide sequence of the reverse primer is shown in SEQ ID No. 78; SNP27, the nucleotide sequences of the forward primer are shown in SEQ ID No. 79 and SEQ ID No. 80, and the nucleotide sequence of the reverse primer is shown in SEQ ID No.

81.

2. The application of the SNP primer combination described in claim 1 in the genotyping of taro germplasm resources.

3. The application of the SNP primer combination described in claim 1 in constructing a taro DNA fingerprint.

4. The application of the SNP primer combination described in claim 1 in the identification of taro resources.

5. A method for identifying taro resources, characterized in that, include: Total DNA was extracted from the sample to be identified and detected using the SNP primer combination described in claim 1. The variety of the sample to be identified is determined based on the test results.