SNP, primers, fingerprint maps for identifying pitaya germplasm resources and their applications
The SNP molecular marker and fingerprint map developed through multiple PCR and high-throughput sequencing technology solves the problems of high accuracy and cost in the identification of dragon fruit varieties, and achieves efficient and accurate variety identification, with a distinction of 99.77%.
Patent Information
- Application Number
- CN202411484710.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-23
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-10-23
AI Technical Summary
The existing methods for identifying dragon fruit varieties have problems such as accuracy and high cost and complex operation, especially when it is difficult to distinguish between relative varieties.
Through multiple PCR combined with high-throughput sequencing technology, SNP molecular markers can be developed that can be used to identify dragon fruit varieties, and the SNP fingerprint map of dragon fruit is constructed to achieve rapid and accurate variety identification.
The accuracy and reliability of dragon fruit variety identification have been improved, misjudgment and operational complexity problems in the prior art have been overcome, and the distinction has reached 99.77%.
Smart Images

Figure CN119061191B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of biotechnology, and particularly relates to a group of SNPs, primers, fingerprint maps for identifying pitaya germplasm resources and their applications. Background Art
[0002] Pitaya (Hylocereus undatus) has attracted wide attention due to its unique appearance, nutritional value and market potential. In recent years, with the rapid development of the pitaya planting industry, the management of variety diversity and the protection of variety rights have become particularly important. Traditional morphological identification methods are often greatly affected by environmental factors and it is difficult to distinguish closely related varieties. Therefore, finding a more accurate and efficient variety identification method has become the focus of current research.
[0003] At present, there are already some technical means for identifying pitaya varieties in the market, such as classification based on morphological characteristics, detection based on molecular markers, etc. However, these methods generally have the following problems: Morphological identification: greatly affected by environmental factors and prone to misjudgment; Detection based on molecular markers: such as RFLP, AFLP, SSR, etc., although having high accuracy, still need to be improved in terms of cost and operation complexity; Small genetic differences between varieties: The genetic differences between some pitaya varieties are small, and traditional molecular marker methods are difficult to effectively distinguish.
[0004] The purpose of the present invention is to provide a new method for identifying pitaya varieties, using single nucleotide polymorphism (SNP) as a molecular marker to construct an SNP fingerprint map of a specific variety, so as to achieve rapid and accurate identification of pitaya varieties. This method can overcome the problems existing in the prior art and improve the accuracy and reliability of variety identification. Summary of the Invention
[0005] Aiming at the problems existing in the prior art, the present invention develops SNP molecular markers that can be used to identify pitaya varieties through multiplex PCR combined with high-throughput sequencing technology, and draws an SNP fingerprint map of pitaya, which can be used for commercial identification of pitayas sold in the market. At the same time, when applied to marker assisted selection (MAS) and genomic selection (GS), it can accelerate the genetic improvement progress of red pitaya varieties.
[0006] The primary objective is to determine SNP molecular markers for identifying pitaya varieties. The molecular markers are located on the pitaya genome (http: / / www.pitayagenomic.com / download.php).
[0007] The sequence of the SNP molecular marker is any one of the molecular markers SNP1 to SNP19, and the nucleotide sequences of the SNP1 to SNP19 molecular markers are as shown in SEQ ID NO.39 to SEQ ID NO.57. The 201st position of SEQ ID NO.39 to SEQ ID NO.57 is the SNP locus.
[0008] In a preferred embodiment, the SNP molecular marker combination of the present invention is the entire combination of SNP1 to SNP19.
[0009] Another object of the present invention is to provide a fingerprint map constructed by the above molecular marker combination.
[0010] Preferably, the fingerprint map is as shown in the following table, where the numbers S1 to S94 are sample numbers, corresponding to the samples in Table 1. A, T, C, and G represent homozygous genotypes, and R, Y, M, K, S, and W represent degenerate bases, that is, heterozygous genotypes. The genotypes are arranged in the order of SNP1 to SNP19 loci.
[0011]
[0012] In a preferred embodiment, different genotypes are represented by different colors to obtain a fingerprint map as Figure 3 shown.
[0013] Another object of the present invention is to provide primers / primer sets for identifying the above SNP molecular markers, and the primer pairs are primer pairs capable of amplifying the above SNP loci.
[0014] In a preferred embodiment, the primers / primer sets are primer pairs of any combination of primer pairs 1 to 19 in Table 2.
[0015] In a preferred embodiment, the PCR primers / primer sets of the molecular markers or molecular marker combinations of the present invention are all primer pairs of primer pairs 1 to 19.
[0016] Another object of the present invention is to provide a probe for identifying the above SNP molecular markers.
[0017] The present invention provides a kit for detecting the above SNP molecular markers, which is characterized in that the kit comprises the foregoing primer pairs and / or probes.
[0018] Another object of the present invention is to provide the application of the above molecular markers or molecular marker combinations in identifying or assisting in identifying pitaya varieties.
[0019] Another object of the present invention is to provide the application of the foregoing fingerprint map in identifying or assisting in identifying pitaya varieties.
[0020] Another object of the present invention is to provide the application of the above primers / primer sets in identifying or assisting in identifying pitaya varieties.
[0021] In the above application, when the fingerprint of the sample to be identified is consistent with that of any variety in Table 3 or Figure 3 it can be determined that the sample to be identified is the corresponding pitaya variety.
[0022] The present invention also provides the application of the above probe or kit in identifying or assisting in identifying pitaya varieties.
[0023] The present invention also provides the application of the foregoing SNP molecular marker combination, primer and / or probe, kit, chip and / or SNP fingerprint in identifying pitaya varieties.
[0024] The present invention also provides a method for constructing a pitaya fingerprint, which is characterized by comprising the following steps:
[0025] 1) Collect tissue samples of the sample; 2) Extract the total DNA of the sample; 3) Perform PCR amplification using the foregoing primers, and sequence the amplification product. 4) Identify the genotype of the locus where the SNP molecular marker combination described in claim 1 is located on the genome of the sample; 5) Construct a pitaya fingerprint using the genotype identified in step 3.
[0026] Preferably, the PCR amplification in step 3 is multiplex PCR amplification.
[0027] Another object of the present invention is to provide a method for identifying or assisting in identifying pitaya varieties, the method comprising the following steps:
[0028] 1) Collect tissue samples of the sample to be identified; 2) Extract the total DNA of the sample to be identified; 3) Identify the genotype of the core SNP loci (i.e., the 19 core SNP loci in Example 2) on the genome of the sample to be tested; 4) Compare the individual genotype of the sample to be identified with the foregoing fingerprint; 5) Determine the variety of the sample to be identified according to the comparison result in step 4).
[0029] In the above method, step 3 is to identify the genotype of the SNP loci shown in Table 3 on the genome of the sample to be identified by genome sequencing or chip detection;
[0030] Furthermore, the sequencing is first-generation sequencing, and / or second-generation sequencing, and / or third-generation sequencing.
[0031] Furthermore, the chip is a DNA probe chip or an RNA probe chip.
[0032] In the above method, the SNP loci are part or all of the SNP loci in Table 3.
[0033] In some cases, genotypes are represented in base form. For example, when performing sequence identity analysis, phylogenetic tree construction, or PCA analysis, A, T, C, and G represent homozygous sites in base form, R, Y, M, K, S, and W represent heterozygous sites in degenerate base form, and "N" represents that the site was not detected. Therefore, within the scope understandable to those skilled in the art, the genotypes described in the present invention have the same meaning as bases in the above cases.
[0034] The present invention has the following advantages and effects compared with the prior art:
[0035] First, the present invention has determined SNP molecular markers that can be used to identify pitaya varieties at the population level and developed corresponding identification methods.
[0036] Second, the SNP molecular markers determined by the present invention can be used alone or in combination, and can be flexibly selected according to specific usage scenarios.
[0037] Third, the SNP molecular markers and identification methods determined by the present invention have the advantages of strict statistics and high accuracy when used to identify pitaya, and are easy to use. Different sequencing methods or chips can be used for identification.
[0038] Fourth, the SNP sites determined by the present invention are in suitable positions and are easy to design primers.
[0039] Fifth, the present invention provides fingerprint maps of 94 pitaya varieties, and the discrimination degree is extremely high, reaching 99.77%. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The methods and their beneficial effects of the present invention will be described in detail below in conjunction with the drawings and specific embodiments.
[0041] Figure 1 It is a PCA graph among individuals constructed using SNPs. A is the PCA graph of all SNPs, and B is the PCA graph of 19 core SNPs.
[0042] Figure 2 It is a graph for analyzing genetic relationships using the gcta software. Among them, A is the genetic relationship of all SNPs, and B is the genetic relationship of 19 core SNPs.
[0043] Figure 3 It is a fingerprint map constructed using 19 core SNPs. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention. Unless otherwise defined, all technical and scientific terms used in this application have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to the above embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention should not be limited to the embodiments shown herein, but should conform to the widest scope consistent with the principles and novel features disclosed herein.
[0045] Example
[0046] Example 1 Resequencing Analysis
[0047] 1. Experimental Materials
[0048] The sources and information of the materials are shown in Table 1, and all experimental materials are from the resource nursery of the Guangxi Academy of Agricultural Sciences:
[0049]
[0050] 2. Sample DNA Extraction and Library Construction for Sequencing
[0051] 1) First, preserve the leaves of each sample with liquid nitrogen and extract genomic DNA using a kit. 2) Design primers for 60 molecular markers based on the previous reduced-representation genome sequencing results and synthesize the primers. 3) Conduct multiplex PCR experiments. 4) Construct a library for the multiplex PCR products using the Illumina TruSeq™ Nano DNA Sample Prep Kit method and perform sequencing. On average, 157.59M of raw data is obtained for each sample, and the sequencing results are 150bp paired-end data.
[0052] 3. Data Quality Control
[0053] Use the software fastp (version: 0.20.0) to filter out low-quality sequences in the raw data and obtain Clean Data. The parameter settings are: -q 5 -n 5, and the following reads are mainly removed:
[0054] (1) Remove reads with the number of unknown bases N < 5; (2) Remove reads where the base quality value of 50% of the length is less than 5; (3) Remove adapter sequences.
[0055] 4. Population Mutation Detection
[0056] In the present invention, taking the genome of pitaya as the reference genome (http: / / www.pitayagenomic.com / download.php), the obtained sequencing reads need to be re-aligned to the reference genome before subsequent mutation analysis can be carried out. The reference genome is aligned using bwa (software version: 0.7.17-r1188). The reference genome is aligned using BWA (Burrows-Wheeler Aligner, 0.7.17-r1188) with the parameter settings: -M -R. The generated sam file is converted to bam format using the samtools software (version: 1.9). Then, PCR duplications are marked using picard MarkDuplicates (version: 2.21.2). Then, only high-quality proper reads are retained for subsequent analysis.
[0057] 5. SNP Filtering
[0058] (1) Sequencing depth and quality value filtering: The average sequencing depth is greater than or equal to 500X, the minimum quality value is greater than or equal to 30, the integrity is 97%, the number of allele genotypes is 2, and the minor allele frequency (maf) is 0.05. (2) Target region filtering: According to the amplified target region, SNPs located within the target region are screened.
[0059] According to the above steps, 179 marker information is finally obtained.
[0060] 6. PCA Principal Component Analysis
[0061] PCA is a multivariate statistical analysis method in mathematics, which studies how to represent the internal structure among multiple variables through a few principal components, that is, to derive a few principal components from the original variables so that they can retain as much information of the original variables as possible, and it has a very wide range of practical applications. In population re-sequencing, PCA analysis is mainly carried out using population SNP data to calculate some main components (eigenvectors) that can represent all SNP variation structures, and then the analysis results are visually displayed by drawing a graph. The results are as Figure 1 shown in Figure A. Samples with similar genetic backgrounds are clustered together, and the variance contribution rates of PC1 and PC2 of the 179 loci are 36.4% and 13.59% respectively.
[0062] 7. Kinship Analysis
[0063] Based on 179 loci, the genetic relationship analysis was performed using the gcta software to obtain the G matrix (genetic relationship matrix) between pairwise samples. The larger the G value, the closer the relationship between the samples. The results are as Figure 2 shown in Figure A. The darker the color, the larger the G value. The maximum G value is 100, indicating that the genetic backgrounds of the samples are exactly the same.
[0064] Example 2 Development of Fingerprint Map
[0065] 1. Screening of Core SNPs
[0066] For the SNP markers screened above, a script was used to find a set of markers that can distinguish all samples as much as possible, and a total of 19 SNP molecular markers were found.
[0067] The labeled primers are as follows (i.e., 19 pairs of the 60 pairs of multiplex PCR primers in Example 1):
[0068]
[0069] The sequences of the SNP molecular markers are shown in SEQ ID NO.39~SEQ ID NO57.
[0070]
[0071] 2. PCA Principal Component Analysis Using Core SNPs
[0072] PCA analysis was performed using the core SNP data to calculate some main components (eigenvectors) that can represent all SNP variation structures, and then the analysis results were visually displayed by drawing. The results are as Figure 1 shown in Figure B. Compared with the PCA map of all SNPs, there is almost no overlap of all samples in the PCA map of core SNPs, and the variance contribution rates of PC1 and PC2 are both greater than those of the PCA map of all SNPs, which are 37.33 and 18.30 respectively.
[0073] 3. Kinship Analysis
[0074] Based on 19 core loci, the genetic relationship analysis was performed using the gcta software to obtain the G matrix (genetic relationship matrix) between pairwise samples. The larger the G value, the closer the relationship between the samples. The results are as Figure 2 shown in Figure B. The darker the color, the larger the G value. The maximum G value is 100, indicating that the genetic backgrounds of the samples are exactly the same. Compared with the 179 loci, the overall G value of the 19 core loci changes little, and the discrimination ability is equivalent to that of all SNP loci.
[0075] 4. Construction of Fingerprint Map
[0076] Determine the genotypes of 19 SNP loci of the above pitaya germplasm resources to construct a DNA fingerprint map. There are four types of bases in an individual, namely adenine (A), cytosine (C), guanine (G), and thymine (T). According to the principle of SNP, there may be homozygotes and heterozygotes at each SNP locus. Four base pairs such as AA, CC, GG, and TT belong to homozygotes, while the other 12 base pairs such as AC, AT, CT, CG, AG, TG, GA, TC, GT, GC, TA, and CA belong to heterozygotes. For missing loci, they can be filled with NN.
[0077] The construction result ( Figure 3 ) can visually express the differences among 94 pitaya germplasms at the gene level. Among them, 94 pitayas can be accurately divided into 88 clusters, and only 3 clusters contain more than 1 sample (Cluster35: containing samples 68, 55, 54, 65; Cluster41: containing samples 18, 26, 79; Cluster42: containing samples 88, 56). Therefore, the discrimination rate of 19 SNPs (the probability that the fingerprint maps of two randomly selected samples are different) reaches 99.77%.
[0078] Example 3 Identification of Pitaya Varieties
[0079] Identification method 1:
[0080] 1) Collect tissue samples of the samples to be identified; 2) Extract the total DNA of the samples to be identified; 3) Sequence the DNA sequences of the samples to be identified containing the 19 SNP loci described in Example 2; 4) Identify the genotypes of the core SNP loci (i.e., the 19 core SNP loci in Example 2) on the genome of the samples to be tested; 5) Compare the individual genotypes of the samples to be identified with the fingerprint map constructed in Example 2; 6) Judge the varieties of the samples to be identified according to the comparison results in step 5).
[0081] A total of 202 samples of different varieties were detected, including 55 samples in Example 1, and all were accurately identified.
[0082] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to the above embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A SNP molecular marker combination for constructing a pitaya fingerprint, characterized in that: The molecular marker combination includes SNP1 to SNP19, and the nucleotide sequences thereof are shown as SEQ ID NO.39 to SEQ ID NO.57, and the SNPs are located at the 201st position of SEQ ID NO.39 to SEQ ID NO.57, respectively.
2. A primer for detecting the SNP molecular marker combination according to claim 1, characterized in that: The primers include primer pair 1 to primer pair 19, and the sequences of primer pair 1 to primer pair 19 are shown in the following table: 。 3. A kit for detecting the SNP molecular marker combination according to claim 1, characterized in that: The kit comprises the primers according to claim 2.
4. A chip for detecting the SNP molecular marker combination according to claim 1, characterized in that: The chip comprises the primers according to claim 2.
5. A method for constructing a dragon fruit fingerprint, characterized in that: The following steps are included: 1) Collecting tissue samples of samples; 2) Extracting total DNA of samples; 3) Performing PCR amplification using the primers described in claim 2 and sequencing the amplified products; 4) Identifying the genotype of the locus where the SNP molecular marker combination described in claim 1 is located on the genome of the sample; 5) Constructing a pitaya fingerprint map using the genotype identified in step 4).
6. The method according to claim 5, characterized in that The PCR amplification in step 3 is multiplex PCR amplification.
7. A method for identifying dragon fruit varieties, characterized in that: The following steps are included: 1) Collecting tissue samples of samples to be identified; 2) Extracting total DNA of samples to be identified; 3) Identifying the genotype of the locus where the SNP molecular marker combination as described in claim 1 is located on the genome of the samples to be identified; 4) Comparing the individual genotypes of the samples to be identified with the fingerprint map constructed in claim 5; 5) Determining the variety of the samples to be identified based on the comparison results of step 4).
8. Use of the SNP molecular marker combination according to claim 1, the primer according to claim 2, the kit according to claim 3 and / or the chip according to claim 4 in identifying pitaya varieties.
9. Use of the SNP molecular marker combination according to claim 1, the primer according to claim 2, the kit according to claim 3 and / or the chip according to claim 4 in constructing a SNP fingerprint of pitaya.
Citation Information
Patent Citations
Construction method of pitaya germplasm resource SSR fingerprint database and identification system
CN111876515A
Rice whole genome breeding chip and application thereof
WO2014121419A1