SNP (Single Nucleotide Polymorphism) site combination for identifying papaya germplasm resources and application
By using 1200 SNP sites combination and molecular probes, the accuracy of papaya germplasm resource identification was solved, and high-resolution papaya germplasm resource identification and genetic relationship analysis were achieved, supporting the establishment of a molecular ID card system.
Patent Information
- Application Number
- CN202510600085.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-07-22
AI Technical Summary
It is difficult for the prior art to accurately identify the quality of papaya medicinal materials of different origins and varieties, and traditional molecular technology is not accurate enough.
The combination of 1,200 high-quality SNP sites based on the papaya reference genome PRJNA890604 was used to construct molecular ID cards, genetic diversity analysis and cluster analysis, and combined with TaqMan probe, molecular beacon probe, CRISPR-Cas12a/crRNA complex probe, and fluorescently labeled oligonucleotide probe, the SNP site genotype was detected by fluorescence quantitative PCR or gene chip hybridization.
High-resolution identification of papaya germplasm resources is achieved, clear location information and physical location are provided, and the molecular ID card system and genetic relationship analysis of papaya germplasm resources is supported.
Smart Images

Figure CN120350155A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of molecular biology, and particularly relates to a SNP locus combination for identifying papaya germplasm resources and its application. Background Art
[0002] Papaya is a commonly used Chinese medicinal material, which has the effects of relaxing tendons and activating collaterals, and regulating the stomach and resolving dampness. The 2020 edition of the Pharmacopoeia of the People's Republic of China stipulates that papaya is derived from the dried nearly ripe fruits of the Rosaceae plant Chaenomeles speciosa (Sweet) Nakai. There are four major papaya production areas in the country, namely Xuancheng City, Anhui Province, Changyang County, Hubei Province, Qijiang District, Chongqing City, and Chun'an County, Zhejiang Province, which are commonly known as "Xuan papaya", "Ziqiu papaya", "Chuan papaya", and "Chun papaya" respectively. Xuan papaya, Ziqiu papaya, and Chun papaya are all derived from Chaenomeles speciosa, while Chuan papaya is derived from Chaenomeles speciosa and Chaenomeles cathayensis (Hemsl.) C.K.Schneid. During the long-term cultivation and propagation process in each production area, many varieties have emerged. For Chaenomeles speciosa, Xuan papaya has four varieties, Ziqiu papaya has three varieties, Chun papaya has one variety, and Chuan papaya has one variety. The medicinal qualities of papaya from different origins and varieties are different. Therefore, it is of great significance to identify papaya from different sources for the quality of medicinal papaya, while traditional molecular techniques are not accurate enough for identifying medicinal papaya varieties.
[0003] With the development of genomics, single nucleotide polymorphism (SNP) loci can provide precise genetic information for plant variety identification, effectively solving the deficiencies of traditional identification methods. Compared with other genetic markers, SNP loci have a higher frequency of occurrence in the genome, have stable heredity, and can exist in the coding region, occasionally causing changes in the amino acids of the encoded polypeptide. Among different genomic variations, SNP loci are the most widely used markers for revealing crop genetic insights because they are precise, easy to automate, and are a highly reliable marker system at present. However, there is currently no method for identifying medicinal papaya varieties using SNPs. Summary of the Invention
[0004] In view of the above problems, on the first aspect, the present invention provides a SNP molecular locus combination for identifying papaya germplasm resources. The SNP molecular locus combination contains 1200 SNP loci, and the information of the SNP loci is shown in Table 2.
[0005] Furthermore, the physical position information of the SNP loci is determined based on the papaya reference genome PRJNA890604.
[0006] On the second aspect, the present invention provides the application of the SNP molecular locus combination in any one of the following:
[0007] Applications in the identification of papaya germplasm resources;
[0008] Applications in the cluster analysis of papaya;
[0009] Applications in the identification of papaya genetic relationships;
[0010] Applications in constructing DNA fingerprint maps of papaya germplasm resources;
[0011] Applications in identifying the genetic purity of papaya hybrid offspring;
[0012] Applications in establishing a molecular identification card system for papaya varieties.
[0013] Thirdly, the present invention provides a molecular probe combination for identifying papaya germplasm resources. The molecular probe combination includes the probe combination specifically binding to the 1200 SNP loci, and is used for detecting SNP locus combinations. Exemplarily, the probes can be selected from at least one of the following types: TaqMan probes, molecular beacon probes, CRISPR-Cas12a / crRNA complex probes, and fluorescently labeled oligonucleotide probes.
[0014] Fourthly, the present invention provides a gene chip for identifying papaya germplasm resources, and the gene chip is loaded with the above-mentioned molecular probe combination.
[0015] Fifthly, the present invention provides an identification system for papaya germplasm resources, and the system includes the above-mentioned molecular probe combination.
[0016] Furthermore, the identification system further includes: papaya genomic DNA extraction reagents, nucleic acid amplification reagents (including Taq DNA polymerase, dNTPs and buffer), and a fluorescence signal detection device.
[0017] Sixthly, the present invention provides a kit for identifying papaya germplasm resources, which includes at least one of the above-mentioned molecular probe combination, the above-mentioned gene chip, and the above-mentioned identification system.
[0018] Seventhly, the present invention provides a method for identifying papaya germplasm resources. The method uses the above-mentioned molecular probe combination, the above-mentioned gene chip or the above-mentioned identification system to identify a papaya sample to be identified. Specifically, the identification method is as follows:
[0019] (a) Extract genomic DNA from the papaya sample to be tested;
[0020] (b) Using the above-mentioned identification system, detect the genotypes of the 1200 SNP loci by fluorescence quantitative PCR or gene chip hybridization;
[0021] (c) Based on the genotyping results of SNP loci, the genetic relationships of papaya germplasms were determined by cluster analysis or principal component analysis (PCA).
[0022] Advantages of the present invention:
[0023] Based on the whole-genome resequencing data of 32 papaya germplasm resources, 1200 SNP loci were obtained through alignment analysis, and there was no linkage relationship among the 1200 SNP loci, which had a high resolution for the identification of papaya germplasm resources; moreover, the allele variation distribution frequencies of each locus in the SNP locus combination of this application were relatively uniform among the collected papaya germplasm resource materials, and had clear positioning information and physical location information.
[0024] Other features and advantages of the present invention will be described in the following specification, and part of them will be obvious from the specification, or can be understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures pointed out in the specification, claims and drawings. Description of the drawings
[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0026] Figure 1 Shows the phylogenetic tree diagram drawn based on the 1200 SNP loci screened in the embodiments of the present invention. Detailed embodiments
[0027] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0028] It should be noted that the SNPs referred to in the present invention mainly refer to DNA sequence polymorphisms caused by single nucleotide variations at the genomic level, and the single nucleotide variations include variations caused by single base transitions, transversions, insertions or deletions. The molecular markers referred to in the present invention are all heritable and detectable DNA sequences or proteins, including but not limited to molecular markers based on molecular hybridization, such as RFLP, MinisatelliteDNA; molecular markers based on PCR technology, such as RAPD, STS, SSR and SCAR; DNA markers based on restriction enzyme digestion and PCR technology; molecular markers based on DNA chip technology, such as SNP; analysis marker technology developed based on EST databases, etc. The probe referred to in the present invention is a nucleic acid sequence (DNA or RNA) with a detection label, known in sequence, and complementary to the target gene.
[0029] Example 1
[0030] This example presents a method for papaya DNA extraction and genomic library construction:
[0031] (1) Use the Plant Genomic DNA Kit (TIANGEN BIOTECH (BEIJING) CO., LTD.) kit to extract DNA from silica gel-dried papaya leaves. Determine the purity and concentration of DNA by a NanoDrop 2000 spectrophotometer (Thermo Fisher Scientific, USA);
[0032] (2) DNA quality inspection: Use 1% agarose gel electrophoresis to analyze the purity and integrity of DNA; use Qubit to accurately quantify the DNA concentration; use Agilent 2100 to accurately detect the integrity. The standard for qualified quality inspection is DNA with a total amount of not less than 4 μg, a sample concentration of not less than 40 ng / μL, good sample integrity, high purity, and no impurity contamination;
[0033] (3) Library construction: Use an ultrasonic disruptor to randomly fragment the qualified DNA, electrophoretically recover the DNA fragments of the required length, and add adapters to their ends to form a library;
[0034] (4) Sequencing: Use an Illumina sequencer and related reagents for sequencing.
[0035] The specific information of the 32 germplasm resources involved in this example is shown in Table 1:
[0036] Table 1 Papaya germplasm resource information
[0037]
[0038] Example 2
[0039] This example proposes a method for identifying papaya SNP sites, and the specific process is as follows:
[0040] (1) The sequencing reads of each test material obtained by sequencing in Example 1 were aligned to the reference genome PRJNA890604 (NCBI, He S, Weng D, Zhang Y, et al. A telomere-to-telomere reference genome provides genetic insight into the pentacyclictriterpenoid biosynthesis in Chaenomeles speciosa[J]. Horticulture research, 2023, 10(10):uhad183.) using bwa mem software to obtain high-quality clean reads.
[0041] (2) Based on the positioning results of clean reads in the reference genome PRJNA890604, GATK (version: 4.3.0.0) was used to detect single nucleotide polymorphisms (SNPs) and obtain the original SNP sites.
[0042] Example 3
[0043] This example proposes a method for screening papaya SNP sites, and the specific process is as follows:
[0044] (1) Analysis was performed based on the original SNP sites obtained in Example 2. The SNP sites were called using the HaplotypeCaller module in GATK (Genomeanalysis toolkit, 4.3.0.0), and the following parameters were used for preliminary screening and filtering: QD<2.0||MQ<40.0||FS>60.0||SOR>3.0||MQRankSum<-12.5||ReadPosRankSum<-8.0, and preliminary SNP sites were obtained.
[0045] (2) Use bcftools to eliminate variants with double allele numbers and output variant sites of single nucleotide polymorphisms (command -m2-M2), and use bcftools to calculate the number of SNPs in each step (command: +counts).
[0046] (3) Then, use the plink (v1.90b6.21 64-bit) software to perform secondary screening and filtering on the preliminary SNP sites in the obtained preliminary filtered vcf file, and screen out the sites with a minor allele frequency (MAF) > 0.2 (command: --maf 0.2), retain the sites with a missing rate (geno) < 0.05 (command: --geno 0.05), screen out the sites with a heterozygosity rate (het) < 0.05 (command: awk '$7<0.05 {print}'), discard the sites with a polymorphism information content (PIC) less than 0.3 (command: INFO / PIC >= 0.3), and retain the sites with consistent genotypes of repeated samples (command: --extract). Meeting all the above conditions is the screening condition.
[0047] In this example, 1200 SNP sites were screened out, and the specific information of the SNP sites is shown in Table 2:
[0048] Table 2
[0049]
[0050]
[0051]
[0052]
[0053]
[0054] It should be noted that in Table 1, LG17_20982331 represents the position 20982331 on the 17th linkage group of papaya, and the variation information C, T indicates the existence of a C→T variation. Other names and variation information are similar.
[0055] Example 4
[0056] This example presents the application of SNP sites in the cluster analysis of papaya:
[0057] Cluster analysis is a dendrogram or tree that describes the evolutionary order among populations and is used to represent the evolutionary relationships among populations. The degree of genetic relationship among them can be inferred based on the commonalities or differences in aspects such as the physical or genetic characteristics of the populations. Using the sequencing data obtained in Example 1, with Chaenomeles cathayensis in Sichuan papaya as the outgroup (Chaenomeles cathaye nsis_ChongQing_Fenhua in the figure), use the software MEGA X (V10.0.5) to draw the phylogenetic tree (see Figure 1). As can be seen in the figure, all varieties of wrinkled papaya (Chaenomeles speciosa) are clustered into a single branch, indicating that the 1200 SNP loci can be used to identify various varieties of wrinkled papaya. Among them, Chun papaya (Chaenomeles speciosa_Chunan_Fenhua in the picture) is a single large branch; Luohanqi (Chaenomeles speciosa_XuanCheng Fenhua LuoHanQ i1, 2, 3 in the picture; 1, 2, 3 are 3 trees of the same variety) and Zhimadian (Chae nomeles speciosa_XuanCheng Fenhua ZhiMaDian in the picture) are each gathered into a small branch and are sister groups to each other; the apple-shaped (Chaenomeles speciosa_XuanCheng Fenhua PingGuoXing in the picture) and Yao papaya (Chaenomeles speciosa_XuanCheng HonghuaYaoMuGua in the picture) are both single branches; Ziqiu papaya is gathered into a single large branch, among which pink flower-long strip (Chaenomeles speciosa_HuBei_Fenhua_ChangTiaoXing in the picture), pink flower-apple-shaped (Chaenomeles speciosa_HuBei_Fenhua_ChangTiaoXing in the picture) speciosa_HuBei_Fenhua_PingGuoXing) and red flower (Chaenomeles speciosa_HuBei_Honghua in the picture) are clustered into a small branch and are sister groups to each other; the wrinkled papaya (Chaenomeles speciosa_ChongQing_Honghua) germplasm in Sichuan papaya is clustered into a small branch.
[0058] Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent substitutions for some of the technical features therein; and these modifications or substitutions do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. SNP molecular locus combination for identifying papaya germplasm resources, characterized in that, The SNP molecular locus combination described above contains 1200 SNP loci, and the information of the SNP loci is shown in Table 2.
2. The SNP molecular locus combination for identifying papaya germplasm resources according to claim 1, wherein The physical position information of the SNP loci is determined based on the papaya reference genome PRJNA890604.
3. Use of the SNP molecular locus combination according to claim 1 or 2 in the identification of papaya germplasm resources.
4. Use of the SNP molecular locus combination according to claim 1 or 2 in the cluster analysis of papaya.
5. Use of the SNP molecular locus combination according to claim 1 or 2 in the identification of the genetic relationship of papaya.
6. A molecular probe combination for identifying papaya germplasm resources, characterized in that, The molecular probe combination contains probes that specifically bind to the 1200 SNP loci in claim 1 or 2 and is used to detect the SNP locus combination in claim 1 or 2.
7. A gene chip for identifying papaya germplasm resources, characterized in that, The gene chip is loaded with the molecular probe combination according to claim 6.
8. An identification system for papaya germplasm resources, characterized in that, The identification system includes the molecular probe combination according to claim 6 or the gene chip according to claim 7.
9. A papaya germplasm resource identification kit, characterized in that, Comprising at least one of the molecular probe combination according to claim 6, the gene chip according to claim 7, and the identification system according to claim 8.
10. A method for identifying papaya germplasm resources, characterized in that, The papaya sample to be identified is identified using the molecular probe combination according to claim 6 and the gene chip according to claim 7.