A method for identifying varieties of apricot germplasm resources

Through high-throughput sequencing and filtering and screening of SNP markers, the problems of environmental and human factors in apricot variety identification are solved, the scientific classification and management of apricot germplasm resources are realized, and theoretical support for the protection of apricot variety resources is provided.

CN115232886BActive Publication Date: 2025-08-15SOUTH CHINA UNIV OF TECH
View PDF 24 Cites 0 Cited by

Patent Information

Application Number
CN202210969124.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-12
Publication Date
2025-08-15
Estimated Expiration
2042-08-12

AI Technical Summary

Technical Problem

In the prior art, the apricot variety identification method is greatly affected by environmental and human factors, and it is difficult to accurately classify and manage. SSR technology has defects in the identification of apricot germplasm resources.

Method used

The filtering and screening methods of high-throughput sequencing and SNP markers are adopted to screen the optimal SNP markers combination through whole-genome sequencing, chloroplast genome assembly, ITS gene sequence assembly and SNP specific site location to achieve scientific classification and management of apricot germplasm resources.

Benefits of technology

It realizes accurate and rapid identification and classification of apricot germplasm resources, provides theoretical support for the protection of apricot variety resources, can distinguish the kinship relationship of apricot cultivars and build a unique molecular identity card for the variety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115232886B_ABST
    Figure CN115232886B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for identifying varieties of apricot germplasm resources. The present invention provides a detailed method for identifying varieties of apricot germplasm resources, and provides a large amount of theoretical support for enriching the information on the protection of apricot variety resources in my country. The present invention has clarified the most suitable molecular marker method for apricot plants through a large number of experimental studies. By determining the SNP sites related to the species of apricot plants, the specific species of unknown apricot germplasm samples can be accurately and quickly determined, and if there are new varieties added, new SNP marker combinations can be more conveniently filtered and screened out. The identification method provided by the present invention can facilitate the classification and management of apricot plants in an efficient manner. In the field of variety identification or breeding based on molecular biology and bioinformatics, the optimal and simplest combination is found mainly by screening the SNP marker combination to achieve the purpose of quickly and accurately identifying Xinjiang apricot cultivars and distinguishing the kinship of Xinjiang apricot cultivars.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of plant variety identification, and particularly relates to a variety identification method for apricot germplasm resources. Background Art

[0002] Apricot varieties are diverse, and their unique growing environment and long-term evolution have resulted in a rich and valuable resource. Fully exploring the nutritional value of different apricot varieties is an urgent task, hoping to provide a scientific basis for the selection of improved varieties, species diversity, and sustainable development of my country's apricot industry.

[0003] With the continuous acceleration of apricot variety breeding in my country, the number of cultivated apricot varieties is increasing. Therefore, the effective and accurate classification and identification of different cultivated apricot varieties is an urgent problem to be solved. Currently, the identification of cultivated apricot varieties in Xinjiang is still limited to traditional methods such as tree morphology, flower form, fruit appearance, and fruit maturity, which are greatly affected by environmental and human factors.

[0004] Most of the apricots cultivated worldwide originated in China. my country has a large planting area for apricot trees and a rich variety of resources.

[0005] Simple sequence repeats (SSRs), also known as microsatellites, have a tandemly repeated core sequence of 1-6 bp. The most common are dinucleotide repeats, namely (CA)n and (TG)n. Each microsatellite DNA core sequence has an identical structure, with 10-60 repeat units. Their high polymorphism stems primarily from the varying number of tandem repeats. The basic principle of SSR markers is to amplify microsatellite fragments using PCR using primers designed to complement the sequences at either end of the microsatellite sequence. The varying number of tandem repeats in the core sequence allows for the generation of PCR products of varying lengths. The amplified products are then subjected to gel electrophoresis, and the size of the separated fragments is used to determine genotype and calculate allele frequencies. However, SSR technology, when analyzing microsatellite DNA polymorphism, requires knowledge of the DNA sequences at both ends of the repeat sequence. Therefore, it presents significant limitations for the identification and management of apricot germplasm resources, which are diverse and complex. Summary of the Invention

[0006] The purpose of the present invention is to solve the problems existing in the prior art of apricot variety identification, so as to obtain the optimal combination of SNP markers through high-throughput sequencing and filtering and screening of SNP markers, so as to more efficiently and accurately distinguish and identify different apricot cultivars and realize the scientific classification and management of apricot resources.

[0007] In order to solve the above technical problems, the present invention is achieved through the following technical solutions.

[0008] The present invention provides a method for identifying varieties of apricot germplasm resources, comprising the following steps:

[0009] (1) Collect apricot samples and extract total DNA;

[0010] (2) Whole genome sequencing of the extracted total DNA;

[0011] (3) Analyzing the sequencing results to identify the apricot variety, wherein the analysis is selected from one or more of the following methods:

[0012] a. Based on chloroplast genome assembly;

[0013] b. Assembly based on ITS gene sequences; these two sequences were not used

[0014] c. Based on SNP-specific site positioning.

[0015] Preferably, the apricots in step (1) are selected from Xinjiang apricots.

[0016] Preferably, the sample in step (1) is selected from one or more of apricot leaves, fruits, roots, and stems; most preferably, the sample is selected from apricot leaves.

[0017] Preferably, in step (1), the CTBA method is used to extract total DNA.

[0018] As a preference, after the total DNA is extracted in step (1), the DNA concentration and purity should be determined to control the DNA OD 260 / OD 280 It is 1.7-2.0.

[0019] Preferably, the analysis in step (3) is most preferably based on SNP-specific site mapping.

[0020] Preferably, the chloroplast genome assembly in step (3) specifically includes the following steps:

[0021] ① Based on the whole genome sequencing data in step (2), filter out the chloroplast genome data, and use Velvet software to assemble a complete circular whole chloroplast genome, and correct some incompletely assembled regions;

[0022] ② Use Plastome annotator to annotate the target genome and correct the gene start and stop codons and intron / exon boundaries reported by the software;

[0023] ③ Using Geneious software, the total length of the chloroplast genome and the length of each region including LSC, SSC, and IR, gene composition, base composition, and GC or AT content are counted and compared; the genes include one or more of protein-coding genes, transcribed RNA genes, ribosomal RNA genes, introns, and exons;

[0024] ④ Use the online software Reputer (https: / / bibiserv.cebitec.uni-bielefeld.de / repuer) to identify four types of repetitive sequences in the chloroplast genome: forward repeats, reverse repeats, complementary repeats, and palindrome repeats;

[0025] ⑤Use mVISTA software to perform sequence alignment and variation analysis of the chloroplast genome, and use IRscope online software to map and analyze the contraction and expansion of the inverted repeat regions (IRs) of the chloroplast genome;

[0026] ⑥ Using MrBayes v.3.2.1 software, we selected assembled and annotated chloroplast whole genome sequences, took almond (Prunus pedunculata) as the outgroup, performed Bayesian analysis, constructed a Bayesian (BI) phylogenetic tree, and analyzed the phylogeny among different accessions. At the same time, we used Dnasp software to calculate the nucleotide diversity of each variety and detect the highly variable regions in the apricot chloroplast genome as the basis for variety identification.

[0027] Preferably, the ITS gene sequence-based assembly in step (3) specifically includes the following steps:

[0028] ① Based on the whole genome sequencing data in step (2), the ITS sequences of all germplasms were extracted, and sequence alignment and comparative analysis were performed using MEGA7.0 software to count the effective informative sites;

[0029] ② Construct a Bayesian phylogenetic tree, analyze the phylogeny and genetic relationships of different germplasms, and reconstruct the phylogenetic tree of all varieties of apricot germplasms.

[0030] Preferably, the SNP-specific site positioning in step (3) specifically includes the following steps:

[0031] ① Using the whole-genome sequenced apricot (Prunus armeniaca L.) genome (gene accession number GCA_903112645.1, https: / / www.ncbi.nlm.nih.gov / assembly / GCA_903112645.1# / def) as the reference sequence, extract nuclear genome single nucleotide polymorphisms (SNPs) based on the whole-genome sequencing data in step (2);

[0032] ②BWA (Burrows-Wheeler Aligner) software was used for sequencing, Picard software was used for deduplication, and SAMtools software was used for identification of SNP sites;

[0033] ③Compare the SNP-specific site results to identify the apricot variety.

[0034] Preferably, the SNP specific site is selected from CAEKDK010000001.1_10440254, CAEKDK010000001.1_13568267, CAEKDK010000001.1_20639545, CAEKDK010000001.1_21909800, CAEKDK010000001.1_25411117, CAEKDK010000001.1_36367780, CAEKDK010000001.1_412 51119, CAEKDK010000002.1_18290549, CAEKDK010000002.1_18928945, CAEKDK010000002.1_22722929, CAEKDK0100000 02.1_28105504, CAEKDK010000003.1_675221, CAEKDK010000003.1_15070082, CAEKDK010000003.1_21123450, CAEKDK0 10000003.1_25439844、CAEKDK010000004.1_7586379、CAEKDK010000004.1_19240527、CAEKDK010000004.1_22531282、 CAEKDK010000005.1_12669636, CAEKDK010000005.1_13466650, CAEKDK010000005.1_17572886, CAEKDK010000006.1_4 One or more of 798304, CAEKDK010000006.1_5750123, CAEKDK010000006.1_20454043, CAEKDK010000006.1_21304403, CAEKDK010000007.1_10365019, CAEKDK010000008.1_5584589, CAEKDK010000008.1_7399902, and CAEKDK010000008.1_15579729.

[0035] It should be understood that in the context of the present invention, unless otherwise specified, the software used, such as Velvet, Plastome, Reputer, mVISTA, MrBayes, IRscope, Dnasp, MEGA7.0, Burrows-WheelerAligner, Picard, SAMtools, etc., are all software or tools known or commonly used in the art. Those skilled in the art are capable of using them normally according to the instructions or user guides provided by the software suppliers, and screening for relevant genes or sequences according to the parameters or conditions provided by the software.

[0036] Single nucleotide polymorphisms (SNPs) refer to variations in a single nucleotide in the genome, including substitutions, transversions, deletions, and insertions. By detecting differences in single nucleotides at the molecular level, SNP markers can help distinguish genetic differences between two individuals. SNPs are considered the third generation of DNA molecular marker technology and are expected to become the most important and effective molecular marker.

[0037] Given the stability and effectiveness of the method, the International Union for the Protection of Plant Varieties Rights (UPOV) has identified SSR and SNP as marker methods for constructing DNA fingerprint databases in its draft BMT testing guidelines. Compared to SSR markers, SNPs are more targeted, have a richer source of variation, and have a potentially large number of potential markers. To date, the total number of SNP markers in cultivated apricot has reached over 19 million. Because most SNPs only have two base forms, they are considered diallelic markers, which has the advantage of simplifying genotyping methods and subsequent digital encoding. Moreover, SNPs are variations in the nucleotides themselves, and there is no need to read the molecular size. Sequencing methods can directly obtain allelic variant genotypes, which have high genetic stability and are more suitable for digital database construction. The use of SNP molecular markers to identify the molecular identity of crops will undoubtedly play a positive role in the management of crop variety resources, the protection, utilization, and evaluation of crop variety resources.

[0038] Compared with the prior art, the present invention has the following technical effects:

[0039] (1) This invention provides a comprehensive method for identifying apricot germplasm resources, which provides a lot of theoretical support for enriching the protection information of apricot variety resources in my country.

[0040] (2) The present invention has identified the most suitable molecular marker method for Apricot plants through a large number of experimental studies. By determining the SNP sites associated with Apricot plant species, the specific species of unknown apricot germplasm samples can be accurately and quickly determined. If new varieties are added, new SNP marker combinations can be more conveniently filtered and screened.

[0041] (3) The identification method provided by the present invention can facilitate the efficient classification and management of Apricot plants. In the field of variety identification or breeding based on molecular biology and bioinformatics, the screening of SNP marker combinations is the main method to find the optimal and simplest combination to achieve the purpose of quickly and accurately identifying Xinjiang apricot cultivars and distinguishing the phylogenetic relationships of Xinjiang apricot cultivars. At the same time, these SNP marker combinations can also be used to construct a unique molecular identity card for each variety. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 This is a schematic diagram of the information results of 29 SNP sites in 20 apricot varieties including Kantaile apricot.

[0043] Figure 2 This is a schematic diagram of the information results of 29 SNP sites in 20 apricot varieties including Danxing.

[0044] Figure 3 This is a schematic diagram of the information results of 29 SNP sites in 20 apricot varieties including Wang Enmao No. 1 apricot.

[0045] Figure 4 This is a schematic diagram of the information results of 29 SNP sites in 20 apricot varieties including Mayimake Yulike apricot.

[0046] Figure 5 This is a schematic diagram of the information results of 29 SNP sites in 20 apricot varieties including Akdarez apricot.

[0047] Figure 6 This is a schematic diagram of the information results of 29 SNP sites in 20 apricot varieties including Asibi Ke apricot.

[0048] Figure 7 This is a schematic diagram of the information results of 29 SNP sites in 9 apricot varieties including Akexing. DETAILED DESCRIPTION

[0049] In order to make the purpose, technical solution and effect of the present invention clearer and more specific, the present invention is further described in detail with reference to the following examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0050] Unless otherwise specified, the samples, consumables, and reagents listed or used in the present invention are all commercially available. The experimental methods used in the present invention, such as DNA extraction and whole genome sequencing, are conventional methods and techniques in the art.

[0051] Example 1

[0052] The determination of SNP loci for known apricot varieties includes the following steps:

[0053] (1) Healthy, clean, and fresh apricot leaves of 129 different varieties were collected from the germplasm resource garden. 1-2 g of each sample was taken for total DNA extraction (Plant Genomic DNA Extraction Kit, Beijing Tiangen Biotechnology Co., Ltd.):

[0054] ① Take 200 mg of leaf tissue and place it in a crushing tube, soak it in liquid nitrogen, and grind it into powder using a machine;

[0055] ② Add appropriate amount of CTAB lysis buffer and mix well with a vortex mixer;

[0056] ③ Place in a constant temperature mixer and incubate at 65°C for 60 min;

[0057] ④ After lysis is complete, cool the tube and centrifuge at 10,000 rpm for 5 minutes at room temperature. Aspirate the remaining liquid and transfer it to a 2.0 mL centrifuge tube. Add an equal volume of Tris-saturated phenol / chloroform / isoamyl alcohol mixed solvent (the volume ratio of the three is 25:24:1) and mix by inverting 5-10 times.

[0058] ⑤ Centrifuge at 10,000 rpm for 10 min at room temperature, aspirate the supernatant, and transfer to a 1.5 mL centrifuge tube;

[0059] ⑥ Add 2 / 3 volume of -20℃ pre-cooled isopropanol, mix well by inverting 3-5 times, and place in a -20℃ refrigerator to settle for more than 2 hours;

[0060] ⑦ Centrifuge at 12000 rpm for 15 min at room temperature, aspirate the supernatant, add 750 μL 75% ethanol, and wash the precipitate with a pipette;

[0061] ⑧ Centrifuge at 12000 rpm for 3 minutes at room temperature, aspirate the supernatant, centrifuge briefly to remove the residual liquid, dry for 3-5 minutes, and dissolve in an appropriate amount of TE solution;

[0062] ⑨ DNA concentration and purity were measured using a NanoDrop 1000 spectrophotometer (Themo Scientific) UV spectrophotometer. 260 / OD 280 1.7-2.0 meets the requirements;

[0063] (2) Whole genome sequencing was performed on the extracted total DNA of each sample using an MGISEQ-2000 sequencing instrument with a sequencing depth of 20×. The specific steps included the following:

[0064] ① Sample testing: Sample testing includes sample concentration, integrity, and purity; concentration is determined by fluorescence quantification or microplate reader; and sample integrity and purity are tested using agarose gel (agarose gel concentration: 1%, voltage: 150V, electrophoresis time: 40min);

[0065] ② Sample shearing: Take 1 μg of genomic DNA sample and perform ultrasonic shearing;

[0066] ③ Fragment size selection: Fragment selection is performed on the sheared sample magnetic beads so that the sample bands are concentrated around 200-400bp;

[0067] ④ End repair, adding "A", and adapter ligation: Prepare the reaction system, react at room temperature for a certain period of time, repair the double-stranded cDNA ends, add the A base at the 3' end, prepare the adapter ligation reaction system, react at room temperature for a certain period of time, and then connect the adapter to the DNA;

[0068] ⑤PCR reaction and product recovery: prepare the PCR reaction system and set the reaction program to amplify the ligation product; the amplified product is purified and recovered using magnetic beads;

[0069] ⑥ PCR product cyclization: After the PCR product is denatured into a single strand, a cyclization reaction system is prepared. After sufficient mixing and reaction at room temperature for a period of time, a single-stranded circular product is obtained, and the uncyclized linear DNA molecules are digested;

[0070] ⑦ Library testing and sequencing: The circularized product undergoes a concentration test before being loaded onto the sequencing machine. Qualified libraries are then scheduled for sequencing (DEBSEQ): Single-stranded circular DNA molecules are replicated via roller rings to form a DNA nanoball (DNB) containing over 300 copies. The resulting DNBs are loaded into the mesh pores of a high-density DNA nanochip and sequenced using combined probe anchor polymerization (cPAS). The resulting data are then quality-controlled using FASTQC software.

[0071] (3) Perform SNP analysis on sequencing results:

[0072] ① Using the whole-genome sequenced apricot (Prunus armeniaca L.) genome (gene accession number GCA_903112645.1, https: / / www.ncbi.nlm.nih.gov / assembly / GCA_903112645.1# / def) as the reference sequence, extract nuclear genome single nucleotide polymorphisms (SNPs) based on the whole-genome sequencing data in step (2);

[0073] ②BWA (Burrows-Wheeler Aligner) software was used for sequencing, Picard software was used for deduplication, and SAMtools software was used for identification of SNP sites;

[0074] ③ Compare the SNP site results of each sample to clarify the number and type of SNP site mutations, so as to determine the apricot variety. There are 29 SNP sites related to apricot variety identification, which are as follows:

[0075] 1. CAEKDK010000001.1_10440254, whose deoxynucleotide sequence is A / T;

[0076] 2. CAEKDK010000001.1_13568267, whose deoxynucleotide sequence is A / T;

[0077] 3. CAEKDK010000001.1_20639545, whose deoxynucleotide sequence is C / T;

[0078] 4. CAEKDK010000001.1_21909800, whose deoxynucleotide sequence is A / G;

[0079] 5. CAEKDK010000001.1_25411117, whose deoxynucleotide sequence is C / T;

[0080] 6. CAEKDK010000001.1_36367780, whose deoxynucleotide sequence is A / G;

[0081] 7. CAEKDK010000001.1_41251119, whose deoxynucleotide sequence is A / G;

[0082] 8. CAEKDK010000002.1_18290549, whose deoxynucleotide sequence is C / T;

[0083] 9. CAEKDK010000002.1_18928945, whose deoxynucleotide sequence is A / G;

[0084] 10. CAEKDK010000002.1_22722929, whose deoxynucleotide sequence is C / T;

[0085] 11. CAEKDK010000002.1_28105504, whose deoxynucleotide sequence is C / T;

[0086] 12. CAEKDK010000003.1_675221, whose deoxynucleotide sequence is A / G;

[0087] 13. CAEKDK010000003.1_15070082, whose deoxynucleotide sequence is C / T;

[0088] 14. CAEKDK010000003.1_21123450, whose deoxynucleotide sequence is A / C;

[0089] 15. CAEKDK010000003.1_25439844, whose deoxynucleotide sequence is A / T;

[0090] 16. CAEKDK010000004.1_7586379, whose deoxynucleotide sequence is A / G;

[0091] 17. CAEKDK010000004.1_19240527, whose deoxynucleotide sequence is C / T;

[0092] 18. CAEKDK010000004.1_22531282, whose deoxynucleotide sequence is A / G;

[0093] 19. CAEKDK010000005.1_12669636, whose deoxynucleotide sequence is C / T;

[0094] 20. CAEKDK010000005.1_13466650, whose deoxynucleotide sequence is A / G;

[0095] 21. CAEKDK010000005.1_17572886, whose deoxynucleotide sequence is C / T;

[0096] 22. CAEKDK010000006.1_4798304, whose deoxynucleotide sequence is A / C;

[0097] 23. CAEKDK010000006.1_5750123, whose deoxynucleotide sequence is C / T;

[0098] 24. CAEKDK010000006.1_20454043, whose deoxynucleotide sequence is C / G;

[0099] 25. CAEKDK010000006.1_21304403, whose deoxynucleotide sequence is C / G;

[0100] 26. CAEKDK010000007.1_10365019, whose deoxynucleotide sequence is A / G;

[0101] 27. CAEKDK010000008.1_5584589, whose deoxynucleotide sequence is C / T;

[0102] 28. CAEKDK010000008.1_7399902, whose deoxynucleotide sequence is A / G;

[0103] 29. CAEKDK010000008.1_15579729, whose deoxynucleotide sequence is C / T.

[0104] The comparative analysis data of SNP sites of 129 apricot samples are as follows Figure 1-7 shown.

[0105] Example 2

[0106] A method for identifying apricot varieties comprises the following steps:

[0107] (1) Obtain a healthy, clean and fresh apricot leaf sample, numbered 4-1, and take 2 g of the sample for total DNA extraction (Plant Genomic DNA Extraction Kit, Beijing Tiangen Biotechnology Co., Ltd.):

[0108] ① Take 200 mg of leaf tissue and place it in a crushing tube, soak it in liquid nitrogen, and grind it into powder using a machine;

[0109] ② Add appropriate amount of CTAB lysis buffer and mix well with a vortex mixer;

[0110] ③ Place in a constant temperature mixer and incubate at 65°C for 60 min;

[0111] ④ After lysis is complete, cool the tube and centrifuge at 10,000 rpm for 5 minutes at room temperature. Aspirate the remaining liquid and transfer it to a 2.0 mL centrifuge tube. Add an equal volume of Tris-saturated phenol / chloroform / isoamyl alcohol mixed solvent (the volume ratio of the three is 25:24:1) and mix by inverting 5-10 times.

[0112] ⑤ Centrifuge at 10,000 rpm for 10 min at room temperature, aspirate the supernatant, and transfer to a 1.5 mL centrifuge tube;

[0113] ⑥ Add 2 / 3 volume of -20℃ pre-cooled isopropanol, mix well by inverting 3-5 times, and place in a -20℃ refrigerator to settle for more than 2 hours;

[0114] ⑦ Centrifuge at 12000 rpm for 15 min at room temperature, aspirate the supernatant, add 750 μL 75% ethanol, and wash the precipitate with a pipette;

[0115] ⑧ Centrifuge at 12000 rpm for 3 minutes at room temperature, aspirate the supernatant, centrifuge briefly to remove the residual liquid, dry for 3-5 minutes, and dissolve in an appropriate amount of TE solution;

[0116] ⑨ DNA concentration and purity were measured using a NanoDrop 1000 spectrophotometer (Themo Scientific) UV spectrophotometer. 260 / OD 280 is 1.8;

[0117] (2) Whole genome sequencing was performed on the extracted total DNA of the sample using the MGISEQ-2000 sequencing instrument with a sequencing depth of 20×. The specific steps included:

[0118] ① Sample testing: Sample testing includes sample concentration, integrity, and purity; concentration is determined by fluorescence quantification or microplate reader; and sample integrity and purity are tested using agarose gel (agarose gel concentration: 1%, voltage: 150V, electrophoresis time: 40min);

[0119] ② Sample shearing: Take 1 μg of genomic DNA sample and perform ultrasonic shearing;

[0120] ③ Fragment size selection: Fragment selection is performed on the sheared sample magnetic beads so that the sample bands are concentrated around 200-400bp;

[0121] ④ End repair, adding "A", and adapter ligation: Prepare the reaction system, react at room temperature for a certain period of time, repair the double-stranded cDNA ends, add the A base at the 3' end, prepare the adapter ligation reaction system, react at room temperature for a certain period of time, and then connect the adapter to the DNA;

[0122] ⑤PCR reaction and product recovery: prepare the PCR reaction system and set the reaction program to amplify the ligation product; the amplified product is purified and recovered using magnetic beads;

[0123] ⑥ PCR product cyclization: After the PCR product is denatured into a single strand, a cyclization reaction system is prepared. After sufficient mixing and reaction at room temperature for a period of time, a single-stranded circular product is obtained, and the uncyclized linear DNA molecules are digested;

[0124] ⑦ Library testing and sequencing: The circularized product undergoes a concentration test before being loaded onto the sequencing machine. Qualified libraries are then scheduled for sequencing (DEBSEQ): Single-stranded circular DNA molecules are replicated via roller rings to form a DNA nanoball (DNB) containing over 300 copies. The resulting DNBs are loaded into the mesh pores of a high-density DNA nanochip and sequenced using combined probe anchor polymerization (cPAS). The resulting data are then quality-controlled using FASTQC software.

[0125] (3) Perform SNP analysis on sequencing results:

[0126] ① Using the whole-genome sequenced apricot (Prunus armeniaca L.) genome (gene accession number GCA_903112645.1, https: / / www.ncbi.nlm.nih.gov / assembly / GCA_903112645.1# / def) as the reference sequence, extract nuclear genome single nucleotide polymorphisms (SNPs) based on the whole-genome sequencing data in step (2);

[0127] ②BWA (Burrows-Wheeler Aligner) software was used for sequencing, Picard software was used for deduplication, and SAMtools software was used for identification of SNP sites;

[0128] ③The SNP site information of this sample is as follows:

[0129] 1. CAEKDK010000001.1_10440254, whose deoxynucleotide sequence is A / A;

[0130] 2. CAEKDK010000001.1_13568267, whose deoxynucleotide sequence is T / A;

[0131] 3. CAEKDK010000001.1_20639545, whose deoxynucleotide sequence is C / T;

[0132] 4. CAEKDK010000001.1_21909800, whose deoxynucleotide sequence is A / G;

[0133] 5. CAEKDK010000001.1_25411117, whose deoxynucleotide sequence is T / T;

[0134] 6. CAEKDK010000001.1_36367780, whose deoxynucleotide sequence is G / G;

[0135] 7. CAEKDK010000001.1_41251119, whose deoxynucleotide sequence is A / A;

[0136] 8. CAEKDK010000002.1_18290549, whose deoxynucleotide sequence is C / T;

[0137] 9. CAEKDK010000002.1_18928945, whose deoxynucleotide sequence is A / G;

[0138] 10. CAEKDK010000002.1_22722929, whose deoxynucleotide sequence is T / C;

[0139] 11. CAEKDK010000002.1_28105504, whose deoxynucleotide sequence is T / T;

[0140] 12. CAEKDK010000003.1_675221, whose deoxynucleotide sequence is G / A;

[0141] 13. CAEKDK010000003.1_15070082, whose deoxynucleotide sequence is T / T;

[0142] 14. CAEKDK010000003.1_21123450, whose deoxynucleotide sequence is C / A;

[0143] 15. CAEKDK010000003.1_25439844, whose deoxynucleotide sequence is A / A;

[0144] 16. CAEKDK010000004.1_7586379, whose deoxynucleotide sequence is A / G;

[0145] 17. CAEKDK010000004.1_19240527, whose deoxynucleotide sequence is C / T;

[0146] 18. CAEKDK010000004.1_22531282, whose deoxynucleotide sequence is G / G;

[0147] 19. CAEKDK010000005.1_12669636, whose deoxynucleotide sequence is C / T;

[0148] 20. CAEKDK010000005.1_13466650, whose deoxynucleotide sequence is A / G;

[0149] 21. CAEKDK010000005.1_17572886, whose deoxynucleotide sequence is C / T;

[0150] 22. CAEKDK010000006.1_4798304, whose deoxynucleotide sequence is C / A;

[0151] 23. CAEKDK010000006.1_5750123, whose deoxynucleotide sequence is C / T;

[0152] 24, CAEKDK010000006.1_20454043, whose deoxynucleotide sequence is G / C;

[0153] 25. CAEKDK010000006.1_21304403, whose deoxynucleotide sequence is C / G;

[0154] 26. CAEKDK010000007.1_10365019, whose deoxynucleotide sequence is A / A;

[0155] 27. CAEKDK010000008.1_5584589, whose deoxynucleotide sequence is C / C;

[0156] 28. CAEKDK010000008.1_7399902, whose deoxynucleotide sequence is G / A;

[0157] 29. CAEKDK010000008.1_15579729, whose deoxynucleotide sequence is C / C.

[0158] Comparing the above results with those in Example 1, it can be clearly seen that the species of the apricot sample is: Mayisaimu apricot.

[0159] Example 3

[0160] A method for identifying apricot varieties comprises the following steps:

[0161] (1) Obtain healthy, clean and fresh apricot leaf samples, sample number: 10-4, sample 2 g, and perform total DNA extraction (Plant Genomic DNA Extraction Kit, Beijing Tiangen Biotechnology Co., Ltd.):

[0162] ① Take 200 mg of leaf tissue and place it in a crushing tube, soak it in liquid nitrogen, and grind it into powder using a machine;

[0163] ② Add appropriate amount of CTAB lysis buffer and mix well with a vortex mixer;

[0164] ③ Place in a constant temperature mixer and incubate at 65°C for 60 min;

[0165] ④ After lysis is complete, cool the tube and centrifuge at 10,000 rpm for 5 minutes at room temperature. Aspirate the remaining liquid and transfer it to a 2.0 mL centrifuge tube. Add an equal volume of Tris-saturated phenol / chloroform / isoamyl alcohol mixed solvent (the volume ratio of the three is 25:24:1) and mix by inverting 5-10 times.

[0166] ⑤ Centrifuge at 10,000 rpm for 10 min at room temperature, aspirate the supernatant, and transfer to a 1.5 mL centrifuge tube;

[0167] ⑥ Add 2 / 3 volume of -20℃ pre-cooled isopropanol, mix well by inverting 3-5 times, and place in a -20℃ refrigerator to settle for more than 2 hours;

[0168] ⑦ Centrifuge at 12000 rpm for 15 min at room temperature, aspirate the supernatant, add 750 μL 75% ethanol, and wash the precipitate with a pipette;

[0169] ⑧ Centrifuge at 12000 rpm for 3 minutes at room temperature, aspirate the supernatant, centrifuge briefly to remove the residual liquid, dry for 3-5 minutes, and dissolve in an appropriate amount of TE solution;

[0170] ⑨ DNA concentration and purity were measured using a NanoDrop 1000 spectrophotometer (Themo Scientific) UV spectrophotometer. 260 / OD 280 is 2.0;

[0171] (2) Whole genome sequencing was performed on the extracted total DNA of the sample using the MGISEQ-2000 sequencing instrument with a sequencing depth of 20×. The specific steps included:

[0172] ① Sample testing: Sample testing includes sample concentration, integrity, and purity; concentration is determined by fluorescence quantification or microplate reader; and sample integrity and purity are tested using agarose gel (agarose gel concentration: 1%, voltage: 150V, electrophoresis time: 40min);

[0173] ② Sample shearing: Take 1 μg of genomic DNA sample and perform ultrasonic shearing;

[0174] ③ Fragment size selection: Fragment selection is performed on the sheared sample magnetic beads so that the sample bands are concentrated around 200-400bp;

[0175] ④ End repair, adding "A", and adapter ligation: Prepare the reaction system, react at room temperature for a certain period of time, repair the double-stranded cDNA ends, add the A base at the 3' end, prepare the adapter ligation reaction system, react at room temperature for a certain period of time, and then connect the adapter to the DNA;

[0176] ⑤PCR reaction and product recovery: prepare the PCR reaction system and set the reaction program to amplify the ligation product; the amplified product is purified and recovered using magnetic beads;

[0177] ⑥ PCR product cyclization: After the PCR product is denatured into a single strand, a cyclization reaction system is prepared. After sufficient mixing and reaction at room temperature for a period of time, a single-stranded circular product is obtained, and the uncyclized linear DNA molecules are digested;

[0178] ⑦ Library testing and sequencing: The circularized product undergoes a concentration test before being loaded onto the sequencing machine. Qualified libraries are then scheduled for sequencing (DEBSEQ): Single-stranded circular DNA molecules are replicated via roller rings to form a DNA nanoball (DNB) containing over 300 copies. The resulting DNBs are loaded into the mesh pores of a high-density DNA nanochip and sequenced using combined probe anchor polymerization (cPAS). The resulting data are then quality-controlled using FASTQC software.

[0179] (3) Perform SNP analysis on sequencing results:

[0180] ① Using the whole-genome sequenced apricot (Prunus armeniaca L.) genome (gene accession number GCA_903112645.1, https: / / www.ncbi.nlm.nih.gov / assembly / GCA_903112645.1# / def) as the reference sequence, extract nuclear genome single nucleotide polymorphisms (SNPs) based on the whole-genome sequencing data in step (2);

[0181] ②BWA (Burrows-Wheeler Aligner) software was used for sequencing, Picard software was used for deduplication, and SAMtools software was used for identification of SNP sites;

[0182] ③The SNP site information of this sample is as follows:

[0183] 1. CAEKDK010000001.1_10440254, whose deoxynucleotide sequence is T / A;

[0184] 2. CAEKDK010000001.1_13568267, whose deoxynucleotide sequence is T / A;

[0185] 3. CAEKDK010000001.1_20639545, whose deoxynucleotide sequence is C / C;

[0186] 4. CAEKDK010000001.1_21909800, whose deoxynucleotide sequence is A / A;

[0187] 5. CAEKDK010000001.1_25411117, whose deoxynucleotide sequence is C / T;

[0188] 6. CAEKDK010000001.1_36367780, whose deoxynucleotide sequence is G / A;

[0189] 7. CAEKDK010000001.1_41251119, whose deoxynucleotide sequence is A / A;

[0190] 8. CAEKDK010000002.1_18290549, whose deoxynucleotide sequence is C / T;

[0191] 9. CAEKDK010000002.1_18928945, whose deoxynucleotide sequence is N / N;

[0192] 10. CAEKDK010000002.1_22722929, whose deoxynucleotide sequence is T / C;

[0193] 11. CAEKDK010000002.1_28105504, whose deoxynucleotide sequence is T / T;

[0194] 12. CAEKDK010000003.1_675221, whose deoxynucleotide sequence is A / A;

[0195] 13. CAEKDK010000003.1_15070082, whose deoxynucleotide sequence is T / C;

[0196] 14. CAEKDK010000003.1_21123450, whose deoxynucleotide sequence is C / A;

[0197] 15. CAEKDK010000003.1_25439844, whose deoxynucleotide sequence is A / A;

[0198] 16. CAEKDK010000004.1_7586379, whose deoxynucleotide sequence is N / N;

[0199] 17. CAEKDK010000004.1_19240527, whose deoxynucleotide sequence is C / T;

[0200] 18. CAEKDK010000004.1_22531282, whose deoxynucleotide sequence is G / A;

[0201] 19. CAEKDK010000005.1_12669636, whose deoxynucleotide sequence is C / C;

[0202] 20. CAEKDK010000005.1_13466650, whose deoxynucleotide sequence is G / G;

[0203] 21. CAEKDK010000005.1_17572886, whose deoxynucleotide sequence is C / T;

[0204] 22. CAEKDK010000006.1_4798304, whose deoxynucleotide sequence is C / A;

[0205] 23. CAEKDK010000006.1_5750123, whose deoxynucleotide sequence is C / T;

[0206] 24. CAEKDK010000006.1_20454043, whose deoxynucleotide sequence is G / G;

[0207] 25. CAEKDK010000006.1_21304403, whose deoxynucleotide sequence is C / G;

[0208] 26. CAEKDK010000007.1_10365019, whose deoxynucleotide sequence is G / G;

[0209] 27. CAEKDK010000008.1_5584589, whose deoxynucleotide sequence is T / T;

[0210] 28. CAEKDK010000008.1_7399902, whose deoxynucleotide sequence is G / A;

[0211] 29. CAEKDK010000008.1_15579729, whose deoxynucleotide sequence is T / T.

[0212] By comparing the above results with those in Example 1, it can be clearly seen that the species of the apricot sample is: Keqiketuman apricot.

[0213] The above detailed description of the analytical methods involved in the present invention provides a detailed introduction. It should be noted that the above description is intended solely to help those skilled in the art better understand the methods and concepts of the present invention, and is not intended to limit the relevant content. Without departing from the principles of the present invention, those skilled in the art may make appropriate adjustments or modifications to the present invention, and such adjustments and modifications shall also fall within the scope of protection of the present invention.

Claims

1. A method for identifying varieties of apricot germplasm resources, characterized in that: The steps include: (1) Collect apricot samples and extract total DNA; (2) Whole genome sequencing of the extracted total DNA; (3) Using the Prunus armeniaca L. apricot genome, which has been fully sequenced, as a reference sequence, the sequencing results were analyzed based on SNP-specific site mapping to identify apricot varieties; The gene accession number of the Prunus armeniaca L. apricot genome is GCA_903112645.1; There are 29 SNP-specific sites, as follows:

1. CAEKDK010000001.1_10440254, whose deoxynucleotide sequence is A / T; 2. CAEKDK010000001.1_13568267, whose deoxynucleotide sequence is A / T; 3. CAEKDK010000001.1_20639545, whose deoxynucleotide sequence is C / T; 4. CAEKDK010000001.1_21909800, whose deoxynucleotide sequence is A / G; 5. CAEKDK010000001.1_25411117, whose deoxynucleotide sequence is C / T; 6. CAEKDK010000001.1_36367780, whose deoxynucleotide sequence is A / G; 7. CAEKDK010000001.1_41251119, whose deoxynucleotide sequence is A / G; 8. CAEKDK010000002.1_18290549, whose deoxynucleotide sequence is C / T; 9. CAEKDK010000002.1_18928945, whose deoxynucleotide sequence is A / G; 10. CAEKDK010000002.1_22722929, whose deoxynucleotide sequence is C / T; 11. CAEKDK010000002.1_28105504, whose deoxynucleotide sequence is C / T; 12. CAEKDK010000003.1_675221, whose deoxynucleotide sequence is A / G; 13. CAEKDK010000003.1_15070082, whose deoxynucleotide sequence is C / T; 14. CAEKDK010000003.1_21123450, whose deoxynucleotide sequence is A / C; 15. CAEKDK010000003.1_25439844, whose deoxynucleotide sequence is A / T; 16. CAEKDK010000004.1_7586379, whose deoxynucleotide sequence is A / G; 17. CAEKDK010000004.1_19240527, whose deoxynucleotide sequence is C / T; 18. CAEKDK010000004.1_22531282, whose deoxynucleotide sequence is A / G; 19. CAEKDK010000005.1_12669636, whose deoxynucleotide sequence is C / T; 20. CAEKDK010000005.1_13466650, whose deoxynucleotide sequence is A / G; 21. CAEKDK010000005.1_17572886, whose deoxynucleotide sequence is C / T; 22. CAEKDK010000006.1_4798304, whose deoxynucleotide sequence is A / C; 23. CAEKDK010000006.1_5750123, whose deoxynucleotide sequence is C / T; 24. CAEKDK010000006.1_20454043, whose deoxynucleotide sequence is C / G; 25. CAEKDK010000006.1_21304403, whose deoxynucleotide sequence is C / G; 26. CAEKDK010000007.1_10365019, whose deoxynucleotide sequence is A / G; 27. CAEKDK010000008.1_5584589, whose deoxynucleotide sequence is C / T; 28. CAEKDK010000008.1_7399902, whose deoxynucleotide sequence is A / G; 29. CAEKDK010000008.1_15579729, whose deoxynucleotide sequence is C / T.

2. The identification method according to claim 1, wherein The apricots described in step (1) are selected from Xinjiang apricots.

3. The identification method according to claim 1, wherein The sample in step (1) is selected from one or more of leaves, fruits, roots and stems of apricot.

4. The identification method according to claim 1, wherein In step (1), total DNA was extracted using the CTAB method.

5. The identification method according to claim 1, wherein Step (1) After total DNA extraction, DNA concentration and purity must be determined to control DNA OD 260 / OD 280 It is 1.7-2.

0.

6. The identification method according to claim 1, wherein The SNP-specific site positioning described in step (3) specifically includes the following steps: ① Using the apricot genome that has been fully sequenced as a reference sequence, extract nuclear genome single nucleotide polymorphism sites based on the data of whole genome sequencing in step (2); ② Use BWA software for sequencing, Picard software for deduplication, and SAMtools software for identification of SNP sites; ③Compare the SNP-specific site results to identify the apricot variety.

Citation Information

Patent Citations

  • Improvements on lubricators

    CA13911A

  • Improvements in oil lamps

    CA14012A

  • Improvements in the "boss washing machine"

    CA14113A

  • Machine for sawing kindling wood

    CA14315A

  • Improvements on telephones

    CA14416A