SSR molecular marker primer set for identifying prunus mume germplasm resources and application thereof

By developing 38 pairs of SSR molecular marker primer sets and utilizing transcriptome sequencing and capillary electrophoresis techniques, the problem of insufficient markers for almond germplasm resources was solved, enabling efficient diversity analysis and identification, and providing a foundation for breeding.

CN115820911BActive Publication Date: 2026-04-24RES INST OF NON TIMBER FORESTRY CHINESE ACAD OF FORESTRY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
RES INST OF NON TIMBER FORESTRY CHINESE ACAD OF FORESTRY
Filing Date
2022-11-11
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

In the existing technology, there are few SSR molecular markers for almond germplasm resources, which limits their application in germplasm resource research.

Method used

A primer set of 38 primer pairs for SSR molecular markers was developed. Suitable primers were screened by transcriptome sequencing and bioinformatics methods to amplify SSR markers of almond germplasm resources. Polymorphism analysis was performed using capillary electrophoresis.

Benefits of technology

This study enabled efficient diversity analysis of almond germplasm resources, provided effective identification markers, laid the foundation for germplasm resource evaluation and assisted breeding, and detected 191 allele loci, enriching the analysis of genetic diversity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115820911B_ABST
    Figure CN115820911B_ABST
Patent Text Reader

Abstract

The application discloses a set of SSR molecular marker primers for identifying almond germplasm resources and application of the set of SSR molecular marker primers. The set of SSR molecular marker primers is composed of 38 pairs of primers, and the 38 pairs of primers are used for analyzing genetic diversity of almonds. Based on transcriptome sequencing data of young fruits of the almond, microsatellite sites are mined, analyzed and evaluated, a batch of microsatellite primers is designed and verified, and 38 pairs of polymorphic primers are developed, wherein 27 pairs of primers are high polymorphic primers, 8 pairs of primers are medium polymorphic primers, and 3 pairs of primers are low polymorphic primers. It is indicated that the development of the primers by using the transcriptome data of the almond has strong feasibility, and diversity of almond germplasm resources in main production areas in Xinjiang is analyzed, thereby providing a basis for deep evaluation of the almond germplasm resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of molecular detection technology. Specifically, it relates to SSR molecular marker primer sets for identifying almond germplasm resources and their applications. Background Technology

[0002] Almond (Prunus dulcis (Mill.) DAWebb, Amygdalus communis L.), also known as almond apricot or almond tree, is one of the world's four major dried fruit tree species and a traditional woody oil tree species in my country. It has been cultivated in my country for more than 1,300 years and is loved by people for its large kernel, thin shell, good processing characteristics, and rich content of various amino acids, vitamins and physiologically active substances needed by the human body.

[0003] my country's traditional almond resources were mainly introduced via the ancient Silk Road. Through a long process of introduction, exchange, seed propagation, and artificial selection, distinctive resource types and local varieties adapted to my country's climate have emerged, including bitter almonds, sweet almonds, and soft-shelled sweet almonds. Leveraging regional advantages, these varieties have fostered a distinctive industry, driving regional economic development. They have also been widely introduced and cultivated in Shanxi, Shaanxi, Gansu, Henan, and other regions, resulting in the selection and breeding of over 80 different resource types and more than 10 superior varieties (Lan Yanping et al., 2004; Zhang Qianru et al., 2016). While developing traditional resources, my country has also continuously introduced foreign resources, importing a large number of varieties and breeding materials from the United States, Italy, France, Algeria, Iran, and other countries for domestication trials (Pan Xiaoyun, 2002; Han Hongwei et al., 2003; Zhang Wenyue et al., 2011), greatly enriching domestic almond germplasm resources. With the growth of social demand, the expansion of cultivation area, and the continuous increase in the cultivation and introduction of resources, people are also deepening the work of collecting, preserving, evaluating, and utilizing almond germplasm resources; therefore, basic work such as the identification, evaluation, and diversity analysis of almond germplasm resources is particularly important.

[0004] Microsatellite (SSR) markers are widely used genetic markers due to their rich polymorphism, good reproducibility, codominant inheritance, and high detection efficiency. They are widely used in germplasm identification, genetic map construction, genetic diversity, and phylogenetic analysis. However, due to limitations in development methods, the number of available markers is still relatively small, which restricts their application in germplasm resource research. Summary of the Invention

[0005] Therefore, the technical problem to be solved by the present invention is to provide an SSR molecular marker primer set for identifying almond germplasm resources that analyzes the diversity of almond germplasm resources in the main producing areas of Xinjiang and its application.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0007] A primer set for identifying SSR molecular markers in almond germplasm resources was developed. The primer set consists of 38 primer pairs, targeting the following SSR molecular markers in almonds: S8, S20, S22, S26, S37, S54, S56, S70, S83, S89, S97, S102, S107, S111, S116, S117, S119, and S121. S123, S127, S136, S168, S178, S179, S180, S182, S184, S200, S203, S210, S213, S216, S219, S223, S225, S227, S228, S536, (1) Primers for amplifying the SSR molecular marker S8, the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO. 1~2;

[0008] (2) Primers for amplifying the SSR molecular marker S20, the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO. 3-4;

[0009] (3) Primers for amplifying the SSR molecular marker S22, the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO. 5-6;

[0010] (4) Primers for amplifying the SSR molecular marker S26, the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO. 7-8;

[0011] (5) Primers for amplifying the SSR molecular marker S37, the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO. 9-10;

[0012] (6) Primers for amplifying the SSR molecular marker S54, the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO. 11-12;

[0013] (7) Primers for amplifying the SSR molecular marker S56: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO. 13-14;

[0014] (8) Primers for amplifying the SSR molecular marker S70: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO. 15-16;

[0015] (9) Primers for amplifying the SSR molecular marker S83: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO. 17-18;

[0016] (10) Primers for amplifying the SSR molecular marker S89: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.19-20;

[0017] (11) Primers for amplifying the SSR molecular marker S97: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.21-22;

[0018] (12) Primers for amplifying the SSR molecular marker S102: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO. 23-24;

[0019] (13) Primers for amplifying the SSR molecular marker S107: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO. 25-26;

[0020] (14) Primers for amplifying the SSR molecular marker S111: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO. 27-28;

[0021] (15) Primers for amplifying the SSR molecular marker S116: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO. 29-30;

[0022] (16) Primers for amplifying the SSR molecular marker S117: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO. 31-32;

[0023] (17) Primers for amplifying the SSR molecular marker S119: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO. 33-34;

[0024] (18) Primers for amplifying the SSR molecular marker S121: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO. 35-36;

[0025] (19) Primers for amplifying the SSR molecular marker S123: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO. 37-38;

[0026] (20) Primers for amplifying the SSR molecular marker S127: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO. 39-40;

[0027] (21) Primers for amplifying the SSR molecular marker S136: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO. 41-42;

[0028] (22) Primers for amplifying the SSR molecular marker S168: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO. 43-44;

[0029] (23) Primers for amplifying the SSR molecular marker S178: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO. 45-46;

[0030] (24) Primers for amplifying the SSR molecular marker S179: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO. 47-48;

[0031] (25) Primers for amplifying the SSR molecular marker S180: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO. 49-50;

[0032] (26) Primers for amplifying the SSR molecular marker S182: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO. 51-52;

[0033] (27) Primers for amplifying the SSR molecular marker S184: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO. 53-54;

[0034] (28) Primers for amplifying the SSR molecular marker S200: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO. 55-56;

[0035] (29) Primers for amplifying the SSR molecular marker S203: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO. 57-58;

[0036] (30) Primers for amplifying the SSR molecular marker S210: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO. 59-60;

[0037] (31) Primers for amplifying the SSR molecular marker S213: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO. 61-62;

[0038] (32) Primers for amplifying the SSR molecular marker S216: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO. 63-64;

[0039] (33) Primers for amplifying the SSR molecular marker S219: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO. 65-66;

[0040] (34) Primers for amplifying the SSR molecular marker S223: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO. 67-68;

[0041] (35) Primers for amplifying the SSR molecular marker S225: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO. 69-70;

[0042] (36) Primers for amplifying the SSR molecular marker S227: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO. 71-72;

[0043] (37) Primers for amplifying the SSR molecular marker S228: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO. 73-74;

[0044] (38) Primers for amplifying the SSR molecular marker S536: The forward primer and the reverse nucleotide sequence are shown in SEQ ID NO. 75-76.

[0045] The above-mentioned SSR molecular marker primer set for identifying almond germplasm resources, the method for obtaining 38 primer pairs, includes the following steps:

[0046] (1) Obtaining the gene database: Transcriptome sequencing samples were obtained from normally developing young fruits 30 days after pollination. During sequencing, samples were taken from 3 plants, with 2 young fruits collected from each plant, and the results were repeated 3 times. The RNA-Seq transcriptome sequencing experiment was commissioned to Lianchuan Biotechnology Co., Ltd. The sequencing results were assembled using the De Novo method. A total of 29,591 gene databases were obtained after assembly, which served as the basic data for the development of microsatellite markers.

[0047] (2) Based on transcriptome sequencing data, microsatellite loci were searched using MISA software. The search criteria were a minimum single nucleotide repeat of 10 times and a 2-6 nucleotide repeat of 4 times or more. Primers were designed in batches using Primer 5.0 software. The design criteria were: flanking sequence length of microsatellite loci ≥ 50 bp; PCR product size of 100–350 bp; primer length of 18–26 bp; primer Tm value of 55–70℃; Tm of upstream and downstream primers ≤ 2℃; GC content of 40%–60%.

[0048] (3) Primer screening: 536 primer pairs with more than 7 repeats and microsatellite repeat units of 2 to 6 nucleotides were selected from the 5545 primers designed in batches and synthesized for PCR amplification. 536 primer pairs with high repeat numbers were initially selected for further experimental verification. The results showed that 358 of the 536 primer pairs could amplify clear agarose bands. Further capillary electrophoresis showed that 38 primer pairs could amplify polymorphic bands.

[0049] Of the 38 primer pairs used to identify almond germplasm resources, 27 pairs were marked with highly polymorphic sites (0.5 ≤ PIC), namely S8, S22, S26, S37, S89, S102, S111, S116, S119, S121, S123, S127, S136, S168, S178, S179, S180, S182, S184, S203, S210, S213, S219, S225, S227, S228, and S536.

[0050] The application of SSR molecular marker primer sets for identifying almond germplasm resources in almond genetic diversity analysis includes the following steps:

[0051] (1) Extract DNA from the almond sample to be tested;

[0052] (2) PCR amplification was performed using the above SSR molecular marker primer set, and the results were detected by capillary electrophoresis.

[0053] (3) The genetic diversity of almond germplasm was analyzed by amplifying polymorphic bands.

[0054] The above-mentioned SSR molecular marker primer set for identifying almond germplasm resources was used in the analysis of almond genetic diversity. In step (1), the almond samples to be tested came from the Second Forest Farm of Shache County, including the local main varieties Paper Skin, Double Kernel, Late Flour, Double Fruit, Small Soft Shell and 19 other local variants collected from the resource nursery.

[0055] The above-mentioned SSR molecular marker primer set for identifying almond germplasm resources was applied in the analysis of almond genetic diversity. In step (1), DNA was extracted from almond leaves using the TIANGEN kit DP305. The concentration and quality of the extracted DNA were detected by agarose gel electrophoresis and Nanodrop. The extracted DNA was stored at -4℃ for later use.

[0056] The above-mentioned SSR molecular marker primer set for identifying almond germplasm resources was used in the analysis of almond genetic diversity. In step (2), the PCR amplification system was 20 μL, including 10 μL of 2×Taq Master Mix and 20-30 ng·L⁻¹. -1 1 μL of DNA template, 10 μmol·L -1 0.5 μL each of the forward and reverse primers, and 8 μL of dd H2O;

[0057] The amplification program was as follows: pre-denaturation at 95°C for 5 min; then 35 cycles were performed, each cycle consisting of denaturation at 95°C for 40 s, annealing at 55°C for 35 s, extension at 72°C for 40 s; and finally extension at 72°C for 10 min.

[0058] The above-mentioned SSR molecular marker primer set for identifying almond germplasm resources was applied to the analysis of almond genetic diversity. In step (3), capillary electrophoresis was used for polymorphism screening and polymorphism analysis.

[0059] The above-mentioned SSR molecular marker primer set for identifying almond germplasm resources was applied to the analysis of almond genetic diversity. Each primer pair was used as one pair of allele loci. The number of alleles was counted and the primer polymorphism information content (PIC) was calculated using Power Marker V3.25 software. The effective number of alleles in the population (Ne), Shannon's information index (I), observed heterozygosity (Ho), expected heterozygosity (He), and Nei's expected heterozygosity (Nei's) diversity index were calculated using POPGENE32 software.

[0060] The technical solution of the present invention achieves the following beneficial technical effects:

[0061] To develop microsatellite (SSR) primers for almonds and provide effective identification markers for the evaluation of almond germplasm resources and assisted breeding, this application used young almond fruits as experimental material for transcriptome sequencing and assembly. Bioinformatics methods were used to statistically analyze the number, frequency, and distribution characteristics of SSR loci in the transcriptome of young almond fruits. SSR loci were screened and primers were developed. The developed primers were then used to evaluate the diversity of 24 almond germplasm accessions. The results showed that the distribution frequency of EST-SSRs in young almond fruits was 25.8%, with each motif type repeating 1-24 times. The dominant motifs were 1-3 nucleotide repeats, accounting for 19.9%, 49.1%, and 24.8% of the total SSRs, respectively. Capillary electrophoresis identified 38 polymorphic primer pairs out of 536 tested primer pairs, including 27 highly polymorphic primer pairs (PIC ≥ 0.5). Thirty-eight primer pairs detected 191 allele loci in 24 almond germplasm accessions. The effective allele count (Ne), Shannon's information index (I), observed heterozygosity (Ho), expected heterozygosity (He), and Nei's expected heterozygosity (Nei's) were 2.99, 1.16, 0.59, 0.63, and 0.61, respectively. UPGMA cluster analysis showed that the 24 accessions could be divided into two large groups. This study demonstrated the feasibility of developing SSR primers for almonds and conducted a diversity analysis of domestic almond germplasm resources, providing a foundation for the identification, evaluation, and assisted breeding of almond germplasm resources.

[0062] Based on transcriptome sequencing data of young almond fruits, this paper explores, analyzes, and evaluates microsatellite loci, designs and validates microsatellite primers in batches, and analyzes the diversity of almond germplasm resources in the main producing areas of Xinjiang, providing a foundation for in-depth evaluation of almond germplasm resources.

[0063] Current research on amygdala genetic diversity is still limited, and most studies rely on polypropylene gel electrophoresis, which is difficult to prepare, involves numerous steps, and is cumbersome. Identifying the relative positions of different bands and standardizing reaction data across different batches remain challenging. Capillary electrophoresis, with its high resolution and automation, offers significant advantages in speed and accuracy, and boasts high reliability.

[0064] Diversity analysis results showed that the average effective allele count (Ne), Shannon's information index (I), observed heterozygosity (Ho), expected heterozygosity (He), and Nei's expected heterozygosity (Nei's) of the 24 almond varieties and clones were 2.99, 1.16, 0.59, 0.63, and 0.61, respectively. These values ​​were close to the diversity levels of closely related tree species such as peach, apricot, and cherry, indicating that Xinjiang almonds have rich genetic diversity and a strong genetic foundation, and possess high potential for variety breeding. Attached Figure Description

[0065] Figure 1 Different sequence repeat types and proportions in amygdala microsatellites;

[0066] Figure 2 Agarose gel electrophoresis bands of the test primers;

[0067] Figure 3a Primers for S184 appear in the capillary electrophoresis bands of variety B10;

[0068] Figure 3b The primers for S180 were used in the capillary electrophoresis bands of variety F09;

[0069] Figure 3c Primers for S184 appear as bands on capillary electrophoresis of variety B02;

[0070] Figure 3d The primers for S180 are shown in the capillary electrophoresis bands of variety F01;

[0071] Figure 4 Cluster analysis diagram of germplasm phylogenetic relationships in almonds.

[0072] 1. Materials and Methods

[0073] 1.1 Test Materials

[0074] Transcriptome sequencing samples were obtained from normally developing young fruits 30 days after pollination. Samples were taken from three plants, with two young fruits collected from each plant, in triplicate. The RNA-Seq transcriptome sequencing experiment was commissioned to Lianchuan Biotechnology Co., Ltd. Sequencing results were assembled using the De Novo method (Grabherr et al., 2011), yielding 29,591 Unigenes (gene database), which served as the basis for microsatellite marker development.

[0075] The samples used for the genetic diversity analysis of almonds came from the Second Forest Farm of Shache County, including the local main varieties Paper Skin, Double Kernel, Late Flour, Double Fruit, Small Soft Shell, and 19 other local variants collected from the resource nursery.

[0076] 1.2 Transcriptome microsatellite locus mining and primer design

[0077] Based on transcriptome sequencing data, microsatellite loci were identified using MISA software. The criteria for identification were a minimum single nucleotide repeat count of 10 and a 2-6 nucleotide repeat count of at least 4. Primers were designed in batches using Primer 5.0 software. The design criteria were: flanking sequence length of microsatellite loci ≥ 50 bp; PCR product size 100–350 bp; primer length 18–26 bp; primer Tm value between 55–70℃; Tm of forward and reverse primers ≤ 2℃; GC content between 40% and 60%. Mismatches and primer dimers were avoided as much as possible during primer design. After primer design, BLAST validation was performed on the primers in a database.

[0078] 1.3 Primer selection and PCR amplification

[0079] From the 5545 primers designed in batches, 536 primer pairs with more than 7 repeats and microsatellite repeat units of 2 to 6 nucleotides were selected for synthesis and used for PCR amplification.

[0080] The PCR reaction volume was 20 μL, containing 10 μL of 2×Taq Master Mix and 20-30 ng·L⁻¹. -1 1 μL of DNA template, 0.5 μL each of 10 μmol·L⁻¹ forward and reverse primers, and 8 μL of dd H₂O.

[0081] The amplification program was as follows: pre-denaturation at 95°C for 5 min; then 35 cycles were performed, each cycle consisting of denaturation at 95°C for 40 s, annealing at 55°C for 35 s, extension at 72°C for 40 s; and finally extension at 72°C for 10 min.

[0082] After PCR product amplification, the primers were first detected by 2% agarose gel electrophoresis to screen primers that had no bands or poor amplification effect.

[0083] Then, capillary electrophoresis was used for polymorphism screening and polymorphism analysis.

[0084] 1.4 DNA extraction from 24 types of samples

[0085] DNA was extracted from almond leaves using the TIANGEN DP305 kit. The concentration and quality of the extracted DNA were detected by agarose gel electrophoresis and Nanodrop. The extracted DNA was stored at -4℃ for later use.

[0086] 1.5 Polymorphism Analysis (Data Analysis)

[0087] Each primer pair was used as one pair of allele loci. The number of alleles was counted and primer polymorphism information content (PIC) was calculated using Power Marker V3.25 software. The population effective allele count (Ne), Shannon's information index (I), observed heterozygosity (Ho), expected heterozygosity (He), Nei's expected heterozygosity (Nei's) and other diversity indices were calculated using POPGENE32 software.

[0088] 2. Results and Analysis

[0089] 2.1 Number and Distribution of Microsatellite Sites

[0090] Sequence analysis and data mining identified 7644 SSR sequences from almonds, accounting for 25.8% of the transcriptome sequences. The length of SSR repeat motifs ranged from 1 to 6 nucleotides, primarily consisting of single nucleotide and 2- or 3-nucleotide repeats. 2- or 3-nucleotide repeats, suitable for microsatellite marker development, accounted for 73.9% of all repeat types. The number of repeats in each microsatellite motif was mainly below 20, accounting for 93.2% of the total repeats. The proportion of repeats with 5-8 repeats was the largest, exceeding 10%. Both the number of repeats and the number of repeat types decreased with increasing quantity (Table 1).

[0091] Table 1. Types, Quantities, and Distribution Frequency of Almond Microsatellites

[0092]

[0093] Among different repeat types, single nucleotide repeats were predominantly T and A, accounting for 59.1% and 40.3% of the total repeat types, respectively. Two-nucleotide repeat types were predominantly AG / CT and GA / TC, accounting for 45.2% and 39.2% of the total repeat types, respectively. Three-nucleotide repeat types were predominantly AAG / CTT, GAA / TTC, and AGA / TCT, accounting for 11.7%, 11.4%, and 10.2% of the total repeats, respectively. Figure 1 ).

[0094] 2.2. Screening and Validation of Microsatellite Primers

[0095] Through batch design, a total of 5,545 pairs of eligible microsatellite primers were obtained. Among them, 536 pairs of primers with a relatively high number of repeats were initially selected for further experimental verification. The results showed that 358 pairs of the 536 pairs of primers tested could amplify clear agarose bands( Figure 2 ), accounting for 66.8% of the primers tested. Further capillary electrophoresis detection showed that 38 pairs of primers could amplify polymorphic bands (Table 2, Figures 3a-3d ).

[0096] Figures 3a-3d The capillary electrophoresis detection results of primers S184 and S180 in different samples are shown respectively. Different varieties can be distinguished by different bands.

[0097] Among them, 27 pairs were marked as highly polymorphic loci (0.5 ≤ PIC), accounting for 71.1% of the polymorphic primers; 8 pairs were marked as moderately polymorphic loci (0.25 < PIC < 0.5), accounting for 21.1% of the polymorphic primers; 3 pairs were marked as lowly polymorphic loci (PIC < 0.25), accounting for 7.9% of the polymorphic primers.

[0098] Table 2 Information Table of Polymorphic Primers

[0099]

[0100]

[0101] 2.3. Analysis of Genetic Diversity of Almond

[0102] Twenty-seven pairs of developed highly polymorphic primers were used to analyze the genetic diversity of almond germplasm in the main production areas. A total of 191 alleles were detected in 24 almond samples by 27 microsatellite primers, with an average of 7.07 alleles detected per pair of primers.

[0103] The diversity analysis showed that the effective number of alleles (Ne), Shannon's information index (I), observed heterozygosity (Ho), expected heterozygosity (He), and Nei's expected heterozygosity (Nei's) of each primer in 24 almond resources were respectively between 1.62 - 6.98, 0.66 - 2.04, 0.16 - 0.87, 0.39 - 0.88, and 0.38 - 0.86, and the average values were 2.99, 1.16, 0.59, 0.63, and 0.61 respectively. The phylogenetic clustering analysis showed that the 24 germplasm resources could be divided into four large groups at a genetic distance of 0.65. Most of them (22) were grouped into one group, and the other two were grouped into one group respectively, as Figure 4 shown.

[0104] 3. Discussion

[0105] Molecular markers are an important supplement to morphological markers. Because the detection process is unaffected by environmental factors and the organism's own growth and development cycle, they offer higher detection efficiency, especially for forestry resources with variable cultivation environments, high heterozygosity, and long growth cycles, where they have significant advantages in early trait identification and stability. Based on their origin, microsatellites can be divided into two main categories: genomic microsatellites (Genomic-SSRs) and expression sequence tag microsatellites (EST-SSRs). Genomic microsatellite development is based on genome sequences and involves a series of steps such as library construction, repetitive sequence screening, and cloning and sequencing. This process is complex, labor-intensive, and uses a limited number of available microsatellite sequences. Expression sequence tag microsatellites, on the other hand, originate from microsatellite sequences contained in gene transcription regions. The application of high-throughput transcriptome sequencing technology provides abundant resources for the development of expression sequence tag microsatellites, offering lower costs and higher efficiency.

[0106] In the transcriptome of young almond fruits, 2- and 3-nucleotide repeats were the most common SSR types, similar to findings in most tree species, indicating that almond transcriptome data is suitable for the development of SSR molecular markers. Analysis revealed a significant increase in the number of repeats of all SSR types in the transcriptome of almond compared to other species within the genus and closely related species, especially 2-nucleotide repeats, with repeats exceeding 10 times accounting for 60.8% of the total. In contrast, closely related species such as peach (3.1%), apricot (8.5%), and Chinese cherry (3.4%) all had repeat counts below 10%. Since microsatellites can influence genome evolution through various mechanisms such as expression regulation, gene conversion, and chromosome organization, this type of highly repetitive dinucleotide adaptive trait warrants further research.

[0107] Regarding the base composition of repeat types, the 2-nucleotide repeat motifs of amygdala are mostly AG / CT and GA / TC, while the 3-nucleotide repeat motifs are mostly AAG / CTT and GAA / TTC. This result is similar to the studies on transcriptome development of closely related species such as peach, apricot, plum, Chinese cherry, and other dicotyledonous plants such as eucommia, soapberry, and wintersweet. It differs from some pine and monocotyledonous plants, indicating that these base combinations have a high mutation probability in dicotyledonous plants. This is consistent with the statement by Morgante (1993) that "AAG / CTT is the main trinucleotide repeat type in dicotyledonous plants".

[0108] Through the mining and analysis of microsatellite sequences in the transcriptome, this study designed 5545 pairs of SSR candidate primers that met the development criteria. Of these, 536 pairs were experimentally validated, resulting in the development of 38 polymorphic primer pairs. Among these, 27 pairs were highly polymorphic, 8 pairs were moderately polymorphic, and 3 pairs were low-polymorphic, demonstrating the strong feasibility of primer development using almond transcriptome data. Diversity analysis of 24 almond varieties and clones showed that the 27 primer pairs could detect 191 microsatellite loci, with an average of 7.07 polymorphic loci detected per primer pair, indicating that the developed primers have a high ability to detect variable loci.

[0109] Current research on the genetic diversity of almonds is still limited, and most studies rely on polypropylene gel electrophoresis, which is difficult to prepare, involves many steps, and is cumbersome. Identifying the relative positions of different bands and standardizing reaction data across different batches remain challenging. Capillary electrophoresis, with its high resolution and automation, offers significant advantages in speed and accuracy, and boasts high reliability. Diversity analysis results showed that the average values ​​of the effective allele count (Ne), Shannon's information index (I), observed heterozygosity (Ho), expected heterozygosity (He), and Nei's expected heterozygosity (Nei's) for 24 almond varieties and clones were 2.99, 1.16, 0.59, 0.63, and 0.61, respectively. These results are similar to those found by Xie et al. (2006); Zeng et al. (2009), and are comparable to the diversity levels of closely related species such as peach, apricot, and cherry. This indicates that Xinjiang almonds possess rich genetic diversity and a strong genetic foundation, demonstrating high potential for variety breeding. Cluster analysis showed that the 24 accessions could be divided into three major groups. Most of them were related and could be classified into one large group, which may be attributed to the long history of almond cultivation in China, which has led to the formation of a cultivated population with certain kinship through long-term natural hybridization and offspring selection. The other two accessions were more distantly related, suggesting that they may have been introduced from areas outside the original introduction area recently.

[0110] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of the claims of this patent application.

Claims

1. An SSR molecular marker primer set for identifying almond germplasm resources, characterized in that, The SSR molecular marker primer set consists of 38 primer pairs. These 38 primer pairs target the almond SSR molecular marker as follows: S8, S20, S22, S26, S37, S54, S56, S70, S83, S89, S97, S102, S107, S111, S116, S117, S119, S121, S123, S127, S136, S168, S178, S179, S180, S182, S184, S200, S203, S210, S213, S216, S219, S223, S225, S227, S228, and S536. (1) Primers for amplifying the SSR molecular marker S8, the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.1~2; (2) Primers for amplifying the SSR molecular marker S20, the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.3~4; (3) Primers for amplifying the SSR molecular marker S22, the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.5~6; (4) Primers for amplifying the SSR molecular marker S26, the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.7-8; (5) Primers for amplifying the SSR molecular marker S37, the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.9-10; (6) Primers for amplifying the SSR molecular marker S54, the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.11-12; (7) Primers for amplifying the SSR molecular marker S56: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.13-14; (8) Primers for amplifying the SSR molecular marker S70: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.15-16; (9) Primers for amplifying the SSR molecular marker S83: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.17-18; (10) Primers for amplifying the SSR molecular marker S89: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.19-20; (11) Primers for amplifying the SSR molecular marker S97: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.21-22; (12) Primers for amplifying the SSR molecular marker S102: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.23~24; (13) Primers for amplifying the SSR molecular marker S107: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.25~26; (14) Primers for amplifying the SSR molecular marker S111: The forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.27-28; (15) Primers for amplifying the SSR molecular marker S116: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.29-30; (16) Primers for amplifying the SSR molecular marker S117: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.31-32; (17) Primers for amplifying the SSR molecular marker S119: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.33-34; (18) Primers for amplifying the SSR molecular marker S121: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.35-36; (19) Primers for amplifying the SSR molecular marker S123: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.37-38; (20) Primers for amplifying the SSR molecular marker S127: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.39~40; (21) Primers for amplifying the SSR molecular marker S136: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.41~42; (22) Primers for amplifying the SSR molecular marker S168: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.43~44; (23) Primers for amplifying the SSR molecular marker S178: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.45~46; (24) Primers for amplifying the SSR molecular marker S179: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.47~48; (25) Primers for amplifying the SSR molecular marker S180: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.49~50; (26) Primers for amplifying the SSR molecular marker S182: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.51-52; (27) Primers for amplifying the SSR molecular marker S184: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.53~54; (28) Primers for amplifying the SSR molecular marker S200: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.55~56; (29) Primers for amplifying the SSR molecular marker S203: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.57-58; (30) Primers for amplifying the SSR molecular marker S210: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.59~60; (31) Primers for amplifying the SSR molecular marker S213: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.61~62; (32) Primers for amplifying the SSR molecular marker S216: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.63~64; (33) Primers for amplifying the SSR molecular marker S219: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.65~66; (34) Primers for amplifying the SSR molecular marker S223: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.67~68; (35) Primers for amplifying the SSR molecular marker S225: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.69~70; (36) Primers for amplifying the SSR molecular marker S227: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.71-72; (37) Primers for amplifying the SSR molecular marker S228: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.73~74; (38) Primers for amplifying the SSR molecular marker S536: the forward primer and the reverse nucleotide sequence are shown in SEQ ID NO.75~76.

2. The SSR molecular marker primer set for identifying almond germplasm resources according to claim 1, characterized in that, The method for obtaining 38 primer pairs includes the following steps: (1) Obtaining the gene database: The transcriptome sequencing samples were obtained from the normally developing young fruits 30 days after pollination. A total of 3 fruits were sampled during sequencing, with 2 young fruits collected from each plant, and the results were repeated 3 times. The RNA-Seq transcriptome sequencing experiment was commissioned to Lianchuan Biotechnology Co., Ltd. The sequencing results were assembled using the De Novo method. A total of 29,591 gene databases were obtained after assembly, which served as the basic data for the development of microsatellite markers. (2) Based on transcriptome sequencing data, microsatellite loci were searched using MISA software. The search criteria were a minimum single nucleotide repeat of 10 times and a 2-6 nucleotide repeat of 4 times or more. Primers were designed in batches using Primer 5.0 software. The design criteria were: flanking sequence length of microsatellite loci ≥ 50 bp. PCR product size ranges from 100 to 350 bp; Primer length is 18–26 bp; Primer Tm values ​​should be between 55 and 70°C; forward and reverse primer Tm ≤ 2°C; GC content should be between 40% and 60%. (3) Primer screening: 536 primer pairs with repeat times of more than 7 times and microsatellite repeat units of 2 to 6 nucleotides were selected from the 5545 primers designed in batches and synthesized for PCR amplification; 536 primer pairs with high repeat times were initially selected for further experimental verification. The results showed that 358 of the 536 primer pairs could amplify clear agarose bands. Further capillary electrophoresis showed that 38 primer pairs could amplify polymorphic bands.

3. The SSR molecular marker primer set for identifying almond germplasm resources according to claim 1, characterized in that, in, Of the 38 primer pairs, 27 pairs were marked as highly polymorphic sites (0.5 ≤ PIC), namely S8, S22, S26, S37, S89, S102, S111, S116, S119, S121, S123, S127, S136, S168, S178, S179, S180, S182, S184, S203, S210, S213, S219, S225, S227, S228, and S536.

4. The application of the SSR molecular marker primer set for identifying almond germplasm resources according to any one of claims 1-3 in the analysis of almond genetic diversity, characterized in that, Includes the following steps: (1) Extract DNA from the almond sample to be tested; (2) PCR amplification was performed using the SSR molecular marker primer set described in any one of claims 1-3, and the results were detected by capillary electrophoresis; (3) The genetic diversity of almond germplasm was analyzed by amplifying polymorphic bands.

5. The application of the SSR molecular marker primer set for identifying almond germplasm resources according to claim 4 in the analysis of almond genetic diversity, characterized in that, In step (1), the almond samples to be tested came from the Second Forest Farm of Shache County, including the local main varieties Paper Skin, Double Kernel, Late Flour, Double Fruit, Small Soft Shell and 19 other local variants collected from the resource nursery.

6. The application of the SSR molecular marker primer set for identifying almond germplasm resources according to claim 4 in the analysis of almond genetic diversity, characterized in that, In step (1), DNA was extracted from almond leaves using the TIANGEN kit DP305. The concentration and quality of the extracted DNA were detected by agarose gel electrophoresis and Nanodrop. The extracted DNA was stored at -4℃ for later use.

7. The application of the SSR molecular marker primer set for identifying almond germplasm resources according to claim 4 in the analysis of almond genetic diversity, characterized in that, In step (2), the PCR amplification system is 20 µL, including 10 µL of 2×Taq Master Mix and 20-30 ng·L⁻¹. -1 DNA template 1 µL, 10 µmol·L -1 0.5 µL each of forward and reverse primers, and 8 µL of dd H2O; The amplification program was as follows: pre-denaturation at 95 °C for 5 min; then 35 cycles were performed, each cycle consisting of denaturation at 95 °C for 40 s, annealing at 55 °C for 35 s, extension at 72 °C for 40 s; and finally extension at 72 °C for 10 min.

8. The application of the SSR molecular marker primer set for identifying almond germplasm resources according to claim 4 in the analysis of almond genetic diversity, characterized in that, In step (3), capillary electrophoresis is used for polymorphism screening and polymorphism analysis.

9. The application of the SSR molecular marker primer set for identifying almond germplasm resources according to claim 7 in the analysis of almond genetic diversity, characterized in that, Each primer pair was used as one pair of allele loci. The number of alleles was counted and the primer polymorphism information content was calculated using Power Marker V3.25 software. The effective number of alleles in the population, Shannon's information index, observed heterozygosity, expected heterozygosity, and Nei's expected heterozygosity diversity index were calculated using POPGENE32 software.