DNA (Deoxyribose Nucleic Acid) fingerprint spectrum of foreign pine improved variety and construction method of DNA fingerprint spectrum
Through SNP site capture and genotyping, the DNA fingerprint of foreign pine varieties was established, which solved the problem of difficult to distinguish and identify foreign pine species in the existing technology, achieved efficient identification and identification of foreign pine varieties, and improved the efficiency of breeding work.
Patent Information
- Application Number
- CN202510194198.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-13
AI Technical Summary
It is difficult for the existing technology to efficiently distinguish and identify foreign pine species with similar morphology but different varieties, and the growth cycle of trees is long. Waiting for breeding or identification until the seedlings become adults will reduce the efficiency of breeding work or cause production losses.
Through SNP site capture and genotyping, the target region of DNA was captured using a 51K liquid phase probe chip, and the SNP site was detected through second-generation sequencing. Combined with strict screening standards and PIC value reduction, the DNA fingerprint map of foreign Songliang Chronicles was successfully established.
It has achieved efficient identification and identification of foreign pine seeds, provided technical support for the management of germ quality and genetic diversity analysis of good varieties, and significantly improved the efficiency of breeding work.
Smart Images

Figure CN120138196A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of molecular biology, and particularly relates to a DNA fingerprint map of improved foreign pine varieties and a method for constructing the same. Background Art
[0002] Since the introduction of Pinus elliottii, P. taeda, and P. caribaea, they have been cultivated in China for a hundred years. They are characterized by strong adaptability and high resin production, and are important timber and resin-producing tree species in southern China. P. elliottii and P. caribaea, as well as P. taeda and P. caribaea, have some complementary traits in terms of growth rate, soil adaptability, pest and disease resistance, and drought resistance, and are ideal materials for cross-breeding. Their interspecific hybrid offspring, P. elliottii × P. caribaea and P. taeda × P. caribaea, have shown better growth characteristics than their parental species in the tropical and subtropical regions of China and have become important tree species for constructing resin-wood dual-purpose and large-diameter forests in suitable regions of southern China.
[0003] The differences in morphological characteristics of foreign pine are not obvious. It is difficult to distinguish varieties with similar traits only relying on traditional morphological identification. Moreover, the growth cycle of forest trees is long. If the selection or identification is carried out after the seedlings grow into adults, the efficiency of breeding work will be greatly reduced or production losses will be caused.
[0004] Therefore, the present invention aims to provide a DNA fingerprint map of improved foreign pine varieties and a method for constructing the same to solve the above problems. Summary of the Invention
[0005] The object of the present invention is to solve the above problems and provide a DNA fingerprint map of improved foreign pine varieties and a method for constructing the same. By capturing SNP sites and genotyping, SNPs sites are obtained. By strictly controlling parameters such as the minor allele frequency, missing rate, and heterozygosity, core SNPs are screened, and further refined to 20 SNPs with the ability to identify improved varieties using the PIC value. A DNA fingerprint map of improved foreign pine varieties is successfully established, providing a reference for the management of improved variety germplasm and the analysis of genetic diversity.
[0006] In order to achieve the above object, the technical solution of the present invention is as follows:
[0007] The present invention provides a DNA fingerprint map of improved foreign pine varieties and a method for constructing the same, and the method comprises the following steps:
[0008] S1. Collect samples and extract DNA from the samples;
[0009] S2. Capture the target region of DNA using a 51K liquid-phase probe chip and perform next-generation sequencing on the target region;
[0010] S3. Perform single nucleotide polymorphism detection through local re-alignment and base recalibration to obtain SNP sites;
[0011] S4. Conduct initial filtering on the captured SNP sites, and the remaining SNPs are used for population genetic structure and phylogenetic analysis;
[0012] S5. Screen the SNP sites in the original data using plink software to obtain core SNPs;
[0013] S6. Use the PIC value of each SNP to reduce the number of SNPs used for constructing the fingerprint map. The formula for calculating PIC is as follows:
[0014]
[0015] p i and p j are the frequencies of the i-th and j-th alleles respectively, and n is the number of alleles;
[0016] S7. Perform PCA calculation using the screened SNPs and core SNPs to evaluate the effectiveness of the screening process.
[0017] The specific steps of step S2 are as follows:
[0018] S201. Fragment the DNA of the sample and repair the ends of the DNA fragments through Pre-PCR to amplify the library;
[0019] S202. Hybridize the 51K liquid-phase probe with the target region;
[0020] S203. Use streptavidin-labeled magnetic beads to capture the hybridization probe;
[0021] S204. After capture, elute and enrich the target fragments and amplify them through Post-PCR;
[0022] S205. Perform next-generation sequencing on the target region.
[0023] The next-generation sequencing mentioned above is: sequencing the target region through the Illumina Xten high-throughput sequencing platform, storing the target region data in FASTQ format, and performing data filtering;
[0024] The steps of the data filtering are as follows:
[0025] S211. Remove the reads with adapters;
[0026] S212. Remove the readings with N content exceeding 10%;
[0027] S213. Remove the readings with more than 50% of the base quality scores lower than 10.
[0028] The specific steps of step S3 are as follows: Use the GATK software toolkit to detect single nucleotide polymorphisms (SNPs) in the target region data after data filtering; use SAMtools to remove duplicates (Mark Duplication), and use GATK for local realignment and base recalibration to obtain SNP sites.
[0029] The screening criteria in step S5 are as follows:
[0030] S501. Screen out the sites with minor allele frequency (MAF) > 0.3;
[0031] S502. Remove the sites with deletions, and only retain the sites with a deletion rate of 0;
[0032] S503. Remove the sites that do not conform to the Hardy - Weinberg equilibrium;
[0033] S504. Filter out the sites with linkage disequilibrium (LD) value < 0.2;
[0034] S505. Calculate the PIC value according to the allele frequency, and remove the sites with PIC < 0.35;
[0035] S506. Remove the sites with heterozygosity rate (Het) > 0.25.
[0036] Compared with the prior art, the beneficial effects of this solution are as follows:
[0037] The present invention uses 38 improved varieties of foreign pine species and their hybrid pine improved variety materials, and uses the 51K liquid - phase probe chip of slash pine and loblolly pine for SNP capture, and successfully genotypes 38 improved varieties of foreign pine; through SNP site capture and genotyping, a total of 5,60,567 SNPs sites are obtained. By strictly controlling parameters such as minor allele frequency, deletion rate, and heterozygosity, 344 core SNPs are finally screened, and further refined to 20 SNPs with the ability to identify improved varieties using the PIC value, and successfully establish the DNA fingerprint maps of 38 pine improved varieties; providing a reference for the management of improved variety germplasm and the analysis of genetic diversity. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a flow chart of the fingerprint map construction method in the embodiment of the present invention;
[0039] Figure 2It is a schematic diagram of the phylogenetic and population genetic analysis of 38 pine resources in the embodiments of the present invention. A is the phylogenetic analysis; B is the genetic structure analysis of 38 pines when K = 2; C is the Delta K line graph, reaching an inflection point when K = 2; D is the principal component analysis. The black circles represent Pinus elliottii, the red circles represent Pinus taeda, the green circles represent Pinus elliottii × Pinus caribaea, and the blue circles represent Pinus taeda × Pinus caribaea;
[0040] Figure 3 It is a schematic diagram of the genetic parameter information of 344 SNPs and 28 SNPs in the embodiments of the present invention. A - D: MAF (A), PIC (B), observed heterozygosity and expected heterozygosity (C) of 344 SNPs; E - H: MAF (D), PIC (E), observed heterozygosity and expected heterozygosity (F) of 20 SNPs;
[0041] Figure 4 It is a schematic diagram of the principal component analysis based on 344 SNPs (A) and 20 SNPs (B) in the embodiments of the present invention. The black circles represent Pinus elliottii (PEE) samples, the red circles represent Pinus taeda (PTA) samples, the green circles represent Pinus elliottii × Pinus caribaea (PEC) samples, and the blue circles represent Pinus taeda × Pinus caribaea (PAC) samples;
[0042] Figure 5 It is the phylogenetic relationship analysis of 38 improved pine varieties based on 344 SNPs (A) and 20 SNPs (B) in the embodiments of the present invention. The black lines represent Pinus elliottii (PEE) samples, the red lines represent Pinus taeda (PTA) samples, the green lines represent Pinus elliottii × Pinus caribaea (PEC) samples, and the blue lines represent Pinus taeda × Pinus caribaea (PAC) samples;
[0043] Figure 6 It is the DNA fingerprint map of 38 improved exotic pine varieties in the embodiments of the present invention. The genotype 1 / 1 represents the homozygote of the major allele, the genotype 0 / 0 represents the homozygote of the minor allele, and the genotype 0 / 1 represents the heterozygote. The different colors on the left represent different exotic pine tree species. Detailed implementation manners
[0044] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solution of the present invention will be further described in detail below in conjunction with the embodiments and drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0045] It should be noted that, without conflict, the embodiments and features in the embodiments of the present invention can be combined with each other. The present invention will be described in detail below in conjunction with the embodiments.
[0046] Embodiment:
[0047] 1. Materials and Methods
[0048] 1.1. Extraction of Materials and Their DNA
[0049] All samples were from Changle Forest Farm in Hangzhou, Zhejiang Province. A total of 38 foreign pine germplasm resources (grafted clones, mainly the parents or mother plants of improved varieties approved and identified in recent years) were selected, including 29 slash pine parents, 2 loblolly pine parents, 2 slash x loblolly pine mother plants, and 5 loblolly x slash pine mother plants. These resources have excellent traits such as fast growth rate, good wood properties, or high resin yield (Table 1).
[0050] Collect their needles and store them frozen. Use the M5 HiPer Universal DNA Mini Kit (Beijing Polymer Beauty Biotechnology Co., Ltd., Beijing) to extract the DNA of pine needle samples. Use a 51K liquid-phase probe chip to capture the target regions of the genome for second-generation sequencing. The specific steps are as follows: (1) Segment the DNA and repair the ends of the DNA fragments by Pre-PCR to amplify the library; (2) Hybridize the probes with the target regions; (3) Use streptavidin-labeled magnetic beads to capture the hybridized probes; (4) After capture, elute the enriched target fragments and amplify them by Post-PCR; (5) Perform second-generation sequencing on the target regions.
[0051] Table 1 Basic Information of 38 Foreign Pine Resources
[0052]
[0053] 1.2. Data Quality Control and Genotyping
[0054] The captured DNA fragments were sequenced by the Illumina Xten high-throughput sequencing platform. The raw data was stored in FASTQ format, and the adapter and low-quality data were filtered using FastQC software (Wingett and Andrews, 2018). The steps of data filtering were as follows: (1) removing the reads with adapters; (2) removing the reads with more than 10% N content; (3) removing the reads with more than 50% of the base quality scores less than 10. Subsequently, single nucleotide polymorphism (SNP) detection was performed using the GATK software toolkit (https: / / gatk.broadinstitute.org / hc / en-us). According to the mapping results of clean Reads in the reference genome, SAMtools was used to remove duplicates (Mark Duplication), and GATK was used for local realignment and base recalibration. The genome was the reference genome of Pinus taeda and the specific long-read transcripts of Pinus elliottii. To ensure the accuracy of SNP detection, GATK was used to detect and filter single nucleotide polymorphisms, and finally SNP sites were obtained.
[0055] 1.3. Population genetic structure and phylogenetic analysis
[0056] The captured SNP sites were initially filtered to screen out the sites with minor allele frequency less than 0.01 and missing rate higher than 0.8. The remaining SNPs were used for population genetic structure and phylogenetic analysis. The neighbor-joining phylogenetic tree was constructed using MEGA 11.0 software. The population structure was inferred using Admixture software (version 1.3.0). The optimal number of clusters (K) of the sample population was determined based on the Delta K method. Principal component analysis (PCA) was performed using PLINK software (version 1.07) to evaluate population stratification and genetic relationships among individuals. The principal component analysis was performed using the "--noweb" command in PLINK to generate the eigenvalues and eigenvectors representing the principal components of genetic variation. The first two principal components were retained for subsequent analysis. The R software package ggplot2 was used to visualize the population genetic structure and PCA results.
[0057] 1.4. Screening of core SNPs
[0058] Screen the SNP sites in the original data of plink software, and the screening criteria are as follows: (1) Screen out the sites with minor allele frequency (MAF) > 0.3; (2) Remove the sites with missing values, and only retain the sites with a missing rate of 0; (3) Remove the sites that do not conform to the Hardy-Weinberg equilibrium; (4) Filter out the sites with linkage disequilibrium (LD) value < 0.2; (5) Calculate the PIC value according to the allele frequency, and remove the sites with PIC < 0.35; (6) Remove the sites with heterozygosity (Het) > 0.25. After screening by the above steps, the core SNPs are obtained.
[0059] Among them, the formula for calculating PIC is as follows:
[0060]
[0061] p i and p j are the frequencies of the i-th and j-th alleles respectively, and n is the number of alleles.
[0062] 1.5. Evaluation of SNP Sites and Construction of DNA Fingerprint Maps
[0063] In order to identify as many samples as possible with a small number of SNPs, we use the PIC value of each SNP to reduce the number of SNPs used to construct the fingerprint map. The PIC value reflects the ability of the marker to distinguish different genotypes. The higher the PIC value, the greater the polymorphism of the marker and the stronger its ability to identify genotypes. Therefore, the SNPs are sorted in descending order according to the PIC value, and the SNPs with the highest PIC value are selected.
[0064] After completing the reduction of the number, PCA calculation is performed using the selected SNPs and core SNPs to evaluate the effectiveness of the screening process. The PCA analysis is performed using the pcadapt v4.3.2 software (https: / / bcm-uga.github.io / pcadapt / articles / pcadapt.html).
[0065] 2. Results
[0066] 2.1. Genetic Structure and Phylogenetic Relationships of 38 Pine Improved Varieties
[0067] For the 38 pine germplasm resource samples in this experiment, after DNA extraction and liquid-phase probe capture of SNPs, a total of 5,60,567 SNPs were obtained, with a genotyping rate of 94.61%. After initial filtering, loci with a minor allele frequency lower than 0.01 and a missing rate higher than 0.8 were screened out, resulting in 183,849 SNPs as the original data for this experiment. Based on the 183,849 SNPs, the software Admixture was used to calculate the optimal number of population clusters. When the K value was 2, the cross-validation error value (CV error) was the smallest ( Figure 2 C), so the pine population selected in this experiment could be divided into two subgroups ( Figure 2 A). Subgroup I had a total of 29 slash pine samples, and subgroup II had a total of 8 samples, including loblolly pine, fire plus pine, and wet plus pine. A neighbor-joining method was used to construct a phylogenetic tree ( Figure 2 B), and the results showed that all slash pine samples clustered together, and loblolly pine, fire plus pine, and wet plus pine samples clustered together, which was consistent with the clustering results of admixture. The number of slash pine (PEE) samples was relatively large, and the materials were from 3 different families with a wide source. The genetic distance between PEE samples was relatively far, and there was no obvious association with the source of the samples; loblolly pine (PTA), fire plus pine (PAC), and wet plus pine (PEC) all had genes of the paternal Caribbean pine, and the number of samples was small, and the genetic relationship between the samples was relatively close. The female parents of PTA and PAC samples were both from the first-generation seed orchard of loblolly pine, and the genetic relationship was relatively close. In contrast, the female trees of PEC samples were the same, and the genetic relationship was relatively far.
[0068] 2.2. Screening of core SNPs
[0069] Based on genetic parameters such as MAF, missing rate, LD, heterozygosity rate, and PIC of SNP loci, a strict screening process was developed (Table 2), and 344 high-quality core SNPs were screened out. The MAF range of the 344 core SNPs was between 0.3026 and 0.5, with an average value of 0.3524; the PIC value range was between 0.4221 and 0.5, with an average value of 0.4507; the expected heterozygosity (He) of the 344 SNPs in 38 samples was 0.5433, and the observed heterozygosity (Ho) range was between 0.3692 and 0.9826, with an average value of 0.7811 ( Figure 3 A-C).
[0070] Table 2 Parameter settings for screening core SNPs
[0071]
[0072]
[0073] 2.3. Evaluation of Core SNPs and Construction of DNA Fingerprint
[0074] To improve the efficiency of elite variety identification, it is necessary to simplify the core SNP loci, aiming to identify more individuals with as few loci as possible. We sorted the 344 core SNPs in descending order according to the PIC value and selected the SNPs with higher PIC values to ensure that the selected SNPs could fully distinguish 38 samples. Finally, 20 loci were selected as the representative loci for 38 samples.
[0075] To evaluate whether the selected SNPs could effectively detect the diversity of the population, principal component analysis and phylogenetic analysis were performed using 344 core SNPs and 20 SNPs respectively. The PCA results showed that the clustering effects of 344 SNPs and 20 SNPs were consistent with those of 183,849 SNPs. As the number of SNPs decreased, the proportion of different loci among samples increased. Finally, at the 20 SNP loci, the samples showed obvious differences, not only retaining the characteristics of 4 tree species but also being able to distinguish different individuals within the same species.
[0076] The phylogenetic tree results showed that based on 344 core SNPs, the 4 exotic pine tree species could be correctly clustered independently. This indicates that the 344 core SNPs screened through the set genetic parameters can effectively assist in the preliminary species identification of exotic pines with similar phenotypic characteristics in the seed orchard. However, when the number of SNPs decreased to 20, the four exotic pine tree species could not be correctly clustered, and the interspecies clustering relationship also changed significantly. Although these 20 loci had the highest polymorphic information content, due to the limited sample source, these loci were not sufficient to support the classification and verification of a wider exotic pine population. Therefore, the 20 SNP markers mainly revealed the different loci among different elite varieties of exotic pines and were suitable for the identification of the offspring and clones of elite varieties within the same seed orchard. The 344 core SNP markers had higher applicability and could be used for the preliminary identification of exotic pine species in the seed orchard.
[0077] Finally, a DNA fingerprint of exotic pine was constructed using 20 SNP loci. The MAF range of the 20 SNPs was between 0.3026 and 0.50, with an average of 0.3144; the PIC value range was between 0.4221 and 0.5, with an average of 0.4448; the expected heterozygosity (He) of the 20 SNPs in 38 samples was 0.55, and the observed heterozygosity (Ho) range was from 0.428571 to 1, with an average of 0.7811( Figure 3 D-F).
[0078] Table 3 Twenty SNPs after quantity reduction and their genetic diversity indices
[0079]
[0080]
[0081] 3. Discussion
[0082] 3.1 Application of 51K liquid-phase probes in exotic pines
[0083] In China, exotic pines refer to several pine species introduced from the United States in the 1930s, including slash pine, loblolly pine, Caribbean pine, and their interspecific hybrids, etc. Since their introduction, exotic pines have been widely planted in southern China and have become important tree species for afforestation. In the process of forest tree breeding, molecular-assisted breeding can accelerate the breeding process. Obtaining molecular markers through genotyping is the key to realizing molecular-assisted breeding. The 51K liquid-phase probe chip has been successfully applied to the breeding programs of slash pine, loblolly pine, and Caribbean pine, proving that this SNP chip can be widely used in the molecular breeding research of exotic pines. Researchers can effectively distinguish the molecular specificities of different exotic pine species, identify genetic relationships, and analyze genetic diversity, providing technical support for the selection of excellent varieties.
[0084] In the present invention, SNPs were captured by 51K liquid-phase probes, and 560,567 SNP loci of 38 improved pine varieties were genotyped, with a genotyping rate of 94.61%. Based on the method process established in the present invention, by integrating more accurate molecular markers, the selected SNP markers were developed into a new set of SNP chips and were actually applied to the management of germplasm resources in seed orchards, accelerating the molecular identification and classification optimization of germplasm resources, significantly improving the scientific nature and efficiency of seed orchard management, and also providing an important direction for the future development of this research.
[0085] 3.2 Construction of DNA fingerprint maps for improved exotic pine varieties
[0086] DNA fingerprint maps are a technology that can effectively identify the genotypes of humans, animals, and plants. Constructing DNA fingerprint maps has important scientific and practical value in the identification and promotion of improved varieties in the forestry industry, can significantly improve the accuracy, efficiency, and reliability of improved variety identification, and at the same time provide solid technical support for forestry breeding and germplasm resource management. In addition, during the process of selecting improved varieties, constructing DNA fingerprint maps helps to protect the intellectual property rights of improved varieties. By clarifying the genetic markers of improved varieties, it can effectively prevent the illegal replication and dissemination of improved varieties and safeguard the legitimate rights and interests of breeding units. The present invention provides a set of core SNP screening processes applicable to improved exotic pine varieties, and constructs DNA fingerprint maps that can distinguish different improved varieties by simplifying the number of SNPs. This fingerprint map represents the molecular specificities of different improved varieties and can be regarded as a standard sample library of improved varieties.
[0087] In the present invention, 344 core SNPs are sorted in descending order according to PIC values, and SNPs with higher PIC values are selected. The 20 selected SNPs have high polymorphism and can be used to identify the improved progeny or clones of exotic pines in the same seed orchard. According to the genetic clustering of different numbers of SNPs ( Figure 5 ), it is shown that at the species identification level, 344 SNPs perform better than 20 SNPs and are applicable to the preliminary identification of exotic pine seedlings with similar phenotypic characteristics in the same seed orchard.
[0088] 4. Conclusions
[0089] In the present invention, a DNA fingerprint map of 38 improved pine germplasm resources including PEE (Pinus elliottii), PTA (Pinus taeda), PEC (Pinus elliottii × Pinus caribaea) and PAC (Pinus taeda × Pinus elliottii) has been successfully established using a 20-SNP marker set constructed by a probe chip. By strictly screening the minor allele frequency, deletion rate, polymorphism information content and heterozygosity rate of SNPs, 344 core SNPs for the preliminary identification of excellent pine varieties are obtained. Subsequently, 20 SNPs are selected as the best locus combination, which can identify all the materials used in this study. However, the identification and analysis effects of this SNP set on other exotic pine germplasm resources still need to be analyzed according to the actual test results.
[0090] The above specific embodiments are only explanations of the present invention, and they are not limitations of the present invention. Those skilled in the art can make modifications without creative contributions to this embodiment according to needs after reading this specification, but as long as they are within the scope of the claims of the present invention, they are protected by the patent law.
Claims
1. A DNA fingerprint of foreign pine varieties and a method for constructing the same, characterized by: The method comprises the following steps: S1, collect samples and extract DNA from samples; S2, using a 51K liquid phase probe chip to capture the target region of DNA, and performing second-generation sequencing on the target region; S3, single nucleotide polymorphism detection is performed through local realignment and base recalibration to obtain SNP sites; S4, initial filtering of captured SNP sites, and the remaining SNPs are used for population genetic structure and phylogenetic analysis; S5. Use plink software to screen the SNP sites in the original data to obtain core SNPs; S6. Use the PIC value of each SNP to further reduce the number of SNPs used to construct the fingerprint map. The formula for calculating PIC is as follows: p i and p j are the frequencies of the ith and jth alleles, respectively, and n is the number of alleles; S7. PCA calculations were performed using the screened SNPs and core SNPs to evaluate the effectiveness of the screening process.
2. The DNA fingerprint of a foreign pine variety and the method for constructing the same as claimed in claim 1, characterized in that: The specific steps of step S2 are: S201, fragmenting the DNA of the sample, and repairing the ends of the DNA fragments by Pre-PCR to amplify the library; S202, hybridizing the 51K liquid phase probe to the target region; S203, using streptavidin affinity labeled magnetic beads to capture the hybridization probe; S204, eluting the enriched target fragments after capture and amplifying them by Post-PCR; S205, performing next-generation sequencing on the target region.
3. The DNA fingerprint of a foreign pine variety and the method for constructing the same as claimed in claim 2, characterized in that: The second generation sequencing is: sequencing the target region through the IlluminaXten high-throughput sequencing platform, storing the target region data in FASTQ format, and performing data filtering; The steps of data filtering are as follows: S211, remove reads with aptamers; S212, remove readings with N content exceeding 10%; S213. Remove more than 50% of the reads with base quality scores lower than 10.
4. The DNA fingerprint of a foreign pine variety and the method for constructing the same as claimed in claim 3, characterized in that: The specific steps of step S3 are to use the GATK software toolkit to perform single nucleotide polymorphism detection on the target region data after data filtering; use SAMtools to remove duplications, and use GATK to perform local rearrangement and base recalibration to obtain SNP sites.
5. The DNA fingerprint of a foreign pine variety and the method for constructing the same as claimed in claim 1, characterized in that: The screening criteria in step S5 are: S501, screen out sites with minor allele frequency > 0.3; S502, removing sites with deletions and retaining only sites with a deletion rate of 0; S503, remove sites that do not conform to the Hardy-Weinberg equilibrium; S504, filter loci with linkage disequilibrium values < 0.2; S505, calculate the PIC value based on the allele frequency and remove sites with PIC < 0.35; S506. Remove sites with heterozygosity > 0.25.