South China tiger individual identification related SNP site combination and individual identification method and application thereof

By screening and preserving high-quality SNP loci, a simple, reliable, and low-cost method for identifying individuals of the South China tiger was established, solving the problems of low reliability of RFID chips and cumbersome and time-consuming molecular identification, and providing strong support for genetic management.

CN119776537BActive Publication Date: 2025-12-30NORTHEAST FORESTRY UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411900829.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-12-30
Estimated Expiration
2044-12-23

AI Technical Summary

Technical Problem

Existing South China tiger individual identification technologies suffer from low reliability of RFID chips and cumbersome and time-consuming experimental processes. Traditional molecular identification methods require the design of a large number of specific primers and repeated optimization.

Method used

By using genome-based SNP loci screening, 300 highly polymorphic SNP loci were selected as genetic markers. Individual identification was achieved through low-depth genome sequencing and database comparison, simplifying the operation process and reducing costs.

Benefits of technology

It enables simple, reliable, and low-cost identification of South China tiger individuals, improving the adaptability and reliability of identification while reducing computational burden and economic costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure QLYQS_1
    Figure QLYQS_1
  • Figure QLYQS_2
    Figure QLYQS_2
  • Figure GDA0005354202700000021
    Figure GDA0005354202700000021
Patent Text Reader

Abstract

The application provides a genetic marker and application thereof in individual identification of South China tigers, and belongs to the technical field of molecular identification.The application provides a genetic marker including 300 SNP molecular markers, which can be applied in individual identification of South China tigers, and provides an individual identification method.The application solves the problems of the traditional RFID chip in individual identification, and also solves the problems that the existing molecular identification mainly depends on STR and SNP site amplification, a large number of specific primers need to be designed in the early stage and repeatedly optimized, the process is complicated and time-consuming, and the like.A convenient, reliable and low-cost individual identification method of South China tigers is established, and strong technical support is provided for genetic management and population protection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of molecular identification technology, specifically relating to the field of animal individual identification technology. Background Technology

[0002] The South China tiger is relatively small. Males have a head-to-tail length of about 2.5 meters and weigh approximately 150 kg. Females have a head-to-tail length of about 2.3 meters, a tail length of 80-100 cm, and weigh approximately 120 kg. The South China tiger has a round head, short ears, strong and powerful limbs, and a relatively long tail. Its chest and abdomen are mixed with a significant amount of milky white, and its entire body is orange-yellow with black horizontal stripes. Its fur has short, narrow stripes, with wider spacing than that of the Bengal tiger or Siberian tiger, and diamond-shaped patterns often appear on its sides. The South China tiger feeds on herbivores such as wild boar, deer, and roe deer, and is one of China's ten most endangered animals and a Class I protected animal in China.

[0003] Accurate individual identification of endangered animal captive populations is more conducive to scientific genetic management of the population, effectively avoiding inbreeding and maintaining the genetic diversity and health of the population.

[0004] Currently, there are two traditional methods for identifying individuals of South China tigers: one relies on RFID (Radio Frequency Identification) chips. This technology requires implanting the chip in the animal during its infancy and allowing it to grow with the animal. In practice, the reliability of identification is limited because the chip may shift position or be damaged by physical collisions during the animal's growth, thus affecting subsequent genetic management. The other technology uses molecular identification methods, which mainly rely on the amplification of STR and SNP loci. This method requires the design of a large number of specific primers and repeated optimization to ensure stability, making the experimental process cumbersome and time-consuming.

[0005] Therefore, there is an urgent need for a simple, reliable, and low-cost method for identifying individual South China tigers. Summary of the Invention

[0006] The purpose of this invention is to propose an individual identification method based on genomic screening of SNP loci, thereby addressing the shortcomings of traditional RFID chips in individual identification. Simultaneously, it also solves the problems of existing molecular identification methods that primarily rely on STR and SNP locus amplification, requiring the design and repeated optimization of numerous specific primers in the early stages, and extensive PCR experiments for locus amplification in the later stages—a process that is cumbersome and time-consuming.

[0007] A genetic marker comprising 300 SNP molecular markers, wherein the 300 SNP molecular markers are:

[0008]

[0009]

[0010] The location and variation information of the loci are represented in the format of chromosome_physical location:reference genotype / variant allele.

[0011] Application of a genetic marker in the identification of South China tiger individuals.

[0012] A method for individual identification of South China tigers, the method being as follows:

[0013] Step 1: Extract DNA from the individual to be identified and perform genome sequencing to obtain sequencing data;

[0014] Step 2: Align the sequencing data to the AmyTig1.0 reference genome using BWA software to generate a BAM file;

[0015] Step 3: Use GATK 4.5 software to generate VCF files for each individual to be identified;

[0016] Step 4: Use the awk command and BCFtools 1.10.1 software to extract genotype data corresponding to the specified SNP loci from the VCF file of the individual to be identified;

[0017] Step 5: Use the pandas library in Python to randomly sample the extracted SNP locus list to obtain the SNP loci and their chromosomal and location information;

[0018] Step 6: Based on the SNP loci and their chromosomal and location information, filter out the corresponding genotype data to obtain the final alignment VCF file for the individual to be identified;

[0019] Step 7: Compare the final VCF file of the individual to be identified with the VCF file containing the 300 SNP molecular markers mentioned above. The target individual corresponding to the SNP molecular marker with the highest similarity is the identity of the individual to be identified.

[0020] Furthermore, the sequencing depth described in step one is 2X.

[0021] Furthermore, the number of SNP sites obtained in step five is 20.

[0022] Furthermore, the comparison method described in step seven is as follows:

[0023] Step 71: Based on the VCF file containing the above 300 SNP molecular markers, create an index using chromosome and location information and store it in a dictionary;

[0024] Step 72: Extract chromosome and location information and their corresponding genotype data from the final VCF file of the individual to be identified and compare them with 300 SNP molecular markers in the dictionary;

[0025] Step 73: During the comparison process, scan the variant sites of the individual to be identified one by one, find the sites that match the chromosome and location information of the 300 SNP molecular markers, and extract the genotype of the individual to be identified. At the same time, update the genotype data of all 300 SNP molecular markers at that site.

[0026] Step 74: Calculate the genotypic similarity between the individual to be identified and 300 SNP molecular markers and calculate the similarity score. The target individual corresponding to the SNP molecular marker with the highest score is the identity of the individual to be identified.

[0027] Beneficial effects

[0028] This invention provides a genetic marker suitable for individual identification of South China tigers, and provides an individual identification method for South China tigers based on this genetic marker, namely: the identification method is based on screening SNP sites in the genome.

[0029] The genetic markers described in this invention involve resequencing the genome of a South China tiger population to obtain all SNP loci in its genome, and then selecting 300 highly polymorphic SNP loci through rigorous quality control to serve as the 'genetic identity markers' of individuals.

[0030] The identification method described in this invention stores the aforementioned locus data in a database for individual identification. During individual identification, only low-depth 2X genome sequencing is required on the target individual. Based on the location information of the 300 stored loci, all successfully sequenced SNP loci are extracted. Twenty loci are randomly selected and compared one by one with known individuals in the database. The true identity of the individual to be identified is confirmed based on the similarity of the comparison results. Compared with traditional STR or SNP amplification methods, this method eliminates the need to design and optimize a large number of specific primers, significantly simplifying the molecular identification process and reducing experimental complexity. Individual identification through low-depth sequencing effectively saves economic costs. High-precision identification can be achieved by randomly selecting only 20 loci from the 300 loci, greatly reducing the computational burden and effectively mitigating the impact of missing data, thus improving the adaptability and reliability of the technology.

[0031] In summary, this invention establishes a convenient, reliable, and low-cost method for identifying individuals of the South China tiger by screening and preserving high-quality SNP loci, providing strong technical support for genetic management and population conservation. Attached Figure Description

[0032] Figure 1 A histogram showing the frequency distribution of SNP loci on chromosomes;

[0033] Figure 2A diagram showing the distribution of the number of SNP sites on each chromosome;

[0034] Figure 3 A probability diagram showing the coupling between the number of SNP sites and individual identification.

[0035] Figure 4 Heatmaps showing the similarity of 20 SNP loci among different individuals;

[0036] Figure 5 A graph showing the number of loci identified at different sequencing depths for different individuals;

[0037] Figure 6 Similarity score maps were generated for 20 SNP sites at different sequencing depths in different individuals. Detailed Implementation

[0038] The sources of DNA used in this experiment are shown in Table 1.

[0039] Table 1

[0040]

[0041]

[0042]

[0043] Example 1: Screening 300 highly polymorphic SNP sites.

[0044] DNA was extracted from 93 South China tigers, and then 30X whole-genome resequencing was performed. The sequencing data was then compared with the AmyTig 1.0 reference genome (https: / / ftp.cngb.org / pub / CNSA / data1 / CNP0001654 / CNS0348048 / CNA0019679) using BWA software to obtain BAM files. Subsequently, GATK 4.5 software (HaplotypeCaller module) was used to construct GVCF files for each individual, and all GVCF files were merged using the CombineGVCFs module to generate VCF files for the 93 samples. To perform quality control on the VCF files, firstly, VariantFiltration was used to filter out sites with a quality score (minQ) less than 30. Next, Bcftools 1.10.1 software was used to filter out sites with excessively high and low sequencing depths (specifically, removing the top and bottom 10% of sites). Then, BCFtools 1.10.1 software (parameter: -m2-M2-vsnps) was used to retain dia alleles. After converting the data into bed, bim, and fam files using PLINK 1.9 software, minor allele filtering was performed (parameter: --maf 0.45). Linkage disequilibrium (LD) filtering was performed using PLINK 1.9 software (parameter: --indep-pairwise 50kb 1 0.3). Finally, Hardy-Weinberg balance filtering was performed (parameter: --hwe0.001) to obtain the quality-controlled SNP dataset. This paper uses the Pandas library in Python to read the SNP dataset, counts the number of selected SNP loci on each chromosome, and plots a histogram of the frequency distribution of SNP loci on the chromosomes. Figure 1 As shown. According to Figure 1 The proportion of SNPs on each chromosome was determined by screening 300 SNP loci from all chromosomes. The distribution of SNPs on each chromosome is shown below. Figure 2 As shown. The physical locations of the selected 300 loci were determined based on the alignment results of the South China tiger reference genome using AmyTig1.0. The location and variation information of the loci are represented in the format of chromosome_physical location:reference genotype / variant allele. The location and variation information of the 300 SNP loci are shown in Table 2.

[0045] Table 2

[0046]

[0047]

[0048] Example 2: Calculation of coupling probability for individual identification with different numbers of SNP loci.

[0049] In this embodiment, to reduce computational costs, we randomly selected different numbers of SNP loci from the 300 highly polymorphic SNP loci to calculate individual coupling probabilities, in order to evaluate the impact of the number of SNP loci on individual recognition ability. Specifically, 13 quantity gradients were set: 2, 5, 10, 20, 30, 40, 50, 80, 100, 150, 200, 250, and 300 SNP loci. For each gradient, a corresponding number of loci were randomly selected from the 300 SNP loci, and 10 repeated random samplings were performed. The Match_probability.py script was used to calculate the individual coupling probability for each group of loci. First, the script iterated through the genotype data of each locus and each sample in the VCF file, counted the frequency of each allele, and calculated the coupling probability (i.e., the sum of squares of allele frequencies) based on the frequency. The script stored the results for each locus (including chromosome, location, allele frequency, and coupling probability) in a list. Then, the product of the coupling probabilities of all loci was calculated to obtain the total coupling probability. The calculation results show a significant negative correlation between the number of SNP sites and the individual identification coupling probability. The relationship between the number of SNP sites and the individual coupling probability is as follows: Figure 3 As shown. Since coupling probability is inversely proportional to accuracy, when using 20 loci for individual identification, the average coupling probability is 1 × 10⁻⁶. -6 The probability of two individuals having the exact same genotype is one in a million, which can significantly distinguish different individuals and reflect their genetic differences. Therefore, 20 SNP loci were selected. Furthermore, a heatmap of inter-individual locus similarity was calculated based on these 20 SNP loci (e.g., ...). Figure 4 As shown in the figure, all individuals exhibit significant differences, and no two individuals have completely identical genotypes. Therefore, using 20 highly polymorphic SNP loci can achieve clear and efficient individual differentiation.

[0050] Example 3: Analysis of the accuracy of individual identification of 20 SNP loci in different individuals and at different sequencing depths.

[0051] To save sequencing costs while improving the accuracy of individual identification, this embodiment verifies the impact of different sequencing depths on the accuracy of individual identification.

[0052] DNA was collected from seven individuals (numbered HNH_677, HNH_367, HNH_437, HNH_514, HNH_552, HNH_555, and HNH_585) and sequenced at four sequencing depths: 1X, 2X, 3X, and 6X.

[0053] Experimental results show that low-depth sequencing cannot cover all 300 known SNP sites; the number of detected SNP sites gradually increases with increasing sequencing depth. The changes in the number of detected sites for each organism at different sequencing depths are shown below. Figure 5 As shown. Although it is impossible to detect all 300 known SNP sites, previous studies have shown that accurate individual identification can be achieved by randomly selecting only 20 sites. Therefore, this experiment collected sequencing data from 7 individuals at different sequencing depths for further validation.

[0054] The sequencing data was processed as follows: First, the sequencing data was aligned to the AmyTig 1.0 reference genome using BWA software to generate a BAM file. Then, the SelectVariants module of GATK 4.5 software was used to generate a VCF file for each individual to be identified. The awk command was used to extract chromosome (CHROM) and position (POS) information from the SNP locus list, and the results were saved to a temporary file. Next, the view module of BCFtools 1.10.1 software was used to extract genotype data corresponding to specified SNP loci from the VCF files of the individuals to be identified. Then, the pandas library in Python was used to randomly sample the extracted SNP locus list, selecting 20 random loci and recording their chromosome and position information. Based on the sampled SNP locus list, the corresponding genotype data was again filtered from the extracted VCF files and saved as a new VCF file. The annotate module of BCFtools 1.10.1 software was used to clean the extracted VCF files, removing all unnecessary formatting fields, ultimately retaining only the VCF files containing genotype data.

[0055] Next, the SNP loci in the VCF file of the individual to be identified are compared with 300 SNP loci of the known individuals in Example 1. The comparison process is as follows: By parsing the VCF files of the known individuals and the individual to be identified, the genotypic similarity between the individual to be identified and each known individual is calculated, and a result report is generated. First, the VCF file of the known individuals is read, the sample name and corresponding genotypic data are extracted, and an index is built using chromosome and position (CHROM, POS), stored in a dictionary for quick lookup. Then, the VCF file of the individual to be identified is read, and the chromosome and position information and its corresponding genotypic data are extracted. During the comparison process, the variant loci of the individual to be identified are scanned one by one to find loci that match the chromosome and position information of the known individuals, and the genotype of the individual to be identified is extracted. At the same time, the genotypic data of all known individuals at that locus are updated synchronously. For each known individual, the genotypic consistency between it and the individual to be identified at the comparison loci is calculated, and the similarity score (i.e., the proportion of similar loci to the total number of comparison loci) is calculated. Finally, the similarity results are written to an output file, reporting the similarity scores between the individual to be identified and all known individuals, and determining the true identity of the individual to be identified based on the similarity scores. Extensive validation has shown that the site similarity score between the individual to be identified and the target individual is usually significantly higher than the similarity score between the individual to be identified and other non-target individuals. Therefore, the target individual with the highest score is the true identity of the individual to be identified.

[0056] In the genotype similarity comparison stage, to reduce the impact of random sampling error, this embodiment performed 10 replicate experiments on 20 randomly selected loci for each individual at each sequencing depth. The final results are as follows: Figure 6 As shown in the figure, the analysis indicates that when the sequencing depth is 1X, the site similarity scores between the individual to be identified and the target individual may overlap with the scores of non-target individuals, resulting in some individuals not achieving a 100% identification success rate. For example, the identification success rate for individual HNH_367 was 80%, meaning only 8 out of 10 repeated experiments were correctly identified; while the identification success rate for individual HNH_677 was only 60%. This suggests that at a sequencing depth of 1X, due to insufficient coverage, there is a certain probability of error in individual identification.

[0057] When the sequencing depth was increased to 2X, the similarity score between the individual to be identified and the target individual was significantly higher than the similarity score with non-target individuals, and the identification success rate for all individuals reached 100% in 10 replicate experiments. Further increasing the sequencing depth to 3X or 6X, a 100% identification success rate was achieved for any randomly selected 20 loci. This indicates that a minimum sequencing depth of 2X is sufficient to meet the accuracy requirements for individual identification.

[0058] Although low-depth sequencing may affect the accuracy of genotyping and lead to some genotype errors, we cannot guarantee that the individual genotypes obtained from low-depth sequencing are completely consistent with the actual genotypes in the database. However, the site similarity score between the individual to be identified and the target individual is still significantly different from the similarity score between the individual to be identified and non-target individuals, which is sufficient to support the accurate determination of individual identity.

[0059] In summary, this embodiment demonstrates through experiments that an efficient and low-cost individual identification method can be achieved at a 2X sequencing depth using 20 randomly selected loci.

[0060] Example 4. Setting up a unique "genomic ID card" through SNP sites.

[0061] Twenty SNP loci were used to assign a unique "genomic ID card" to each individual. The rules for assigning this unique "genomic ID card" are as follows: since SNPs are usually dimorphic sites, involving only two different bases, the genotype of each SNP locus in the genome can only have 16 possible values: AA, AT, AC, AG, TA, TT, TC, TG, CC, CA, CT, CG, GA, GT, GC, and GG. The genotype coding rules are defined in Table 3. Twenty loci were extracted from 300 high-quality loci for each individual, and the genotypes of these 20 loci were encoded using the above method to obtain each individual's unique "genomic ID card."

[0062] Table 3

[0063]

[0064]

[0065] We extracted 20 SNP loci from 7 individuals (numbered HNH_367, HNH_437, HNH_514, HNH_552, HNH_555, HNH_585, and HNH_677), and the individual and locus information is shown in Table 4. The location and variation information of the loci are represented in the format of chromosome_physical location:reference genotype / variant allele. According to the encoding method in Table 2, a unique "genomic ID card" was generated for each individual, as shown in Table 5. The encoded genomic ID card is equivalent to a unique identifier based on genetic information, which can effectively distinguish individuals, facilitate display, and hide the individual's locus and genotype information.

[0066] Table 4

[0067]

[0068]

[0069] Table 5

[0070] Individual ID Genomic ID Card HNH_367 HIIMDDJFIGFAMIPFIPMP HNH_437 FFIMDDJGFIECMGPFOMMP HNH_514 FFLPDDIFGGECAGPGOMPA HNH_552 HKLPDPIFGGFCAGDFIMPP HNH_555 PKIAAAJFGIFAPIDGIMAP HNH_585 HFPPPPAIFFAIMFAIIMMA HNH_677 PFIPPPJIFFECPGDGIMMA

Claims

1. Application of 300 SNP molecular markers combination in individual identification of South China tigers, characterized in that, The 300 SNP molecular markers are combined as follows: Among them, the position of the site and the variation information are represented in the format of Chromosome _ Physical Position: Reference Genotype / Variation Allele Genotype, and the physical positions of the selected 300 sites are determined based on the AmyTig1.0 reference genome alignment result.

Citation Information

Patent Citations

  • Genetic marker for individual recognition and / or paternity test of South China tiger and application

    CN115961054A