No-reference genome species individual identification method based on simple genome assembly
By assembling a simplified genome and constructing a population SNP library, combined with low-depth sequencing and SNP locus alignment, the problem of individual identification in species without a reference genome has been solved, realizing an efficient and low-cost individual identification method that is applicable to the genetic management of endangered animals and non-model species.
Patent Information
- Application Number
- CN202511030508.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2025-11-11
AI Technical Summary
Existing technologies for individual identification in species lacking high-quality reference genomes are cumbersome and time-consuming, making accurate individual identification difficult.
By assembling a simplified genome and constructing a population SNP library, low-depth sequencing is performed, combined with a small number of SNP site alignments, to achieve individual identification, reducing dependence on the reference genome and the requirements for sequencing depth.
It simplifies the operation process, reduces economic costs, and improves the adaptability and reliability of individual identification, making it suitable for genetic monitoring and research of endangered animals and non-model species.
Smart Images

Figure CN120932735A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of bioinformatics and genetics, and in particular to a method and application for identifying individuals of species without a reference genome based on simplified genome assembly. Background Technology
[0002] For endangered species, whether in captivity in zoos or in the wild, accurate individual identification is fundamental to the scientific genetic management of populations. From a molecular biology perspective, current molecular identification methods primarily rely on the amplification of STR and SNP loci. This requires the design and repeated optimization of numerous specific primers to ensure stability, and the subsequent individual identification process still necessitates extensive PCR experiments for locus amplification—a cumbersome and time-consuming process. With the development of genomics, genomic data-based individual identification technologies have played a crucial role in endangered species conservation, genetic resource management, and non-model species research. However, this often depends on high-quality reference genomes. Many non-model species still lack reference genomes, significantly limiting the application of individual identification technologies in these species. Therefore, developing an individual identification method for species lacking high-quality reference genomes and operating at low sequencing depth is key to solving these challenges in genetic monitoring of endangered animals. Summary of the Invention
[0003] To address the problems existing in current technologies, this invention proposes an individual identification method based on simplified genome assembly. By assembling a simplified reference genome and constructing a population SNP library, low-depth sequencing is performed, combined with alignment of a small number of SNP sites, to achieve accurate individual identification. This method not only reduces dependence on a reference genome but also significantly reduces the requirements for sequencing depth and data processing, exhibiting broad adaptability and practicality.
[0004] In order to achieve the purpose of this invention, after research, the technical solution provided by this invention is as follows:
[0005] In a first aspect, the present invention provides a method for identifying individuals of a species without a reference genome based on simplified genome assembly, the method comprising the following steps:
[0006] (1) Assembly of a simplified genome: This includes extracting DNA from any individual sample in a population of a species without a reference genome, performing whole-genome de novo sequencing, assembling the sequencing depth to 30X or higher to obtain assembled sequencing data; and calculating the average insert length after trimming and filtering the sequencing reads; then performing simplified genome assembly, setting the K-mer length to 31-63; and extracting discontinuous long scaffold sequences with a length exceeding 1000bp as the simplified reference genome of the species without a reference genome.
[0007] (2) Construction of population SNP library: including mutation detection, extracting SNP mutation sites in the population, filtering out SNP sites within 50bp at both ends of each discontinuous long scaffold sequence, then filtering out all SNP sites with biallelic genes, then filtering out SNP sites with secondary allele frequency maf < 0.4, and finally deleting all missing SNP sites in the population to obtain population SNP library;
[0008] (3) Using low-depth sequencing authentication data for individual identification: including extracting DNA from the individual to be identified, performing low-depth genome resequencing to obtain authentication data, wherein the depth of the low-depth sequencing is 2X or higher; comparing the authentication data with the simplified reference genome; comparing the discontinuous long scaffolds and location information of all SNP sites in the population SNP library with the discontinuous long scaffolds and location information of the SNP variant sites extracted from the individual to be identified, to obtain the SNP sites common to the population SNP library and the SNPs of the individual to be identified; arbitrarily selecting more than 100 SNP sites from the SNP sites of the individual to be identified and comparing their genotypes with the SNP sites of all individuals in the population SNP site library.
[0009] In this invention, the population of a species without a reference genome includes at least 2, at least 5, at least 8, at least 10, at least 12, at least 15, at least 18, at least 20, at least 25, at least 30, at least 40, at least 50, at least 60, at least 80, at least 100, at least 500, at least 1000, at least 2000, at least 5000, at least 10000, and at least 20000 individuals.
[0010] In some preferred embodiments, the reference genome-free species is a mammal. In some more preferred embodiments, the reference genome-free species is a feline.
[0011] Further, step (1) includes the following steps:
[0012] Step (1-1): Extract DNA from any individual sample in a population of a species without a reference genome, perform whole-genome de novo sequencing, and obtain assembled sequencing data with an assembly sequencing depth of 30X or higher.
[0013] Steps (1-2): After trimming and filtering the sequencing reads using trimmomatic software, the average insert length of the data is calculated using fastp software; then, a simplified genome assembly is performed using SOAPdenovo software, with the K-mer length set to 31-63 to generate a genome file; sequences with discontinuous long scaffold lengths exceeding 1000 bp are extracted from the generated genome file and used as a simplified reference genome for the species without a reference genome.
[0014] Further, in step (1-1), the assembly sequencing depth includes 30X, 40X, or 50X.
[0015] Furthermore, step (2) includes the following steps:
[0016] Step (2-1): Extract DNA from at least two individuals to be tested from the population of the species without a reference genome, and perform whole-genome resequencing to obtain sequencing data;
[0017] Step (2-2): Use BWA software to align the sequencing data with the simplified reference genome to generate the first BAM file;
[0018] Steps (2-3): Use Samtools software to sort the data in the first BAM file according to the simplified reference genome coordinates, and use GATK software to mark and remove PCR repeats and optical repeats;
[0019] Step (2-4): Use the Samtools software to filter out records with MAPQ < 20, and at the same time filter out read segments in the flag bits that are not matched, are not primary matches, fail quality control, or are supplemented matches;
[0020] Steps (2-5): Use GATK software to perform mutation detection, generate gvcf files, and merge the files of each individual into a population file. Extract the SNP mutation sites present in the population file, use bcftools software to filter out SNP sites within 50bp at both ends of each discontinuous long scaffold sequence, then filter out all SNP sites with biallelic alleles, then filter out SNP sites with a secondary allele frequency maf < 0.4, and finally delete all missing SNP sites in the population VCF file to obtain the population SNP library.
[0021] Furthermore, in step (2-1), the sequencing depth of the whole genome resequencing is 20X or higher.
[0022] Furthermore, step (3) includes the following steps:
[0023] Step (3-1): Extract DNA from the individual to be identified, perform low-depth genome sequencing, and obtain identity verification data;
[0024] Step (3-2): Use BWA software to align the authentication data with the simplified reference genome to generate a second BAM file;
[0025] Step (3-3): Use Samtools software to sort the data in the second BAM file according to the simplified reference genome sequence order, and use GATK software to mark and remove PCR repeats and optical repeats;
[0026] Steps (3-4): Use GATK software to generate VCF files for each individual to be identified, and use BCFtools software to extract SNP variant sites from the VCF files of the individuals to be identified;
[0027] Step (3-5): Use the pandas library in Python to compare the discontinuous long scaffolds and location information of all SNP sites in the population SNP library with the discontinuous long scaffolds and location information of the SNP variant sites extracted from the VCF file of the individual to be identified, and obtain the common SNP sites with the same discontinuous long scaffolds and locations as the SNPs in the population SNP library and the SNPs in the individual to be identified.
[0028] Step (3-6): Randomly select more than 100 SNP loci from the SNP loci of the individual to be identified and compare them with the SNP loci of all individuals in the population SNP locus database for genotyping; the comparison information includes the discontinuous long scaffold of the SNP locus, its location coordinates and genotype.
[0029] Furthermore, in steps (3-6), the depth of the low-depth resequencing of the authentication data includes 2X, 3X, 4X, 5X or 6X.
[0030] Further, in steps (3-6), when the depth of low-depth sequencing of the authentication data is 2X, at least 200 SNP sites are randomly selected from the SNP sites of the individual to be identified; when the depth of low-depth sequencing of the authentication data is 3X, at least 200 SNP sites are randomly selected from the SNP sites of the individual to be identified; when the depth of low-depth sequencing of the authentication data is 4X, 5X or 6X, at least 100 SNP sites are randomly selected from the SNP sites of the individual to be identified.
[0031] Further, in steps (3-6), the comparison method includes:
[0032] Step (3-6-1): Extract the discontinuous long scaffold and location information of more than 100 SNP sites from the VCF file of the individual to be identified, build an index using a Python script and store it in a dictionary;
[0033] Step (3-6-2): Extract the SNP loci at the corresponding positions from the population SNP database, and compare the population SNP database with the genotype data of the individual SNP to be identified;
[0034] Step (3-6-3): During the comparison process, the SNP variant sites of the individual to be identified are scanned one by one to find the genotypic differences with the SNP sites at the same position in each individual in the population, the genotypic consistency of all SNP sites is statistically analyzed and the similarity score is calculated.
[0035] Secondly, the present invention provides the application of the above-described method in the identification of individuals in species without a reference genome.
[0036] The beneficial effects of the present invention include at least the following:
[0037] This invention provides a method for individual identification of species without a reference genome based on simplified genome assembly. This method eliminates the reliance on high-quality reference genomes, solving the technical challenge of individual identification in many non-model species due to the lack of a reference genome. It is applicable to genetic monitoring and research of endangered animals and other non-model species. In some implementation schemes, only low-depth 2X genome sequencing is required for the target individual. More than 200 SNP loci are compared with genotypes in a population SNP database, and the target individual's identity is confirmed based on the similarity of the comparison results. Furthermore, in some preferred implementation schemes, only low-depth 6X genome sequencing is required for the target individual. More than 100 SNP loci are compared with genotypes in a population SNP database, and the target individual's identity is confirmed based on the similarity of the comparison results. Compared with traditional STR or SNP amplification methods, this invention eliminates the need for designing and optimizing a large number of specific primers, significantly simplifying the molecular identification process and reducing experimental complexity. Individual identification through low-depth sequencing effectively saves economic costs. Achieving high-precision identification with only a small number of loci significantly reduces computational burden and effectively mitigates the impact of missing data, thus improving the adaptability and reliability of the technology. In summary, this invention establishes a convenient, reliable, and low-cost method for identifying individuals in species without a reference genome by screening and preserving SNP loci, providing strong technical support for genetic management and population conservation.
[0038] Other features and advantages of the present invention will be described in detail in the following detailed description section. Attached Figure Description
[0039] Figure 1 This shows the quality assessment results of the simplified genome assembly.
[0040] Figure 2 This image shows the genotyping results of the fluorescence signal of the SNP locus SDA02-p10176749 identified in Example 2. (The impact of sequencing depth on the applicability of the simplified reference genome in individual identification.)
[0041] Figure 3 This image shows the genotyping results of the fluorescence signal of the SNP locus SDA02-p10668400 identified in Example 2. (Feasibility analysis of individual identification using different low-depth sequencing data)
[0042] Figure 4 This image shows the genotyping results of the fluorescent signal of the SNP locus SDA02-p8849037 identified in Example 2. (The impact of different numbers of SNP loci on the feasibility of individual identification)
[0043] Figure 5 This image shows the genotyping results of the fluorescence signal of the SNP locus SDA02-p9195997 identified in Example 2. (Analysis of the optimal individual identification threshold under different sequencing depths and different numbers of SNP loci).
[0044] Figure 6 shows the results of the individual identification accuracy verification analysis for the four individuals. Detailed Implementation
[0045] To make the objectives, technical solutions, and beneficial effects of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Examples of the embodiments are shown in the accompanying drawings. It should be understood that the specific embodiments described in the following embodiments of the invention are merely illustrative examples of specific implementations of the invention and are intended to explain the invention, but do not constitute a limitation thereof.
[0046] The endpoints and any values of the ranges disclosed in this invention are not limited to the precise ranges or values, and these ranges or values should be understood to include values close to these ranges or values. For numerical ranges, the endpoint values of the various ranges, the endpoint values of the various ranges and individual point values, and individual point values can be combined with each other to obtain one or more new numerical ranges, which should be considered as specifically disclosed herein. In the description of this application, unless otherwise stated, terms such as "multiple" mean two / agents or more.
[0047] The sources of the individual DNA used in the following examples are shown in Table 1.
[0048] Table 1 shows the sources of individual DNA in the examples.
[0049]
[0050]
[0051] Assuming that the groups and individuals in Table 1 above are unknown, the method of the present invention will be verified through the following embodiments.
[0052] Example 1
[0053] I. Quality of simplified reference genomes assembled from sequencing data at different depths and their impact on individual identification analysis
[0054] In order to minimize the assembly sequencing depth while ensuring the quality of the simplified reference genome assembly, this embodiment uses whole genome de novo sequencing data with different assembly sequencing depths (20X, 30X, 40X, 50X) to assemble a simplified genome, and analyzes the quality of the simplified assembled genome and its performance in individual identification.
[0055] Assembly of the simplified genome: First, DNA was extracted from the HEB-010-2 individual sample in the population of species without a reference genome listed in Table 1. The raw sequencing data of this individual was randomly sampled proportionally using seqtk (v1.3-r106; https: / / github.com / lh3 / seqtk) software for de novo whole-genome sequencing, generating datasets with assembly sequencing depths of 20X, 30X, 40X, and 50X, respectively. Subsequently, SOAPdenovo2 (v2.04-r241) was used to assemble the reference genome from the datasets at different assembly sequencing depths, with the K-mer length set to 63 and all other parameters set to default values. Sequences with discontinuous long scaffold lengths exceeding 1000 bp were extracted and used as the simplified reference genome for the species without a reference genome. Specifically, the following steps were included:
[0056] Step (1-1): Extract DNA from individual samples of HEB-010-2 in the population of species without reference genomes in Table 1, and perform whole genome de novo sequencing. The assembly sequencing depth is 20X, 30X, 40X and 50X to obtain assembly sequencing data.
[0057] Steps (1-2): After trimming and filtering the sequencing reads using trimmomatic software, the average insert length of the data is calculated using fastp software; then, a simplified genome assembly is performed using SOAPdenovo software, with the K-mer length set to 63, to generate a genome file; sequences with non-contiguous long scaffold lengths exceeding 1000 bp are extracted from the generated genome file and used as a simplified reference genome for the species without a reference genome.
[0058] The quality assessment results of the simplified genome assembly are as follows: Figure 1 As shown, with increasing assembly sequencing depth, the N50 value of the simplified assembled genome continuously increases, indicating that the degree of genome fragmentation gradually decreases. Simultaneously, the number of discontinuous long scaffolds exceeding 1 Kbp in length decreases with increasing sequencing depth, from 707,708 scaffolds assembled at 20X sequencing depth to 573,852 scaffolds assembled at 50X sequencing depth; while the number of discontinuous long scaffolds exceeding 10 Kbp in length increases significantly, from 6,514 at 20X to 27,913 at 50X. This suggests that high-depth sequencing data at assembly sequencing depth can better utilize pairing information, effectively merging short discontinuous long scaffolds, thereby generating more and longer discontinuous long scaffolds.
[0059] To evaluate the applicability of simplified reference genomes assembled at different sequencing depths for individual identification, sequencing data from samples with a low sequencing depth of 6X were used as the test set. Reference genomes assembled from sequencing data at depths of 20X, 30X, 40X, and 50X were used as comparison reference genomes. Finally, 1000 SNP loci of the individuals to be identified were selected for individual identification testing. The results are as follows: Figure 2 As shown, the reference genome assembled using 20X depth data failed to accurately identify individuals. The genotype matching rate of the target real individuals and the false matching rate of the non-target real individuals (where the real individuals were assumed to be unknown individuals to be identified) highly overlapped across all samples, with a success rate of 0%. In contrast, when using a reference genome assembled at a depth of 30X or higher, the genotype matching rate of the target real individuals and the false matching rate of the non-target real individuals were completely separated, achieving a success rate of 100%. Under these three conditions (30X, 40X, and 50X), the target individuals significantly outperformed other individuals in all repeated tests, and their score distribution was clearly distinguishable from that of non-target individuals, demonstrating excellent identification performance. This result indicates that a simplified reference genome, at a sequencing depth of 30X or higher, possesses reliable individual identification capabilities, while assembly quality at lower depths may struggle to support effective identity determination.
[0060] Further analysis of the read alignment rates between the raw sequencing data of the simplified assembled reference genome and its assembled reference genome revealed the following results (as shown in Table 2): The most significant improvement in read alignment rate was observed in the 20X to 30X range. However, beyond 30X, the improvement in read alignment rate gradually decreased with increasing sequencing depth. Therefore, for assembled reference genomes, a sequencing depth of around 30X is sufficient to achieve a high alignment rate, and further increases in sequencing depth have limited effect on improving the alignment rate.
[0061] Table 2. Read alignment rates of raw sequencing data from the simplified assembled reference genome and its assembled simplified reference genome.
[0062] Assembly Sequencing Depth Comparison of reading passages Total reading section Comparison rate 20X 453,098,774 508,183,689 89.16% 30X 713,904,755 760,885,325 93.83% 40X 959,334,027 1,014,803,958 94.53% 50X 1,228,765,981 1,295,789,013 94.83%
[0063] II. Feasibility Analysis of Individual Identification Using Different Low-Depth Sequencing Data
[0064] To save sequencing costs while improving the accuracy of individual identification, this embodiment uses a simplified reference genome assembled from 30X depth data to further explore the impact of low-depth sequencing data of different depths on the feasibility of individual identification.
[0065] Construction of the population SNP library: This includes mutation detection, extraction of SNP variant sites in the population, filtering out SNP sites within 50 bp of both ends of each discontinuous long scaffold sequence, filtering out all SNP sites with biallelic alleles, filtering out SNP sites with a secondary allele frequency (maf) < 0.4, and finally deleting all missing SNP sites in the population to obtain the population SNP library. Specifically, it includes the following steps:
[0066] Step (2-1): Blood DNA was collected from 16 Siberian tiger individuals (as shown in Table 1, numbered AT_hd88, AT_hd26, AT_1578, AT_1367, AT_hd36, AT_1731, AT_hd32, AT_he24, AT_1716, dbh_1761, AT_hd70, HEB-1129, HEB-1383, HEB-1712, HEB-1761, HEB-182, HEB-010-2), and 20X whole-genome resequencing was performed to obtain sequencing data;
[0067] Step (2-2): Use BWA software to align the sequencing data with a simplified reference genome assembled from 30X sequencing data to generate the first BAM file;
[0068] Step (2-3): Use Samtools software to sort the data in the first BAM file according to the coordinates of the simplified reference genome, and use GATK software to mark and remove PCR repeats and optical repeats;
[0069] Step (2-4): Use the Samtools software to filter out records with MAPQ < 20, and at the same time filter out read segments in the marker bits that are not matched, non-primary matches, fail quality control, or require supplementary matching;
[0070] Steps (2-5): Use GATK software to perform mutation detection, generate gvcf files, and merge the files of each individual into a population file. Extract the SNP mutation sites present in the population file, use bcftools software to filter out SNP sites within 50bp at both ends of each discontinuous long scaffold sequence, then filter out all SNP sites with biallelic alleles, then filter out SNP sites with a secondary allele frequency maf < 0.4, and finally delete all missing SNP sites in the population VCF file to obtain the population SNP library.
[0071] Next, individual identification is performed using low-depth sequencing authentication data: this includes extracting DNA from the individual to be identified, performing low-depth genome resequencing to obtain authentication data, wherein the low-depth sequencing depth is 2X or higher; comparing the authentication data with the simplified reference genome; comparing the discontinuous long scaffolds and location information of all SNP sites in the population SNP library with the discontinuous long scaffolds and location information of the SNP variant sites extracted from the individual to be identified, to obtain the SNP sites shared by the population SNP library and the SNPs of the individual to be identified; and arbitrarily selecting more than 100 SNP sites from the SNP sites of the individual to be identified and performing genotyping with the SNP sites of all individuals in the population SNP site library. Specifically, this includes the following steps:
[0072] Step (3-1): Fecal DNA from 12 Siberian tigers to be identified (AT_hd88, AT_hd26, AT_1578, AT_1367, AT_hd36, AT_1731, AT_hd32, AT_he24, AT_1716, dbh_1761, AT_hd70, HEB-010-2) was used for low-depth sequencing at six sequencing depths: 1X, 2X, 3X, 4X, 5X, and 6X, to obtain identity verification data.
[0073] Step (3-2): Align the authentication data to the simplified reference genome generated from the 30X sequencing sample using BWA software to generate a second BAM file;
[0074] Step (3-3): Use Samtools software to sort the data in the second BAM file according to the simplified reference genome sequence order, and use GATK software to mark and remove PCR repeats and optical repeat reads.
[0075] Steps (3-4): Use GATK software to generate VCF files for each individual to be identified, and use BCFtools software to extract SNP variant sites from the VCF files of the individuals to be identified;
[0076] Step (3-5): Use the pandas library in Python to compare the discontinuous long scaffolds and location information of all SNP sites in the population SNP library with the discontinuous long scaffolds and location information of the SNP variant sites extracted from the VCF file of the individual to be identified, and obtain the common SNP sites with the same discontinuous long scaffolds and locations as the SNPs in the population SNP library and the individual to be identified.
[0077] Step (3-6): Randomly select more than 100 SNP loci from the SNP loci of the individual to be identified and compare them with the SNP loci of all individuals in the population SNP locus database for genotyping; the comparison information includes the discontinuous long scaffold of the SNP locus, the location coordinates and the genotype.
[0078] In steps (3-6), the comparison method includes:
[0079] Step (3-6-1): Extract the discontinuous long scaffold and location information of more than 100 SNP sites from the VCF file of the individual to be identified, build an index using a Python script and store it in a dictionary;
[0080] Step (3-6-2): Extract the SNP loci at the corresponding positions from the population SNP database, and compare the population SNP database with the genotype data of the individual SNP to be identified;
[0081] Step (3-6-3): During the comparison process, the SNP variant sites of the individual to be identified are scanned one by one to find the genotypic differences with the SNP sites at the same position in each individual in the population, the genotypic consistency of all SNP sites is statistically analyzed and the similarity score is calculated.
[0082] The calculation is performed according to the following formula:
[0083] Similarity score = Number of matching SNP sites / Total number of SNP sites
[0084] The number of matching SNP sites is the number of SNP sites in the individual to be identified that have the same genotype at the same location as each individual in the population; the total number of SNP sites is the total number of all SNP sites in the population SNP database.
[0085] The genotype matching rates of real individuals and irrelevant individuals (where "real individuals" refers to hypothetically unknown individuals to be identified, and "irrelevant individuals" refers to all other individuals within the same population distinct from the real individual) were calculated. A Δ_value was defined, calculated by subtracting the genotype matching rate of irrelevant individuals from the genotype matching rate of real individuals. The Δ_values of all individuals were statistically analyzed and correlated with the depth of low-depth sequencing. The results are as follows: Figure 3 As shown, the Δ_value increases with sequencing depth. From 1X to 6X sequencing depths, the median value gradually increases, indicating a positive correlation between sequencing depth and Δ_value. At a sequencing depth of 1X, Δ_values sometimes fall below 0, indicating that the genotype matching rate of the actual individual is lower than that of unrelated individuals, suggesting potential misidentification at 1X sequencing depth. Increasing the sequencing depth reveals that Δ_values are all greater than 0, indicating that at depths of 2X and above, the genotype matching rate of the actual individual is significantly higher than that of unrelated individuals, making individual identification possible.
[0086] III. The Impact of Different Numbers of SNP Loci on the Feasibility of Individual Identification
[0087] To conserve computational resources and reduce computation time, the impact of varying numbers of SNP loci on the feasibility of individual identification was verified by reducing the number of SNP loci in the sample of the individuals to be identified. Eight SNP locus number gradients were set: 20, 50, 100, 200, 300, 400, 500, and 600. SNP loci of corresponding numbers from the number gradients were randomly extracted from the dataset of common SNP loci of the individuals to be identified, and this process was repeated 10 times. Genotypic similarity scores for different numbers of SNP loci in the individuals to be identified were calculated, and the Δ_value was further calculated. The results are shown below. Figure 4 As shown, when the sequencing depth is 1X, even with 600 loci for identification, there are still cases where the Δ_value is less than 0, indicating that the individual's true identity is being misidentified. However, when the sequencing depth is increased to 2X, using more than 200 loci, the Δ_value is always greater than 0, indicating that individual identification is possible. Therefore, at a minimum sequencing depth of 2X, using at least 200 SNP loci of the individual to be identified is sufficient for individual identification.
[0088] IV. Optimal Threshold Setting for Individual Identification at Different Sequencing Depths and Numbers of SNP Loci
[0089] As discussed in Part III above, the genotype matching rate of real individuals was generally significantly higher than that of non-real individuals, while the matching rate differences between different non-real individuals were relatively small. Therefore, this part further calculated the Δ_threshold (i.e., the difference between the genotype matching rate (TMR) of real individuals and the highest false matching rate (FMR) of irrelevant individuals) and the "secondary difference" (i.e., the difference between the highest FMR and the second highest FMR) under different low-depth sequencing depths (2X to 6X) and the number of SNP loci of the individuals to be identified (50 to 600). ROC analysis was used to evaluate the Δ_threshold and the secondary difference to determine the optimal Δ_threshold. The results are as follows: Figure 5 As shown, the AUC values under all conditions are higher than 0.5, and the ROC curve is close to the upper left corner, indicating that the model performs well under high true positive rate (TPR) and low false positive rate (FPR), suggesting that the classification model has a certain discriminative ability.
[0090] Wherein, the true positive rate (TPR) = TP / (TP + FN);
[0091] TP (True Positive): The number of samples where Δ_threshold ≥ the threshold and are actually real individuals;
[0092] FN(False Negative): Δ_threshold < threshold and the actual number of samples representing real individuals;
[0093] False positive rate (FPR) = FP / (FP + TN);
[0094] FP(False Positive): The number of samples where Δ_threshold ≥ the threshold and are actually irrelevant individuals;
[0095] TN (True Negative): The number of samples where Δ_threshold < threshold and is actually irrelevant individuals.
[0096] When using a 2X low-depth sequencing depth and 50 SNP loci for the individual to be identified, the AUC value is 0.76, showing moderate discriminative ability, but accompanied by a high error rate and limited discriminative accuracy. When the number of SNP loci for the individual to be identified is ≥200, the AUC value is greater than 0.9, indicating that the samples can be distinguished well under this condition. As the low-depth sequencing depth increases, the performance of the classifier also improves. When using low-depth sequencing depths of 2X or higher, the AUC value is above 0.8 under all conditions, indicating strong discriminative ability. In particular, at a 3X low-depth sequencing depth, when the number of SNP loci for the individual to be identified reaches 200 or more, the AUC value reaches 1.00, showing 100% accuracy. When the low-depth sequencing depth is 3X, 4X, 5X, or 6X, with more than 200 SNP loci for the individual to be identified, the classifier's AUC value can reach 1.00, demonstrating extremely excellent classification ability. In comparison, the 2X classifier performs weaker, but its AUC value significantly improves when the number of SNP loci in the individual to be identified is ≥200, demonstrating good recognition performance. Furthermore, the optimal Δ_threshold obtained by the Youden index and Euclidean distance method was calculated (as shown in Table 3). The results show that the thresholds obtained by the two calculation methods are very close. This invention provides the optimal Δ_threshold obtained by the Youden index method, which can be used as a standard for judging the accuracy of individual identification. When determining the identity of an unknown individual, as long as the highest genotype matching rate score minus the second highest genotype matching rate score is ≥ the optimal Δ_threshold, the true identity of the individual can be accurately determined.
[0097] Table 3 Optimal Threshold - Yoden
[0098]
[0099]
[0100] V. Evaluation of the accuracy of individual identification for four individuals with different sequencing depths and SNP loci
[0101] After determining the optimal threshold, the Δ_threshold was further evaluated for four individuals (HEB-1129, HEB-1712, HEB-1761, and HEB-182) under different low-depth sequencing depths and the number of SNP sites in the individuals to be identified, to accurately determine the true identity of the individuals. First, low-depth sequencing was performed on the fecal DNA of the four individuals at five sequencing depths: 2X, 3X, 4X, 5X, and 6X, to obtain identity verification data. The identity verification data was then aligned to a simplified reference genome generated from 30X sequencing samples using BWA software to generate a second BAM file. The data in the second BAM file was sorted according to the simplified reference genome sequence using Samtools software, and GATK software was used to label and remove PCR repeats and optical repeats. A VCF file for each individual was generated using GATK software, and SNP variant sites were extracted from the VCF files of the individuals using BCFtools software. Using the pandas library in Python, the discontinuous long scaffolds and location information of all SNP sites in the population SNP library were compared with the SNP variant sites extracted from the VCF file of the individual to be identified, to obtain the SNP sites shared by the population SNP library and the SNPs of the individual to be identified. Seven SNP site number gradients were set, namely 50, 100, 200, 300, 400, 500 and 600. Sites of the corresponding number gradients were randomly extracted from the dataset of SNP sites shared by the individual to be identified, and this was repeated 10 times. The genotype similarity score of different number of sites was calculated, and the Δ_threshold was further calculated. The results are shown in Figure 6 (A, B, C, D). At the lowest sequencing depth of 2X, at least 200 SNP sites are required to satisfy that all calculated Δ_thresholds are greater than or equal to the optimal threshold, accurately identifying the individual's true identity. At a low sequencing depth of 3X, for some individuals, using at least 100 SNP sites, all calculated Δ_thresholds are greater than or equal to the optimal threshold, accurately identifying the individual's true identity. However, for some individuals, the calculated Δ_threshold is less than the optimal threshold. Increasing the number of SNP sites to at least 200 allows for accurate identification of all individuals. At low sequencing depths of 4X, 5X, and 6X, using at least 100 SNP sites for the individual to be identified, all Δ_thresholds exceed the optimal threshold, enabling accurate individual identification.
[0102] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and do not constitute a limitation on the content of the present invention. Within the scope of the technical concept of the present invention, various simple modifications can be made to the technical solutions of the present invention, including combining various technical features in any other suitable manner. These simple modifications and combinations should also be regarded as the content disclosed in the present invention and all fall within the protection scope of the present invention.
Claims
1. A method for identifying individuals of a species without a reference genome based on simplified genome assembly, characterized in that, The method includes the following steps: (1) Assembly of a simplified genome: This includes extracting DNA from any individual sample in a population of a species without a reference genome, performing whole-genome de novo sequencing, assembling the sequencing depth to 30X or higher to obtain assembled sequencing data; and calculating the average insert length after trimming and filtering the sequencing reads; then performing simplified genome assembly, setting the K-mer length to 31-63; and extracting discontinuous long scaffold sequences with a length exceeding 1000 bp as the simplified reference genome of the species without a reference genome. (2) Construction of the population SNP library: This includes performing mutation detection, extracting SNP mutation sites in the population, filtering out SNP sites within 50 bp at both ends of each discontinuous long scaffold sequence, then filtering out all SNP sites with biallelic genes, then filtering out SNP sites with secondary allele frequency maf < 0.4, and finally deleting all missing SNP sites in the population to obtain the population SNP library. (3) Using low-depth sequencing authentication data for individual identification: including extracting DNA from the individual to be identified, performing low-depth genome resequencing to obtain authentication data, wherein the depth of the low-depth sequencing is 2X or higher; comparing the authentication data with the simplified reference genome; comparing the discontinuous long scaffolds and location information of all SNP sites in the population SNP library with the discontinuous long scaffolds and location information of the SNP variant sites extracted from the individual to be identified, to obtain the SNP sites common to the population SNP library and the SNPs of the individual to be identified; arbitrarily selecting more than 100 SNP sites from the SNP sites of the individual to be identified and comparing their genotypes with the SNP sites of all individuals in the population SNP site library.
2. The method for identifying individuals of a species without a reference genome based on simplified genome assembly according to claim 1, characterized in that, Step (1) further includes the following steps: Step (1-1): Extract DNA from any individual sample in a population of a species without a reference genome, perform whole-genome de novo sequencing, and obtain assembled sequencing data with an assembly sequencing depth of 30X or higher. Steps (1-2): After trimming and filtering the sequencing reads using trimmomatic software, the average insert length of the data is calculated using fastp software; then, a simplified genome assembly is performed using SOAPdenovo software, with the K-mer length set to 31-63 to generate a genome file; sequences with discontinuous long scaffold lengths exceeding 1000 bp are extracted from the generated genome file and used as a simplified reference genome for the species without a reference genome.
3. The method for identifying individuals of a species without a reference genome based on simplified genome assembly according to claim 2, characterized in that, In step (1-1), the assembly sequencing depth includes 30X, 40X or 50X.
4. The method for identifying individuals of a species without a reference genome based on simplified genome assembly according to claim 1 or claim 2, characterized in that, Step (2) includes the following steps: Step (2-1): Extract DNA from at least two individuals to be tested from the population of the species without a reference genome, and perform whole-genome resequencing to obtain sequencing data; Step (2-2): Use BWA software to align the sequencing data with the simplified reference genome to generate the first BAM file; Step (2-3): Use Samtools software to sort the data in the first BAM file according to the coordinates of the simplified reference genome, and use GATK software to mark and remove PCR repeats and optical repeats; Step (2-4): Use the Samtools software to filter out records with MAPQ < 20, and at the same time filter out read segments in the flag bits that are not matched, non-primary matches, fail quality control, or require supplementary matching; Steps (2-5): Use GATK software to perform mutation detection, generate gvcf files, and merge the files of each individual into a population file. Extract the SNP mutation sites present in the population file, use bcftools software to filter out SNP sites within 50 bp at both ends of each discontinuous long scaffold sequence, then filter out all SNP sites with biallelic alleles, then filter out SNP sites with a secondary allele frequency maf < 0.4, and finally delete all missing SNP sites in the population VCF file to obtain the population SNP library.
5. The method for identifying individuals of a species without a reference genome based on simplified genome assembly according to claim 1, characterized in that, In step (2-1), the sequencing depth of the whole genome resequencing is 20X or higher.
6. The method for identifying individuals of a species without a reference genome based on simplified genome assembly according to claim 1, characterized in that, Step (3) includes the following steps: Step (3-1): Extract DNA from the individual to be identified, perform low-depth genome sequencing, and obtain identity verification data; Step (3-2): Use BWA software to align the authentication data with the simplified reference genome to generate a second BAM file; Step (3-3): Use Samtools software to sort the data in the second BAM file according to the simplified reference genome sequence order, and use GATK software to mark and remove PCR repeats and optical repeats; Steps (3-4): Use GATK software to generate VCF files for each individual to be identified, and use BCFtools software to extract SNP variant sites from the VCF files of the individuals to be identified; Step (3-5): Use the pandas library in Python to compare the discontinuous long scaffolds and location information of all SNP sites in the population SNP library with the discontinuous long scaffolds and location information of the SNP variant sites extracted from the VCF file of the individual to be identified, and obtain the common SNP sites with the same discontinuous long scaffolds and locations as the SNPs in the population SNP library and the SNPs in the individual to be identified. Step (3-6): Randomly select more than 100 SNP loci from the SNP loci of the individual to be identified and compare them with the SNP loci of all individuals in the population SNP locus database for genotyping; the comparison information includes the discontinuous long scaffold of the SNP locus, its location coordinates and genotype.
7. The method for identifying individuals of a species without a reference genome based on simplified genome assembly according to claim 6, characterized in that, In steps (3-6), the depth of the low-depth resequencing of the authentication data includes 2X, 3X, 4X, 5X or 6X.
8. The method for identifying individuals of a species without a reference genome based on simplified genome assembly according to claim 5, characterized in that, In steps (3-6), the comparison method includes: Step (3-6-1): Extract the discontinuous long scaffold and location information of more than 100 SNP sites from the VCF file of the individual to be identified, build an index using a Python script and store it in a dictionary; Step (3-6-2): Extract the SNP loci at the corresponding positions from the population SNP database, and compare the population SNP database with the genotype data of the individual SNP to be identified; Step (3-6-3): During the comparison process, the SNP variant sites of the individual to be identified are scanned one by one to find the genotypic differences with the SNP sites at the same position in each individual in the population, the genotypic consistency of all SNP sites is statistically analyzed and the similarity score is calculated.
9. The method for identifying individuals of a species without a reference genome based on simplified genome assembly according to claim 1, characterized in that, In steps (3-6), When the depth of low-depth sequencing of the authentication data is 2X, at least 200 SNP sites are randomly selected from the SNP sites of the individual to be identified. When the depth of low-depth sequencing of the authentication data is 3X, at least 200 SNP sites are randomly selected from the SNP sites of the individual to be identified. When the depth of low-depth sequencing of the identity verification data is 4X, 5X or 6X, at least 100 SNP sites are randomly selected from the SNP sites of the individual to be identified.
10. The application of the method according to any one of claims 1-9 in the identification of individuals in species without a reference genome.