Individual Identification Methods and Systems Based on Whole-Genome Single Nucleotide Polymorphism Data

By screening for linkage-balanced SNP loci and calculating the total likelihood ratio, the dynamics and linkage disequilibrium issues of WGA-WGS data were resolved, ensuring the fairness and reliability of the evidence interpretation process for whole-genome sequencing data, and making it suitable for individual identification in forensic genetics.

CN121459920BActive Publication Date: 2026-05-26SICHUAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SICHUAN UNIV
Filing Date
2025-10-16
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing STR typing technology systems are not efficient enough when processing trace amounts of biodegradable biological samples. WGA-WGS data suffer from dynamism, error specificity, and linkage disequilibrium, making it difficult to effectively interpret evidence.

Method used

Based on whole-genome single nucleotide polymorphism data, by screening for linkage equilibrium SNP sites, calculating the total likelihood ratio, constructing an automated evidence interpretation process, integrating typing error rates, avoiding subjective selection bias, and ensuring the impartiality and reproducibility of evidence.

Benefits of technology

It has achieved high-throughput and compliant site preparation of WGA-WGS data, scientifically handled typing errors, ensured the fairness of evidence interpretation and the reproducibility of results, and solved key technical problems of WGA-WGS technology in legal science practice.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121459920B_ABST
    Figure CN121459920B_ABST
Patent Text Reader

Abstract

This invention discloses an individual identification method and system based on whole-genome single nucleotide polymorphism (SNP) data, belonging to the field of forensic genetics technology. The specific steps are as follows: obtaining the intersection SNP sites of the sample pairs to be identified, constructing a linkage-equilibrium SNP site set; evaluating the individual identification ability of the linkage-equilibrium SNP site set based on a power evaluation index, generating an optimal SNP site set; under the set mutual exclusion hypothesis, calculating the probability value of each optimal SNP site based on the population genotype frequency and the prior genotyping error rate, and calculating the total likelihood ratio based on the probability values ​​of each optimal SNP site; selecting a supporting hypothesis or hypothesis based on the relationship between the total likelihood ratio and a preset threshold. This invention, by introducing a semi-continuous likelihood ratio calculation model that integrates the genotyping error rate, can scientifically handle and explain the unavoidable genotyping errors in whole-genome sequencing data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of forensic genetics technology, and more specifically to an individual identification method and system based on whole-genome single nucleotide polymorphism data. Background Technology

[0002] Currently, the field of forensic genetics faces technical bottlenecks in handling trace amounts and degraded biological samples, including insufficient efficiency of existing STR typing technology and irreversible DNA template consumption. While emerging whole-genome amplification combined with whole-genome sequencing (WGA-WGS) technology can provide massive amounts of SNP data, it cannot directly adopt evidence interpretation methods developed based on fixed STR loci panels. This is because WGA-WGS data has inherent characteristics such as the dynamic nature of locus detection, the universality of typing errors, the independence of error patterns from sequencing depth, and the widespread linkage disequilibrium among massive loci. Therefore, how to construct an evidence interpretation method that can solve the problems of dynamic nature, error specificity, linkage disequilibrium, and computational fairness of WGA-WGS data is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0003] In view of this, the present invention provides an individual identification method and system based on whole-genome single nucleotide polymorphism data, which overcomes the above-mentioned defects.

[0004] To achieve the above objectives, the present invention provides the following technical solution:

[0005] An individual identification method based on whole-genome single nucleotide polymorphism data, comprising the following steps:

[0006] Obtain the intersection SNP sites of the sample pairs to be identified, and filter the intersection SNP sites based on physical location and heterozygosity to construct a set of linkage-balanced SNP sites;

[0007] Based on the effectiveness evaluation index, the individual identification ability of the linkage equilibrium SNP locus set is evaluated to generate the optimal SNP locus set.

[0008] In setting mutual exclusion assumptions , The probability values ​​of each optimal SNP locus are calculated based on the population genotype frequency and the prior genotyping error rate, and the total likelihood ratio is calculated based on the probability values ​​of each optimal SNP locus. The two samples in the sample pair to be identified originate from the same individual; The two samples in the pair of samples to be identified are from two unrelated individuals;

[0009] Selecting supporting hypotheses based on the relationship between the total likelihood ratio and a preset threshold. Or assume .

[0010] Optionally, the steps for obtaining the linkage-balanced SNP site set are as follows:

[0011] The whole genome of the reference sample is divided into multiple continuous genomic regions according to a preset physical distance;

[0012] Within a portion of the genome, the heterozygosity of each SNP locus is calculated based on allele frequency data from a reference population.

[0013] Based on the heterozygosity, representative SNP sites for each of the genomic regions are determined, and the SNP site set is constructed based on the representative SNP sites.

[0014] Optionally, the steps for assessing the individual identification capability of the linkage-balanced SNP locus set are as follows:

[0015] Calculate the individual identification probability value of each SNP site, multiply the individual identification probability value of each SNP site to obtain the cumulative individual identification probability value, and determine whether the identification efficiency is met based on the cumulative individual identification probability value and a preset efficiency threshold. If the efficiency is met, the optimal SNP site set is generated; if the efficiency is not met, the linkage equilibrium SNP site set is determined to have no identification ability.

[0016] Optionally, the calculation steps for the total likelihood ratio are as follows:

[0017] Obtain the observed genotype combinations of the sample pairs to be identified at each of the optimal SNP loci;

[0018] In the assumption Next, obtain all real genotype combinations under each observed genotype combination, calculate the probability of occurrence of each real genotype combination based on the genotype frequency of the population and the prior genotyping error rate, and obtain the first conditional probability based on the probability of occurrence.

[0019] In the assumption Next, obtain all real genotype combinations under each observed genotype combination, calculate the probability of occurrence of each real genotype combination based on the genotype frequency of the population and the prior genotyping error rate, and obtain the second conditional probability based on the probability of occurrence.

[0020] The total likelihood ratio is obtained based on the ratio of the first conditional probability to the second conditional probability.

[0021] Alternatively, assume All true genotype combinations below are those that are consistent with the observed genotype combinations and those that are inconsistent with the observed genotype combinations due to genotyping errors.

[0022] Alternatively, assume All true genotype combinations are those generated from samples from two unrelated individuals.

[0023] Optionally, the calculation steps for the first conditional probability are as follows:

[0024] The conditional probability of each optimal SNP site is obtained by summing the multiple occurrence probabilities at each optimal SNP site.

[0025] The conditional probabilities of each of the optimal SNP sites are multiplied together to obtain the first conditional probability.

[0026] An individual identification system based on whole-genome single nucleotide polymorphism data includes:

[0027] The SNP locus set construction module is used to obtain the intersection SNP loci of the sample pairs to be identified, and to filter the intersection SNP loci based on physical location and heterozygosity to construct a linkage-balanced SNP locus set.

[0028] The efficacy evaluation module is used to evaluate the individual identification ability of the linkage equilibrium SNP locus set based on efficacy evaluation indicators, and generate the optimal SNP locus set.

[0029] The likelihood ratio calculation module is used to calculate the likelihood ratio based on the set mutually exclusive assumptions. , The probability values ​​of each optimal SNP locus are calculated based on the population genotype frequency and the prior genotyping error rate, and the total likelihood ratio is calculated based on the probability values ​​of each optimal SNP locus. The two samples in the sample pair to be identified originate from the same individual; In the sample pair to be identified, the two samples originate from two unrelated individuals.

[0030] The result judgment module is used to select supporting hypotheses based on the relationship between the total likelihood ratio and a preset threshold. Or assume .

[0031] As can be seen from the above technical solution, the present invention provides an individual identification method and system based on whole-genome single nucleotide polymorphism data, which has the following advantages compared with the prior art:

[0032] 1. This invention constructs a complete evidence interpretation process for the characteristics of whole genome amplification and sequencing data, systematically solves the key technical problems from data generation to evidence interpretation, and provides an indispensable methodological foundation and standard for the reliable application of WGA-WGS technology in forensic practice;

[0033] 2. This invention proposes a dynamic and automated screening mechanism based on physical distance, which eliminates the dependence on preset fixed sites and can make full use of the advantage of whole genome coverage of WGA-WGS data to adapt to the unique detection sites of each pair of samples; and achieves high-throughput and compliant site preparation of massive WGA-WGS data, laying a solid foundation for subsequent analysis.

[0034] 3. This invention introduces a semi-continuous likelihood ratio calculation model that integrates genotyping error rate, which can scientifically handle and explain the unavoidable genotyping errors in whole genome sequencing data;

[0035] 4. The entire evidence interpretation process, from site selection to likelihood ratio calculation, is based entirely on pre-set objective conditions and methods, fundamentally avoiding subjective selection bias and the methodological risks of "presumption of guilt," and effectively ensuring the impartiality of scientific evidence and the repeatability of results. Attached Figure Description

[0036] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0037] Figure 1 This is a schematic diagram of the method flow of the present invention;

[0038] Figure 2 This is a schematic diagram of the SNP locus set construction process of the present invention. Detailed Implementation

[0039] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0040] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0041] One embodiment of the present invention discloses an individual identification method based on whole-genome single nucleotide polymorphism data, the steps of which are as follows: Figure 1 As shown, specifically:

[0042] Step 1: Obtain the intersection SNP sites of the sample pairs to be identified, and filter the intersection SNP sites based on physical location and heterozygosity to construct a set of linkage-balanced SNP sites.

[0043] Step 2: Evaluate the individual identification ability of the linkage equilibrium SNP locus set based on the efficacy evaluation index, and generate the optimal SNP locus set.

[0044] Step 3: Set the mutual exclusion assumption , Based on the population genotype frequency and prior genotyping error rate, the probability value of each optimal SNP locus is calculated, and the total likelihood is calculated based on the probability value of each optimal SNP locus. The two samples in the sample pair to be identified originate from the same individual; The two samples in the pair to be identified are from two unrelated individuals;

[0045] Step 4: Select supporting hypotheses based on the relationship between the total likelihood ratio and the preset threshold. Or assume .

[0046] In one embodiment, step 1, the step of obtaining the linkage-balanced SNP site set, is as follows:

[0047] Step 11: Divide the whole genome of the reference sample into multiple continuous genomic regions according to a preset physical distance;

[0048] Step 12: Within a portion of the genome, calculate the heterozygosity of each SNP locus based on allele frequency data from the reference population.

[0049] Step 13: Determine the representative SNP sites for each genomic region based on heterozygosity, and construct a set of SNP sites based on the representative SNP sites.

[0050] Furthermore, in step 1, to address the issue of non-fixed loci and potential linkage disequilibrium (LD) in whole-genome amplification-whole-genome sequencing (WGA-WGS) data, a dynamic screening strategy based on preset rules, independent of genotyping results, was designed, such as... Figure 2 As shown, specifically:

[0051] Step 11, Interval Division: Based on the reference sample (i.e., the human reference genome GRCh38), the entire genome is divided into 2,730 consecutive intervals with 1Mb as the basic unit.

[0052] Step 12, Interval Sampling: To maximize the physical distance between sites and reduce the risk of LD, only all odd-numbered intervals (a total of 1,365) are selected for subsequent site selection.

[0053] Step 13, Representative Locus Selection: Within each selected odd-numbered interval, calculate the heterozygosity of each SNP locus based on the allele frequency data of the reference population, and select the SNP locus with the highest heterozygosity in that interval as the representative locus of that region; construct an SNP locus set based on the representative locus of each region.

[0054] Its advantage lies in the fact that this strategy ensures that the final set of SNP sites used for LR calculation (up to 1,365) maintains sufficient physical spacing, effectively avoiding the LD effect. The entire screening process is based solely on two objectively preset parameters: physical location and heterozygosity, and is unrelated to the genotyping results of specific samples. This completely avoids the problem of presumption of guilt based on the plaintiff's or defendant's assumptions, ensuring the objectivity and impartiality of the evidence.

[0055] In one embodiment, the steps for assessing the individual identification capability of a linkage-balanced SNP locus set are as follows:

[0056] Calculate the individual identification probability value of each SNP locus, multiply the individual identification probability value of each SNP locus to obtain the cumulative individual identification probability value, and determine whether the identification efficiency is met based on the cumulative individual identification probability value and the preset efficiency threshold. If it is met, the optimal SNP locus set is generated; if it is not met, the SNP locus set in linkage equilibrium is determined to have no identification ability.

[0057] Furthermore, in step 2, after dynamically generating a set of SNP loci for each pair of samples (i.e., the sample pairs to be identified), it is necessary to quantitatively evaluate the individual identification ability of the loci set.

[0058] The cumulative individual identification probability (CDP) is used as the core evaluation index. A CDP closer to 1 indicates a stronger ability of the locus system to distinguish between different individuals. The calculation formula is as follows:

[0059] ;

[0060] In the formula, The number of phenotypes for a given genetic marker; For the first in this group The frequency of each phenotype; This is the probability that two unrelated individuals randomly selected from a population will have the same phenotype at a certain locus purely by chance.

[0061] To more intuitively display high CDP values, the CDP value is transformed using the logarithm.10 (1-CDP) is used to express this.

[0062] In one embodiment, the recognition capability is evaluated by: setting -log 10 The criterion is (1-CDP) ≥ 10. This threshold means that CDP ≥ 0.9999999999. Its theoretical recognition capability has exceeded the reciprocal requirement of the total global population (approximately 8 billion), indicating that this dynamic locus set has an extremely high level of individual recognition efficiency, which is sufficient for individual identification on a global scale.

[0063] In one embodiment, the steps for calculating the total likelihood ratio are as follows:

[0064] Step 31: Obtain the observed genotype combinations of the sample pairs to be identified at each optimal SNP locus;

[0065] Step 32, under the assumption Under each observed genotype combination, all real genotype combinations are obtained. The probability of occurrence of each real genotype combination is calculated based on the population genotype frequency and the prior genotyping error rate. The first conditional probability is obtained based on the probability of occurrence.

[0066] Step 33, under the assumption Under each observed genotype combination, all real genotype combinations are obtained. The probability of occurrence of each real genotype combination is calculated based on the population genotype frequency and the prior genotyping error rate. The second conditional probability is obtained based on the probability of occurrence.

[0067] Step 34: Obtain the total likelihood ratio based on the ratio of the first conditional probability to the second conditional probability.

[0068] Furthermore, in step 3, the genotyping error rate present in whole-genome sequencing ( This is directly integrated into the probability model for LR calculation. The calculation steps are as follows: For each pair of samples, calculate the probability at each dynamically selected locus. and Then, based on the assumption of site independence, the probability values ​​of all sites are multiplied together to obtain the total conditional probability, and finally the LR value is obtained.

[0069] Furthermore, the mutual exclusion assumption set in step 3... , Specifically:

[0070] Plaintiff's assumption ( ): The two samples came from the same individual.

[0071] The defendant assumes ( ): The two samples came from two unrelated individuals.

[0072] In one embodiment, the probability model includes: The calculation formula is: , This indicates that the subtyping was obtained under the plaintiff's assumptions. The probability, This represents the probability of obtaining a fractal under the defendant hypothesis. In the calculation... and The calculation comprehensively considers all possible true genotype combinations and all genotyping error patterns that lead to the observed current genotyping results. The calculation formulas used in each case are shown in Table 1. Here, 0 / 0 represents a homozygote consistent with the genomic reference allele, 1 / 1 represents a homozygote inconsistent with the reference allele, and 0 / 1 represents a heterozygote. , , These represent the genotype frequencies for 0 / 0, 0 / 1, and 1 / 1 subtypes, respectively. This represents the classification error rate, which is obtained from empirical data.

[0073] Table 1 Plaintiff's Hypothesis Considering Classification Error Rate and the defendant's assumption Calculation formula

[0074]

[0075] Taking an example where both sides are observed to be 0 / 0: In Under these conditions, the true genotype combinations could be 0 / 0 (all correct typing), 0 / 1 (one error in each), or 1 / 1 (all errors). In this case, it is necessary to consider all nine genotype combinations of two unrelated individuals (simplified to six scenarios) and the corresponding genotyping error paths (such as errors occurring in one or both individuals). The probability of each scenario is determined by the population genotype frequency (…). ) and classification error rate ( This was calculated together.

[0076] Its advantage lies in embedding the fractal error rate (ω) as a core parameter into the semi-continuous... In the computing framework, The computation can tolerate and quantify phenotypic inconsistencies, and can still calculate correct support even with a large number of erroneous sites. highly persuasive value.

[0077] In one embodiment, to empirically evaluate the ability of this embodiment to distinguish between real contributors and non-contributors, this control group test specifically includes:

[0078] 1. Control group construction: Individuals unrelated to the test samples were selected as the irrelevant control group.

[0079] 2. Pairing and calculation: Pair non-contributors with WGA-WGS samples.

[0080] 3. Process Application: For each pair of samples, perform steps 1 to 3 above in their entirety: obtain the intersection SNP, dynamically screen linkage equilibrium sites, evaluate systemic efficacy (CDP), and calculate... value.

[0081] 4. Results Analysis: The large number of "non-contributor LR" and "real contributor LR" obtained were compared and analyzed to verify the effectiveness and reliability of the method in this embodiment in excluding irrelevant individuals, thereby providing a solid experimental basis for the interpretation of the strength of evidence.

[0082] In one embodiment, this embodiment aims to compare its own samples (which meets the plaintiff's assumption). Comparison with non-contributor samples (conforms to the defendant's hypothesis) The results comprehensively verify the reliability, accuracy, and practicality of the above process, specifically as follows:

[0083] 1. Experimental Objective

[0084] Verify whether the evidence interpretation process based on WGS data constructed in this embodiment can: (1) provide an extremely high number of results when the samples originate from the same individual. It deserves strong support (2) When the sample comes from irrelevant individuals, give a very low number of samples. The values ​​accurately exclude non-contributors, and the results in the two cases are clearly separated with no erroneous overlap.

[0085] 2. Experimental Materials and Setup

[0086] Test sample:

[0087] Contributor sample: WGA-WGS data of 12 NA12878 samples.

[0088] Non-contributor sample: Genome typing data of 98 unrelated individuals from the 1,000 Genomes Project (1KGP) database (GRCh38) of Utah with Nordic and Western European ancestry (CEU population, homologous to NA12878).

[0089] Comparison Design:

[0090] Self-comparison ( (Established): 12 NA12878 samples were paired with themselves, forming a total of 66 sample pairs.

[0091] Non-contributor comparison ( Establishment): The 12 test samples were paired with 98 non-contributors, resulting in a total of 1176 sample pairs.

[0092] Analysis Procedure: For each pair of samples, the above procedure was strictly followed: (a) Obtain the intersection SNP; (b) Dynamically screen linkage equilibrium sites; (c) Calculate the system performance (-log) 10 (1-CDP)); (d) Calculate the likelihood ratio ( ).

[0093] 3. Experimental Results

[0094] 3.1 Self-comparison ( (Established) Result

[0095] The analysis results of all 66 self-aligned sample pairs strongly support the conclusion that "the two samples originated from the same person".

[0096] Available loci and system performance: As shown in Table 2, all sample pairs have a sufficient number of independent loci (minimum 1,297, maximum 1,365), and the system performance of all sample pairs is (-log) 10 (1-CDP) all exceeded the threshold (10), indicating that the system has extremely strong distinguishing ability.

[0097] Strength of evidence ( This embodiment calculates the error rate using a semi-continuous model that integrates the classification error rate, ultimately yielding an extremely high number of results for all sample pairs. value (log) 10 The minimum LR value is 217, and without exception, it provides a value for... There is extremely strong supporting evidence.

[0098] Table 2 Self-comparison ( Key Indicator Statistics Table

[0099]

[0100] 3.2 Non-contributor comparison ( (Established) Result

[0101] All 1,960 non-contributor pairs met the system efficiency requirements, and the analysis results correctly supported the conclusion that "the two samples came from different people," successfully excluding irrelevant individuals.

[0102] 4. Experimental Conclusions

[0103] The overall results of this embodiment indicate that:

[0104] Validity: Under the above procedure, WGA-WGS data can reliably obtain a large number of loci, and the system efficiency far exceeds the discrimination threshold. For samples from the same source, even with sequencing errors, an extremely high number can be calculated. This value strongly supports the determination of identity.

[0105] Specificity: For non-contributors, the process of this invention can provide 100% accurate support for exclusion. value (log) 10 The results showed that LR < -4 and there were no false positives, demonstrating that the method has extremely high reliability.

[0106] This embodiment also discloses an individual identification system based on whole-genome single nucleotide polymorphism data, including:

[0107] The SNP site set construction module is used to obtain the intersection SNP sites of the sample pairs to be identified, and to filter the intersection SNP sites based on physical location and heterozygosity to construct a linkage-balanced SNP site set.

[0108] The efficacy evaluation module is used to evaluate the individual identification ability of the linkage equilibrium SNP locus set based on efficacy evaluation indicators, and generate the optimal SNP locus set.

[0109] The likelihood ratio calculation module is used to calculate the likelihood ratio based on the set mutually exclusive assumptions. , The probability values ​​of each optimal SNP locus are calculated based on the population genotype frequency and the prior genotyping error rate, and the total likelihood ratio is calculated based on the probability values ​​of each optimal SNP locus. The two samples in the sample pair to be identified originate from the same individual; In the sample pair to be identified, the two samples come from two unrelated individuals.

[0110] The result evaluation module is used to select supporting hypotheses based on the relationship between the total likelihood ratio and a preset threshold. Or assume .

[0111] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0112] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for individual identification based on whole genome single nucleotide polymorphism data, characterized by, The specific steps are as follows: The process involves obtaining the intersection SNP sites of the sample pairs to be identified, filtering the intersection SNP sites based on physical location and heterozygosity, and constructing a linkage-balanced SNP site set. The steps for obtaining the linkage-balanced SNP site set are as follows: dividing the whole genome of the reference sample into multiple continuous genomic regions according to a preset physical distance; calculating the heterozygosity of each SNP site within a portion of the genomic regions based on allele frequency data from the reference population; determining representative SNP sites for each genomic region based on the heterozygosity, and constructing the SNP site set based on the representative SNP sites. The individual identification ability of the linkage-balanced SNP locus set is evaluated based on the effectiveness evaluation index to generate the optimal SNP locus set. The individual identification ability evaluation steps of the linkage-balanced SNP locus set are as follows: calculate the individual identification probability value of each SNP locus, multiply the individual identification probability value of each SNP locus to obtain the cumulative individual identification probability value, and determine whether the identification effectiveness is met based on the cumulative individual identification probability value and a preset effectiveness threshold. If the recognition effectiveness is met, the optimal SNP locus set is generated; if the recognition effectiveness is not met, the linkage-balanced SNP locus set is determined to have no identification ability. Under the setting of mutual exclusive assumptions , , calculating the probability value of each optimal SNP site based on the genotype frequency of the population and the prior typing error rate, and calculating the total likelihood ratio based on the probability value of each optimal SNP site; The two samples in the pair of samples to be identified come from the same individual; The two samples in the pair of samples to be identified come from two unrelated individuals; selecting the support hypothesis based on a relationship between the total likelihood ratio and a preset threshold or hypothesis .

2. The individual identification method based on whole-genome single nucleotide polymorphism data according to claim 1, characterized in that, The steps for calculating the total likelihood ratio are as follows: Obtain the observed genotype combinations of the sample pairs to be identified at each of the optimal SNP loci; In the assumption Next, obtain all real genotype combinations under each observed genotype combination, calculate the probability of occurrence of each real genotype combination based on the genotype frequency of the population and the prior genotyping error rate, and obtain the first conditional probability based on the probability of occurrence. In the assumption Next, obtain all real genotype combinations under each observed genotype combination, calculate the probability of occurrence of each real genotype combination based on the genotype frequency of the population and the prior genotyping error rate, and obtain the second conditional probability based on the probability of occurrence. The total likelihood ratio is obtained based on the ratio of the first conditional probability to the second conditional probability.

3. The individual identification method based on whole-genome single nucleotide polymorphism data according to claim 2, characterized in that, Assumption All true genotype combinations below are those that are consistent with the observed genotype combinations and those that are inconsistent with the observed genotype combinations due to genotyping errors.

4. The individual identification method based on whole-genome single nucleotide polymorphism data according to claim 2, characterized in that, Assumption All true genotype combinations are those generated from samples from two unrelated individuals.

5. The individual identification method based on whole-genome single nucleotide polymorphism data according to claim 2, characterized in that, The steps for calculating the first conditional probability are as follows: The conditional probability of each optimal SNP site is obtained by summing the multiple occurrence probabilities at each optimal SNP site. The conditional probabilities of each of the optimal SNP sites are multiplied together to obtain the first conditional probability.

6. An individual identification system based on whole-genome single nucleotide polymorphism (WNP) data, employing the individual identification method based on whole-genome single nucleotide polymorphism (WNP) data as described in any one of claims 1-5, characterized in that, include: The SNP locus set construction module is used to obtain the intersection SNP loci of the sample pairs to be identified, and to filter the intersection SNP loci based on physical location and heterozygosity to construct a linkage-balanced SNP locus set. The efficacy evaluation module is used to evaluate the individual identification ability of the linkage equilibrium SNP locus set based on efficacy evaluation indicators, and generate the optimal SNP locus set. The likelihood ratio calculation module is used to calculate the likelihood ratio based on the set mutually exclusive assumptions. , The probability values ​​of each optimal SNP locus are calculated based on the population genotype frequency and the prior genotyping error rate, and the total likelihood ratio is calculated based on the probability values ​​of each optimal SNP locus. The two samples in the sample pair to be identified originate from the same individual; The two samples in the sample pair to be identified are from two unrelated individuals; The result judgment module is used to select supporting hypotheses based on the relationship between the total likelihood ratio and a preset threshold. Or assume .