Method and system for interpreting hybrid DNA biallelic SNP data based on WGA-WGS
By employing a WGA-WGS-based method for interpreting mixed DNA biseleural SNP data, and utilizing dynamic linkage equilibrium screening and an integrated genotyping error model, the complexity of mixed DNA evidence was addressed, enabling accurate calculation of evidence strength and reliable analysis results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SICHUAN UNIV
- Filing Date
- 2025-10-16
- Publication Date
- 2026-06-02
AI Technical Summary
Existing technologies have several drawbacks when processing mixed DNA evidence, including poor adaptability to trace amounts of degraded samples, low utilization of SNP sites across the entire genome, high typing error rate, and the assumption that the model assumes the reference sample is correctly typed, leading to deviations in LR calculation results.
A WGA-WGS-based method for interpreting mixed DNA biseleural SNP data was adopted. SNP sites were screened using a dynamic linkage equilibrium screening strategy, and the total likelihood ratio was calculated to interpret the mixed samples by combining an observation probability model that integrates genotyping errors.
It achieves precise processing of WGA-WGS data, accurately calculates the likelihood value of evidence strength, is applicable to complex and difficult biological samples, reduces human intervention bias, and ensures the scientific rigor and reproducibility of analysis results.
Smart Images

Figure CN121459919B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of forensic genetics technology, and more specifically to a method and system for interpreting mixed DNA diallelic SNP data based on WGA-WGS. Background Technology
[0002] Currently, probabilistic genotyping (PG) based on short tandem repeat (STR) capillary electrophoresis (CE) systems or targeted amplification next-generation sequencing (NGS) panels is the closest existing technology for interpreting mixed DNA evidence. This technology typically employs a fixed combination of genetic markers and relies on continuous or semi-continuous PG models. Its core assumption is that the genotyping of the reference sample is absolutely accurate, and only allele drop-out and drop-in are probabilistically assessed in the evidence sample. However, this technology has significant limitations when processing whole-genome sequencing (WGS) data: First, STR-CE and targeted NGS panels rely on intact DNA templates and specific primer binding, making them poorly adaptable to trace amounts of degraded samples and prone to amplification failure due to primer binding site loss. Second, fixed-site panels cannot dynamically adapt to the massive number of SNP sites across the entire genome, resulting in low utilization of WGS data information. Third, existing PG models optimized for PCR bias in STR processing cannot effectively handle the high systematic genotyping error rate in WGA-WGS, which is independent of sequencing depth. Fourth, the model assumes that the reference sample is correctly genotyped, but WGA-WGS can produce non-negligible genotyping errors even with high-quality reference samples. Directly using incorrect genotyping as the "gold standard" input will lead to serious deviations in LR calculation results. Therefore, how to construct a hybrid DNA interpretation method that can adapt to the characteristics of whole-genome biselequential SNP data while comprehensively considering the genotyping error probability of evidence and reference samples is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0003] In view of this, the present invention provides a method and system for interpreting mixed DNA diallelic SNP data based on WGA-WGS, overcoming the above-mentioned defects.
[0004] To achieve the above objectives, the present invention provides the following technical solution:
[0005] A method for interpreting mixed DNA diallelic SNP data based on WGA-WGS, the specific steps of which are as follows:
[0006] Intersecting SNP sites were obtained from WGA-WGS genotyping data of mixed samples and target individual samples, and a dynamic linkage equilibrium screening strategy was used to screen the intersecting SNP sites to generate a set of SNP sites.
[0007] Determine the number of mixed contributors N and the error rate of individual allele typing;
[0008] Based on the application scenario, a preset observation probability model for integrated genotyping errors is invoked. Based on the SNP locus set, the total likelihood ratio is calculated based on the number of mixed contributors N and the genotyping error rate of a single allele.
[0009] The mixed samples are interpreted based on the total likelihood ratio.
[0010] Optionally, the dynamic chain balance screening strategy is:
[0011] The whole genome is divided into multiple regions according to a preset physical interval. Within the multiple regions, the SNP sites that intersect within the regions are screened based on heterozygosity to form the SNP site set.
[0012] Optionally, the error rate of individual allele typing is obtained based on previous Bayesian estimation.
[0013] Optionally, the observation probability model for integrating classification errors includes an observation probability model with known contributors and an observation probability model without known contributors.
[0014] Alternatively, the expression for the observation probability model without known contributors is:
[0015] ;
[0016] in, ;
[0017] ;
[0018] In the formula, The observed genotype of the target individual; The observed genotypes for the mixed sample; , The assumption is mutual exclusion; In order to be in the true genotype Observed at time The probability of; For all the true genotypes of the target individual; The prior probability of the true genotype of the target individual; Given N unknown contributors, this represents the sum of the number of 0 alleles in all individuals. For a total of 0 alleles Observed at time The probability of; The number of 0 alleles in the true genotype of the target individual; The total number of 0 alleles from all contributors in the mixed sample.
[0019] Optionally, the expression for the observed probability model of known contributors is:
[0020] ;
[0021] in,
[0022] ;
[0023] ;
[0024] ;
[0025] In the formula, The observed genotype of the target individual; Observed genotypes of known contributors; The observed genotypes for the mixed sample; , The assumption is mutual exclusion; , These are the actual genotypes of the known contributors and the target individuals, respectively. for The total number of 0 alleles from unknown contributors; , These represent the number of 0 alleles in the true genotypes of known contributors and target individuals, respectively. for The total number of 0 alleles from unknown contributors.
[0026] Optionally, evidence interpretation is performed based on the total likelihood ratio and a preset judgment threshold, and a judgment result is output, the judgment result including supporting hypotheses. Supporting hypothesis And the evidence lacks discriminatory power.
[0027] An interpretation system for mixed DNA diallelic SNO data based on WGA-WGS includes:
[0028] The locus screening module is used to obtain the intersection SNP loci based on WGA-WGS genotyping data of mixed samples and target individual samples, and to screen the intersection SNP loci using a dynamic linkage balance screening strategy to generate a set of SNP loci.
[0029] The parameter setting module is used to determine the number of mixed contributors N and the error rate of individual allele typing;
[0030] The probability calculation module is used to call a preset observation probability model for integrated genotyping errors based on the application scenario. Based on the SNP locus set, it calculates the total likelihood ratio based on the number of mixed contributors N and the genotyping error rate of a single allele.
[0031] The sample interpretation module is used to interpret the mixed samples based on the total likelihood ratio.
[0032] As can be seen from the above technical solution, the present invention provides a method and system for interpreting mixed DNA diallelic SNP data based on WGA-WGS, which has the following advantages compared with the prior art:
[0033] 1. A complete and specific evidence interpretation process for WGA-WGS mixed DNA diallelic SNP data was constructed, providing a new and reliable technical path for processing such complex data and promoting the advancement of forensic genetic analysis technology;
[0034] 2. By integrating a computational model for typing errors, this method can accurately address the common SNP typing inconsistency problem in WGA-WGS data, enabling the calculation of a likelihood (LR) value that accurately reflects the strength of evidence, whether for real contributors or non-contributors, thereby eliminating misjudgments.
[0035] 3. The dynamic site screening strategy adopted enables it to adaptively adjust according to the quality and sequencing depth of different samples, making it particularly suitable for complex and difficult biological samples such as trace amounts, degradation, and extremely low proportion mixing that are difficult to handle by traditional analysis methods, greatly expanding the range of sources of effective DNA evidence;
[0036] 4. The entire process operates based on a pre-defined mathematical model and objective parameters (such as ω), minimizing potential biases introduced by human intervention and subjective judgment. This not only ensures the scientific rigor of the evidence interpretation but also guarantees the reproducibility and verifiability of the analysis results. Attached Figure Description
[0037] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0038] Figure 1 This is a schematic diagram of the method flow of the present invention. Detailed Implementation
[0039] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0040] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0041] One embodiment of the present invention discloses a method for interpreting mixed DNA diallelic SNP data based on WGA-WGS, such as... Figure 1 As shown, the specific steps are as follows:
[0042] Step 1: Obtain the intersection SNP sites based on the WGA-WGS genotyping data of the mixed sample and the target individual sample, and use the dynamic linkage equilibrium screening strategy to screen the intersection SNP sites to generate a set of SNP sites.
[0043] Step 2: Determine the number of mixed contributors N and the error rate of individual allele typing;
[0044] Step 3: Based on the application scenario, call the preset observation probability model of integrated typing error, and calculate the total likelihood ratio based on the number of mixed contributors N and the typing error rate of a single allele on the basis of the SNP locus set;
[0045] Step 4: Interpret the mixed samples based on the total likelihood ratio.
[0046] In one embodiment, the dynamic chain equilibrium screening strategy is as follows:
[0047] The whole genome is divided into multiple regions according to a preset physical interval. Within these regions, SNP sites that intersect within the regions are screened based on heterozygosity to form a set of SNP sites.
[0048] Furthermore, for each mixed sample-target individual sample pair, the SNP sites that are commonly detected across the whole genome are first obtained, and a dynamic linkage equilibrium screening strategy is applied (e.g., screening the SNPs with the highest heterozygosity in odd intervals at 1Mb intervals) to form the final set of SNP sites used for calculation.
[0049] In one embodiment, the error rate of individual allele typing is obtained based on prior Bayesian estimation.
[0050] In one embodiment, all parameters and symbols required during the calculation process are defined:
[0051] : The observed genotype of the target individual (such as a suspect) (values: 00, 01, 11).
[0052] : Observed genotypes of the mixed sample (values: 00, 01, 11).
[0053] : (If available) Observed genotypes of known contributors (such as victims).
[0054] The genotyping error rate of a single allele, obtained through prior Bayesian estimation, is a key input to the model.
[0055] : Total number of contributors in the mixed sample.
[0056] , : The frequency of alleles 0 and 1 in the population.
[0057] , , The frequency of genotypes 00, 01, and 11 in the population.
[0058] The steps for constructing an observation probability model that integrates classification errors are as follows:
[0059] Construct a 3x3 probability matrix, defined given the true genotype ( Under the condition of ), a specific genotype was observed ( The probability of each allele occurring independently is given in Table 1. This model is constructed based on the assumption that each allele independently makes an error.
[0060] Table 1 shows the given true genotypes ( Under the condition of ), a specific genotype was observed ( The probability of )
[0061]
[0062] The core formula for the calculation is: Depending on whether there are known contributors, there are two scenarios:
[0063] Scenario A: No known contributors The calculation is as follows:
[0064] Define two mutually exclusive assumptions:
[0065] (Plaintiff's assumption): The target individuals are a mixed sample. One of the contributors.
[0066] (Defendant's assumption): The target individual was not a contributor to the pooled sample; the pool was composed of... It consists of other unrelated individuals.
[0067] Calculate the numerator :
[0068] ;
[0069] In the formula, This means iterating through all possible true genotypes (00, 01, 11) of the target individual, calculating the joint probability for each case, and summing the results. This indicates that the true genotype is Observed at time The probability is calculated using the sequencing error model; The prior probability representing the true genotype of the target individual is derived from the population genotype frequency. , , ); Given N individuals with no known contributors, this represents the sum of the number of 0 alleles. For a total of 0 alleles Observed at time The probability is calculated using the sequencing error model; This indicates the number of 0 alleles in the true genotype of the target individual (00, 01, and 11 correspond to 2, 1, and 0, respectively). The total number of 0 alleles from all contributors in the mixed sample.
[0070] The total number of 0 alleles is At that time, observation The probability of is calculated as follows:
[0071] ;
[0072] In the formula, For the number of contributors to the mixture, a biselequential SNP locus has a total of One allele; This represents the total number of alleles that were actually "0" before the error occurred, with values ranging from 0 to ... ; A single allele was misclassified as an opposing allele. The probability is that each allele makes an independent error. This indicates that only the "0" allele was observed. This indicates that only the "1" allele was observed. This indicates that both "0" and "1" are observed simultaneously. The mixture is considered to be composed of... Each independent allele composition; if an allele misinterpretation occurs, if this If all the observed isopleths are "0", it is recorded as 00; if all are "1", it is recorded as 11; otherwise, it is recorded as 01.
[0073] when When the actual value is "0" Given 10 alleles, to maintain the correct sequence (0→0), the probability of each allele is (1- ). ), merged into (1- ) T ; The real value is "1" There are 1 alleles, and an error occurs during the flipping (1→0). The probability of each allele is... , merged into (2N-T) The two cases are independent of each other, therefore, = .
[0074] when When the actual value is "0" Given a set of alleles, the probability of an incorrect flip (0→1) is... ; The real value is "1" Given 1 to 1 alleles, the probability of maintaining the correct sequence (1→1) is _____. ;therefore, = .
[0075] when When both "0" and "1" are present, using the complement event calculation, the result is "neither 00 nor 11", which is 01. Therefore, = .
[0076] The joint probability calculation process can be understood as follows: First, iterate through the true genotypes of the target individual and calculate their probability in the observed values. The probability is calculated and weighted by combining it with the population frequency; then, the possible allelic composition of the remaining unknown contributors is considered, through... and Calculate the observed genotypes of the mixed sample The probability is calculated by summing the results for each true genotype; finally, the joint probability is obtained by summing the results for each true genotype. .
[0077] Calculate the denominator :
[0078] exist Below, the target individual is independent of the mixture, therefore the observed genotype of the target individual is... Observed genotypes of mixed samples Given the group parameters, the groups are independent, and their joint probability decomposes as follows:
[0079] ;
[0080] In the formula, Marginal probability of the target individual; This represents the marginal probability of the mixture.
[0081] Among them, the marginal probability of the target individual for:
[0082] The true genotype of the target individual Perform a full probability expansion:
[0083] ;
[0084] In the formula, From the sequencing / genotyping error model (see observation probability matrix); The prior probability of a population genotype (which can be derived from allele frequencies) , According to HWE: , , (or directly use empirical genotype frequencies).
[0085] Marginal probability of mixture for:
[0086] exist Below, the mixture is made from It consists of 1 unknown contributor. Let the count of contributions made by the i-th contributor to the "0 allele" be denoted as . ,but:
[0087] , , ;
[0088] make: Let be the total number of "0 alleles" in the mixture. Then:
[0089] ;
[0090] in, The total number of "0 alleles" is given, and:
[0091] ;
[0092] The solution is the same part.
[0093] The calculation process of the joint probability under the following conditions can be understood as: Below, the mixture and the target individual are independent. For the mixture, first determine the distribution of "how many zero allotments are there in total". Then, under the condition that "the total number is k", the observed values are calculated using the single equipotential error model. The probability is calculated, and finally a weighted sum is taken over all possible k.
[0094] Seeking Value: Divide the numerator by the denominator to obtain the value of that site. value.
[0095] Scenario B: There are known contributors (such as victims). The calculation is as follows:
[0096] Definition and assumptions:
[0097] The target individual and known contributors are both contributors to the mixed sample.
[0098] It is known that the contributor is a contributor, but the target individual is not; the mixture consists of the target individual and others. It consists of unrelated individuals.
[0099] ;
[0100] Calculate the numerator :
[0101] Both the target individuals and the victims are contributors to the mixed sample:
[0102] ;
[0103] In the formula, , These are the actual genotypes of the known contributors and the target individuals, respectively. , These represent the number of 0 alleles in the true genotypes of known contributors and target individuals, respectively. for The total number of 0 alleles from unknown contributors (which can be calculated using dynamic programming); The total number of alleles is 0. The probability of observing Om at that time:
[0104] ;
[0105] Sum the probabilities of all scenarios using a weighted average.
[0106] Calculate the denominator :
[0107] The target individual is unrelated to the pooled sample; the victim remains a contributor.
[0108] ;
[0109] ;
[0110] for The total number of 0 alleles from unknown contributors (including the original target individual location).
[0111] Seeking Value: Divide the numerator by the denominator to obtain the value of that site. value.
[0112] Furthermore, the overall Calculation: Assuming each SNP site is independent, multiply the probability values calculated in step 3 above at each site to obtain the overall likelihood ratio at the genome level. ).
[0113] Interpretation of Results: Objective interpretation of the strength of evidence based on the final LR value:
[0114] like If the number is ≥10,000, then the evidence supports the plaintiff's hypothesis. This means that the target individual is a contributor to the mixture.
[0115] If 0.0001 < If the number is less than 10,000, the evidence has no discriminatory power.
[0116] like If the value is ≤0.0001, then the evidence supports the defendant's hypothesis ( This means that the target individual is not a contributor to the mixture.
[0117] In one embodiment, the validation of the WGA-WGS-based hybrid DNA evidence interpretation method in a scenario without known contributors is as follows:
[0118] 1. Experimental Objective
[0119] The effectiveness of the mixed DNA evidence interpretation process proposed in this embodiment in the absence of known contributors was verified, with a focus on evaluating its ability to distinguish between primary and secondary contributors under complex mixed conditions, as well as its specificity for excluding non-contributors.
[0120] 2. Experimental Materials and Setup
[0121] Hybrid Sample Design: Using samples with a total template size of 10ng, a hybrid system containing samples of varying complexity is constructed.
[0122] Two-person mixing: Set four mixing ratios: 1:1, 1:5, 1:10, and 1:20;
[0123] Three-person mix: Set four mixing ratios: 1:1:1, 1:1:10, 1:10:10, and 1:10:20;
[0124] A total of 8 mixed conditions were constructed, and each condition was repeated multiple times.
[0125] Reference sample:
[0126] WGA-WGS typing data from real contributors 533, 534 and NA12878 were used;
[0127] 533 and 534 each have 3 sequencing replicates, and NA12878 has 6 sequencing replicates;
[0128] Non-contributor control group: Nine real samples unrelated to the mixed sample were selected as the non-contributor control group.
[0129] 3. Implementation process
[0130] For each pooled sample-target individual sample pair, perform the following standardized analysis procedure:
[0131] Step 1: Data Preprocessing and Dynamic Site Selection
[0132] Input mixed samples ( ) and target individual samples ( WGA-WGS classification data (VCF format);
[0133] Obtain all SNP sites detected jointly by both methods;
[0134] A dynamic linkage balance screening strategy was applied: odd-numbered intervals were screened at 1 Mb physical intervals, and the SNP with the highest heterozygosity in each interval was selected as the representative site to ensure linkage balance between sites.
[0135] Step 2: Parameter Settings and Model Configuration
[0136] Set the number of mixed contributors (2 or 3);
[0137] The input is estimated using Bayesian methods to determine the pre-determined error rate of individual allele typing. ( =0.017573);
[0138] Step 3: Likelihood ratio calculation:
[0139] Calling for no known contributors Calculation module:
[0140] Based on the formula perform the calculation;
[0141] wherein : the target individual is one of the true contributors to the mixture;
[0142] : the target individual is not a contributor to the mixture;
[0143] During the calculation process:
[0144] Traverse all possible true genotypes (00, 01, 11) of the target individual;
[0145] Use the dynamic programming algorithm to calculate the allele combination probability of the remaining unknown contributors;
[0146] Integrate the genotyping error rate ω and calculate the joint probability under each hypothesis respectively.
[0147] Step 4: Result recording and statistical analysis, specifically:
[0148] Record the lgLR value of each sample pair;
[0149] According to the international general standard:
[0150] lgLR ≥ 4: Support the contributor hypothesis;
[0151] -4 < lgLR < 4: The evidence has no discriminative power;
[0152] lgLR ≤ -4: Support the exclusion hypothesis.
[0153] Statistically analyze the support rate of true contributors and the false positive rate of non - contributors.
[0154] 4. Experimental Results
[0155] 4.1 Analysis Results of True Contributors
[0156] Table 2 Statistical Results of lgLR of True Contributors under Different Mixing Conditions
[0157]
[0158] It can be seen from Table 2 that: Robustness of major contributor identification: The major contributors (534 and NA12878) received extremely strong support (lgLR ≥ 4) under the vast majority of mixing conditions, and the support rate reached 100%;
[0159] Excellent proportion sensitivity: The support rate of the minor contributor (533) decreased significantly as the mixing ratio decreased, and the support rate dropped to 0% at a ratio of 1:10, accurately reflecting its minor contribution in the mixture;
[0160] Applicability to complex hybrid systems: The method maintains good performance in three-person hybrid systems and can accurately identify the main contributors.
[0161] 4.2 Non-contributor exclusion specificity
[0162] Table 3. Statistics on the effect of excluding non-contributors
[0163]
[0164] As can be seen from Table 3:
[0165] Highly efficient exclusion capability: Under most mixed conditions, the exclusion rate of non-contributors exceeds 90%;
[0166] The false positive rate is controllable: the highest false positive rate was 14.8% (three-person mixed 1:1:1 condition), and all other conditions were below 12%;
[0167] The method exhibits good specificity, indicating that the present invention can effectively distinguish between contributors and non-contributors.
[0168] 5. Experimental Conclusions
[0169] This embodiment verifies the performance of the method with a template amount of 10 ng. The results show that:
[0170] 1. Excellent ability to identify major contributors: This method can effectively identify major contributors in mixtures and provides strong evidence support (lgLR ≥ 4) under various mixing ratios, with a support rate of 100%.
[0171] 2. Precise proportional sensitivity: The method is highly sensitive to the proportion of contributors and can accurately reflect the objective fact of changes in the proportion of minor contributors in the mixture.
[0172] 3. Good specificity: Under the condition of 10ng template amount, the method has good ability to exclude non-contributors, and the false positive rate is generally less than 15%, which proves that the method has high specificity.
[0173] 4. Applicability to complex mixed systems: The method is stable in both two-person and three-person mixed systems, demonstrating its good applicability under complex mixing conditions.
[0174] The evidence interpretation process provided in this embodiment can provide objective, accurate, and quantitative statistical evidence for mixed DNA analysis in scenarios where there are no known contributors. It performs particularly well in identifying major contributors and in terms of proportional sensitivity, providing a reliable technical means for the analysis of complex mixed samples in forensic practice.
[0175] This embodiment also discloses an interpretation system for mixed DNA diallelic SNP data based on WGA-WGS, including:
[0176] The locus screening module is used to obtain the intersection SNP loci based on WGA-WGS genotyping data of mixed samples and target individual samples, and to screen the intersection SNP loci using a dynamic linkage equilibrium screening strategy to generate a set of SNP loci.
[0177] The parameter setting module is used to determine the number of mixed contributors N and the error rate of individual allele typing;
[0178] The probability calculation module is used to call the preset observation probability model of integrated genotyping error based on the application scenario. Based on the SNP locus set, it calculates the total likelihood ratio based on the number of mixed contributors N and the genotyping error rate of a single allele.
[0179] The sample interpretation module is used to interpret mixed samples based on the total likelihood ratio.
[0180] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0181] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for interpreting mixed DNA diallelic SNP data based on WGA-WGS, characterized in that, The specific steps are as follows: Intersecting SNP sites were obtained from WGA-WGS genotyping data of mixed samples and target individual samples, and a dynamic linkage equilibrium screening strategy was used to screen the intersecting SNP sites to generate a set of SNP sites. The dynamic linkage equilibrium screening strategy is as follows: the whole genome is divided into multiple regions according to a preset physical interval, and the intersection SNP sites in the multiple regions are screened according to heterozygosity to form the SNP site set. Determine the number of mixed contributors N and the error rate of individual allele typing; Based on the application scenario, a preset observation probability model for integrated genotyping errors is invoked. Based on the SNP locus set, the total likelihood ratio is calculated based on the number of mixed contributors N and the genotyping error rate of a single allele. The mixed sample is interpreted based on the total likelihood ratio; The observation probability model for integrating classification errors includes an observation probability model with known contributors and an observation probability model without known contributors. The expression for the observation probability model with no known contributors is: ; in, ; ; The expression for the observed probability model of known contributors is: ; in, ; ; ; In the formula, In order to be in the true genotype Observed at time The probability of; The prior probability of the true genotype of the target individual; In order to be in In the case of an unknown contributor, the sum of the number of 0 alleles in all individuals; For a total of 0 alleles, Observed at time The probability of; The total number of zero alleles from all contributors in the pooled sample; The observed genotype of the target individual; Observed genotypes of known contributors; The observed genotypes for the mixed sample; , The assumption is mutual exclusion; , These are the actual genotypes of the known contributors and the target individuals, respectively. for The total number of zero alleles from unknown contributors; , These represent the number of 0 alleles in the true genotypes of known contributors and target individuals, respectively. for The total number of 0 alleles from unknown contributors.
2. The method for interpreting mixed DNA diallelic SNP data based on WGA-WGS according to claim 1, characterized in that, The error rate of individual allele typing was obtained based on previous Bayesian estimation.
3. The method for interpreting mixed DNA diallelic SNP data based on WGA-WGS according to claim 1, characterized in that, Evidence interpretation is performed based on the total likelihood ratio and a preset judgment threshold, and a judgment result is output, which includes supporting hypotheses. Supporting hypothesis And the evidence lacks discriminatory power.
4. A system for interpreting mixed DNA biseleural SNP data based on WGA-WGS, comprising an interpretation method for mixed DNA biseleural SNP data based on WGA-WGS as described in any one of claims 1-3, characterized in that, include: The locus screening module is used to obtain the intersection SNP loci based on WGA-WGS genotyping data of mixed samples and target individual samples, and to screen the intersection SNP loci using a dynamic linkage balance screening strategy to generate a set of SNP loci. The parameter setting module is used to determine the number of mixed contributors N and the error rate of individual allele typing; The probability calculation module is used to call a preset observation probability model for integrated genotyping errors based on the application scenario. Based on the SNP locus set, it calculates the total likelihood ratio based on the number of mixed contributors N and the genotyping error rate of a single allele. The sample interpretation module is used to interpret the mixed samples based on the total likelihood ratio.