Method for mixed DNA profile association analysis based on STR genotyping and probabilistic models

By employing a hybrid DNA profile correlation analysis method based on STR genotyping and probability models, the challenges of DNA analysis across samples and cases were solved, enabling joint analysis of multiple cases, improving the resolution capability of low-information profiles, supporting correlation analysis of massive hybrid STR data, and providing reliable quantitative evidence support.

CN120564849BActive Publication Date: 2025-12-26CHENGDU SAIXI BIOTECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510709239.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-12-26
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

Existing forensic DNA analysis techniques are ineffective in handling mixed DNA samples with low template quantity or low quality, especially in mixed pattern association analysis across samples and cases. They cannot integrate biological sample data from different cases, resulting in the failure of information complementarity and the inability to reveal criminal networks or behavioral patterns.

Method used

A hybrid DNA profile correlation analysis method based on STR genotyping and probability models was adopted. By setting mutually exclusive propositions, calculating likelihood functions and likelihood ratios, hybrid DNA profiles from different cases were integrated to identify common contributors and quantify their probability of existence, thus achieving joint analysis across samples and cases.

Benefits of technology

It breaks through the limitations of traditional PG models in analyzing single cases, realizes joint analysis of multiple cases, improves the resolution capability of low-information-content graphs, accurately identifies the DNA characteristics of mobile perpetrators, supports the correlation analysis of massive mixed STR data, expands the application scenarios of DNA databases, and provides reliable evidence for case investigation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564849B_ABST
    Figure CN120564849B_ABST
Patent Text Reader

Abstract

The application discloses a mixed DNA map correlation analysis method based on STR gene typing and a probability model, and comprises the following steps: step 1, determining the number of STR map contributors according to N map data characteristic information; step 2, setting mutually exclusive propositions Hp and Hd, Hp is that there is a common contributor among the N mixed maps, and Hd is that there is no common contributor among the N mixed maps; step 3, calculating the frequency of a candidate genotype combination and the probability of the occurrence of the peak height combination; step 4, calculating the likelihood ratio LRcommon of the common contributor among the maps; step 5, setting mutually exclusive propositions H1 and H2, H1 is that a reference individual is a common contributor among the N mixed maps, and H2 is that an unknown unrelated individual is a common contributor among the N mixed maps; and step 6, calculating the likelihood ratio LR, and obtaining a joint analysis result according to the LR. The method is suitable for the joint analysis of different biological samples in the same case or among different cases, and improves the systematicness and scientificity of the mixed DNA map analysis.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of forensic DNA analysis, and in particular to a mixed DNA profile correlation analysis method based on STR genotyping and probability model. BACKGROUND

[0002] In the field of forensic DNA analysis, the analysis of mixed DNA samples (containing DNA of two or more contributors) has always been an important and extremely challenging task. With the complexity of criminal methods and the improvement of detection technology sensitivity, complex mixed samples such as low template amount and multiple contributors have gradually become the mainstream sample type. To cope with this trend, probabilistic genotyping (PG) has become the core solution of the existing technology. PG establishes a biological statistical model (covering random effects such as stutter, drop-out, drop-in), combines statistical theory and computer algorithm to analyze mixed DNA profile, and quantifies the evidence weight based on likelihood ratio (LR).

[0003] For low template amount or low quality mixed samples, part of the PG mathematical model focuses on the joint analysis of multiple genotyping data of the same sample. Through this technical repetition, the genotyping data of multiple repeated samples are integrated to constrain the parameter estimation range of the model, reduce the interference of random effects, thereby improving the accuracy of the inference of the contributor genotype combination and outputting a higher evidence power LR value. Compared with single analysis, this method can enhance the ability to determine whether the Person of Interest (POI) is a contributor. However, technical repetition relies on multiple genotyping of the same sample, and only when the sample size is sufficient for multiple detection can it be performed. However, in actual cases, technical repetition samples cannot be obtained due to trace amount or degradation of the sample. On the other hand, technical repetition needs to strictly meet the consistency of the number of contributors and the composition of individuals, resulting in different samples (such as multiple evidence in a crime scene) being unable to improve the evidence power through joint analysis. Existing PG software usually processes the DNA profile of a single scene sample independently. Even if there are multiple samples in a case, analysts still analyze each evidence profile separately and report LR values respectively, resulting in the fragmentation of potential correlation information across samples. When a single mixed profile cannot confirm the contributor due to insufficient information, case analysis will be at a standstill.

[0004] Serial and parallel case analysis can establish a connection between cases occurring at different times and spaces, providing more clues for crime characteristic description and case detection. The existing PG mathematical model is limited to the integration of data of technical repeat samples of the same detection material within a single case, without establishing a joint probability model across detection materials, making it difficult to support mixed graph correlation analysis across cases and unable to provide scientific basis for serial and parallel case investigation. This limitation leads to the failure of information complementation, which cannot improve the resolution of low-information graphs through multi-detection material data fusion, and makes it difficult to directly correlate DNA evidence with mixed graphs across cases, and unable to reveal criminal networks or behavioral patterns. Although the number of shared alleles between mixed graphs across detection materials and cases can be counted by allele counting method to preliminarily judge the correlation of the graphs, the binary logic cannot be compatible with random effects such as drop-out and drop-in, making it difficult to deal with complex mixed graphs. Although the existing technology has optimized the analysis capability of a single detection material through technical repetition, there is still a technical gap in the fusion of cross-detection material data, multi-graph common contributor inference and serial and parallel case correlation analysis. SUMMARY

[0005] The present application provides a mixed DNA graph correlation analysis method based on STR genotyping and probability model to solve the problems in the prior art.

[0006] The technical scheme adopted by the present application is: a mixed DNA graph correlation analysis method based on STR genotyping and probability model, comprising the following steps:

[0007] Step 1: obtaining STR graphs of N mixed DNA samples to obtain data characteristic information, N ≥2; determining the number of contributors of each STR graph according to the data characteristic information;

[0008] Step 2: setting mutually exclusive propositions Hp and Hd, Hp is N There is a common contributor between N mixed graphs, and Hd is There is no common contributor between

[0009] Step 3: according to the number of contributors of each graph in step 1, the candidate genotype combination set of each graph at the locus is obtained based on the mutually exclusive propositions in step 2; the frequency of each candidate genotype combination and the probability of occurrence of the peak height combination are calculated according to the candidate genotype combination set;

[0010] Step 4: constructing a likelihood function based on the mutually exclusive propositions in step 2 according to the frequency and the probability of occurrence of the peak height combination obtained in step 3; calculating the maximum likelihood value and the corresponding parameters under Hp and Hd according to the likelihood function to obtain the likelihood ratio LRcommon of the common contributor between the graphs. If LRcommon>1, it supportsN If there are common contributors among the mixed maps, proceed to step 5; if LRcommon≤1, exit.

[0011] Step 5: For each possible co-contributor, set mutually exclusive propositions H1 and H2; H1 states that the individual is a co-contributor among N mixed graphs, and H2 states that an unknown, unrelated individual is a co-contributor among N mixed graphs. N Common contributors among the mixed maps;

[0012] Step 6: Based on the mutually exclusive propositions and the likelihood functions under the corresponding parameters from Step 5, obtain the maximum likelihood value and the corresponding parameter values ​​under the conditions of propositional hypotheses H1 and H2. N The deconvolution results of the genotypes of the co-contributors of the mixed map were used to calculate the likelihood ratio (LR) under the corresponding hypothesis, and the association analysis results were obtained based on the LR.

[0013] Furthermore, the data feature information includes locus vectors, allele vectors, and peak height vectors.

[0014] Furthermore, the calculation process for the set of candidate genotype combinations in step 3 is as follows:

[0015] right N Taking the intersection of the locus vectors of each map, we get... N The set of intersection loci of the graphs Mcommon ;

[0016] right Mcommon Each intersecting locus m For loci m The union of allele vectors is taken to include potentially missing alleles. Q ,get N A map at the intersection locus m The union of allele vectors;

[0017] The intersection loci are obtained by sampling alleles twice from the union allele vectors. m The union of genotype vectors on;

[0018] For each mixture map, genotypes in the union genotype vector are repeatedly sampled according to the number of map contributors to obtain the intersection loci of the mixture map. m By using the union-type combination vector on the vector, we can obtain the set of candidate genotype combinations.

[0019] Furthermore, the frequency calculation method for candidate genotype combinations in step 3 is as follows:

[0020] If all alleles in the STR map are present in the population allele frequency table, the genotype frequency and genotype combination frequency can be obtained based on the allele frequency information.

[0021] If the STR profile includes alleles not contained in the population allele frequency table, preset allele frequency information is selected to obtain genotype frequency and genotype combination frequency;

[0022] The genotype combination of a certain genotype combination is expressed as , wherein i is the sequence number of the STR profile, m is the intersection locus sequence number, p is the sequence number of the genotype combination in the union genotype combination vector, and H represents H1 or H2, is the subpopulation correction coefficient of the first i profile.

[0023] Further, the probability calculation method of the peak height combination is as follows:

[0024] The fraction of each allele in the genotype combination is calculated;

[0025] According to the alleles in the union genotype combination and the alleles detected in the profile, each allele is classified:

[0026] The alleles present in the union genotype combination but not in the profile are dropout alleles, the alleles present in the profile but not in the union genotype combination are dropin alleles, and the alleles present in both the profile and the union genotype combination are real alleles;

[0027] According to the classification of alleles, the probability of the occurrence of the corresponding peak height of the alleles is obtained;

[0028] The probability of the occurrence of the corresponding peak height of each allele peak of each profile at a given autosomal locus under a specified genotype combination is calculated, and the probabilities are multiplied to obtain the probability of the occurrence of the peak height combination.

[0029] Further, the process of obtaining the probability of the occurrence of the corresponding peak height of the alleles according to the classification of alleles is as follows:

[0030] The probability of the real alleles m in the intersection locus al of the profile is:

[0031]

[0032] In the formula: is the fraction of the alleles al in the profile, is the peak height expectation of the profile,​ denoted as the coefficient of variation of peak height in the spectrum. Represents the computational model;

[0033] Considering the peak height when the stutter occurs probability for:

[0034]

[0035] In the formula: The peak height of the stutter peak in the spectrum is the proportion of the peak height of the true allele peak that produces the stutter.

[0036] Considering the peak height of true alleles when the map degrades. probability for:

[0037]

[0038] In the formula: The segment length of gene a1 at the intersection gene locus in the map. This represents the fragment length of the allele with the smallest fragment length among all loci in the population in the gene map. The degradation parameters of the spectrum; the probability of the corresponding stutter peak. as follows:

[0039]

[0040] In the formula: stutter peak The length of the segment;

[0041] When a specified genotype combination contains the missing allele Q, its peak height probability for:

[0042]

[0043] In the formula: Union loci in the map m upper allele Q The number of portions;

[0044] When degradation occurs, the peak height probability is calculated as follows:

[0045]

[0046] In the formula: Union loci in the map m The length of the fragment corresponding to the allele with the longest fragment length;

[0047] When the dropin occurs at the intersection locus in the profile m The peak height of the dropin The probability of the dropin is:

[0048]

[0049] In the formula: is the probability of the dropin in the profile, is the parameter of the exponential distribution to which the dropin in the profile is subjected, is the probability of the dropin allele.

[0050] Further, in the step 4, the process of constructing the likelihood function based on the mutual exclusive propositions in the step 2 is as follows:

[0051] Under the Hp hypothesis, the likelihood functions corresponding to different genotype combination sets (in the frequency calculation part, N The genotype frequencies of the common contributors of the profiles are calculated only once) are added to obtain the likelihood function of the Hp hypothesis N Under the Hp hypothesis, the likelihood function of the common contributors of the profiles at the intersection loci m is obtained by multiplying the likelihood functions of the loci, and the likelihood function under the Hp hypothesis is obtained.

[0052] Under the Hd hypothesis, the likelihood functions of each profile at all loci are multiplied to obtain the joint N The likelihood function of all intersection loci in the profiles, i.e. the likelihood function under the Hd hypothesis is obtained.

[0053] Further, in the step 4, the likelihood ratio of the common contributors of the profiles LRcommon is calculated as follows:

[0054]

[0055] In the formula: is the maximum value of the likelihood function under the Hp hypothesis, is the maximum value of the likelihood function under the Hd hypothesis.

[0056] Further, the process of obtaining the joint analysis result according to the LR is as follows:

[0057] If LR>1, the H1 hypothesis is supported; if LR<1, the H2 hypothesis is supported, and if LR=1, the process is exited.

[0058] Further, in the step 1, the maximum allele count method or the maximum likelihood estimation method is used to determine the number of contributors of each STR profile.

[0059] The present application has the following advantages:

[0060] (1) The method of this invention breaks through the limitation of the traditional PG model that can only handle a single case, and can be used to analyze multiple cases in combination; it can integrate biological sample data from different cases and accurately identify the DNA characteristics of itinerant criminals; it achieves a breakthrough from single-point analysis to multi-point linkage.

[0061] (2) This invention can quantitatively assess the correlation between the number of contributors and mixed DNA maps with different individual compositions by calculating the joint likelihood ratio across maps; it realizes information complementarity between any samples or specimens and improves the resolution capability of low-information maps by using multi-map data.

[0062] (3) This invention not only realizes the scientific correlation analysis between mixed DNA maps, and can objectively assess the probability of the existence of common contributors, providing reliable quantitative evidence support for the work of linking cases; but also effectively improves the map resolution by combining multiple sample data, providing a solution for the accurate determination and analysis of the genotypes of mixed map contributors.

[0063] (4) This invention supports the correlation analysis of massive mixed STR data, expands the application scenarios of DNA databases, enhances their application value in the investigation of serial cases, and provides technical means for the rapid detection of complex series of cases. Attached Figure Description

[0064] Figure 1 This is a schematic diagram of the method flow of the present invention. Detailed Implementation

[0065] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0066] like Figure 1 As shown, a hybrid DNA mapping association analysis method based on STR genotyping and probability models includes the following steps:

[0067] Step 1 to obtain N STR profiles of mixed DNA samples were obtained to acquire data feature information. N ≥2; Determine the number of contributors for each STR map based on data feature information;

[0068] Data feature information includes N A mixed DNA map locus vectors Allele vectors and peak height vector .in E N For the first N A map. , For the first N Locus vectors of each map; allele vectors , For the first N Allele vectors of each map. Peak height vector. , For the first N Allele peak height vectors of each spectrum.

[0069] The number of contributors to each STR map was determined using the maximum allele count method or the maximum likelihood estimation method based on the data characteristics.

[0070] Step 2: Set mutually exclusive propositions Hp and Hd, where Hp is... N There are common contributors among the mixed maps. Hd is N There are no common contributors among the mixed maps; among them For the first i DNA map E i The first contributor, .

[0071] In this invention, it is assumed that the common contributor is the first contributor in each DNA map, and this assumption has no effect on the calculation results.

[0072] Step 3: Based on the number of contributors to each graph in Step 1, obtain the set of candidate genotype combinations for each graph at the intersection locus based on the mutual exclusion proposition in Step 2; calculate the frequency of each candidate genotype combination and the probability of that peak combination appearing based on the set of candidate genotype combinations.

[0073] The calculation process for candidate genotype combinations is as follows:

[0074] right N Locus vectors of each map Taking the intersection, we get N The set of intersection loci of the graphs Mcommon .

[0075] right Mcommon Each intersecting locus m In the map E i loci m Allele vectors on For each Take the union of the sets, and add any potentially missing alleles. Q ,get N A map at the intersection locus m Union of allele vectors .

[0076] Union of allele vectors Alleles were sampled twice to obtain the intersection loci. m Union genotype vectors on ;exist There is x In the case of one allele, There is Different genotypes.

[0077] union genotype vectors Genotype duplication sampling K i Next, the map was obtained. E i At the locus m Union-type combined vectors on , , For the first i A map E i At the locus m The first p There are 2 union genotype combinations, containing a total of 2 K i One allele. There is By combining different genotypes, a set of candidate genotype combinations can be obtained.

[0078] The calculation process for the frequencies of candidate genotype combinations is as follows:

[0079] Genotype frequencies and genotype combination frequencies are calculated based on the frequency information in the population allele frequency table. The heterozygous genotype frequency is the product of the two allele frequencies multiplied by 2, and the homozygous genotype frequency is the square of the allele frequency.

[0080] If an allele not included in the frequency table appears in the graph, the frequency of the rare allele is used instead of the frequency of that allele for calculation.

[0081] If the genotype contains potentially lost alleles Alleles The frequency is 1 minus the sum of the frequencies of all alleles detected at that locus in the map;

[0082] When considering the existence of subpopulation structures between corresponding populations within a map, the allele frequencies of the map need to be corrected using the Fst value. express;

[0083] Atlas E i At the locusm The specified union of genotypes The genotype combination frequency is expressed as In the formula i This represents the sequence number of the sample in the STR map. m The intersection locus number, p Here, H represents the index in the union-type combined vector, where H denotes H1 or H2. For the first i Subgroup correction coefficients for each spectrum.

[0084] The probability of this peak height combination occurring is calculated as follows:

[0085] Calculate the number of copies of each allele in the genotype combination;

[0086] Known map E i At the locus m Union Genotype Combination Vector ,for Each specified genotype combination ,calculate Number of copies of each allele C Alleles a Number of portions for:

[0087]

[0088] In the formula: K i For the map E i The number of contributors, k The contributor's serial number. For the map E i The Middle k One contributor at the locus m Alleles contained in the specified genotype a The number of For the map E i The Middle k The combined percentage of contributors.

[0089] Based on genotype combination Alleles and maps E i At the locus m Allele vectors on Classify each allele:

[0090] Existing in the union of genotype combinations dropout; present in allele vector dropin; present in both allele vector dropin; present in both allele vector

[0091] According to the classification of alleles, the probability of the peak height corresponding to the occurrence of alleles is obtained;

[0092] Specifically, the following cases can be divided:

[0093] 1) For the map E i The true allele m at the locus al , the probability of the peak height is calculated as follows:

[0094]

[0095] In the formula: is the fraction of alleles in the map, al is the peak height expectation of the map, is the peak height coefficient of variation of the map, denotes the calculation model;

[0096] 2) Considering the occurrence of stutter, for the specified genotype combination , where each true allele al produces stutter, denoted as stutter set , where, is the map E i The probability of the peak height m of the stutter peak produced by the i-th true allele of the i-th genotype combination at the locus p is calculated as follows: t

[0097]

[0098] In the formula: is the proportion of stutter peaks in the map to the peak height of the true allele peak that produces stutter;

[0099] Corresponding to it, the true allele peak that produces the stutter peak​​​​​​​​al peak height probability for:

[0100]

[0101] 3) When considering map degradation, the true allele peak al peak height probability The calculation method is as follows:

[0102]

[0103] In the formula: The segment length of gene a1 at the intersection gene locus in the map. This represents the fragment length of the allele with the smallest fragment length among all loci in the population in the gene map. The degradation parameters of the spectrum; the probability of the corresponding stutter peak. as follows:

[0104]

[0105] In the formula: stutter peak The length of the segment;

[0106] 4) When a specified genotype combination contains a missing allele Q At that time, the diagram E i At the locus m Dropout alleles Q Without considering degradation, its peak height The probability calculation method is as follows:

[0107]

[0108]

[0109] In the formula: For the map The analysis threshold, Indicates alleles peak height Less than The conditional probability, For the map Middle gene locus upper allele The number of portions.

[0110] When degradation occurs, the peak height probability is calculated as follows:

[0111]

[0112] wherein: is the profile is the locus is the fragment length corresponding to the allele with the longest fragment length.

[0113] 5) When the profile E i is the locus m drop in, the peak height of the drop in allele is calculated as follows:

[0114]

[0115]

[0116] wherein: is the profile is the probability of drop in, is the profile is the parameter (constant value) of the exponential distribution to which the drop in complies in the profile is the frequency of the drop in allele, represents the conditional probability of the occurrence of the drop in allele given the parameters .

[0117] The probability of the occurrence of each allele peak with its corresponding peak height is calculated for each profile at a given autosomal locus for a given genotype combination. The probabilities are multiplied to obtain the probability of the occurrence of the peak height combination.

[0118] For the profile E i is the locus m is the genotype combination , the peak height probabilities of the alleles in are multiplied to obtain the probability of the peak height combination .

[0119] is the profile E i is the locus m is the genotype combination corresponding to the peak height combination probability of the specified union genotype combination at the locus , which can be expressed as the conditional probability .

[0120] Step 4: Based on the frequency and probability of the peak height combination obtained in Step 3, construct a likelihood function based on the mutually exclusive proposition in Step 2; calculate the maximum likelihood value and corresponding parameters under Hp and Hd using the likelihood function, and obtain the likelihood ratio LRcommon of common contributors between the spectra. If LRcommon > 1, then it is supported. N If there are common contributors among the mixed maps, proceed to step 5; if LRcommon≤1, exit.

[0121] The calculation process of the likelihood function is as follows:

[0122] 1) Under the Hd hypothesis, individual For unrelated individuals, their genotype frequencies need to be calculated separately, and the genotype map needs to be created. E i At the locus m The likelihood function on is:

[0123]

[0124] In the formula: For the map E i At the locus m The specified genotype combination The frequency. for E i At the locus m superior The corresponding peak height combination probability, for E i At the locus m Union Genotype Combination Vector Each of them Calculate the product of the genotype combination frequency and the peak height combination probability, and sum the products corresponding to different genotype combinations to obtain the result. .

[0125] joint N A graph, combining all intersecting loci Mcommon The likelihood function under the Hd assumption is:

[0126]

[0127] Since the mixture maps are independent of each other and the intersection loci are independent of each other, the likelihood function of each map on all loci is calculated. Multiplication yields union N Likelihood function of all intersecting loci in a given map .

[0128] 2) Under the Hp hypothesis, individual ,N Shared contributors among the graphs At the locus m Possible genotype vectors For co-contributors at the locus m Upper q One possible genotype.

[0129] according to exist m The genotype on Divided into q Group ,in, Representation of the spectrum E i medium-sized individuals At the locus m The genotype is The set of all genotype combinations, .

[0130] joint N A map, at the locus m Above, when the genotypes of co-contributors are The likelihood function at time is:

[0131]

[0132] In the formula: Representation of the spectrum At the locus The specified genotype combination frequency, express exist superior The corresponding peak height combination probability;

[0133] Traversal Each of them Calculate the product of the frequency of genotype combinations and the probability of peak height combinations, and then... Adding the corresponding products together, we get exist superior The genotype is The likelihood function;

[0134] joint When each graph is displayed, Multiply the likelihood functions. When hour, The frequency of the spectrum in The product of the genotype frequencies of all individuals; when hour, The frequency of the spectrum in The product of the genotype frequencies of all individuals except co-contributors, i.e., the individual Genotype frequency It is still calculated based on frequency information, and No need for repeated calculations;

[0135] joint A graph, combining all intersecting loci The likelihood function under the Hp assumption is:

[0136]

[0137] Set up different genotype combinations The corresponding likelihood functions are added together to obtain A map at the gene locus The likelihood function of a common contributor exists, and the product of the likelihood functions of each point is obtained. .

[0138] The likelihood ratio of common contributors among graphs LRcommon The calculation process is as follows:

[0139]

[0140] In the formula: Let be the maximum value of the likelihood function under the Hp assumption. Let be the maximum value of the likelihood function under the Hd assumption.

[0141] If the calculation yields It supports the Hp hypothesis, that is, it supports There are common contributors among the mixed maps; if It supports the Hd hypothesis, that is, it supports There are no common contributors among the mixed maps; Then it is considered that there is not enough information to make a judgment. Are there common contributors among the mixed maps?

[0142] Step 5: For each possible individual, set mutually exclusive propositions H1 and H2, where H1 is the reference individual. N Common contributors among the hybrid maps H2 is an unknown, unrelated individual that is a common contributor among N mixed graphs. ;

[0143] in, Indicates the first i DNA map The first contributor in, ( The first contributor in each DNA map is a co-contributor.

[0144] Step 6: Based on the mutual exclusion propositions and the likelihood functions under the corresponding parameters in Step 5, obtain the maximum likelihood value and the corresponding parameter value under the conditions of propositional hypotheses H1 and H2, obtain the likelihood ratio LR under this condition, and obtain the correlation analysis results based on LR and the set threshold.

[0145] The likelihood ratio (LR) calculation process is as described above. LRcommon The calculation process:

[0146]

[0147]

[0148] Representation of the spectrum medium-sized individuals At the locus When the genotype on the upper part is the Ref genotype, exist The set of all possible genotype combinations. ; Indicates the H1 hypothesis and parameters Under these conditions, the map At the locus The above observed a specific genotype combination The conditional probability is calculated using the method described above.

[0149] Indicates in the parameter Under the conditions, exist superior The corresponding peak height combination is The conditional probability is calculated using the method described above. The product of the genotype combination frequency and the peak height combination probability will China and other countries Adding the corresponding products together, we get exist superior The genotype of is the likelihood function of the genotype of Ref. Multiplying the likelihood functions of the mixture maps ( (No need for repeated calculations) Multiplying the likelihood functions at each point yields the result. The calculation method is the same as step 5.

[0150] The maximum values ​​of the likelihood functions under hypotheses H1 and H2 are obtained respectively, and then the likelihood ratio LR and the spectrum are obtained. The unknown parameters corresponding to the maximum values ​​of the likelihood functions under H1 and H2 are respectively and .

[0151] If LR > 1, support H1 hypothesis, that is, support Ref is If LR < 1, support H2 hypothesis, that is, support unknown independent individual is If LR = 1, exit, that is, it is considered that there is not enough information to judge.

[0152] Embodiment

[0153] The application will be further described below in combination with specific embodiments.

[0154] A mixed DNA profile correlation analysis method based on STR genotyping and probability model, comprising the following steps:

[0155] Step 1: Simulate preparation of two mixed DNAs from different scenes, and there is a common contributor between the two mixed DNAs, use GoldenEye 25AC kit combined with polymerase chain reaction for DNA amplification, capillary electrophoresis counting for STR typing, and obtain the locus vector , allele vector and peak height vector of two mixed DNA profiles.

[0156] According to the data information of two STR profiles, use maximum likelihood estimation method to determine the number of contributors of two STR mixed DNA profiles , wherein , .

[0157] Step 2: Set mutually exclusive propositions Hp and Hd, Hp is that there is a common contributor between two mixed profiles , Hd is that there is no common contributor between mixed profiles;

[0158] wherein, represents the first contributor in the i th DNA profile , .

[0159] In order to facilitate expression and calculation, it is assumed in the method that the common contributor is the first contributor in each DNA profile, and this assumption has no effect on the calculation result.

[0160] Step 3: According to the number of contributors of each profile in step 1, on the intersection locus of the profiles, based on the mutually exclusive propositions in step 2, obtain the set of candidate genotype combinations of each profile at the locus; according to the set of candidate genotype combinations, calculate the frequency of each candidate genotype combination and the probability of appearing the peak height combination. ​​​​

[0161] The process of obtaining the genotype combination vector is as follows:

[0162] right Taking the intersection yields the set of intersection loci of the two maps. Mcommon , Mcommon The CCP contains 23 autosomal loci.

[0163] for Mcommon Each intersecting locus m In the map E i loci m Allele vectors are present For each Take the union and add any missing alleles. Q ,get N A map at the intersection locus m Union of allele vectors Taking the first intersecting locus as an example, , Known , ,but .

[0164] Will Alleles in the sample were sampled twice to obtain loci. Union genotype vectors on ;exist There is In the case of one allele, There is Different genotypes; taking the first intersecting locus as an example. There are 4 alleles. There are 10 different genotypes;

[0165] Will Genotype duplication sampling Next, the map was obtained. At the locus Union Genotype Combination Vector ; , For the first A map At the locus The first The union of genotype combinations contains a total of 100,000 genotype combinations. One allele; There is Different combinations of genotypes; taking the first intersecting locus as an example. There are 100 different genotype combinations, each of which contains 4 alleles, There are 1000 different genotype combinations, each of which contains 6 alleles.

[0166] Then, the frequency of the generated genotype combination is calculated:

[0167] The genotype frequency and genotype combination frequency are calculated according to the frequency information in the population allele frequency table, wherein the heterozygous genotype frequency is the product of the frequencies of two alleles multiplied by 2, and the homozygous genotype frequency is the square of the allele frequency; The present application adopts the STR frequency of the national population;

[0168] If an allele not included in the frequency table appears in the map, the frequency of the rare allele is used instead of the frequency of the allele for calculation;

[0169] If there is a possible missing allele in the genotype , the frequency of the allele is 1 minus the sum of the frequencies of all detected alleles at this locus in the map;

[0170] When considering the existence of sub-population structure corresponding to the population in the map, the allele frequency of the map needs to be corrected by the Fst value, denoted as The Fst value used in the present application is 0.0108;

[0171] Then, the genotype combination frequency of the specified union genotype combination at the locus can be expressed as the conditional probability , wherein is Hp or Hd.

[0172] For each map, at each intersection locus, for each genotype combination, the probability of the occurrence of the peak height combination is calculated using the gamma model, and the calculation steps are as follows:

[0173] S1: Calculate the fraction of each allele in the genotype combination;

[0174] Given the map , the union genotype combination vector at the locus , for each specified genotype combination in , the fraction of each allele in is calculated , then the fraction of the allele is: ​​

[0175]

[0176] wherein: the number of contributors of the map is the serial number of the contributor is the number of alleles contained in the genotype designated by the th contributor of the map at the locus is the proportion of the mixture of the th contributor of the map ; taking the first designated genotype combination of the first intersection locus of the first map as an example, , . .

[0177] S2: classify each allele by comparing the alleles in the genotype combination with the alleles detected by the map;

[0178] According to the allele vector of the alleles in the genotype combination and the alleles in the map at the locus , , the alleles present in but not in are dropout alleles; the alleles present in but not in are dropin alleles; the alleles present in both and are real alleles; taking the first intersection locus of the first map and a certain designated genotype combination thereon as an example, , dropout alleles are dropin alleles are real alleles.

[0179] S3: calculate the probability of each allele in the genotype combination appearing at the corresponding peak height, and the specific calculation process is as described above.

[0180] S4: calculate the probability of each allele peak appearing at its corresponding peak height for each genotype combination of a given autosomal locus in each map, and then multiply the probabilities to obtain the probability of appearing at the peak height combination;

[0181] For the map at the locus ​​​The specified genotype combination ,Will The peak height probabilities of each allele are multiplied together to obtain the peak height combination. The probability of;

[0182] Then, the map At the locus The specified union of genotypes The probability of peak height combinations corresponding to genotype combinations can be expressed as conditional probability. ,in .

[0183] Step 4: Based on the frequency and probability of the peak height combination obtained in Step 3, construct a likelihood function based on the mutually exclusive proposition in Step 2; calculate the maximum likelihood value and corresponding parameters under Hp and Hd using the likelihood function, and obtain the likelihood ratio LRcommon of common contributors between the spectra. If LRcommon > 1, then it is supported. N If there are common contributors among the mixed maps, proceed to step 5; if LRcommon≤1, exit.

[0184] By combining two maps and all intersecting loci, write the likelihood function given the hypothesis and corresponding parameters, which can be divided into the following two cases:

[0185] 1) Under the Hd hypothesis, individual For unrelated individuals, their genotype frequencies need to be calculated separately, and the genotype map needs to be created. At the locus The likelihood function on is:

[0186]

[0187] in, Representation of the spectrum At the locus The specified genotype combination frequency, express exist superior The corresponding peak height combination probability, for exist Union Genotype Combination Vector Each of them Calculate the product of the genotype combination frequency and the peak height combination probability, and sum the products corresponding to different genotype combinations to obtain the result. ;

[0188] Combine the two maps and combine all 23 intersecting loci. The likelihood function under the Hd assumption is:

[0189]

[0190] Multiply the likelihood function of each profile at all loci Get the likelihood function of the joint of all 23 intersection loci of two profiles .

[0191] 2) Under the Hp hypothesis, the common contributor between two profiles The possible genotype vector at locus , , is the possible genotype of the common contributor at locus ;

[0192] According to the genotype at , , is divided into groups , where, represents the set of all genotype combinations of individuals in profile , at locus , , ;

[0193] Take the first intersection locus as an example, the common contributor of two profiles is , The possible genotype vector of the common contributor at the first intersection locus is , where , , , , , ; According to the genotype at the first intersection locus , is divided into 10 groups , where, represents the set of all genotype combinations of individuals in profile , at the first intersection locus , , where , , , .

[0194] Joint two profiles, loci​ The likelihood function of the genotype of the common contributor when the genotypes of the contributors are is:

[0195]

[0196] where, represents the profile of the genotypes of the individuals in the locus , the frequency of the specified genotype combination , represents the corresponding peak height combination probability.

[0197] Traverse each of the , calculate the product of its genotype combination frequency and peak height combination probability, add the corresponding products of different to obtain the likelihood function of the genotype of the common contributor in .

[0198] When combining two profiles, the likelihood functions of each are multiplied. When , the frequency of is the product of the genotype frequencies of all individuals in the profile at ; when , the frequency of is the product of the genotype frequencies of all individuals in the profile except the common contributor at , i.e., the genotype frequency of individual is still calculated according to the frequency information, and there is no need to repeat the calculation; taking the first intersection locus as an example, , the frequency of which is , and , the frequency of which is .

[0199] When combining two profiles, the likelihood functions of all 23 intersection loci are multiplied, and the likelihood function under the Hp hypothesis is:

[0200]

[0201] Add the likelihood functions corresponding to different genotype combination sets to obtain the likelihood function of the existence of a common contributor at locus of the two profiles, and the likelihood functions of each intersection locus are multiplied to obtain .​​​​​

[0202] The calculation method of LRcommon is as follows:

[0203]

[0204] The maximum value of the likelihood function under the assumption of Hp and Hd is obtained respectively, , and the likelihood ratio is obtained , , supporting the assumption of Hp, that is, supporting the existence of common contributors between the two mixed profiles;

[0205] The profile The unknown parameters corresponding to the maximum value of the likelihood function under the assumption of Hp and Hd are and , The parameters include the mixed ratio parameter , , the peak height expectation parameter , the peak height coefficient of variation parameter , the degradation parameter , and the stutter ratio parameter ; The parameters include , , , , , , ; similarly, , .

[0206] Step 5: For each possible individual, set the mutually exclusive propositions H1 and H2; H1: ref2 is a common contributor between the two mixed profiles, ;

[0207] H2: unknown unrelated individual is a common contributor between the two mixed profiles, ;

[0208] wherein represents the first contributor in the th DNA profile , , and the first contributor in each DNA profile is the common contributor.

[0209] Step 6: According to the likelihood function under the corresponding parameters of the mutually exclusive propositions in step 5, the maximum likelihood value and the corresponding parameter value under the assumption of the corresponding propositions H1 and H2 are obtained, the likelihood ratio LR under the condition is obtained, and the correlation analysis result is obtained according to the threshold set according to LR.

[0210] maximize the likelihood function under H1 and H2 respectively, , and then get the likelihood ratio , support H1 hypothesis, i.e. support ref2 is the common contributor between the two mixed profiles;

[0211] profile the unknown parameters corresponding to the maximum likelihood function under H1 and H2 are and respectively; the parameters in the parameter set include the mixing ratio parameter , the peak height expectation parameter , the peak height coefficient of variation parameter , the degradation parameter , the stutter ratio parameter ; the parameters in the parameter set include , , , , , , ; similarly, , .

[0212] if LR>1, support H1 hypothesis; if LR<1, support H2 hypothesis, if LR=1, exit.

[0213] The application breaks through the limitation of the prior art that can only process a single sample or technically repeated samples. By establishing a cross-mapping joint likelihood ratio (LRcommon) calculation system, the quantitative evaluation of the correlation between mixed DNA maps with different numbers of contributors and individual compositions is realized. This technology not only supports the joint analysis of multiple maps derived from cases occurring at different times and spaces, but also can be used for the joint analysis of different biological samples of the same case. The technical bottleneck of the traditional PG mathematical model in the analysis of serial and parallel cases is solved. A double-layer hypothesis testing system is adopted to evaluate the probabilities of "whether there is a common contributor between maps" and "whether a specific individual is a common contributor". This hierarchical architecture forms a complete evidence chain: first confirm the cross-map correlation, and then verify the specific individual identity. Compared with the single-layer comparison mode of the prior art, the systematicness and scientificity of the analysis are significantly improved. The cross-map constraint condition is introduced, which requires the first contributor to be the common contributor. This optimization design greatly reduces the number of invalid combinations, significantly improves the calculation efficiency while ensuring the analysis accuracy, and provides feasibility for handling complex serial and parallel case scenarios. The application breaks through the limitation of the traditional database that only supports individual comparison, develops a joint likelihood ratio algorithm for mixed maps, and can be used for comparison between mixed maps and mixed maps in the database, providing directly adoptable DNA evidence for serial and parallel cases. By outputting the LRcommon value and the probability type retrieval index between mixed maps, the correlation between mixed maps is quantified, the direct correlation analysis between mixed maps is realized, and a new way for the serial and parallel application of massive DNA data is opened up.

Claims

1. A method for STR genotyping and probabilistic model-based mixed DNA profile association analysis, characterized in that, The method comprises the following steps: Step 1 acquisition N STR profiles of the mixed DNA sample, obtaining data characteristic information, N ≥2; determining the number of STR profile contributors according to the data characteristic information; Step 2: Set up exclusive propositions Hp, Hd, Hp is N There is a common contributor between the two mixed maps, Hd is N There is no common contributor between the two mixed maps; Step 3: according to the number of contributors of each profile in step 1, the set of candidate genotype combinations of each profile at the locus of the intersection of the profiles is obtained based on the exclusive proposition in step 2; the frequency of each candidate genotype combination and the probability of the occurrence of the peak height combination are calculated according to the set of candidate genotype combinations; Step 4: according to the frequency and the probability of the occurrence of the peak height combination obtained in step 3, a likelihood function based on the exclusive proposition in step 2 is constructed; the maximum likelihood value and the corresponding parameters under Hp and Hd are calculated according to the likelihood function, and a likelihood ratio LRcommon of the common contributors between the profiles is obtained; If LRcommon > 1, then support N a common contributor between the two mixed profiles exists and go to step 5; if LRcommon < 1, then exit. Step 5: For each possible co-contributor, set the mutually exclusive propositions H1, H2; H1 is the possible co-contributor is N a co-contributor between the two mixed profiles, H2 is the unknown unrelated individual is N a co-contributor between the two mixed profiles; Step 6: According to the exclusive propositions of step 5 and the likelihood function under the corresponding parameters, the maximum likelihood value and the corresponding parameter value under the corresponding proposition hypothesis H1, H2 are obtained, and N The genotype deconvolution results of the common contributors of the mixed profile are obtained, and the likelihood ratio LR under the corresponding hypothesis is calculated. The correlation analysis result is obtained according to the LR.

2. The method according to claim 1, wherein the STR genotyping and probabilistic model-based mixed DNA profile relatedness analysis method is characterized by, The data feature information comprises a locus vector, an allele vector and a peak height vector.

3. The method of claim 2, wherein the method is based on STR genotyping and a probabilistic model. The calculation process of the set of candidate genotype combinations in step 3 is as follows: right N Taking the intersection of the locus vectors of each map, we get... N The set of intersection loci of the graphs Mcommon ; For each of the intersection loci Mcommon m , the allele vectors for the loci m are taken in union, adding any alleles that can have been lost Q , resulting in N a union allele vector for the intersection loci m ;​ The alleles in the union allele vector are resampled twice to obtain the intersection genotype vector on the union loci m ​ For each mixed profile, the genotypes in the union genotype vector are resampled according to the number of profile contributors to obtain the union genotype combination vector of the mixed profile at the intersection loci m , i.e. the set of desired candidate genotype combinations is obtained.

4. The method of claim 3, wherein the method is characterized by, The calculation method of the frequency of the candidate genotype combination in step 3 is as follows: If all the alleles in the STR profile exist in the population allele frequency table, the genotype frequency and the genotype combination frequency are obtained according to the allele frequency information; If the STR profile includes an allele not contained in the population allele frequency table, preset allele frequency information is selected to obtain the genotype frequency and the genotype combination frequency; Genotype combination The genotype combination frequency is represented as , where i is the serial number of the STR profile, m is the serial number of the intersection locus, p is the serial number of the genotype combination in the union genotype combination vector, and H represents H1 or H2, is the subpopulation correction coefficient of the i th profile.

5. The method for STR-based genotyping and probabilistic model-based mixed DNA profile relatedness analysis according to claim 4, wherein, The calculation method of the probability of the occurrence of the peak height combination is as follows: The number of parts of each allele in the genotype combination is calculated; According to the alleles in the union genotype combination and the alleles detected by the profile, each allele is classified: The allele that exists in the union genotype combination but does not exist in the profile is a dropout allele; The allele that exists in the profile but does not exist in the union genotype combination is a dropin allele; the allele that exists in both the profile and the union genotype combination is a real allele; The probability of the occurrence of the corresponding peak height of the allele is obtained according to the classification of the alleles; The probability of the occurrence of the corresponding peak height of the allele is obtained according to the classification of the alleles; 6. The method for STR-based genotyping and probabilistic model-based mixed DNA profile relatedness analysis according to claim 5, wherein, The probability of the occurrence of the corresponding peak height of the allele is obtained according to the classification of the alleles; The map is on the true alleles at the intersection loci m al , the probability of peak height is:​​ In the formula: Alleles in the map al The number of portions, The expected peak height of the spectrum. denoted as the coefficient of variation of peak height in the spectrum. Represent the computational model; Considering the occurrence of a stutter, the probability of its peak height is: is: In the formula: is the proportion of stutter peaks in the profile to the peak height of the true allele peak that generates the stutter. Considering the degradation of the profile, the probability of the peak height of the true allele is: :​ where: is the fragment length of the allele al at the locus of interest in the profile, is the fragment length of the smallest allele in the population at all loci in the profile, is the degradation parameter of the profile; the probability of stutter peaks as follows: wherein: is the stutter peak fragment length; When the specified genotype combination contains a missing allele Q the probability of its peak height is: ​ In the formulae: Union loci in the atlas m Upper allele Q Fraction; The probability of the occurrence of the corresponding peak height of the allele is obtained according to the classification of the alleles; In the formula: is the union locus in the map m the fragment length of the allele with the longest fragment length in the upper When the union locus in the profile m occurs a drop in, its peak height is the probability that: wherein: is the probability of a dropin, is the parameter of the exponential distribution to which dropins in the profile are subject, is the probability of a dropin allele.

7. The method according to claim 6, wherein the method is characterized by, The calculation method of the peak height probability when degradation occurs is as follows: Under the Hp hypothesis, the likelihood functions corresponding to different genotype combinations are added to obtain N The likelihood function of the common contributor at the union locus m The likelihood function of the common contributor at the union locus ; Under the Hd assumption, the likelihood function of each profile is multiplied over all loci to obtain the joint N likelihood function of all intersecting loci in the individual profiles, i.e., the likelihood function under the Hd assumption .

8. The method according to claim 7, wherein the STR genotyping and probabilistic model-based mixed DNA profile relatedness analysis method is characterized by, The likelihood ratio of the common contributors between the profiles in step 4 In step 4, the process of constructing the likelihood function based on the exclusive proposition in step 2 is as follows: The calculation process is as follows: where: is the maximum of the likelihood function under the Hp hypothesis, is the maximum of the likelihood function under the Hd hypothesis.

9. The method of claim 8, wherein the method is based on STR genotyping and a probabilistic model. LRcommon The process of obtaining the association analysis result according to LR is as follows:

10. The method according to claim 1, wherein the method is characterized by, If LR>1, H1 hypothesis is supported; if LR<1, H2 hypothesis is supported; if LR=1, exit. In step 1, the maximum allele counting method or the maximum likelihood estimation method is used to determine the number of contributors of each STR profile.

Citation Information

Patent Citations

  • Mixed DNA evidence analysis method based on STR genotyping map

    CN119028440A

  • Mixed DNA sample donor proportion analysis method based on STR typing and residual matrix analysis

    CN119811502A