20838 snp locus combinations for inferring human kinship ranks and applications thereof

By combining 20,838 SNP loci and a computational model, and integrating the genetic similarity coefficient GISC and the zero-shared genetic index GSI0, the problem of accuracy in predicting complex kinship levels in forensic identification was solved, achieving highly accurate prediction of kinship levels.

CN120913638BActive Publication Date: 2026-05-05INST OF FORENSIC SCI OF MIN OF PUBLIC SECURITY
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INST OF FORENSIC SCI OF MIN OF PUBLIC SECURITY
Filing Date
2025-06-26
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing technologies are insufficient to accurately predict complex and diverse kinship levels in forensic identification, often leading to difficulties in practical applications.

Method used

Using 20,838 SNP loci combinations, combined with the genetic similarity coefficient GISC and the zero-shared genetic index GSI0, the degree of kinship was inferred through a computational model. Primers, probes, or gene chips were used for detection. Specific amplification primer pairs were designed or gene chips were used for the detection of SNP loci combinations.

Benefits of technology

It achieved accurate prediction of forensic genealogy relationships within the fourth degree, with a confidence interval accuracy greater than 99.77% and a low false negative rate. The accuracy rate for fourth-degree kinship was 95.51%, with a false negative rate of only 0.83%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913638B_ABST
    Figure CN120913638B_ABST
Patent Text Reader

Abstract

The application provides a 20838-SNP-site combination for inferring human kinship grades and application thereof. The application provides a SNP-site combination for inferring human kinship grades, which comprises 20838 SNP sites. Based on the 20838-SNP-site combination, the application establishes an average moment method as a pedigree relationship inference algorithm, which is used for predicting kinship within four grades (inclusive) in forensic genealogy. The confidence interval accuracy rate of the system for predicting first to third grades is greater than 99.77% and there is no false negative, and the confidence interval accuracy rate of the four-grade kinship is 95.51%, and the false negative rate is only 0.83%.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bioinformatics, and in particular to a combination of 20,838 SNP loci for inferring the degree of human kinship and its application. Background Technology

[0002] Kinship refers to the biological lineage of relatives, including parent-child and sibling relationships. Kinship ranking refers to the quantitative classification of kinship based on the degree of genetic similarity or the closeness of blood ties between individuals. In the field of forensic identification, researching and inferring kinship ranking is of paramount importance, with a wide-ranging and profound impact, providing core support for solving many complex cases. In the process of solving criminal cases, kinship identification can be a powerful tool for identifying suspects. When biological evidence left at the scene (such as bloodstains, hair, and skin flakes) is associated with known relatives of the suspect, analyzing the kinship between the two, using genetic clues, can greatly narrow down the scope of the investigation and improve the efficiency of case solving. In some crime scenes, if it is not possible to directly obtain biological samples from the suspect, analyzing the kinship between those related to the case and their relatives may uncover potential criminal clues, thus providing a crucial path to breakthroughs in the case. In handling cases of missing persons and identifying unidentified bodies, kinship identification is the core means of achieving accurate identification. By identifying the kinship between DNA samples of missing persons' relatives and biological samples of suspected missing persons, the connection between the two can be accurately established, and identity information can be determined.

[0003] Single nucleotide polymorphisms (SNPs) are third-generation genetic markers in forensic medicine, characterized by their wide distribution, low mutation rate, and high genetic stability. They are crucial genetic markers for individual identification and kinship analysis. Techniques for inferring kinship based on SNPs are also known as forensic SNP pedigree techniques. These typically utilize whole-genome SNP microarrays or whole-genome resequencing to obtain SNP genotyping data, followed by computational models to infer kinship. However, when faced with complex and diverse kinship levels, accurate differentiation is difficult, often leading to challenges in practical applications. Therefore, providing a combination of SNP loci with higher accuracy for inferring kinship levels has become an urgent need in this field. Summary of the Invention

[0004] This invention provides a combination of SNP loci for inferring the degree of human kinship and its application, thereby improving the accuracy of kinship degree inference.

[0005] In a first aspect, the present invention provides an application for detecting a substance for detecting combinations of SNP sites, the combinations of SNP sites comprising 20,838 SNP sites, the information of which is shown in Table 2.

[0006] The application is selected from at least one of A1)-A4):

[0007] A1) Application in inferring the degree of kinship;

[0008] A2) Application in the preparation of products for inferring kinship levels;

[0009] A3) Application in genetic analysis of kinship;

[0010] A4) Application in the preparation of products for genetic analysis of kinship.

[0011] As described above, the substance used to detect SNP site combinations is selected from at least one of primers, probes, or gene chips. In one specific embodiment, those skilled in the art can obtain the genomic sequence containing the target SNP site based on the SNP site information shown in Table 2, and design specific amplification primer pairs based on the genomic sequence; using the genomic DNA of the individual to be tested as a template, a PCR amplification reaction is performed using the specific amplification primer pairs, and the genotype of the above-mentioned SNP site is determined based on the amplification results; in another specific embodiment, a gene chip capable of detecting the genotype of the above-mentioned SNP site combination can also be used for detection, and the specific type of chip can be determined according to conventional techniques in the art.

[0012] As described above, the SNP locus combination is used to infer the kinship between the first and second individuals from an Asian population; further, it is used to infer the kinship between the first and second individuals from an East Asian population.

[0013] Secondly, the present invention provides a product for inferring the degree of kinship, including the aforementioned substance for detecting SNP locus combinations.

[0014] The product described above can be a kit.

[0015] Thirdly, the present invention provides a method for inferring the degree of kinship, comprising:

[0016] Obtain genomic DNA samples from the first and second individuals whose kinship level is to be inferred;

[0017] Genomic DNA samples from the first and second individuals were tested to obtain genotyping data of the SNP site combinations described above for the first and second individuals.

[0018] Based on the genotyping data of the SNP locus combinations of the first and second individuals, the genetic similarity coefficient GISC and the zero-shared genetic index GSI0 were calculated.

[0019] The kinship level of the first and second individuals was determined based on the genetic similarity coefficient GISC and the zero-shared genetic index GSI0.

[0020] The formula for calculating the genetic similarity coefficient GISC using the method described above is shown in Equation 1:

[0021]

[0022] In Equation 1, M Aa,Aa M represents the marker number indicating that both the first and second individuals are heterozygous. AA,aa The marker number indicates that both the first and second individuals are homozygous for their genotypes. The number of markers indicating that the first individual's genotype is heterozygous. The marker number indicates that the second individual's genotype is heterozygous.

[0023] The zero-shared genetic index GSI0 is calculated using the method described above, as shown in Equation 2:

[0024]

[0025] In Equation 2, M AA,aa The markers representing the number of homozygous genotypes in both the first and second individuals, m representing the SNP locus number, and p... m denoted as the allele frequency at the m-th SNP locus;

[0026] p m The calculation formula is shown in Equation 3:

[0027]

[0028] In Equation 3, M AA M represents the total number of individuals with genotype AA at the m-th SNP locus. Aa M represents the total number of individuals with genotype Aa at the m-th SNP locus. aa This represents the total number of individuals with genotype aa at the m-th SNP locus.

[0029] The method described above, determining the kinship level between the first and second individuals based on the genetic similarity coefficient GISC and the zero-shared genetic index GSI0, includes:

[0030] When the genetic similarity coefficient If so, it is inferred that the first individual and the second individual were twins;

[0031] when Genetic similarity coefficient Furthermore, if the zero-shared genetic index GSI0 < 0.001, it is inferred that the first individual and the second individual are parent-child.

[0032] when Genetic similarity coefficient If GSI0 > 0.001, then it is inferred that the first and second individuals are full siblings.

[0033] when Genetic similarity coefficient If so, it is inferred that the first individual and the second individual are second-degree related;

[0034] when Genetic similarity coefficient If so, it is inferred that the first individual and the second individual are related by a third degree;

[0035] when Genetic similarity coefficient If so, it is inferred that the first individual and the second individual are related at the fourth degree of kinship;

[0036] when Genetic similarity coefficient If so, it is inferred that the first individual and the second individual are related at the fifth degree of kinship;

[0037] when Genetic similarity coefficient If so, it is inferred that the first individual and the second individual are related at the sixth degree of kinship;

[0038] when Genetic similarity coefficient If so, it is inferred that the first individual and the second individual are related by a seventh degree of kinship;

[0039] When the genetic similarity coefficient GISC ≤ 0, it is inferred that the first individual and the second individual are not related.

[0040] Fourthly, the present invention provides an apparatus for inferring the degree of kinship, comprising:

[0041] Data acquisition module: used to acquire genotyping data of the first and second individuals based on the SNP locus combinations described above;

[0042] Data processing module: Calculates the genetic similarity coefficient GISC and the zero-shared genetic index GSI0 based on the genotyping data of the SNP locus combinations of the first and second individuals;

[0043] Data Judgment Module: Based on the genetic similarity coefficient GISC and zero-shared genetic index GSI0 calculated by the data processing module, the kinship level between the first individual and the second individual is inferred.

[0044] Fifthly, the present invention provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method for inferring kinship levels as described above.

[0045] In a sixth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for inferring kinship levels as described above.

[0046] In the above text, the degree of kinship is classified according to the conventional classification method in this field. Specifically, the kinship between parents and children, and siblings with the same father and mother, is classified as Level 1. The kinship between oneself and one's grandparents, maternal grandparents, nephews, grandchildren, uncles, aunts, and cousins ​​is classified as Level 2. The kinship between oneself and one's great-grandparents, maternal great-grandparents, uncles, maternal uncles, paternal aunts, maternal aunts, maternal uncles, maternal aunts, maternal uncles, maternal aunts, cousins, paternal cousins, and paternal cousins ​​is classified as Level 2. The kinship level between myself and my great-great-grandparents, maternal great-great-grandparents, great-uncles, paternal uncles, paternal aunts, paternal nephews, paternal nieces, and cousins ​​is level four. The kinship level between myself and my great-great-grandparents, maternal great-great-grandparents, great-uncles, paternal uncles, paternal aunts, paternal nephews, and cousins ​​is level five. For details on the kinship levels, please refer to the literature "Forensic SNP Genealogy Inference Technology Helps Solve a 14-Year-Old Cold Case", DOI: 10.16467 / j.1008-3650.2021.0028.

[0047] This invention provides a combination of SNP loci for predicting the degree of kinship. Based on this combination of SNP loci and combined with the genetic similarity coefficient GISC and the zero-shared genetic index GSI0, it is possible to predict kinship at the fourth degree (inclusive) in forensic genealogy. This system achieves a confidence interval accuracy of over 99.77% for first to third degree kinship with no false negatives, and a confidence interval accuracy of 95.51% for fourth degree kinship with a false negative rate of only 0.83%. Attached Figure Description

[0048] Figure 1 This is a distribution map of the 20,838 SNP loci screened in Example 1 of the present invention on autosomes;

[0049] Figure 2 This is a histogram of the molar distances between the 20,838 SNP loci selected in Example 1 of the present invention on autosomes.

[0050] Figure 3 This is a histogram showing the physical distances between the 20,838 SNP loci selected in Example 1 of the present invention on autosomes.

[0051] Figure 4 This is a histogram of the minimum allele frequency (MAF) of the 20,838 SNP loci screened in Example 1 of the present invention;

[0052] Figure 5 The distribution of genetic similarity coefficients (GISC) calculated for different kinship levels from 20,838 SNP loci screened according to Example 1 of the present invention;

[0053] Figure 6 The distribution of the zero-shared genetic index GSI0, calculated from the parent-child (PO) and full-sibling (FS) relationships of the 20,838 SNP loci screened according to Example 1 of the present invention. Detailed Implementation

[0054] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, embodiments of this invention, and should not be construed as limiting the invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention. In the description of this invention, it should be understood that the terminology used is for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0055] Unless otherwise specified, the experimental methods used in the following examples are conventional methods, performed according to the techniques or conditions described in the literature in this field or according to the product instructions. Unless otherwise specified, the materials and reagents used in the following examples are commercially available.

[0056] Example 1: Screening and Confirmation of SNP Locus Combinations

[0057] The entire human genome was analyzed, and a combination of SNP loci, comprising 20,838 SNP loci, was obtained through the following screening method:

[0058] 1.1 Preliminary Screening:

[0059] a) 553,049 SNPs were found at the intersection of the Infinium Global Screening Array (GSA) and the Infinium Chinese Genotyping Array (CGA). b) The intersection of these 553,049 SNPs with the Affymetrix GeneChip (Affy) was found, and sex chromosome SNPs were removed, leaving 69,496 SNPs. c) Overlapping sites with The Single Nucleotide Polymorphism Database (dbSNP151) were selected, leaving 69,490 SNPs. d) Dialleles were identified, and after deleting multiallele sites, 58,113 SNPs remained. e) Sites with other mutations at the same location were removed, leaving 58,107 SNPs. f) Intersection sites were found with the 1000 Genomes Database, leaving 57,030 SNPs. g. After deleting loci with a MAF (Minor Allele Frequency) of 0 and loci with a microarray genotyping detection rate of less than 99.9% in the East Asian population, 50,516 SNPs remain. h. Using a window size of 1000 and a step size of 1, remove r... 2 With SNPs having a value >0.2, there are 39,526 SNPs remaining.

[0060] 1.2 Detailed Screening:

[0061] a. The target number of loci is set at 15,000; b. Loci are allocated to each chromosome according to the proportion of chromosome centromere length; c. The centromere length of the chromosome is divided by the number of allocated loci to obtain the fragment centromere length; d. Loci with the highest MAF are selected from each fragment, resulting in 15,000 SNPs. The number of SNP loci allocated to each chromosome is shown in Table 1; e. The union of the pedigree SNPs included in the 9K locus combination disclosed in Chinese Invention Patent Application No. 202310586302.0 is taken, and r is removed. 2 SNPs with a value >0.2 were ultimately identified as 20,838 SNPs.

[0062] The final 20,838 SNP locus combinations are shown in Table 2, the number of SNP loci assigned to each chromosome is shown in Table 1, and their distribution on the chromosomes is shown in Table 2. Figure 1 As shown, Figure 1 The missing positions are mainly due to the fact that the chip design does not include the region. Possible reasons include: the region is the centromere region; the region has virtually no gene frequency information, etc.

[0063] Table 1. Distribution of SNP loci on various chromosomes.

[0064] Chromosome numbering Centimolar length of each chromosome percentage Combining the number of chromosome allocation sites before GISNP 9k Final number of allocation sites on each chromosome 1 292.7716 0.0809 1213 1673 2 274.2171 0.0757 1136 1609 3 227.8490 0.0629 944 1357 4 219.7976 0.0607 911 1275 5 208.9552 0.0577 866 1231 6 198.2419 0.0548 821 1184 7 190.3783 0.0526 789 1141 8 178.1477 0.0492 738 1020 9 180.2760 0.0498 747 969 10 182.4558 0.0504 756 1072 11 161.8505 0.0447 671 974 12 174.9621 0.0483 725 1015 13 129.5776 0.0358 537 775 14 116.7649 0.0323 484 688 15 150.7651 0.0416 625 713 16 131.1084 0.0362 543 755 17 128.5343 0.0355 533 742 18 120.0761 0.0332 497 668 19 106.8500 0.0295 443 605 20 110.2054 0.0304 457 628 21 63.7514 0.0176 264 357 22 72.9868 0.0202 302 387 total 3620.5229 1.0000 15000 20838

[0065] Table 2 Combinations of 20,838 SNP sites

[0066]

[0067]

[0068]

[0069]

[0070]

[0071]

[0072]

[0073]

[0074]

[0075]

[0076]

[0077]

[0078]

[0079]

[0080]

[0081]

[0082]

[0083]

[0084]

[0085]

[0086]

[0087]

[0088]

[0089]

[0090]

[0091]

[0092]

[0093]

[0094]

[0095]

[0096]

[0097]

[0098]

[0099]

[0100]

[0101]

[0102]

[0103]

[0104]

[0105]

[0106]

[0107]

[0108]

[0109]

[0110]

[0111]

[0112]

[0113]

[0114]

[0115]

[0116]

[0117]

[0118]

[0119]

[0120]

[0121]

[0122]

[0123]

[0124]

[0125]

[0126]

[0127]

[0128]

[0129]

[0130]

[0131]

[0132]

[0133]

[0134]

[0135]

[0136]

[0137]

[0138]

[0139]

[0140]

[0141]

[0142]

[0143]

[0144]

[0145]

[0146]

[0147]

[0148]

[0149]

[0150]

[0151]

[0152]

[0153]

[0154]

[0155]

[0156]

[0157]

[0158]

[0159]

[0160]

[0161]

[0162]

[0163]

[0164]

[0165]

[0166]

[0167]

[0168]

[0169]

[0170]

[0171]

[0172]

[0173]

[0174]

[0175]

[0176]

[0177]

[0178]

[0179]

[0180]

[0181]

[0182]

[0183] Where chr represents the chromosome number where the SNP site is located, pos represents the position of the SNP site on the chromosome, and id represents the identification number of the SNP site.

[0184] By plotting the selected 20,838 SNP loci on autosomes based on their molar and physical distances, and then creating histograms based on the statistical counts of pairwise molar and physical distances, as shown in the figure... Figure 2-3 As shown, the selected sites are basically evenly distributed in terms of centomotor distance, but also relatively dispersed in terms of physical distance.

[0185] The minimum allele frequency (MAF) was calculated for 20,838 SNP loci combinations, and a histogram was plotted. Figure 4 As shown, the MAF value of most points is close to 0.5, indicating that these points contain more frequency information and can help distinguish whether sample pairs are related.

[0186] Example 2: Inferring kinship level based on 20,838 SNP locus combinations and assessing accuracy.

[0187] 2.1 DNA Extraction and Detection of Actual Kinship Sample Sets

[0188] A total of 304 samples were collected from seven volunteer families in East Asia, including 244 parent-child (PO) pairs, 131 full sibling (FS) pairs, 333 second-degree (2nd) pairs, 439 third-degree (3rd) pairs, 602 fourth-degree (4th) pairs, 915 fifth-degree (5th) pairs, 976 sixth-degree (6th) pairs, 885 seventh-degree (7th) pairs, and 1000 unrelated pairs (UN). All participants signed informed consent forms, and the samples were approved by the Ethics Committee of the Forensic Science Center of the Ministry of Public Security (No.: 2022-017).

[0189] Saliva samples were collected from 304 samples, and DNA was extracted using the MagAttract M48 DNA Manual Kit (Qiagen, Hilden, Germany). DNA concentration was determined using the Qubit dsDNA HS Assay Kit (Invitrogen, Carlsbad, CA). SNPs were analyzed in 253 samples from this family using the Illumina Infinium Global Screening Array (GSA) chip, and in the remaining 51 samples using the Infinium Chinese Genotyping Array (CGA) chip. All samples were analyzed using Illumina... The Genotyping Module v2.0 software determines the genotype (Anlan Intelligent, Shenzhen, China), and then extracts the locus information of 20,838 selected SNPs for subsequent kinship inference tests on real family samples.

[0190] 2.2 Calculation of Kinship Level

[0191] First, genotypic data at each SNP locus between each pair of individuals are analyzed to calculate the number of SNPs detected in common between each pair and their allele sharing patterns, resulting in three key genetic sharing indices: zero-sharing index, single-allelic sharing index, and bialic sharing index. Then, based on the Hardy-Weinberg equilibrium law, the genetic similarity coefficient (GISC) is calculated by analyzing genotype and allele frequency distributions. In practical applications, due to factors such as samples potentially coming from populations with different genetic backgrounds or genotyping errors, GISC may be overestimated. Therefore, a lower heterozygosity rate between individuals is used to correct the GISC value to improve the accuracy of the inference. Finally, the calculated GISC and GSI0 values ​​are compared with preset standard ranges to infer the specific kinship level between individuals. GSI0 is specifically used to distinguish between parent-child and full-sibling relationships, as these two relationships have similar GISC values, but the GSI0 value for parent-child relationships is lower than that for full-sibling relationships.

[0192] The mean moment method refers to inferring the degree of kinship between individuals by calculating the Genetic Sharing Index (GSI) and Genetic Inheritance Similarity Coefficient (GISC).

[0193] Specifically, the Genetic Sharing Index (GSI) is a quantitative indicator of allele sharing between two individuals, comprising three variables: GSI0, GSI1, and GSI2. GSI0 represents the probability that two individuals share zero alleles from a common ancestor (zero-sharing index); GSI1 represents the probability that two individuals share one allele from a common ancestor (single-allelic sharing index); and GSI2 ​​represents the probability that two individuals share two alleles from a common ancestor (bialic sharing index). Let p be the frequency of a reference allele (labeled A) at a given SNP locus, and IBS... ij IBD represents the number of shared alleles between individuals i and j. ij This represents the number of alleles shared by individuals i and j from a common ancestor. Since only when IBD... ij IBS can only be enabled when = 0. ij =0. Therefore, according to the Hardy-Weinberg Equilibrium (HWE), the proportion of SNPs sharing zero common alleles between two individuals can be expressed as:

[0194] p(IBS ij =0)=p(AA,aa|IBD ij=0)p(IBD ij =0)=2p 2 (1-p) 2 GSI0

[0195] Therefore, the zero-shared genetic index GSI0 can be expressed as:

[0196]

[0197] in Indicates whether individuals i and j do not share any alleles at the m-th SNP marker, M AA,aa p represents the number of SNP markers for individuals i and j who are relatively homozygous for their genotypes. m Let be the allele frequency of the m-th SNP marker, which is estimated from the genotype frequencies of the entire sample:

[0198]

[0199] M AA M Aa M aa Let M represent the total number of individuals with genotypes AA, Aa, and aa on the m-th SNP marker, respectively. The other two shared genetic indices, GSI1 and GSI2, can be determined based on the number M of SNP markers with IBS=1 in two individuals. IBS=1 The number M of SNP markers with IBS=2 IBS=2 p m And GSI0 estimation, and the three shared genetic indices satisfy the following conditions:

[0200] GSI0 + GSI1 + GSI2 ​​= 1

[0201] The genetic similarity coefficient GISC is defined as the probability that a pair of alleles randomly sampled from two individuals at a SNP locus expresses a common ancestral origin, and it is related to the estimated shared genetic index mentioned above as follows:

[0202]

[0203] Assume that the reference allele frequencies of both individuals are p, and the number of reference alleles in individual i is X. (i) According to the HWE equilibrium law, the genetic distance between individuals i and j can be modeled as a function of their allele frequencies and kinship coefficients:

[0204] (X (i) -X (j) ) 2 =4p(1-p)(1-2GISC)

[0205] Therefore, the genetic similarity coefficient GISC can be expressed as:

[0206]

[0207] Where M Aa,Aa The number of markers indicating that individuals i and j are both heterozygous. This is the number of markers indicating that the genotype of individual x (where x represents i or j) is heterozygous. The genetic similarity coefficient GISC calculated in the above formula is based on the assumption that SNP loci satisfy Hardy-Weinberg equilibrium. However, in practical applications, due to genotyping errors and the inclusion of individuals from diverse genetic backgrounds in the sample, the genotype distribution of some individuals deviates from the expected HWE, thus the estimator in the above formula overestimates the genetic similarity coefficient. To prevent overestimation of the genetic similarity coefficient due to deviations from individual-level HWE equilibrium, a value with lower heterozygosity among individuals is used for estimation. Assuming that the heterozygosity of individual i is lower than that of individual j, the genetic similarity coefficient GISC can be expressed as:

[0208]

[0209] After calculating the genetic similarity coefficient (GISC) and the zero-shared genetic index (GSI0), the kinship level between individuals can be inferred by combining them with the inference criteria in Table 3. Specifically, the genetic similarity coefficient (GISC) predicted by the algorithm for all individuals is compared with the range of the inference criteria in this table to determine the kinship level between individuals. It should be noted that since the inference range of the genetic similarity coefficient (GISC) for parent-child and full-sibling relationships is the same, the genetic similarity coefficient (GISC) alone cannot distinguish between the two. Therefore, the zero-shared genetic index (GSI0) can be used in conjunction to distinguish them. When GSI0 ≤ 0.001, it is judged as parent-child (PO); when GSI0 > 0.001, it is judged as full-sibling (FS).

[0210] Table 3. Criteria for Inferring Kinship

[0211]

[0212] 2.3 Accuracy under the Assumption of Kinship Level in Known Scenarios

[0213] The distribution of kinship similarity coefficients (GISC) calculated based on 304 real family samples is as follows: Figure 5As shown, the genetic similarity coefficient distribution curves of related pairs, including third-degree related pairs, are completely separated from those of unrelated pairs. Starting from fourth-degree related pairs, the genetic similarity coefficient distribution curves of related pairs begin to overlap with those of unrelated pairs. Specifically, the genetic similarity coefficients for PO (proximal genital kinship) ranged from 0.2369 to 0.2547 (Median = 0.2490, IQR = 0.0035), for FS (proximal genital kinship) from 0.1957 to 0.2913 (Median = 0.2515, IQR = 0.0267), for second-degree relatives from 0.0470 to 0.1692 (Median = 0.1232, IQR = 0.0210), for third-degree relatives from 0.0218 to 0.0990 (Median = 0.0601, IQR = 0.0181), and for fourth-degree relatives from -0.0316 to 0.0601. The genetic similarity coefficients for fifth-degree relatives ranged from -0.0260 to 0.0423 (Median = 0.0113, IQR = 0.0117), sixth-degree relatives from -0.0149 to 0.0321 (Median = 0.0033, IQR = 0.0100), seventh-degree relatives from -0.0684 to 0.0235 (Median = -0.0009, IQR = 0.0086), and UN relatives from -0.0803 to 0.0173 (Median = -0.0031, IQR = 0.0088). The distribution of the zero-shared genetic index GSI0 between PO and FS is shown below. Figure 6 As shown. Specifically, the zero-shared genetic index GSI0 of PO is distributed between 0.0000 and 0.0003 (Median = 0.0000, IQR = 0.0001), while the zero-shared genetic index GSI0 of FS is distributed between 0.0122 and 0.0323 (Median = 0.0230, IQR = 0.0051), and there is no overlap between the two.

[0214] The predicted kinship levels were compared with the actual kinship levels found in the survey to evaluate the inference power of 20,838 SNP locus combinations for unknown kinship. Table 4 shows the predicted kinship levels and surveyed kinship levels for pairwise relationships of 304 individuals, and statistically analyzes the absolute accuracy (AC), confidence interval accuracy (CIA), false negatives (FN), and false positives (FP) used to evaluate the inference power of kinship.

[0215] Where, absolute accuracy (AC) = the number of pairs of predicted kinship results that are consistent with the surveyed kinship results in this level / the total number of surveyed kinship pairs in this level;

[0216] Confidence interval accuracy (CIA) = Number of relationships in this level whose predicted kinship outcome is ±1 level of the investigated kinship outcome / Number of investigated kinship relationships in this level;

[0217] False negative (FN) = Relationship pairs whose predicted kinship result is "irrelevant" at this level / All investigated kinship pairs at this level;

[0218] False positive (FP) = Relationship pairs that predict "related" kinship among those that indicate "unrelated" kinship / All relationship pairs that indicate "unrelated" kinship.

[0219] Table 4. Parameters for Kinship Inference Assessment

[0220]

[0221] As shown in Table 4, the accuracy of the confidence intervals, including those for third-degree kinship, is higher than 99.77%, and there are no false negatives. The accuracy of the confidence intervals for fourth-degree kinship is 95.51%, with a low proportion of false negatives (0.83%). Starting from fifth-degree kinship, the false negative rate increases with the increase of the kinship level, while the absolute accuracy and the accuracy of the confidence interval decrease.

[0222] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An application of a substance for detecting SNP site combinations, characterized in that, The SNP locus combination includes 20,838 SNP loci, and the information of these 20,838 SNP loci is as follows: Where chr represents the chromosome number where the SNP site is located, pos represents the position of the SNP site on the chromosome, and id represents the identification number of the SNP site. The application is selected from at least one of A1)-A2): A1) Application in inferring the degree of kinship; A2) Application in genetic analysis of kinship.

2. The application according to claim 1, characterized in that, The substance used to detect SNP site combinations is selected from at least one of primers, probes, or gene chips.

3. Products for inferring the degree of kinship, characterized in that, Includes the substance for detecting SNP site combinations as described in claim 1 or 2.

4. A method for inferring the degree of kinship, characterized in that, include: Obtain genomic DNA samples from the first and second individuals whose kinship level is to be inferred; The genomic DNA samples of the first individual and the second individual are tested to obtain the genotyping data of the SNP site combinations of the first individual and the second individual as described in claim 1 or 2. Based on the genotyping data of the SNP locus combinations of the first and second individuals, the genetic similarity coefficient GISC and the zero-shared genetic index GSI0 were calculated. The kinship level of the first and second individuals was determined based on the genetic similarity coefficient GISC and the zero-shared genetic index GSI0.

5. The method according to claim 4, characterized in that, The formula for calculating the genetic similarity coefficient GISC is shown in Equation 1: Formula 1 In Equation 1, The marker number indicates that both the first and second individuals are heterozygous. The marker number indicates that both the first and second individuals are homozygous for their genotypes. The number of markers indicating that the first individual's genotype is heterozygous. The marker number indicates that the second individual's genotype is heterozygous.

6. The method according to claim 4 or 5, characterized in that, The formula for calculating the zero-shared genetic index GSI0 is shown in Equation 2: Formula 2 In Equation 2, The marker number indicates that both the first and second individuals are homozygous for their genotypes, and m represents the SNP locus number. denoted as the allele frequency at the m-th SNP locus; The calculation formula is shown in Equation 3: Formula 3 In Equation 3, This represents the total number of individuals with genotype AA at the m-th SNP locus. This represents the total number of individuals with genotype Aa at the m-th SNP locus. This represents the total number of individuals with genotype aa at the m-th SNP locus.

7. The method according to claim 4 or 5, characterized in that, The kinship level between the first and second individuals is determined based on the genetic similarity coefficient GISC and the zero-shared genetic index GSI0, including: When the genetic similarity coefficient GISC > If so, it is inferred that the first individual and the second individual were twins; when <Genetic similarity coefficient GISC≤ If the zero-shared genetic index GSI0 ≤ 0.001, then it is inferred that the first individual and the second individual are parent-child. when <Genetic similarity coefficient GISC≤ If the zero-shared genetic index GSI0 > 0.001, then it is inferred that the first individual and the second individual are full siblings. when <Genetic similarity coefficient GISC≤ If so, it is inferred that the first individual and the second individual are second-degree related; when <Genetic similarity coefficient GISC≤ If so, it is inferred that the first individual and the second individual are related by a third degree; when <Genetic similarity coefficient GISC≤ If so, it is inferred that the first individual and the second individual are related at the fourth degree of kinship; when <Genetic similarity coefficient GISC≤ If so, it is inferred that the first individual and the second individual are related at the fifth degree of kinship; when <Genetic similarity coefficient GISC≤ If so, it is inferred that the first individual and the second individual are related at the sixth degree of kinship; when <Genetic similarity coefficient GISC≤ If so, it is inferred that the first individual and the second individual are related by a seventh degree of kinship; When the genetic similarity coefficient GISC ≤ 0, it is inferred that the first individual and the second individual are not related.

8. A device for inferring the degree of kinship, characterized in that, include: Data acquisition module: used to acquire genotyping data of the SNP locus combinations as described in claim 1 or 2 for the first and second individuals; Data processing module: Calculates the genetic similarity coefficient GISC and the zero-shared genetic index GSI0 based on the genotyping data of the SNP locus combinations of the first and second individuals; Data Judgment Module: Based on the genetic similarity coefficient GISC and zero-shared genetic index GSI0 calculated by the data processing module, the kinship level between the first individual and the second individual is inferred.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method for inferring kinship levels as described in any one of claims 4-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method for inferring kinship levels as described in any one of claims 4-7.

Citation Information

Patent Citations

  • SNP locus combination and its application for estimating the degree of human kinship

    CN117524308B

  • Genetic relationship identification method with SNP as genetic marker

    CN111091869A

  • Method for judging genetic relationship through SNP (Single Nucleotide Polymorphism) mismatch rate

    CN115572770A