20838 SNP site combination for speculating genetic relationship grade of human and application thereof
By using a combination of 20,838 SNP loci and a computational model, combined with the genetic similarity coefficient GISC and the zero-shared genetic index GSI0, the problem of inaccurate kinship prediction in existing technologies has been solved, and highly accurate kinship level prediction has been achieved.
Patent Information
- Application Number
- CN202510868584.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-06-26
Smart Images

Figure CN120913638A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of bioinformatics, and particularly relates to a 20838 SNP site combination for inferring human kinship grades and application thereof. BACKGROUND
[0002] Kinship refers to the relationship formed through biological descent, including parent-child relationship, sibling relationship, etc. Kinship grade refers to the quantitative classification of the relationship between individuals according to the similarity of their genes or the distance of their blood relationship. In the field of forensic identification, studying and inferring kinship grades is of great significance, and has a wide and far-reaching impact, providing core support for the solution of many complex cases. In the process of solving criminal cases, kinship identification can be a powerful tool for locking in suspects. When biological evidence (such as bloodstains, hair, and skin flakes) left at the scene is related to a known suspect's relative, by analyzing the kinship between the two, and using genetic clues, the scope of the investigation can be greatly reduced, improving the efficiency of case solving. In some crime scenes, if it is not possible to directly obtain the biological sample of the suspect, but by analyzing the kinship of the people related to the case and their relatives, it is possible to discover potential criminal clues and thus provide a key path for the breakthrough of the case. In dealing with missing person cases and unidentified corpse identity confirmation work, kinship identification is a core means of accurate identity recognition. By identifying the kinship between the DNA sample of the missing person's relatives and the biological sample of the suspected missing person, the connection between the two can be accurately established, and the identity information can be determined.
[0003] Single nucleotide polymorphism (SNP) is the third generation of genetic markers in the field of forensic science, with the characteristics of wide distribution, low mutation rate, and high genetic stability, and can be used for individual identification and kinship analysis, and is an important genetic marker for individual identification. The technology of inferring kinship based on SNP is also known as forensic SNP pedigree technology, which usually uses whole genome SNP chips or whole genome resequencing to obtain SNP typing data information, and then infers the kinship through a calculation model. However, when faced with complex and diverse kinship grades, it is difficult to accurately distinguish them, which often leads to difficulties in practical applications. Therefore, how to provide a SNP site combination with higher accuracy for inferring human kinship grades has become an urgent need in the field. SUMMARY
[0004] The present application provides a SNP site combination for inferring human kinship grades and application thereof, which can improve the accuracy of kinship grade inference.
[0005] In a first aspect, the present application provides an application of a substance for detecting a SNP site combination, wherein the SNP site combination comprises 20838 SNP sites, and the information of the 20838 SNP sites is shown in Table 2.
[0006] The application is at least one selected from A1) to A4):
[0007] A1) an application in inferring a degree of kinship;
[0008] A2) an application in preparing a product for inferring a degree of kinship;
[0009] A3) an application in genetic analysis of kinship;
[0010] A4) an application in preparing a product for genetic analysis of kinship.
[0011] The substance for detecting the SNP site combination in the above-mentioned application is at least one selected from a primer, a probe or a gene chip. In a specific embodiment, a person skilled in the art can obtain a genomic sequence containing a target SNP site according to the information of the SNP site shown in Table 2, and design a specific amplification primer pair according to the genomic sequence; and then perform a PCR amplification reaction using the specific amplification primer pair and the genomic DNA of a subject to be detected as a template, so as to determine the genotype of the above-mentioned SNP site according to the amplification result; in another specific embodiment, a gene chip capable of detecting the genotype of the above-mentioned SNP site combination can also be used for detection, and the specific chip type can be determined according to the conventional technical means in the art.
[0012] The application in the above-mentioned application, wherein the SNP site combination is used for inferring the kinship between a first individual and a second individual derived from an Asian population; and further, the kinship between a first individual and a second individual derived from East Asians.
[0013] In a second aspect, the present application provides a product for inferring a degree of kinship, comprising the above-mentioned substance for detecting a SNP site combination.
[0014] The product in the above-mentioned application can be a kit.
[0015] In a third aspect, the present application provides a method for inferring a degree of kinship, comprising:
[0016] obtaining genomic DNA samples of a first individual and a second individual to be inferred for a degree of kinship;
[0017] detecting the genomic DNA samples of the first individual and the second individual to obtain the typing data of the SNP site combination of the first individual and the second individual as described above;
[0018] Based on the genotyping data of the SNP locus combinations of the first and second individuals, the genetic similarity coefficient GISC and the zero-shared genetic index GSI0 were calculated.
[0019] The kinship level of the first and second individuals was determined based on the genetic similarity coefficient GISC and the zero-shared genetic index GSI0.
[0020] The formula for calculating the genetic similarity coefficient GISC using the method described above is shown in Equation 1:
[0021]
[0022] In Equation 1, M Aa,Aa M represents the marker number indicating that both the first and second individuals are heterozygous. AA,aa The marker number indicates that both the first and second individuals are homozygous for their genotypes. The number of markers indicating that the first individual's genotype is heterozygous. The marker number indicates that the second individual's genotype is heterozygous.
[0023] The zero-shared genetic index GSI0 is calculated using the method described above, as shown in Equation 2:
[0024]
[0025] In Equation 2, M AA,aa The markers representing the number of homozygous genotypes in both the first and second individuals, m representing the SNP locus number, and p... m denoted as the allele frequency at the m-th SNP locus;
[0026] p m The calculation formula is shown in Equation 3:
[0027]
[0028] In Equation 3, M AA M represents the total number of individuals with genotype AA at the m-th SNP locus. Aa M represents the total number of individuals with genotype Aa at the m-th SNP locus. aa This represents the total number of individuals with genotype aa at the m-th SNP locus.
[0029] The method described above, determining the kinship level between the first and second individuals based on the genetic similarity coefficient GISC and the zero-shared genetic index GSI0, includes:
[0030] When the genetic similarity coefficient If so, it is inferred that the first individual and the second individual were twins;
[0031] When Genetic similarity coefficient and zero-sharing genetic index GSI0<0.001, then the first individual and the second individual are presumed to be parent-offspring relationship;
[0032] When Genetic similarity coefficient and GSI0>0.001, then the first individual and the second individual are presumed to be full siblings;
[0033] When Genetic similarity coefficient , then the first individual and the second individual are presumed to be second-degree relatives;
[0034] When Genetic similarity coefficient , then the first individual and the second individual are presumed to be third-degree relatives;
[0035] When Genetic similarity coefficient , then the first individual and the second individual are presumed to be fourth-degree relatives;
[0036] When Genetic similarity coefficient , then the first individual and the second individual are presumed to be fifth-degree relatives;
[0037] When Genetic similarity coefficient , then the first individual and the second individual are presumed to be sixth-degree relatives;
[0038] When Genetic similarity coefficient , then the first individual and the second individual are presumed to be seventh-degree relatives;
[0039] When the genetic similarity coefficient GISC≤0, then the first individual and the second individual are presumed to have no genetic relationship.
[0040] In a fourth aspect, the present application provides a device for inferring the degree of genetic relationship, comprising:
[0041] A data acquisition module for acquiring the typing data of the SNP site combination of the first individual and the second individual as described above;
[0042] A data processing module for calculating the genetic similarity coefficient GISC and the zero-sharing genetic index GSI0 from the typing data of the SNP site combination of the first individual and the second individual;
[0043] Data Judgment Module: Based on the genetic similarity coefficient GISC and zero-shared genetic index GSI0 calculated by the data processing module, the kinship level between the first individual and the second individual is inferred.
[0044] Fifthly, the present invention provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method for inferring kinship levels as described above.
[0045] In a sixth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for inferring kinship levels as described above.
[0046] In the above text, the degree of kinship is classified according to the conventional classification method in this field. Specifically, the kinship between parents and children, and siblings with the same father and mother, is classified as Level 1. The kinship between oneself and one's grandparents, maternal grandparents, nephews, grandchildren, uncles, aunts, and cousins is classified as Level 2. The kinship between oneself and one's great-grandparents, maternal great-grandparents, uncles, maternal uncles, paternal aunts, maternal aunts, maternal uncles, maternal aunts, maternal uncles, maternal aunts, cousins, paternal cousins, and paternal cousins is classified as Level 2. The kinship level between myself and my great-great-grandparents, maternal great-great-grandparents, great-uncles, paternal uncles, paternal aunts, paternal nephews, paternal nieces, and cousins is level four. The kinship level between myself and my great-great-grandparents, maternal great-great-grandparents, great-uncles, paternal uncles, paternal aunts, paternal nephews, and cousins is level five. For details on the kinship levels, please refer to the literature "Forensic SNP Genealogy Inference Technology Helps Solve a 14-Year-Old Cold Case", DOI: 10.16467 / j.1008-3650.2021.0028.
[0047] This invention provides a combination of SNP loci for predicting the degree of kinship. Based on this combination of SNP loci and combined with the genetic similarity coefficient GISC and the zero-shared genetic index GSI0, it is possible to predict kinship at the fourth degree (inclusive) in forensic genealogy. This system achieves a confidence interval accuracy of over 99.77% for first to third degree kinship with no false negatives, and a confidence interval accuracy of 95.51% for fourth degree kinship with a false negative rate of only 0.83%. Attached Figure Description
[0048] Figure 1 This is a distribution map of the 20,838 SNP loci screened in Example 1 of the present invention on autosomes;
[0049] Figure 2 Histogram of the molar distance between loci on autosomes for the 20838 SNP loci screened in Example 1 of the present application;
[0050] Figure 3 Histogram of the physical distance between loci on autosomes for the 20838 SNP loci screened in Example 1 of the present application;
[0051] Figure 4 Histogram of the minimum allele frequency (MAF) for the 20838 SNP loci screened in Example 1 of the present application;
[0052] Figure 5 Distribution of the genetic identity similarity coefficient GISC calculated for different levels of genetic relationship for the 20838 SNP loci screened in Example 1 of the present application;
[0053] Figure 6 Distribution of the zero-sharing genetic index GSI0 calculated for parent-offspring (PO) and full-sibling (FS) relationships for the 20838 SNP loci screened in Example 1 of the present application. DETAILED DESCRIPTION
[0054] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below with reference to the drawings in the present application. Obviously, the described embodiments are some but not all embodiments of the present application, and they should not be interpreted as limiting the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work should fall within the scope of the present application. In the description of the present application, it should be understood that the terms used are only for the purpose of describing and should not be interpreted as indicating or implying the relative importance.
[0055] In the following examples, the experimental methods are all conventional methods, and are performed according to the techniques or conditions described in the literature in the art or according to the product instructions, unless otherwise specified. The materials, reagents and the like used in the following examples can be obtained from commercial channels, unless otherwise specified.
[0056] Example 1, Screening and Confirmation of SNP Locus Combination
[0057] The human whole genome is detected, and a SNP locus combination including 20838 SNP loci is screened by the following screening method. The screening method includes:
[0058] 1.1, Preliminary screening:
[0059] a. Take the intersection of Infinium Global Screening Array (GSA) and Infinium Chinese Genotyping Array (CGA) sites, a total of 553049 SNPs. b. Take the intersection of the 553049 SNPs in step a with Affymetrix GeneChip (Affy) and remove sex chromosome SNPs, leaving 69496 SNPs. c. Select sites that coincide with The Single Nucleotide Polymorphism Database (dbSNP151), leaving 69490 SNPs. d. Select biallelic sites, and after removing sites with multiple alleles, there are 58113 SNPs left. e. Remove sites with other mutations at the same location, leaving 58107 SNPs. f. Take the intersection of the sites with 1000 Genomes data, leaving 57030 SNPs. g. Remove sites with a MAF (Minor Allele Frequency) of 0 in the East Asian population and sites with a chip typing detection rate of less than 99.9%, leaving 50516 SNPs. h. Take out SNPs with r 2 > 0.2 with a window of 1000 and a step of 1, leaving 39526 SNPs.
[0060] 1.2. Fine screening:
[0061] a. Set the target number of sites to 15000; b. Distribute the number of sites to each chromosome according to the centiMorgan length of the chromosome; c. Divide the centiMorgan length of the chromosome by the number of sites allocated to obtain the segment centiMorgan length; d. Select the site with the highest MAF in each segment, and finally obtain 15000 SNPs, the number of SNP sites allocated to each chromosome is shown in Table 1; e. Take the union of the pedigree SNPs contained in the 9K site combination disclosed in the Chinese invention patent with application number 202310586302.0, and remove SNPs with r 2 > 0.2, finally obtaining 20838 SNPs.
[0062] The final combination of 20838 SNP sites is shown in Table 2, the number of SNP sites allocated to each chromosome is shown in Table 1, and the distribution on the chromosome is shown in Figure 1 . Figure 1 The positions shown as blanks are mainly due to the fact that the chip design does not contain sites in this region, and the possible reasons are: the region is a centromere region; the region has little gene frequency information.
[0063] Table 1 Distribution of SNP sites on each chromosome
[0064] Chromosome number Chromosome centimorgan length Percentage Number of allocated sites per chromosome before incorporation of GISNP 9k Final number of allocated sites per chromosome 1 292.7716 0.0809 1213 1673 2 274.2171 0.0757 1136 1609 3 227.8490 0.0629 944 1357 4 219.7976 0.0607 911 1275 5 208.9552 0.0577 866 1231 6 198.2419 0.0548 821 1184 7 190.3783 0.0526 789 1141 8 178.1477 0.0492 738 1020 9 180.2760 0.0498 747 969 10 182.4558 0.0504 756 1072 11 161.8505 0.0447 671 974 12 174.9621 0.0483 725 1015 13 129.5776 0.0358 537 775 14 116.7649 0.0323 484 688 15 150.7651 0.0416 625 713 16 131.1084 0.0362 543 755 17 128.5343 0.0355 533 742 18 120.0761 0.0332 497 668 19 106.8500 0.0295 443 605 20 110.2054 0.0304 457 628 21 63.7514 0.0176 264 357 22 72.9868 0.0202 302 387 Total 3620.5229 1.0000 15000 20838
[0065] Table 2 Combinations of 20838 SNP loci
[0066]
[0067]
[0068]
[0069]
[0070]
[0071]
[0072]
[0073]
[0074]
[0075]
[0076]
[0077]
[0078]
[0079]
[0080]
[0081]
[0082]
[0083]
[0084]
[0085]
[0086]
[0087]
[0088]
[0089]
[0090]
[0091]
[0092]
[0093]
[0094]
[0095]
[0096]
[0097]
[0098]
[0099]
[0100]
[0101]
[0102]
[0103]
[0104]
[0105]
[0106]
[0107]
[0108]
[0109]
[0110]
[0111]
[0112]
[0113]
[0114]
[0115]
[0116]
[0117]
[0118]
[0119]
[0120]
[0121]
[0122]
[0123]
[0124]
[0125]
[0126]
[0127]
[0128]
[0129]
[0130]
[0131]
[0132]
[0133]
[0134]
[0135]
[0136]
[0137]
[0138]
[0139]
[0140]
[0141]
[0142]
[0143]
[0144]
[0145]
[0146]
[0147]
[0148]
[0149]
[0150]
[0151]
[0152]
[0153]
[0154]
[0155]
[0156]
[0157]
[0158]
[0159]
[0160]
[0161]
[0162]
[0163]
[0164]
[0165]
[0166]
[0167]
[0168]
[0169]
[0170]
[0171]
[0172]
[0173]
[0174]
[0175]
[0176]
[0177]
[0178]
[0179]
[0180]
[0181]
[0182]
[0183] wherein, chr represents the number of the chromosome where the SNP site locates, pos represents the position of the chromosome where the SNP site locates, and id represents the identification number of the SNP site.
[0184] By mapping the selected 20838 SNP sites on autosomes according to their centiMorgan distance and physical distance, a histogram was drawn according to the number of Morgan distance and physical distance between each other, as shown in Figures 2-3 The selected sites are basically evenly distributed in centiMorgan distance, and are also relatively dispersed in physical distance.
[0185] The minimum allele frequency (MAF) of the combination of 20838 SNP sites was calculated and a histogram was drawn, as shown in Figure 4 The MAF value of most points is close to 0.5, indicating that these points contain more frequency information and can help distinguish whether the sample pairs have a kinship.
[0186] Example 2, Estimation of kinship levels based on 20838 SNP loci combinations and accuracy assessment
[0187] 2.1 DNA extraction and detection of the actual kinship sample set
[0188] Seven families of volunteers from East Asia, a total of 304 samples, including 244 pairs of parent-offspring (PO), 131 pairs of full siblings (FS), 333 pairs of 2nd (2nd), 439 pairs of 3rd (3rd), 602 pairs of 4th (4th), 915 pairs of 5th (5th), 976 pairs of 6th (6th), 885 pairs of 7th (7th) kinship, and 1000 pairs of unrelated pairs (UN). All participants signed the informed consent form and passed the approval of the Ethical Committee of the Ministry of Public Security Forensic Center (No: 2022-017).
[0189] Saliva of 304 samples was collected, and DNA was extracted using MagAttract M48 DNA Manual Kit (Qiagen, Hilden, Germany), and the concentration of DNA was detected using Qubit dsDNA HS Assay Kit (Invitrogen, Carlsbad, CA). Among them, 253 samples in the family were detected by SNP using Illumina Infinium Global Screening Array (GSA) chip, and the remaining 51 samples were detected by SNP through Infinium Chinese Genotyping Array (CGA) chip. All samples were genotyped by Illumina Genotyping Module v2.0 software (Anlan Intelligence, Shenzhen, China), and then the site information of 20838 SNPs screened out was extracted for subsequent real family sample kinship inference test.
[0190] 2.2 Kinship level calculation
[0191] Firstly, genotype data at each SNP site between two individuals are analyzed to calculate the number of SNPs and their allele sharing patterns that are commonly detected between each pair of individuals, resulting in three key genetic sharing indices: zero sharing index, single allele sharing index, and double allele sharing index. Then, based on the Hardy-Weinberg equilibrium law, by analyzing the genotype frequency and allele frequency distribution, the genetic inheritance similarity coefficient GISC is calculated. In practical applications, due to factors such as samples from different genetic backgrounds or genotyping errors, GISC may be overestimated, so the value of GISC is corrected by a lower heterozygosity rate between individuals to improve the accuracy of inference. Finally, the calculated GISC and GSI0 values are compared with the preset standard range, and the specific degree of kinship between individuals can be inferred, where GSI0 is particularly used to distinguish parent-child relationships and full sibling relationships, because these two relationships have similar GISC values, but the GSI0 value of the parent-child relationship is lower than that of the full sibling relationship.
[0192] The average moment method refers to inferring the degree of kinship between individuals by calculating the genetic sharing index (Genetic Sharing Index, GSI) and the genetic inheritance similarity coefficient (Genetic Inheritance Similarity Coefficient, GISC) between two individuals.
[0193] Specifically, the genetic sharing index GSI is a quantitative indicator of the allele sharing state between two individuals, including GSI0, GSI1, and GSI2. Among them, GSI0 represents the probability of two individuals sharing zero common ancestor alleles (zero sharing genetic index); GSI1 represents the probability of two individuals sharing one common ancestor allele (single allele sharing genetic index); GSI2 represents the probability of two individuals sharing two common ancestor alleles (double allele sharing genetic index). Assuming p is the frequency of the reference allele (labeled A) at a certain SNP site, IBS ij represents the number of common alleles between individuals i, j, IBD ij represents the number of common ancestor common alleles between individuals i, j. Since IBS ij = 0 only when IBD ij = 0. Therefore, according to the Hardy-Weinberg equilibrium law (Hardy-Weinberg Equilibrium, HWE), the proportion of SNPs that share zero common alleles between two individuals can be represented as:
[0194] p(IBS ij = 0) = p(AA, aa | IBD ij= 0) = 2p ij = 0) = 2p 2 (1 - p) 2 GSI0
[0195] Therefore, the zero shared genetic index GSI0 can be expressed as:
[0196]
[0197] where denotes whether the i, j individuals do not share any alleles at the mth SNP marker, M AA,aa is the number of SNP markers for which the genotypes of individuals i, j are homozygous, p m is the allele frequency of the mth SNP marker, which is estimated from the genotype frequencies of the entire sample:
[0198]
[0199] M AA , M Aa , M aa denote the total number of individuals with genotypes AA, Aa, and aa at the mth SNP marker, respectively. The other two shared genetic indices GSI1 and GSI2 can be estimated according to the number of SNP markers with IBS = 1 M IBS=1 , the number of SNP markers with IBS = 2 M IBS=2 , p m and GSI0, and the three shared genetic indices satisfy the following conditions:
[0200] GSI0 + GSI1 + GSI2 = 1
[0201] The genetic similarity coefficient GISC is defined as the probability that a pair of alleles randomly sampled from two individuals at SNP sites exhibit the same ancestral origin, and has the following relationship with the estimated shared genetic indices:
[0202]
[0203] Assuming that the reference allele frequencies of the two individuals are both p, and the number of reference alleles of individual i is X (i) , according to the HWE balance law, the genetic distance between individuals i, j can be modeled as a function of their allele frequencies and the genetic similarity coefficient:
[0204] (X (i) -X (j) ) 2 = 4p(1 - p)(1 - 2GISC)
[0205] Therefore, the genetic similarity coefficient GISC can be expressed as:
[0206]
[0207] where M Aa,Aa is the number of markers with heterozygous genotype for individual i, j, is the number of markers with heterozygous genotype for individual x (x represents i or j). The GISC calculation in the above formula is based on the assumption that the SNP loci satisfy Hardy-Weinberg equilibrium. However, in practical applications, due to genotyping errors and the inclusion of people from different genetic backgrounds in the sample, the genotype distribution of some individuals deviates from the expected HWE, so the above formula estimate will overestimate the genetic similarity coefficient. In order to prevent the overestimation of the genetic similarity coefficient due to the deviation of the HWE balance at the individual level, the smaller value of the heterozygosity between individuals is used for estimation. Assuming that the heterozygosity of individual i is lower than that of individual j, the genetic similarity coefficient GISC can be expressed as:
[0208]
[0209] After calculating the genetic similarity coefficient GISC and the zero sharing genetic index GSI0, the kinship level between individuals can be inferred by combining the inference criteria in Table 3. Specifically, according to the genetic similarity coefficient GISC between all individuals predicted by the algorithm and the range of inference criteria in this table, the kinship level between individuals can be determined. It should be noted that since the genetic similarity coefficient GISC of parent-child and full siblings is consistent in the inference range, only the genetic similarity coefficient GISC cannot be used to distinguish between the two, so the zero sharing genetic index GSI0 can be used for distinction. When GSI0≤0.001, it is judged to be a parent-child (PO); when GSI0>0.001, it is judged to be a full sibling (FS).
[0210] Table 3, kinship inference criteria
[0211]
[0212] 2.3, accuracy of kinship level under the assumption of known scenarios
[0213] The distribution of the genetic similarity coefficient GISC calculated based on the 304 real sample families is as follows: Figure 5The distribution of the kinship genetic similarity coefficient of unrelated pairs of individuals was completely separated from the distribution of the kinship genetic similarity coefficient of pairs of individuals with the first three degrees of kinship, and the distribution of the kinship genetic similarity coefficient of unrelated pairs of individuals began to overlap from the fourth degree of kinship. Specifically, the distribution of the PO kinship genetic similarity coefficient was between 0.2369-0.2547 (Median=0.2490, IQR=0.0035), the distribution of the FS kinship genetic similarity coefficient was between 0.1957-0.2913 (Median=0.2515, IQR=0.0267), the distribution of the second degree of kinship genetic similarity coefficient was between 0.0470-0.1692 (Median=0.1232, IQR=0.0210), the distribution of the third degree of kinship genetic similarity coefficient was between 0.0218-0.0990 (Median=0.0601, IQR=0.0181), the distribution of the fourth degree of kinship genetic similarity coefficient was between -0.0316-0.0601 (Median=0.0286, IQR=0.0130), the distribution of the fifth degree of kinship genetic similarity coefficient was between -0.0260-0.0423 (Median=0.0113, IQR=0.0117), the distribution of the sixth degree of kinship genetic similarity coefficient was between -0.0149-0.0321 (Median=0.0033, IQR=0.0100), the distribution of the seventh degree of kinship genetic similarity coefficient was between -0.0684-0.0235 (Median=-0.0009, IQR=0.0086), and the distribution of the UN kinship genetic similarity coefficient was between -0.0803-0.0173 (Median=-0.0031, IQR=0.0088). The distribution of the zero-sharing genetic index GSI0 of PO and FS is shown in FIG. 4. Specifically, the distribution of the zero-sharing genetic index GSI0 of PO was between 0.0000-0.0003 (Median=0.0000, IQR=0.0001), and the distribution of the zero-sharing genetic index GSI0 of FS was between 0.0122-0.0323 (Median=0.0230, IQR=0.0051), and there was no intersection between the two. Figure 6
[0214] The predicted degree of kinship was compared with the actually investigated degree of kinship, so as to evaluate the inference efficiency of the 20838 SNP site combinations for unknown kinship. The predicted degree of kinship of the pairwise relationship pairs of 304 individuals and the investigated degree of kinship are shown in Table 4, and the absolute accuracy (Accuracy, AC), the confidence interval accuracy (Confidence interval accuracy, CIA), and the false negative (False negative, FN) and the false positive (False positive, FP) are counted to evaluate the index of the inference efficiency of the kinship. Figure 6
[0214] The predicted degree of kinship was compared with the actually investigated degree of kinship, so as to evaluate the inference efficiency of the 20838 SNP site combinations for unknown kinship. The predicted degree of kinship of the pairwise relationship pairs of 304 individuals and the investigated degree of kinship are shown in Table 4, and the absolute accuracy (Accuracy, AC), the confidence interval accuracy (Confidence interval accuracy, CIA), and the false negative (False negative, FN) and the false positive (False positive, FP) are counted to evaluate the index of the inference efficiency of the kinship.
[0215] Accuracy (AC) = the number of pairs of relationships in which the predicted kinship result is consistent with the investigated kinship result / the total number of pairs of investigated kinship relationships in this grade;
[0216] Confidence interval accuracy (CIA) = the number of pairs of relationships in which the predicted kinship result is within ±1 grade of the investigated kinship result / the total number of pairs of investigated kinship relationships in this grade;
[0217] False negative (FN) = the number of pairs of relationships in which the predicted kinship result is "unrelated" in this grade / the total number of pairs of investigated kinship relationships in this grade;
[0218] False positive (FP) = the number of pairs of relationships in which the predicted kinship result is "related" in the pairs of investigated kinship relationships that are "unrelated" / the total number of pairs of investigated kinship relationships that are "unrelated".
[0219] Table 4: Evaluation parameters of kinship inference
[0220]
[0221] According to Table 4, the confidence interval accuracy of the third grade of kinship and above is higher than 99.77%, and there is no false negative; the confidence interval accuracy of the fourth grade of kinship is 95.51%, and there is a lower proportion of false negative (0.83%); from the fifth grade of kinship, with the increase of the grade of kinship relationship, the false negative rate increases, and the absolute accuracy and the confidence interval accuracy decrease.
[0222] Finally, it should be noted that: the above examples are only used to illustrate the technical solutions of the present application, but not to limit it; although the present application has been described in detail with reference to the foregoing examples, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing examples, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. Use of a substance for detecting a combination of SNP loci, characterized in that, The SNP site combination comprises 20838 SNP sites, and information of the 20838 SNP sites is specifically as follows: Wherein, chr represents the number of the chromosome where the SNP site is located, pos represents the position of the chromosome where the SNP site is located, and id represents the identification number of the SNP site; The application is selected from at least one of A1) to A4): A1) application in the presumed kinship rank; A2) application in the preparation of a product of the presumed kinship rank; A3) application in the genetic analysis of kinship; A4) application in the preparation of a product of the genetic analysis of kinship.
2. Use according to claim 1, characterized in that, The substance for detecting the SNP site combination is selected from at least one of a primer, a probe or a gene chip.
3. A product for inferring a degree of kinship, characterized by The substance for detecting the SNP site combination comprises claim 1 or 2.
4. A method of inferring a rank of relatedness, characterized by It comprises: Obtaining the genomic DNA samples of the first individual and the second individual whose kinship rank is to be inferred; Detecting the genomic DNA samples of the first individual and the second individual to obtain the typing data of the SNP site combination of claim 1 or 2 of the first individual and the second individual; According to the typing data of the SNP site combination of the first individual and the second individual, calculating the genetic similarity coefficient GISC and the zero-sharing genetic index GSI0; and determining the kinship rank of the first individual and the second individual according to the genetic similarity coefficient GISC and the zero-sharing genetic index GSI0.
5. The method of claim 4, wherein, The calculation formula of the genetic similarity coefficient GISC is shown in formula 1: In Formula 1, M Aa,Aa represents the number of markers for which both the first individual and the second individual are heterozygous, AA,aa represents the number of markers for which both the first individual and the second individual are homozygous, represents the number of markers for which the first individual is heterozygous, represents the number of markers for which the second individual is heterozygous.
6. The method according to claim 4 or 5, characterized in that, The calculation formula of the zero-sharing genetic index GSI0 is shown in formula 2: In formula 2, M AA,aa represents the number of markers for which the first individual and the second individual are both homozygous, m represents the number of the SNP site, p m is the allele frequency of the mth SNP site; p m The calculation formula is shown in Equation 3: In formula 3, M AA represents the total number of individuals with genotype AA at the mth SNP site, M Aa represents the total number of individuals with genotype Aa at the mth SNP site, M aa represents the total number of individuals with genotype aa at the mth SNP site.
7. The method according to any one of claims 4-6, characterized in that, According to the genetic similarity coefficient GISC and the zero-sharing genetic index GSI0, determining the kinship rank of the first individual and the second individual comprises: When the coefficient of genetic similarity is greater than 0.35, then the first and second individuals are presumed to be twins. When and zero shared genetic index GSI0≤ 0.001, then the first individual and the second individual are presumed to be in a parent-offspring relationship. When and zero shared genetic index GSI0> 0.001, then the first and second individuals are inferred to be full-sibs; when then the first and second individuals are inferred to be second-degree relatives. When then the first individual and the second individual are presumed to be third degree relatives. When then the first individual and the second individual are presumed to be in a fourth degree of consanguinity. When then the first individual and the second individual are presumed to be in a fifth degree of consanguinity. When then the first individual and the second individual are presumed to be in a sixth degree of consanguinity. When then the first individual and the second individual are presumed to be in a seventh degree of consanguinity. When the genetic similarity coefficient GISC≤0, it is inferred that the first individual and the second individual have no kinship.
8. An apparatus for inferring a rank of kinship, characterized by It comprises: A data acquisition module for acquiring the typing data of the SNP site combination of claim 1 or 2 of the first individual and the second individual; A data processing module for calculating the genetic similarity coefficient GISC and the zero-sharing genetic index GSI0 according to the typing data of the SNP site combination of the first individual and the second individual; A data judgment module for inferring the kinship rank of the first individual and the second individual according to the genetic similarity coefficient GISC and the zero-sharing genetic index GSI0 calculated by the data processing module.
9. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the program to realize the method for inferring the kinship rank according to any one of claims 4 to 7.
10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to realize the method for inferring the kinship rank according to any one of claims 4 to 7.
Citation Information
Patent Citations
SNP locus combination and its application for estimating the degree of human kinship
CN117524308B
Genetic relationship identification method with SNP as genetic marker
CN111091869A
Non-invasive parent-child detection analysis method and device
CN113584178A
SNP (Single Nucleotide Polymorphism)-based genetic relationship identification method
CN115565604A
Method for judging genetic relationship through SNP (Single Nucleotide Polymorphism) mismatch rate
CN115572770A