SNP-based kinship identification method

By using high-throughput sequencing technology and SNP site analysis, the problems of high mutation rate and limited detection sites in kinship identification have been solved, achieving highly accurate kinship identification that is applicable to various sample types, including degraded and old samples.

CN115565604BActive Publication Date: 2026-02-27WUHAN LANSHA MEDICAL LAB CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210995843.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-18
Publication Date
2026-02-27
Estimated Expiration
2042-08-18

AI Technical Summary

Technical Problem

Existing technologies for kinship identification suffer from problems such as high mutation rates, limited detection sites, gender restrictions on test subjects, and poor accuracy, especially in identifying sibling and grandparent-grandchild relationships, where it is difficult to provide a biased opinion.

Method used

Using next-generation high-throughput sequencing technology and thousands of SNP loci as genetic markers, this study simulates the cumulative paternity index under different kinship relationships through bioinformatics analysis and statistical algorithms, achieving the identification of parent-child, sibling, half-sibling/uncle-nephew/grandparent-grandchild and great-grandparent-grandchild relationships with an accuracy rate >99.99%.

Benefits of technology

It reduces the impact of gene mutations, increases the number of detection sites, is suitable for degraded or old samples, and the test results are not limited by gender, with significantly improved accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115565604B_ABST
    Figure CN115565604B_ABST
Patent Text Reader

Abstract

The application discloses a kind of SNP-based kinship identification method, belong to the technical field of parentage identification.Method comprising: obtaining SNP typing data of sample S1 and S2;Suppose S1 is F1, with S2 as benchmark, with SNP typing data to calculate the cumulative paternity index CPI1 of S1 under various assumed kinship;With S2 as benchmark, simulate multiple gene data Sm under each assumed kinship respectively, calculate the cumulative paternity index CPI2 of Sm, obtain the probability distribution of multiple lg (CPI2) of each kinship;Under various assumed kinship, judge whether lg (CPI1) conforms to the probability distribution of lg (CPI2);If only one kind of kinship except great-grandson relationship is consistent, then S1 and S2 are the kinship;If it is consistent in at least two kinds of kinship, then the kinship corresponding to the maximum lg (CPI1) is the kinship between S1 and S2.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of biological information analysis, and particularly relates to a SNP-based kinship identification method. BACKGROUND

[0002] Genetic theory has proved that the genomic DNA of children is half from each of the biological parents. Kinship identification is to determine whether there is a biological, intergenerational or other blood relationship between disputed individuals according to the basic principles of genetics and using modern DNA typing detection technology. Current kinship identification includes the following categories:

[0003] (1) Conventional biological blood relationship identification: This is the largest demand category of paternity relationship identification, including parent-child three-party (also known as triad), father-child (or mother-child) two-party (also known as diad) paternity identification.

[0004] (2) Intergenerational kinship identification: This type of identification refers to confirming the kinship between great-grandparents, grandparents, and great-grandchildren, grandchildren. In addition, it also includes simple paternal kinship identification such as confirming the kinship between great-grandfathers, grandfathers, and great-grandchildren, grandchildren, and simple maternal kinship identification such as confirming the kinship between great-grandmothers, grandmothers, and granddaughters, great-granddaughters.

[0005] (3) Difficult kinship identification: In addition to the above two categories, there are some more difficult kinship identifications, such as identification of siblings (brothers, sisters, brothers and sisters, sisters and brothers) when both parents are suspected (absent), identification of cousins, identification of the kinship between uncles and nieces, between aunts and goddaughters, and between uncles and godsons (goddaughters).

[0006] Parentage identification is a very mature application field of genetic detection technology. Generally, by detecting genetic markers of two test materials, the detection results of the two samples are compared. If the genetic markers of the two samples comply with Mendelian inheritance law, it is considered that the two samples comply with the parent-child relationship. Currently, there are two commonly used genetic markers, short sequence tandem repeats (STR) and single nucleotide polymorphism sites (SNP), in addition to some research using insertion and deletion (InDel) mutation sites as genetic markers. First-generation sequencing technology is the most mature detection technology in the field of parentage identification. Generally, 21 STR sites are used for parentage identification and discrimination. First-generation sequencing has the advantages of fast speed, low cost and simple operation, and is widely used by various identification institutions. It is the main detection technology for current parentage identification.

[0007] The STR marker detection sites used for first-degree sibling relationship identification are generally 19-39. The average mutation frequency of STR markers is 0.2%. STR markers are used to detect the biological parents and offspring, biological siblings, and grandchild identification. Gene mutations often lead to mismatches. Therefore, if 1-2 loci do not comply with the corresponding genetic law (considering the existence of mutations), the "Biological Full Sibling Relationship Identification Implementation Specification" suggests adding autosomal STR, Y chromosome STR or mitochondrial DNA sequence (SF / Z JD0105002-2014, Biological Full Sibling Relationship Identification Implementation Specification). Some studies suggest that other genetic markers, such as SNP markers and / or insertion / deletion polymorphism, can be combined with more stable biallelic genetic markers (Tang Meiyun, Huang Jian, Cai Jinhong, et al. Identification of sibling relationship by STR and Y chromosome biallelic markers [J]. Journal of Forensic Medicine, 2012(03): 190-194. Zhang Suhua, Zhao Shumin, Li Li. Identification of full sibling relationship by STR and SNP genetic markers [J]. Journal of Forensic Medicine, 2010(03): 185-187). Due to the limited number of STR detections and high mutation rate, first-degree sibling relationship identification often cannot give a tendency opinion. When STR is used for grandchild identification, the "Technical Specifications for Paternity Identification" requires the provision of the grandparent's test material. If 1-2 loci do not comply with the corresponding genetic law (considering the existence of mutations), other highly polymorphic and genetically stable STR loci should be added. If the child being tested is a girl, the X-STR test of the tested grandmother and granddaughter should be increased as much as possible. If the child being tested is a boy, the Y-STR test of the tested grandfather and grandson should be increased as much as possible. And only the identification opinion of "not excluding the existence of biological grandchild relationship" can be given.

[0008] The present application adopts a new generation of high-throughput sequencing technology, uses SNP sites as genetic markers, captures and sequences thousands of biallelic autosomal SNP sites in the human genome, detects low-frequency mutations as low as one thousandth at each SNP site, and uses specific bioinformatics analysis and mathematical algorithm to identify sibling, half-sibling / uncle-niece / cousin / uncle-niece / uncle-niece / ancestor-grandchild, and the accuracy is > 99.99%. The detection object is not limited by gender. The method can solve the problems of poor feasibility and low detection rate in traditional first-degree kinship identification. SUMMARY

[0009] In order to solve the above problems, the patent of the application adopts a new generation of high-throughput sequencing technology, takes SNP sites as genetic markers, carries out target region capture sequencing on thousands of SNP sites in the human genome, can detect low-frequency mutations as low as one thousandth at each SNP site, and can at least judge three kinds of kinship through specific bioinformatics analysis and statistical algorithm, with an accuracy of > 99.99%. The technical scheme is as follows:

[0010] The embodiment of the application provides a SNP-based kinship identification method, which comprises the following steps:

[0011] S101: obtaining SNP typing data of two samples S1 and S2 to be identified;

[0012] S102: assuming that S1 is F1, taking S2 as a reference, and calculating the cumulative parentage index CPI1 of S1 under various assumed kinship relationships based on the SNP typing data;

[0013] S103: taking S2 as a reference, simulating a plurality of gene data Sm under each assumed kinship relationship, calculating the cumulative parentage index CPI2 of Sm based on the SNP typing data, and then calculating the probability distribution of a plurality of lg(CPI2) of each kinship relationship;

[0014] S104: under various assumed kinship relationships, determining whether lg(CPI1) conforms to the probability distribution of lg(CPI2) through a statistical algorithm; if it does not conform under various assumed kinship relationships, the assumed kinship relationship between S1 and S2 is not established; if it conforms to only one kinship relationship except the great-grandson relationship, S1 and S2 are in the kinship relationship; if it conforms to at least two kinship relationships, the kinship relationship between S1 and S2 is the kinship relationship corresponding to the maximum lg(CPI1);

[0015] Among them, the assumed kinship relationship is divided into the following categories: category 1: parent-child relationship; category 2: sibling relationship; category 3: half-sibling relationship, uncle-niece relationship and grandparent-grandchild relationship; category 4: great-grandson relationship.

[0016] Further, in step S104, if only the great-grandson relationship conforms, it cannot be determined that S1 and S2 are in the great-grandson relationship.

[0017] Specifically, the step S103 comprises:

[0018] assuming that the kinship relationship between S1 and S2 is category 1, simulating gene data Sm1 of M1 offspring of S2;

[0019] assuming that the kinship relationship between S1 and S2 is category 2, simulating gene data Sm2 of M2 siblings of S2;

[0020] Assuming that the relationship between S1 and S2 is category 3, M3 gene data Sm3 of half-siblings, nephews or grandchildren of S2 are simulated;

[0021] Assuming that the relationship between S1 and S2 is category 4, M4 gene data Sm4 of great-grandchildren of S2 are simulated.

[0022] The M1≥30, M2≥30, M3≥30, M4≥30.

[0023] In step S103, a plurality of gene data Sm are simulated at a theoretical frequency of genotypes.

[0024] The cumulative paternity index CPI=CPI1×PI2×PI3×…×PIn, wherein PIn is a PI value of a SNP site shared by two samples, and the PI value is a ratio of a possibility of F1 being a certain genotype under the condition that the two samples have a relationship to a possibility of F1 being the certain genotype under the condition that a random individual and F1 have the relationship.

[0025] Specifically, when the relationship is category 1, the PI value is as follows:

[0026]

[0027] When the relationship is category 2, the PI value is as follows:

[0028]

[0029] When the relationship is category 3, the PI value is as follows:

[0030]

[0031] When the relationship is category 4, the PI value is as follows:

[0032]

[0033] F1 is a child generation; P1 is a parent generation, and the relationship with F1 is a parent-child relationship; P2 is a grandparent generation, and the relationship with F1 is a great-grandchild relationship; P3 is a great-grandparent generation, and the relationship with F1 is a great-great-grandchild relationship; F1S1 is a sibling of F1; F1S2 is a half-sibling of F1; P1S1 is a sibling of P1, and the relationship with F1 is an uncle-nephew relationship; P(A) is a population frequency of wild type A; P(B) is a population frequency of mutant type B; and μ is a sequencing error rate.

[0034] In step S104, it is determined whether lg(CPI1) conforms to a probability distribution of lg(CPI2) by chi-square test.

[0035] Specifically, when performing the chi-square test, if |Z-score|≤3, lg(CPI1) belongs to the probability distribution of lg(CPI2); if |Z-score|>3, lg(CPI1) does not belong to the probability distribution of lg(CPI2).

[0036] The method has the following advantages:

[0037] 1. Lower mutation rate

[0038] STR markers have an average mutation frequency of 0.2%, and SNPs are referred to as "third-generation genetic markers" with a mutation rate of less than one in ten million, which can reduce the impact of genetic mutations on parentage identification.

[0039] 2. More detection sites

[0040] The STR markers used in the first generation of parentage identification generally have 16-33 detection sites. The method can select more than 200 high-quality SNP sites from the ten thousand sites of human autosomes for analysis, which is more accurate.

[0041] 3. More sample types can be selected.

[0042] STR short tandem repeat sequences are 100-200 bp in length, and DNA degradation can affect the results of the first generation of parentage identification. SNP, as a single-base genetic marker, does not depend on the integrity of the detection sample. Samples with DNA degradation (improper transportation and storage; or easily degradable samples such as fingernails) will not affect the detection results. Therefore, SNP-based kinship identification requires less test material, has high sensitivity and success rate, and is suitable for degraded and old test materials.

[0043] 4. Better feasibility

[0044] Using SNP markers for sibling identification, whether it is parentage, sibling, half-sibling / uncle-niece / ancestor-grandchild, or great-grandchild relationship identification, a set of bioinformatics analysis processes can be used, plus the low mutation frequency of SNPs and the large number of sites. All detection cases can give a tendency result. The first generation of sibling and grandchild relationship identification often needs to be supplemented and verified by increasing autosomal STR, X chromosome STR, Y chromosome STR, or mitochondria, and there is a certain probability that it cannot give a tendency result. The method of the present application can identify kinship, including parentage, sibling, half-sibling / uncle-niece / ancestor-grandchild, and great-grandchild relationship identification. Using the method does not limit by gender factors. BRIEF DESCRIPTION OF DRAWINGS

[0045] Figure 1 is a flowchart of the SNP-based kinship identification method provided by the embodiments of the present application;

[0046] Figure 2 is a relationship diagram of each kinship. DETAILED DESCRIPTION

[0047] In order to make the objects, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings.

[0048] Embodiment 1

[0049] Referring to Figure 1 , embodiment 1 discloses a SNP-based kinship identification method, which comprises the following steps:

[0050] S101: obtaining SNP typing data of two samples S1 and S2 to be identified, wherein the SNP typing data is obtained by using second-generation sequencing technology.

[0051] S102: assuming that S1 is F1, taking S2 as a reference, and calculating a cumulative parentage index CPI1 of S1 under various assumed kinships by using the SNP typing data; generally, the two samples to be submitted are expected to be child generation, and the child generation is considered by the customer as S1.

[0052] S103: taking S2 as a reference, simulating a plurality of gene data Sm under each assumed kinship (four in this embodiment), and calculating a cumulative parentage index CPI2 of Sm by using the SNP typing data, and then calculating a probability distribution of a plurality of lg(CPI2) of each kinship.

[0053] S104: under various assumed kinships, determining whether lg(CPI1) conforms to the probability distribution of lg(CPI2) by using a statistical algorithm; if not conforming under various assumed kinships, the assumed kinship between S1 and S2 is not established; if only one kinship except the great-grandson relationship conforms, S1 and S2 are the kinship; if at least two kinships conform, the kinship corresponding to the maximum lg(CPI1) is the kinship between S1 and S2 (also cannot be the great-grandson relationship).

[0054] In this embodiment, kinship relationships are categorized as follows: Category 1: Parent-child relationship; Category 2: Sibling relationship; Category 3: Half-sibling relationship, uncle-nephew relationship, and grandparent-grandchild relationship; where, in this embodiment, uncle-nephew relationship includes uncle-nephew, aunt-nephew, maternal uncle-niece, and maternal aunt-niece, etc., regardless of gender. This embodiment sets half-sibling relationship, uncle-nephew relationship, and grandparent-grandchild relationship as the same type. Category 4: Great-grandparent-grandchild relationship. All of the above relationships are not limited by gender. For Category 3, if the kinship relationship is determined to be Category 3 and matches the expected kinship relationship of samples S1 and S2, then the expected kinship relationship of S1 and S2 is taken as the kinship relationship between S1 and S2. If a client requests a determination of whether two samples are half-siblings when submitting samples for testing, then if they match the kinship relationship of Category 3, it can be used as the basis for determining the half-sibling relationship (in conjunction with other methods).

[0055] Furthermore, in step S104, if only the great-grandfather-grandson relationship is met, then it cannot be determined that S1 and S2 are great-grandfather-grandson relationships.

[0056] Specifically, step S103 includes:

[0057] Assuming that the kinship between S1 and S2 is category 1, the genetic data Sm1 of M1 offspring of S2 are simulated.

[0058] Assuming that the kinship between S1 and S2 is category 2, the genetic data Sm2 of M2 siblings of S2 are simulated.

[0059] Assuming that the kinship between S1 and S2 is category 3, then the genetic data Sm3 of M3 half-siblings, nephews or grandchildren of S2 are simulated.

[0060] Assuming that the kinship between S1 and S2 is category 4, the genetic data Sm4 of M4 great-grandchildren of S2 are simulated.

[0061] Among them, M1≥30, M2≥30, M3≥30, M4≥30.

[0062] In step S103, multiple gene datasets Sm are simulated using the theoretical frequencies of genotypes. For example, if a locus S2 has the AA genotype, P(B) = 0.4 and P(B) = 0.6, then theoretically, the probabilities of AA, BB, and AB in the first generation F1 of S2 at this SNP locus are 0.2, 0.3, and 0.5. We simulate 300 first-generation F1s of S2; at this SNP locus, there are 60 AA, 90 BB, and 150 AB.

[0063] CPI = PI1x PI2x PI3x...x PIn, PIn is the PI value of the SNP site shared by two samples. The PI value is the ratio of the possibility of a certain genotype of F1 to the possibility of a certain genotype of F1 in the case that F1 is a certain genotype and the random individual and F1 are in a certain kinship.

[0064] Specifically, when the kinship is type 1, the PI value is shown in Table 1:

[0065] Table 1

[0066]

[0067] When the kinship is type 2, the PI value is shown in Table 2:

[0068] Table 2

[0069]

[0070] When the kinship is type 3, the PI value is shown in Table 3:

[0071] Table 3

[0072]

[0073] When the kinship is type 4, the PI value is shown in Table 4:

[0074] Table 4

[0075]

[0076] Wherein, F1 is the child generation; P1 is the parent generation, and the kinship with F1 is parent-child relationship; P2 is the second parent generation, and the kinship with F1 is grandparent-grandchild relationship; P3 is the third parent generation, and the kinship with F1 is great-grandparent-great-grandchild relationship; F1S1 is the sibling of F1; F1S2 is the half sibling of F1; P1S1 is the sibling of P1, and the kinship with F1 is uncle-niece relationship; P(A) is the frequency of wild type A population; P(B) is the frequency of mutant type B population; μ is the sequencing error rate.

[0077] In step S104, it is determined whether lg(CPI1) conforms to the probability distribution of lg(CPI2) by chi-square test.

[0078] Specifically, when performing chi-square test, if |Z-score|≤3, lg(CPI1) belongs to the probability distribution of lg(CPI2); if |Z-score|>3, lg(CPI1) does not belong to the probability distribution of lg(CPI2).

[0079] The number of SNP sites is preferably greater than or equal to 100, and in the embodiment, can be 200, and the autosomal SNP sites are biallelic with a mutation frequency of 0.4-0.6. In step S103, the number of simulation samples of each assumed kinship is preferably greater than or equal to 30, and in the embodiment, is 300.

[0080] Embodiment 2

[0081] Embodiment 2 discloses a SNP-based kinship identification method, which comprises the following steps:

[0082] S101 sequencing: More than 200 autosomal SNP sites on the human genome with a mutation frequency of 0.4-0.6 are selected as genetic markers for kinship identification. After obtaining the test material, the target test material is first subjected to nucleic acid extraction, and whole genome library construction. In the library construction process, the DNA sequence of each sample is added with a barcode sequence representing the number and a sequencing adapter that can be used for high-throughput sequencing and other necessary sequences, and whole genome amplification is performed. After library construction is completed, a set of probe sequences is used to perform liquid hybridization capture on thousands of SNP sites, and high-throughput sequencing and bioinformatics analysis are performed.

[0083] S102 typing: After sequencing and analysis, the SNP typing of the test sample is obtained.

[0084] S103 calculating the PI value and CPI value of the two samples.

[0085] Referring to Figure 2 In the embodiment, the kinship is defined by the following letters:

[0086] F1: first generation;

[0087] P1: first generation, parent of F1, including father and mother;

[0088] P2: second generation, parent of P1, including grandfather and grandmother of F1, etc.

[0089] P3: third generation, parent of P2, including great-grandfather and great-grandmother of F1, etc.

[0090] F1S1: sibling of F1, i.e. brother and sister of F1 from the same father and mother;

[0091] F1S2: half sibling of F1, i.e. brother and sister of F1 from the same father and different mother or from the same mother and different father;

[0092] P1S1: sibling of P1, i.e. brother and sister of P1 from the same father and mother; including uncle, aunt, uncle and aunt of F1, etc.

[0093] PI: Paternity Index, the ratio of the probability of a certain genotype of F1 between the individuals to be tested under the condition of existing kinship and the probability of a certain genotype of F1 between the random individual and F1 under the condition of existing kinship. The PI value is calculated by formula 1 as follows:

[0094]

[0095]

[0096]

[0097]

[0098]

[0099]

[0100] CPI: Combined Paternity Index, the PI value of all genetic marker sites is multiplied, that is, the CPI value.

[0101] Wherein, A: genotype is wild type, that is, the same as the genotype of the reference genome. B: genotype is mutant type, that is, different from the genotype of the reference genome. P(A): population frequency of wild type A. P(B): population frequency of mutant type B. μ: sequencing error rate, the average sequencing error rate is 0.002.

[0102] Let the paternity index of each genetic marker be PI1, PI2, PI3, … PIn, and the paternity index of n genetic markers is multiplied to be CPI. The CPI value is calculated by formula 2 as follows:

[0103] CPI = PI1 × PI2 × PI3 × … × PIn (1, 2, 3, n represent the PI values of the 1st, 2nd, 3rd, and nth SNP sites).

[0104] S104 calculates the PI value and CPI value of the simulation data under various kinship relationships and judges the kinship relationship.

[0105] Fixing the genotype of P1, the genotypic data of 300 theoretical offspring of P1 are simulated, the simulated CPI values of 300 P1 offspring are calculated according to formula 1 and 2, and the probability distribution of lg(CPI) of P1 offspring is obtained. Whether lg[CPI(F1-P1)] belongs to the probability distribution of lg(CPI) of P1 simulated offspring is judged by chi-square test. If |Z-score|≤3, lg[CPI(F1-P1)] belongs to the simulated lg(CPI) distribution; if |Z-score|>3, lg[CPI(F1-P1)] does not belong to the probability distribution of simulated lg(CPI).

[0106] Fixing the genotype of F1S1, the genotypic data of 300 theoretical siblings of F1S1 are simulated, the simulated CPI values of 300 F1S1 siblings are calculated according to formula 1 and 2, and the probability distribution of lg(CPI) of F1S1 siblings is obtained. Whether lg[CPI(F1-F1S1)] belongs to the probability distribution of lg(CPI) of F1S1 simulated siblings is judged by chi-square test. If |Z-score|≤3, lg[CPI(F1-F1S1)] belongs to the simulated lg(CPI) distribution; if |Z-score|>3, lg[CPI(F1-F1S1)] does not belong to the probability distribution of simulated lg(CPI).

[0107] Fixing the genotype of F1S2 / P1S1 / P2, the genotypic data of 300 theoretical half-siblings of F1S2, 300 theoretical grandchildren of P1S1 and / or 300 theoretical grandchildren of P2 are simulated (preferably, all three are simulated), the CPI values of 300 F1S2 half-siblings, 300 P1S1 grandchildren and / or 300 P2 theoretical grandchildren are calculated according to formula 1 and 2; the probability distribution of lg(CPI) of F1S2 half-siblings, the probability distribution of lg(CPI) of P1S1 grandchildren and / or the probability distribution of lg(CPI) of P2 grandchildren are obtained. Whether lg[CPI(F1-F1S2 / P1S1 / P2)] belongs to the probability distribution of lg(CPI) of simulated half-siblings, uncles and grandchildren is judged by chi-square test. If |Z-score|≤3, it belongs to; if |Z-score|>3, it does not belong to.

[0108] Fixed the genotype of P3, simulated the theoretical grandchild genotype data of 300 P3, calculated the simulated CPI value of 300 P3 grandchild according to formula 1 and 2, and obtained the probability distribution of lg(CPI) of P3 grandchild. Whether lg[CPI(F1-P3)] belongs to the probability distribution of lg(CPI) of P3 simulated grandchild is judged by chi-square test. If |Z-score|≤3, lg[CPI(F1-P3)] belongs to the probability distribution of simulated lg(CPI); if |Z-score|>3, lg[CPI(F1-P3)] does not belong to the probability distribution of simulated lg(CPI).

[0109] When the CPI values of two or more kinds of kinship meet the theoretical distribution, it is necessary to determine which kind of kinship probability is large.

[0110] For example, in a certain family, the CPI values of parent-child relationship and sibling relationship meet the theoretical distribution, and the probability of sibling relationship between the individuals to be tested is: the probability of parent-child relationship between the individuals to be tested.

[0111]

[0112] CI = I1x I2x I3x…x In (1, 2, 3, n represent the I value of the 1st, 2nd, 3rd, n SNP site) = CPI(F1-F1S1) / CPI(F1-P1).

[0113] CI>1, i.e. CPI(F1-F1S1)>CPI(F1-P1), indicating that the probability of sibling relationship between the individuals to be tested is greater than the probability of parent-child relationship between the individuals to be tested.

[0114] CI<1, i.e. CPI(F1-F1S1)<CPI(F1-P1), indicating that the probability of sibling relationship between the individuals to be tested is less than the probability of parent-child relationship between the individuals to be tested.

[0115] CI=1, i.e. CPI(F1-F1S1)=CPI(F1-P1), indicating that the probability of sibling relationship between the individuals to be tested is equal to the probability of parent-child relationship between the individuals to be tested.

[0116] Therefore, when the CPI values of two or more kinds of kinship meet the theoretical distribution, the kinship with larger CPI value is supported as the identification result.

[0117] Specifically, if lg[CPI(F1-P1)] belongs to the theoretical offspring CPI distribution of P1, and lg[CPI(F1-P1)] is greater than lg[CPI(F1-F1S1)], lg[CPI(F1-F1S2 / P1S1 / P2)], and lg[CPI(F1-P3)], it supports the existence of parent-child relationship between the two samples. If lg[CPI(F1-F1S1)] belongs to the CPI distribution of the theoretical sibling of F1S1, and lg[CPI(F1-F1S1)] is greater than lg[CPI(F1-P1)], lg[CPI(F1-F1S2 / P1S1 / P2)], and lg[CPI(F1-P3)], it supports the existence of sibling relationship between the two samples. If lg[CPI(F1-F1S2 / P1S1 / P2)] belongs to the CPI distribution of the theoretical half-sibling / uncle / niece / grandson of F1S2 / P1S1 / P2, and lg[CPI(F1-F1S2 / P1S1 / P2)] is greater than lg[CPI(F1-P1)], lg[CPI(F1-F1S1)], and lg[CPI(F1-P3)], it supports the hypothesis that the two samples are half-siblings, uncles and nieces, and grandchildren. If lg[CPI(F1-P3)] belongs to the theoretical great-grandson CPI distribution of P3, and lg[CPI(F1-P3)] is greater than lg[CPI(F1-P1)], lg[CPI(F1-F1S1)], and lg[CPI(F1-F1S2 / P1S1 / P2)]; it does not exclude the existence of great-grandson relationship between the two samples

[0118] If all the calculated relative lg[CPI] values do not belong to the theoretical lg[CPI] distribution, it is excluded that the two samples have parent-child / sibling / half-sibling / uncle / niece / grandson / great-grandson relationship.

[0119] Example 3

[0120] A family numbered 1703 sent a blood sample of an adult and a blood sample of a child for parent-child relationship identification, which are marked as RT1703P1 and RT1703F1 respectively. The SNP typing results of RT1703P1 and RT1703F1 are obtained by sequencing analysis. The CPI values of various relatives and the chi-square test results are shown in Table 5:

[0121] Table 5

[0122] Parent-offspring Siblings Half-siblings Uncle-niece Grandparent-grandchild Great grandparent-great grandchild 1g(CPI) value -∞ -293.78 -51.46 -51.46 -51.46 -0.67 Mean of simulated 1g(CPI) distribution 465.58 355.06 88.90 89.04 87.26 22.3 SD of simulated 1g(CPI) distribution 14.82 22.90 11.24 11.82 12.72 6.37 Kolmogorov-Smirnov test z-score -∞ -28.24 -12.49 -11.89 -10.90 -3.61 Fit to distribution Not fit Not fit Not fit Not fit Not fit Not fit

[0123] Detection conclusion: It is excluded that RT1703P1 and RT1703F1 have parent-child, sibling, half-sibling, uncle, niece, and great-grandson relationship.

[0124] Example 4

[0125] A family numbered 1704 sent a blood sample of an adult and a blood sample of a child for paternity testing, marked as RT1704P1 and RT1704F1 respectively. The SNP typing results of RT1704P1 and RT1704F1 were obtained by sequencing analysis. The CPI values and chi-square test results of various relationships are shown in Table 6:

[0126] Table 6

[0127] Parent-offspring Siblings Half-siblings Uncle-niece Grandparent-grandchild Great grandparent-great grandchild 1g(CPI) value 469.82 359.43 227.06 227.06 227.06 152.88 Mean of simulated 1g(CPI) distribution 462.90 353.23 88.96 89.37 87.32 22.10 SD of simulated 1g(CPI) distribution 12.11 25.05 11.66 14.31 12.79 6.53 Kolmogorov-Smirnov test z-score 0.57 0.24 16.13 13.11 14.84 20.04 Fit to theoretical distribution Fit Fit Not fit Not fit Not fit Not fit

[0128] The detection conclusion is that RT1704P1 and RT1704F1 are both in the paternity relationship.

[0129] Example 5

[0130] A family numbered 1523 sent a blood sample and a nail sample of two children for sibling relationship testing, marked as RT1523F1 and RT1523F1S1 respectively. The SNP typing results of RT1523F1 and RT1523F1S1 were obtained by sequencing analysis. The CPI values and chi-square test results of various relationships are shown in Table 7:

[0131] Table 7

[0132] Parent-offspring Siblings Half-siblings Uncle-niece Grandparent-grandchild Great grandparent-great grandchild 1g(CPI) value 21.64 307.06 228.96 228.96 228.96 134.86 Mean of simulated 1g(CPI) distribution 445.31 346.47 87.02 87.20 87.51 20.65 SD of simulated 1g(CPI) distribution 12.81 24.69 14.63 12.39 12.45 6.81 Kolmogorov-Smirnov test z-score -33.84 -1.60 9.70 11.44 11.36 16.78 Fit to theoretical distribution Not fit Not fit Fit Fit Fit Not fit

[0133] The detection conclusion is that RT1523F1 and RT1523F1S1 are both in the sibling relationship.

[0134] Example 6

[0135] A family numbered 782 sent a buccal swab sample of an adult and a blood sample of a child for grandchild relationship testing, marked as RT782P2 and RT782F1 respectively. The SNP typing results of RT782P2 and RT782F1 were obtained by sequencing analysis. The CPI values and chi-square test results of various relationships are shown in Table 8:

[0136] Table 8

[0137] Parent-offspring Siblings Half-siblings Uncle-niece Grandparent-grandchild Great grandparent-great grandchild 1g(CPI) value -581.31 -19.31 80.81 80.81 80.81 61.31 Mean of simulated 1g(CPI) distribution 449.12 340.29 85.405 85.35 83.47 20.74 SD of simulated 1g(CPI) distribution 13.50 25.12 11.55 13.37 11.46 6.99 Kolmogorov-Smirnov test z-score -76.33 -14.31 -0.40 -0.34 -0.23 5.80 Fit to theoretical distribution Not fit Not fit Fit Fit Fit Not fit

[0138] The detection conclusion is that RT782P2 and RT782F1 are both in the half-sibling, uncle-niece or grandchild relationship.

[0139] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for identifying the kinship based on SNP, characterized in that, The method comprises the following steps: S101: obtaining SNP typing data of two samples S1 and S2 to be identified; S102: assuming S1 as F1, taking S2 as a reference, and calculating a cumulative parentage index CPI1 of S1 under various assumed parentage relationships according to the SNP typing data; S103: taking S2 as a reference, simulating a plurality of genetic data Sm under each assumed parentage relationship, calculating a cumulative parentage index CPI2 of Sm according to the SNP typing data, and then calculating a probability distribution of a plurality of lg(CPI2) of each parentage relationship; S104: determining whether lg(CPI1) conforms to the probability distribution of lg(CPI2) under various assumed parentage relationships by a statistical algorithm; if it does not conform under various assumed parentage relationships, the assumed parentage relationship between S1 and S2 is not established; if it conforms to only one parentage relationship except the great-grandson relationship, S1 and S2 are the parentage relationship; if it conforms to at least two parentage relationships, the parentage relationship corresponding to the maximum lg(CPI1) is the parentage relationship between S1 and S2; if only the great-grandson relationship conforms, it cannot be determined that S1 and S2 are the great-grandson relationship; wherein the assumed parentage relationships are divided into the following categories: category 1: parent-child relationship; category 2: sibling relationship; category 3: half-sibling relationship, uncle-nephew relationship and grandparent-grandchild relationship; and category 4: great-grandparent-great-grandchild relationship; The cumulative parentage index CPI is PI1×PI2×PI3×…×PIn, PIn is a PI value of a SNP site shared by the two samples, and the PI value is a ratio of a possibility of F1 being a certain genotype under the condition that the individuals to be tested have a parentage relationship to a possibility of a random individual and F1 having the certain genotype under the parentage relationship.

2. The SNP-based relationship identification method according to claim 1, characterized in that, The step S103 comprises: assuming that the parentage relationship between S1 and S2 is category 1, and then simulating M1 genetic data Sm1 of the offspring of S2; assuming that the parentage relationship between S1 and S2 is category 2, and then simulating M2 genetic data Sm2 of the siblings of S2; assuming that the parentage relationship between S1 and S2 is category 3, and then simulating M3 genetic data Sm3 of the half-siblings, nephews or grandchildren of S2; assuming that the parentage relationship between S1 and S2 is category 4, and then simulating M4 genetic data Sm4 of the great-grandchildren of S2; M1≥30, M2≥30, M3≥30, and M4≥30.

3. The SNP-based relationship identification method according to claim 2, wherein, In step S103, the plurality of genetic data Sm is simulated according to the theoretical frequency of the genotype.

4. The SNP-based parentage identification method according to claim 1, characterized in that, when the parentage relationship is type 1, the PI value is as shown in the following table: when the parentage relationship is type 2, the PI value is as shown in the following table: when the parentage relationship is type 3, the PI value is as shown in the following table: when the parentage relationship is type 4, the PI value is as shown in the following table: Wherein, F1 is a sub-generation; P1 is a parent generation, and the relationship with F1 is parent-child relationship; P2 is a second-generation, and the relationship with F1 is grandchild relationship; P3 is a third-generation, and the relationship with F1 is great-grandchild relationship; F1S1 is a sibling of F1; F1S2 is a half-sibling of F1; P1S1 is a sibling of P1, and the relationship with F1 is uncle-niece relationship; P(A) is the population frequency of wild type A; P(B) is the population frequency of mutant type B; μ is the sequencing error rate.

5. The SNP-based relationship identification method according to claim 1, wherein, In step S104, it is judged whether lg(CPI1) conforms to the probability distribution of lg(CPI2) by chi-square test.

6. The SNP-based relationship identification method according to claim 5, wherein, When performing chi-square test, if |Z-score|≤3, lg(CPI1) belongs to the probability distribution of lg(CPI2); if |Z-score|>3, lg(CPI1) does not belong to the probability distribution of lg(CPI2).

Citation Information

Patent Citations

  • Genetic relationship identification method with SNP as genetic marker

    CN111091869A

  • Genetic relationship identification method, system and equipment

    CN112053743A