Genetic relationship identification method and system based on homologous co-progenitor fragment IBD

By using a kinship identification method based on the common ancestor fragment IBD, and employing a likelihood ratio test framework and distribution fitting, the method solves the problem of low accuracy in the identification of third-degree and above kinship in existing technologies, and achieves efficient multi-level kinship discrimination.

CN121747684APending Publication Date: 2026-03-27FUDAN UNIVERSITY +2
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-18
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing methods for identifying kinship relationships show a significant decrease in their ability to distinguish between third-degree and higher kinship relationships, increasing the risk of misjudgment, especially in complex kinship relationships where accuracy is reduced.

Method used

A kinship identification method based on common ancestral fragments (IBD) was adopted. By constructing a likelihood ratio test framework and combining the number and length of IBD fragments, a significance test system for determining specific kinship was established. Poisson distribution and exponential distribution were used to fit the IBD characteristics, and the likelihood ratio was calculated to determine kinship.

Benefits of technology

It improves the accuracy of identifying third-degree and above kinship, especially with a 100% efficiency for third- and fourth-degree kinship and a 99.85% efficiency for fifth-degree kinship. It breaks through the limitations of existing strategies that compare IBS scores with preset thresholds and is applicable to the identification of multi-degree kinship.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121747684A_ABST
    Figure CN121747684A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of genetic relationship identification methods, particularly relates to a genetic relationship identification method and system based on homologous co-progenitor fragments IBD, and aims to solve the technical problem of identifying genetic relationships of five levels and within five levels. According to the technical scheme, the method comprises the following steps: detecting IBD fragments between two disputed individuals; constructing a likelihood ratio test framework, taking two dispute individuals as irrelevant individuals as an original hypothesis, and taking a specific genetic relationship between the two dispute individuals as a standby hypothesis; solving the probability of IBD fragment data under the original hypothesis; solving the probability of IBD fragment data under the alternative hypothesis; calculating a likelihood ratio gt by synthesizing the probabilities of IBD fragment data appearing under the original hypothesis and the alternative hypothesis; if 1 is less than 1, determining that the individual is an unrelated individual.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical field of kinship identification methods, specifically relating to a kinship identification method and system based on the common ancestor fragment IBD. Background Technology

[0002] Kinship identification is a crucial scientific issue in forensic medicine, human genetics, and biology. Its core objective is to determine the biological blood relationship between individuals through genetic marker analysis. Traditional kinship identification primarily infers kinship by sharing allele information on genetic markers such as short tandem repeats (STRs), demonstrating high accuracy in identifying first- and second-degree kinship (e.g., pairs, triplets, grandparents, grandchildren). However, as the degree of kinship increases (e.g., third-degree and above), the discriminatory power of these methods significantly declines, and the risk of misjudgment increases. The limitations of existing methods are particularly pronounced in identifying complex kinship relationships (e.g., first cousins, grandparents, etc.).

[0003] Current mainstream analytical frameworks can be divided into two main strategies: Identity-by-State (IBS) and Identity-by-Descent (IBD). In addition, the Likelihood Ratio (LR) is also a widely used statistical method. These methods differ significantly in their technical approaches and applicability. The IBS strategy focuses on allele sharing at the phenotypic level, while the IBD strategy focuses on the dynamic tracing of genetic segments through generations. The Likelihood Ratio assesses relationships by comparing genotype probabilities under different kinship assumptions.

[0004] IBS refers to two individuals sharing the same nucleotide sequence or alleles, but these genes may not originate from a common ancestor. IBS focuses on the superficial similarity of genotypes, regardless of genetic origin. For example, if two individuals both have the same allele A at a certain locus, but these A alleles come from different ancestors, they are still considered IBS. IBD refers to two individuals sharing alleles that originate from a common ancestor; in other words, these genes are directly passed down through heredity. If the shared gene segment between two individuals can be traced back to the same ancestor, rather than occurring randomly, it is identified as IBD.

[0005] (1) IBS-based identification methods: Currently, IBS-based kinship identification mainly includes three representative methods: IBS scoring method and Method of Moment (MoM).

[0006] The IBS scoring method is the most direct method for analyzing kinship. It infers kinship by calculating the number of shared alleles between two individuals. A weighted scoring system based on STR loci distinguishes kinship levels by setting thresholds. my country's Ministry of Justice's "Technical Specifications for Biological Full-Sibling Identification" (SF / T0117—2021) adopts this method, setting IBS scoring thresholds for determining full-sib relationships when testing 19-55 autosomal STR loci. Using the 13 core STRs of the US CODIS system as an example, if two individuals have an IBS score of 17 or higher, they can be identified as full siblings with high confidence. However, the accuracy of the IBS scoring method decreases significantly when identifying second-degree and higher kinship. In the study by Li et al. (Tao et al., 2022a), they found that most second-degree related individuals were difficult to distinguish from unrelated individuals, especially cousins, whose IBS scores significantly overlapped with unrelated individuals, leading to an increased risk of misjudgment. Although the IBS method is simple to calculate and provides intuitive data, it has poor universality and is not suitable for low copy number DNA and mixed samples.

[0007] The method of moments (MoM) typically quantitatively characterizes the strength of kinship by estimating the kinship coefficient and correlation indices between individuals (Dou et al., 2017). Typical implementations include PLINK (Purcell et al., 2007) and KING (Manichaikul et al., 2010) software; the former is suitable for large-scale population genetic analysis, while the latter focuses on efficient family pedigree testing and individual relationship classification. The MoM method estimates kinship based on the proportion of homologous sharing across multiple loci in sample pairs, making it suitable for large-sample genetic structure analysis. This method is computationally efficient but lacks the ability to significantly distinguish between different kinship categories (such as full siblings versus half siblings, uncles and nephews versus cousins). Furthermore, MoM estimates are sensitive to population background structure, often exhibiting systematic bias in populations with subpopulation structures, inbreeding, or heterozygous bias, sometimes even yielding negative results, leading to difficulties in practical interpretation.

[0008] In summary, IBS-based methods offer advantages such as high computational speed and simple model construction, making them effective for initial screening of kinship or identification of lower-level kinship. However, most of these methods rely on autosomal STRs or independent SNP loci for identification, exhibiting high accuracy in some first- and second-degree kinship identifications. However, their accuracy rapidly declines in identifying third-degree and more distant relatives, especially in cases of sample contamination, DNA degradation, or mixed samples, where accuracy becomes difficult to guarantee. In recent years, with the continuous development of forensic eccentricity, the scope of kinship identification has expanded from routine parentage testing to complex kinship identification such as uncles and nephews, half-siblings, and cousins. These kinship identification cases often require the participation of grandparents, siblings, or collateral relatives due to the absence of key individuals. The application of complex kinship identification is widespread, and the demand for such identification is increasing annually. However, existing technologies and methods for complex kinship identification cannot adequately meet practical needs, and many aspects require improvement. In addition, traditional scoring methods and likelihood ratio models suffer from low computational efficiency and model simplification when dealing with complex family structures and high-throughput data, which limits the depth and breadth of their practical applications.

[0009] (2) IBD-based identification methods: kinship identification based on bloodline consistency is one of the core technologies in forensic medicine and population genetics for studying distant kinship. An IBD fragment refers to a continuous chromosomal region inherited from a common ancestor between two individuals. Therefore, the closer the kinship, the more numerous and longer the shared IBD fragments. For example, in parent-child relationships, individuals almost share an entire chromosome-level IBD region, while cousins ​​only share scattered, shorter IBD fragments. Through systematic analysis of the number, length, and distribution of IBD fragments, kinship relationships from first to sixth degree or even further can be inferred (Huff et al., 2011).

[0010] In recent years, many efficient IBD detection algorithms and supporting software have been developed, improving the accuracy and operability of inference under large sample sizes and distant kinship. GERMLINE, an earlier IBD detection tool, uses a sliding window matching algorithm (Gusev et al., 2009) to achieve efficient detection of IBD fragments. Its core lies in introducing a seed expansion strategy, which can quickly identify shared fragments longer than 2 cM (centimolecular weight) across the entire genome. It is suitable for high-density SNP microarray data, but is relatively sensitive to genotyping errors. The ERSA (Estimation of Recent Shared Ancestry) algorithm proposes a maximum likelihood estimation method by establishing a statistical model of the number of IBD fragments and the number of common ancestral generations (Huff et al., 2011). It can accurately infer second to fourth degree kinship, and its advantage is that it can perform relationship inference without pre-setting specific kinship types (Smith et al., 2022). For large-scale databases, RaPID (Naseri et al., 2019) and its extended version RaPID-Query (Wei et al., 2023) achieve linear scaling of IBD detection by introducing the PBWT (Positional Burrows-WheelerTransform) index (Durbin, 2014) and a random projection strategy, significantly improving computational efficiency. RaPID-Query supports fast IBD queries between external individuals and large sample databases, with an average query time of only about 2.76 seconds. Furthermore, FISHR focuses on improving the accuracy of IBD segment endpoint detection, performing best in the detection of medium-to-long fragments (IBD fragment length > 3cM) and exhibiting higher computational speed compared to Refined IBD and HaploScore (Bjelland et al., 2017). IBIS further breaks through the dependence on phase information, enabling rapid and accurate detection of long-fragment IBD in genotype data without phase information, making it suitable for use in databases with over a million genomic samples (Seidman et al., 2020). In terms of algorithm fusion, RAFFI constructs a kinship inference model based on RaPID results. By correcting IBD segment quality bias, it effectively improves the distant relative identification rate, demonstrating higher accuracy and stability than KING on real datasets such as UK Biobank (Naseri et al., 2021). For low-depth sequencing data, LocalNgsRelate introduces genotype similarity but not fixed genotype values, enabling reliable IBD segment identification even under low coverage conditions (Severson et al., 2022).

[0011] Regarding kinship inference, current research generally agrees that the IBD method can reliably infer kinship relationships up to six generations or even further. Taking ERSA as an example, Huff et al. (2011) demonstrated in empirical data from three major human pedigrees that this method achieved an accuracy of 97% in first to fifth generation kinship, and even for sixth and seventh generation relatives, the accuracy remained within the one-generation error range of 80%. The Rapid-Query method, using IBD segments of 7 cM or larger on UK Biobank data, can accurately distinguish kinship levels up to four generations, with an AUC value as high as 97.28% on its ROC curve. Comprehensive method evaluation studies also support the superiority of the IBD method in multi-level kinship relationships. Ramstetter et al. (2017) systematically evaluated 12 mainstream kinship inference methods, finding that most IBD detection methods (such as GERMLINE, Refined IBD combined with ERSA) achieved accuracy rates of 92%-99% in first and second-degree relationships, but for seventh-degree relatives, the accuracy generally dropped to below 43%. However, even with inference distortion, the IBD method can still accurately locate more than 76% of seventh-degree relatives within adjacent degree intervals, demonstrating its robustness in distant kinship identification. Furthermore, Liu et al. (2023) compared the IBS statistical method (KING) with the IBD fragment algorithm (GERMLINE+ERSA) in a Han Chinese pedigree study, finding that the latter was more accurate in identifying relatives within eight degrees, especially with a significantly lower misclassification rate (less than 2%), although a certain proportion of false negatives (<27.4%) existed.

[0012] Overall, the IBD method has become one of the main methods for inferring kinship, especially demonstrating outstanding practical value in areas such as distant relative identification, family structure reconstruction, and data privacy protection. Although the inference accuracy decreases in kinship relationships beyond six generations, combining multiple methods with high-quality data input can still significantly improve the stability of the inference.

[0013] (3) Likelihood Ratio Method: The likelihood ratio method is a classic statistical method widely used in kinship identification. Its core idea is to compare the genotype probability distributions under the assumption of kinship (such as father-son relationship) and the assumption of no kinship (such as random individuals). If the probability of the kinship assumption is significantly higher than the probability of no kinship, then the kinship is more likely to be established. The Ministry of Justice has issued the following standards, all of which use this method as the main identification method: the Technical Specifications for Paternity Identification (GB / T 37223-2018) (Ministry of Justice, 2018), the Technical Specifications for Biological Full Sibling Relationship Identification (GB / T 43641-2024) (Ministry of Justice of the People's Republic of China, 2024), and the Specifications for Biological Grandparent-Grandchild Relationship Identification (SF / Z JD0105005-2015) (Ministry of Justice of the People's Republic of China, 2015).

[0014] Currently, the mainstream phylogenetic index (LR) calculation method in China is the ITO method (Lu Huiling and Yang Qing'en, 2002; Lu Huiling et al., 2009). This method calculates the kinship index based on the probability of two individuals sharing 0, 1, or 2 IBS alleles. The formula is simple and the calculation is fast (Li & Sacks, 1954). This method performs well in first- and second-degree relationships, but still has some errors in inferring more distant relationships. Other methods include family reconstruction (He et al., 2013), the Elston-Stewart algorithm (Elston & JA, 1971), and the Lander-Green algorithm (Lander & Green, 1987). These methods generally rely on specialized genetic statistical software to perform complex data processing and parameter estimation, such as EasyDNA, ForeStatistics, Familias, and Merlin. The LR method has a clear statistical basis and standardized decision thresholds, but it has several problems: First, in distant relatives, the proportion of shared alleles between individuals is low, and the LR value tends to be close to 1, making it difficult to form a significant judgment; Second, this method is usually based on the assumption of independent gene loci and fails to effectively utilize the linkage information between markers, thus affecting the accuracy of judgment on complex kinship structures.

[0015] While current IBD fragment-based kinship inference methods (such as ERSA, GERMLINE, and Rapid) can accurately infer kinship levels up to seven, these methods primarily focus on estimating the continuity of kinship ranks and classification tasks, suitable for population genetic structure analysis or genealogy construction. In contrast, the core need in forensic medicine and judicial practice is often to determine whether a specific kinship exists between two individuals—that is, "identification" rather than "inference." Existing methods lack a statistically significant and controllable hypothesis testing framework for this task, especially when dealing with distant kinship or high background noise, making it difficult to provide reliable and quantitative support for judgments.

[0016] Relevant patent documents retrieved: Publication country: China, Publication number: CN107609343A, Publication date: 2018-01-19. This document discloses a method, system, computer equipment, and readable storage medium for kinship identification, including the following steps: Step S1: Receiving an input instruction for the type of kinship identification; if it is a two-person kinship identification, proceed to step S21; if it is a three-person kinship identification, proceed to step S22; Step S21: Obtaining the two-person STR typing data of the child and the hypothetical parent, proceed to step S3; Step S22: Obtaining the three-person STR typing data of the child, mother, and hypothetical father, proceed to step S4; Step S3: Calculating the PI value of each STR, proceed to step S5; Step S4: Calculating the PI value of each STR, proceed to step S5; Step S5: Multiplying the PI values ​​of all STRs to obtain the total kinship index. This kinship identification method is based on STR and PI for kinship identification.

[0017] Relevant non-patent literature retrieved: Journal or book title: PLOS Genetics, Article title: Relationship Estimation from Whole-Genome Sequence Data, Volume: January 2014, Volume 10, Issue 1, e1004144, Publication date: January 30, 2014. This article discloses association estimation based on whole-genome sequence data, utilizing homology information (IBD) generated from high-density single nucleotide polymorphism (SNP) data, which can significantly improve the efficiency and accuracy of genetic relationship detection.

[0018] The prior art represented by the aforementioned documents has at least the following unresolved technical problems or defects: Most second-degree related individuals are difficult to distinguish from unrelated individuals, especially cousins, whose IBS scores significantly overlap with those of unrelated individuals, leading to an increased risk of misclassification. Although the IBS method is simple to calculate and provides intuitive data, it has poor universality and is not suitable for low-copy-number DNA and mixed samples. Relevant evidence can be found in the study by Li et al. (Tao et al., 2022a).

[0019] Currently, the mainstream phylogenetic index (LR) calculation method in China is the ITO method (Lu Huiling and Yang Qing'en, 2002; Lu Huiling et al., 2009). This method calculates the kinship index based on the probability of two individuals sharing 0, 1, or 2 IBS alleles. The formula is simple and the calculation is fast (Li & Sacks, 1954). This method performs well in first- and second-degree relationships, but it still has some errors in inferring more distant relationships. Relevant evidence is provided by: Lu Huiling and Yang Qing'en, 2002; Lu Huiling et al., 2009.

[0020] In solving the above problems or overcoming the above defects, the present invention has encountered the following difficulties and obstacles: how to break through the limitations of existing IBS scoring and preset threshold comparison strategies in terms of application scope and discrimination ability. Summary of the Invention

[0021] The purpose of this invention is to provide a method for kinship identification based on the common ancestor fragment IBD, and related technologies, to solve one or a combination of the technical problems of existing identification methods, such as the significant decrease in the distinguishing ability, the increased risk of misjudgment, and the reduced accuracy of existing identification methods as the kinship level increases (e.g., level three and above).

[0022] Terminology Explanation: Unless otherwise defined, all technical terms in this document have the same meanings as commonly understood by one of ordinary skill in the art to which the subject matter of the claims pertains. Unless otherwise stated, all patents, patent inventions, and publications cited in this document are incorporated herein by reference in their entirety. If multiple definitions exist for terms in this document, the definitions in this chapter shall prevail.

[0023] It should be understood that the above brief description and the following detailed description are exemplary and for illustrative purposes only, and do not limit the subject matter of the invention in any way. In this invention, the singular is used in conjunction with the plural unless otherwise specifically stated. It should also be noted that, unless otherwise stated, the use of “or” or “or” means “and / or”. Furthermore, the use of the term “comprising” and other forms such as “including,” “containing,” and “contains” are not limiting.

[0024] Definitions of standard chemical terms can be found in the references “Technical Specification for Parentage Testing GB-T 37223-2018”, “Biological Grandparent-Grandchild Relationship Testing Specification SF-Z JD0105005-2015”, and “Biological Full Sibling Relationship Testing Specification GB / T 43641-2024”.

[0025] Unless specifically defined herein, the use of all commercially available products herein employs standard techniques. For example, it may be carried out using the manufacturer's instructions for use with the kit, or in accordance with methods known in the art or the description of this invention. The techniques and methods described herein can generally be implemented according to conventional methods well known in the art, based on the descriptions in the various summary and more specific documents cited and discussed in this specification.

[0026] The term "parentage testing" as used in this article refers to the activity of determining the parent-child relationship between individuals by detecting human genetic markers and analyzing genetic laws.

[0027] The term “tripartite paternity test” used in this article refers to a paternity test involving the male being tested, the child’s biological mother, and the child, or the female being tested, the child’s biological father, and the child.

[0028] The term “dual parentage testing” used in this article refers to parentage testing between the tested male and child or between the tested female and child.

[0029] The term “STR” used in this article refers to short tandem repeat sequences.

[0030] The term "full sibling" as used in this article refers to multiple offspring individuals who share the same biological father and mother.

[0031] The term "system effectiveness" as used in this article refers to the likelihood that a clear conclusion can be reached when using a given detection system and corresponding criteria for biological kinship identification.

[0032] The term "IBS" used in this article refers to State Consistency Score.

[0033] The term "CGI" used in this article refers to the Cumulative Grandparent-Grandchild Relationship Index.

[0034] The term "CPI" used in this article refers to the Cumulative Parental Relationship Index.

[0035] The term "CFSI" used in this article refers to the Cumulative Brother-Sister Relationship Index.

[0036] In a first aspect, the present invention provides a method for kinship identification based on the common ancestor fragment IBD, comprising: Detect IBD fragments between two disputed individuals, and denote the set of all detected IBD fragments as s, where s = {i1, i2, i3, ..., i n}, where i represents the length of the detected IBD fragment in centimoles, and the set s contains n elements; A likelihood ratio testing framework is constructed. Regarding the existence of kinship, the null hypothesis is that the two disputed individuals are unrelated individuals, and the alternative hypothesis is that there is a specific kinship between the two disputed individuals. This paper models the probability of IBD fragments occurring under the null hypothesis in the likelihood ratio test framework, and calculates the probability of IBD fragments occurring under the null hypothesis. Under the null hypothesis, all fragments in the IBD fragment set come from the population background, denoted as s. P ={i1,i2,i3,…,i np}, n P s represents the number of elements in the IBD fragment set originating from the group background. P =s,n P =n; This paper models the probability of IBD fragments occurring under the alternative hypothesis in the likelihood ratio test framework, and calculates the probability of IBD fragments occurring under the alternative hypothesis. Under the alternative hypothesis, a portion of the IBD fragment set originates from the population background, denoted as s. P Part of it comes from a common ancestor, denoted as s A ;s P and s A Let s be two mutually exclusive subsets of s. A The number of elements in is n A Set set s P The number of elements in is n P And n A +n P =n; The likelihood ratio is calculated by combining the probabilities of IBD fragment data under both the null and alternative hypotheses.

[0037] Furthermore, the likelihood ratio testing framework is as follows:

[0038] Where LR represents the likelihood ratio, Pr represents the probability, E represents the IBD fragment data detected between disputed individuals, H1 represents the hypothesis that there is a specific kinship between the two disputed individuals, H1 is the alternative hypothesis; H0 represents the hypothesis that the two disputed individuals are unrelated individuals, H0 is the null hypothesis. Where Pr(E|H1) represents the probability of the IBD segment data appearing under the H1 hypothesis, and Pr(E|H0) represents the probability of the IBD segment data appearing under the H0 hypothesis.

[0039] Furthermore, the probability of IBD fragment data occurring under the null hypothesis is modeled, and the established model is as follows: , Where t is the minimum detection threshold for IBD fragment length; This represents the number of IBD fragments n under the H0 hypothesis, given the minimum detection threshold t. P and length set s P The joint probability; This indicates that, under the H0 hypothesis, given a minimum detection threshold t, the individuals in dispute share n. P The probability of an IBD fragment; This indicates that, under the H0 hypothesis, given a minimum detection threshold t, the individuals involved in the dispute share a set of IBD fragments s. P The probability of.

[0040] Furthermore, ,

[0041] Where λ represents the average number of IBD fragments with a length exceeding the minimum detection threshold t shared among all unrelated individuals in the population; θ represents the probability that the IBD fragment length is i given a minimum detection threshold t; θ represents the average length of IBD fragments exceeding the minimum detection threshold t among all unrelated pairs of individuals in the population.

[0042] Furthermore, when there is a specific kinship between two disputed individuals, the probability of their IBD fragments appearing is defined as the product of the probability that the IBD fragments originate from the most recent common ancestor and the probability that they originate from the group background.

[0043] Furthermore, the probability of IBD fragment data occurring under the alternative hypothesis is modeled, and the established model is as follows: , This represents the number of IBD fragments n under the H1 hypothesis, given the minimum detection threshold t. A and length set s A The joint probability; This represents the number of IBD fragments n under the H1 hypothesis, given the minimum detection threshold t. p and length set s p The joint probability.

[0044] In the above technical solution, and The calculation method is the same, as shown in the following technical solution.

[0045] Furthermore, ,in,

[0046]

[0047] ,in, ,

[0048] Where: k is the chromosome number; n A k λ represents the number of IBD segments of length greater than t and originating from a common ancestor among individuals on chromosome k in the disputed sequence. k This represents the average number of shared IBD segments of length greater than t among all individuals with a specific kinship in chromosome k; Let j be the set of IBD segments of length greater than t among the individuals in the dispute on chromosome k that originate from a common ancestor; j is the set The length of an IBD segment; θ k t represents the average length of IBD segments with a length greater than t among all individuals with a specific kinship on chromosome k.

[0049] Furthermore, the detected IBD fragments are sorted in ascending order of length, assuming the first n P One line originates from the group's background, while the rest stem from a common ancestor; traversing n... P Calculate the probability from 0 to the total number. The maximum value is taken as the probability under the alternative hypothesis.

[0050] Furthermore, by combining the probabilities of IBD fragment data occurring under both the null and alternative hypotheses, the likelihood ratio is calculated as follows:

[0051] Substitute the probabilities under the null and alternative hypotheses into the likelihood ratio calculation formula; if the likelihood ratio is >1, it is determined that there is a hypothetical kinship; if the likelihood ratio is <1, it is determined that the individuals are unrelated.

[0052] Secondly, This invention provides a kinship identification system based on the common ancestor fragment IBD, used to perform the above-described kinship identification method. The kinship identification system includes: Data reading and format parsing module: used to receive and parse standardized .txt format IBD test result files; IBD Feature Modeling Module: Models the number and length of IBD fragments separately, using Poisson and exponential distributions for fitting; Likelihood Ratio Calculation and Relationship Determination Module: Calculates the likelihood ratio of each IBD segment under two hypotheses: "existence of a specific kinship" and "unrelated individuals"; and makes an overall determination based on the cumulative LR value of the 22 autosomes: cumulative LR > 1: determined to have a specific kinship; cumulative LR < 1: determined to be an unrelated individual. Results output and visualization module: Automatically generates reports that display the LR value, cumulative LR value and final judgment result for each IBD segment.

[0053] The present invention has at least the following beneficial effects: Compared with existing technologies, this invention has better technical performance in terms of accuracy in identifying kinship at the third level and above. According to data simulation tests, the traditional STR-based identification method has an efficiency of 66.62% to 99.995% for identifying first- and second-degree kinship, while this invention improves it to 100%. Furthermore, the efficiency for third- and fourth-degree kinship identification remains 100%, and the efficiency for fifth-degree kinship identification is 99.85%, thus improving the accuracy of kinship identification.

[0054] This invention proposes a novel method for kinship identification based on IBD fragment features. This method innovatively introduces a likelihood ratio statistical framework, jointly modeling the number and length of IBD fragments to establish a significance test system for determining whether a specific kinship exists. This method overcomes the limitations of existing IBS scoring and preset threshold comparison strategies in terms of application scope and discriminative ability. It can adapt to multi-level kinship identification tasks from level one to level five, and is particularly suitable for practical scenarios such as forensic identification where a clear determination of "whether a specific kinship exists" is required. It possesses significant methodological innovation and practical value.

[0055] This invention establishes an efficient, stable, and scalable kinship identification model, widely applicable to various judicial and social service applications such as distant relative identification, historical figure identification, identification of victims of major disasters, and search for missing persons. Through systematic modeling and statistical quantification of IBD fragments across the entire genome, this method significantly improves the identification resolution at the third degree of kinship and above, effectively filling the capability gap of traditional STR methods in such tasks.

[0056] This invention proposes a novel method for kinship identification based on IBD fragment features, representing a paradigm shift in kinship identification technology from "limited genetic locus comparison" to "whole-genome big data modeling." This lays the theoretical and technical foundation for expanding the kinship identification system in terms of accuracy, breadth, and universality. The quantifiable and reusable statistical model it constructs not only possesses strong scalability and adaptability but also provides solid support for subsequent case analysis, judicial practice, population evolution, and other research and practical fields, demonstrating significant theoretical value and promising practical applications.

[0057] The kinship identification system of this invention combines genome-wide IBD fragment data with a likelihood ratio statistical model to achieve high-precision determination of complex kinship relationships from first to fifth degree. The system is applicable to multiple scenarios such as forensic laboratories, genealogical research, and historical individual identification. Attached Figure Description

[0058] Figure 1 The cumulative paternity index density distribution map of triplets based on the mandatory loci tested by the Department of Justice, with STR loci number = 19; Figure 2 The cumulative paternity index density distribution of triplets based on the GlobalFiler™ PCR amplification kit, with STR loci number = 21; Figure 3 The cumulative paternity index density distribution of triplet based on the VeriFiler™ Plus PCR amplification kit, with STR loci number = 23; Figure 4 The cumulative paternity index density distribution map of duos based on the mandatory loci tested by the Department of Justice, with STR loci number = 19; Figure 5 The cumulative paternity index density distribution of the duozones based on the GlobalFiler™ PCR amplification kit, with STR loci number = 21; Figure 6 The cumulative paternity index density distribution of the duozones based on the VeriFiler™ Plus PCR amplification kit, with STR loci number = 23; Figure 7 The distribution map of the cumulative grandparent-grandchild relationship index based on the mandatory gene loci examined by the Ministry of Justice is shown. The number of STR loci is 19. Figure 8 The cumulative grandparent-grandchild relationship index density distribution map based on the GlobalFiler™ PCR amplification kit, with STR loci number = 21; Figure 9 The density distribution of the cumulative grandparent-grandchild relationship index based on the VeriFiler™ Plus PCR amplification kit, with STR loci number = 23; Figure 10 The distribution map of the cumulative full-sibling relationship index density based on the Ministry of Justice's mandatory loci, with STR loci number = 19; Figure 11 This is a density distribution map of the cumulative full-sib relationship index of full siblings based on the GlobalFiler™ PCR amplification kit, with STR loci number = 21; Figure 12 The image shows the density distribution of the cumulative full-sib relationship index of full siblings based on the VeriFiler™ Plus PCR amplification kit, with STR loci number = 23. Figure 13 This is a distribution diagram of the LR calculation simulation results based on whole-genome parent-child kinship obtained by the method of the present invention; Figure 14 This is a distribution diagram of the LR calculation simulation results based on whole-genome full-sibling kinship obtained by the method of this invention; Figure 15The distribution diagram shows the LR calculation simulation results based on whole-genome half-sibling kinship obtained by the method of this invention; Figure 16 This is a distribution diagram of the LR calculation simulation results based on whole-genome uncle-nephew kinship obtained by the method of this invention; Figure 17 This is a distribution diagram of the LR calculation simulation results based on whole-genome grandparent-grandchild kinship obtained by the method of this invention; Figure 18 This is a distribution diagram of the LR calculation simulation results based on whole-genome first cousin kinship obtained by the method of this invention; Figure 19 This is a distribution diagram of the LR calculation simulation results based on the whole genome first-generation first cousin skip-generation kinship obtained by the method of this invention; Figure 20 This is a distribution diagram of the LR calculation simulation results based on whole-genome second-generation cousin kinship obtained by the method of this invention; Figure 21 The distribution diagram shows the LR calculation simulation results based on 800,000 SNPs of parent-child kinship obtained by the method of this invention; Figure 22 The distribution diagram shows the LR calculation simulation results based on the full sibling kinship of 800,000 SNPs obtained by the method of this invention; Figure 23 The distribution diagram shows the LR calculation simulation results based on 800,000 SNP half-sib relationships obtained by the method of this invention. Figure 24 The distribution diagram shows the LR calculation simulation results based on 800,000 SNP uncle-nephew kinship obtained by the method of this invention; Figure 25 The distribution diagram shows the LR calculation simulation results based on 800,000 SNP grandparent-grandchild kinship obtained by the method of this invention; Figure 26 The distribution diagram shows the LR calculation simulation results based on the kinship of 800,000 SNPs of first cousins ​​obtained by the method of this invention. Figure 27 The distribution diagram shows the LR calculation simulation results based on the one-generation skip-generation kinship of 800,000 SNPs for first-generation first cousins ​​obtained by the method of this invention; Figure 28 This is a distribution diagram of the LR calculation simulation results based on the kinship of 800,000 SNPs of second-generation cousins ​​obtained by the method of this invention. Detailed Implementation

[0059] The following non-limiting embodiments are intended to enable those skilled in the art to gain a more comprehensive understanding of the present invention, but do not limit the invention in any way. The following content is merely an exemplary description of the scope of protection claimed by the present invention, and those skilled in the art can make various changes and modifications to the present invention based on the disclosed content, and such changes should also fall within the scope of protection claimed by the present invention.

[0060] The present invention will be further described below by way of specific embodiments. Unless otherwise specified, all instruments, devices, equipment, reagents, products, etc., used in the embodiments of the present invention are obtained through conventional commercial means.

[0061] Example 1 This invention provides a method for kinship identification based on the common ancestor fragment IBD, comprising the following steps: S1. Detect IBD fragments between two disputed individuals, and denote the set of all detected IBD fragments as s, s={i1,i2,i3,…,i n}, where i represents the length of the detected IBD fragment in centimoles, and the set s contains n elements; S2. Construct a likelihood ratio test framework. Regarding whether a kinship exists, take the null hypothesis that the two disputed individuals are unrelated individuals and the alternative hypothesis that a specific kinship exists between the two disputed individuals. The likelihood ratio test framework in step S2 is as follows:

[0062] Where LR represents the likelihood ratio, Pr represents the probability, E represents the IBD fragment data detected between disputed individuals, H1 represents the hypothesis that there is a specific kinship between the two disputed individuals, H1 is the alternative hypothesis; H0 represents the hypothesis that the two disputed individuals are unrelated individuals, H0 is the null hypothesis. Where Pr(E|H1) represents the probability of the IBD segment data appearing under the H1 hypothesis, and Pr(E|H0) represents the probability of the IBD segment data appearing under the H0 hypothesis.

[0063] S3. Model the probability of IBD fragment data occurring under the null hypothesis in the likelihood ratio test framework, and calculate the probability of IBD fragment data occurring under the null hypothesis; under the null hypothesis, all fragments in the IBD fragment set come from the population background, denoted as s. P ={i1,i2,i3,…,i np}, n P s represents the number of elements in the IBD fragment set originating from the group background. P =s,n P =n; In step S3, the probability of IBD fragment data occurring under the null hypothesis is modeled, and the established model is as follows: , Where t is the minimum detection threshold for IBD fragment length; This represents the number of IBD fragments n under the H0 hypothesis, given the minimum detection threshold t. P and length set s P The joint probability; This indicates that, under the H0 hypothesis, given a minimum detection threshold t, the individuals in dispute share n. P The probability of an IBD fragment; This indicates that, under the H0 hypothesis, given a minimum detection threshold t, the individuals involved in the dispute share a set of IBD fragments s. P The probability of.

[0064] Assuming the number of IBD fragments follows a Poisson distribution, then under the H0 assumption, the individuals involved in the dispute share n. P The probability of an IBD segment, i.e., the number of segments shared by unrelated individuals. P The probability of an IBD fragment is: , Where λ represents the average number of IBD fragments with a length exceeding the minimum detection threshold t shared among all unrelated individuals in the population;

[0065] in This represents the probability that the IBD fragment length is i given a minimum detection threshold t; Assuming the length of an IBD segment follows an exponential distribution, the probability that the length of an IBD segment is i is: , θ represents the average length of IBD fragments exceeding the minimum detection threshold t among all irrelevant pairs of individuals in the population.

[0066] S4. Model the probability of IBD fragment data occurring under the alternative hypothesis in the likelihood ratio test framework, and calculate the probability of IBD fragment data occurring under the alternative hypothesis; under the alternative hypothesis, the fragment portion in the IBD fragment set comes from the population background, denoted as s. P Part of it comes from a common ancestor, denoted as s A ;s P and s A Let s be two mutually exclusive subsets of s. A The number of elements in is n A, Set set s P The number of elements in is n P And n A +nP =n; In step S4, In alternative hypothesis H1, it is assumed that if two individuals are related, they share at least one most recent common ancestor. Therefore, the IBD fragments shared between the two individuals primarily originate from their most recent common ancestor, with a portion also derived from their group background. Based on this, when a specific kinship exists between two disputed individuals, the probability of their IBD fragments appearing is defined as the product of the probability that the IBD fragment originates from their most recent common ancestor and the probability that it originates from their group background.

[0067] The probability of IBD fragments occurring under the alternative hypothesis is modeled, and the model established is as follows: , This represents the number of IBD fragments n under the H1 hypothesis, given the minimum detection threshold t. A and length set s A The joint probability; This represents the number of IBD fragments n under the H1 hypothesis, given the minimum detection threshold t. p and length set s p The joint probability.

[0068] The calculation method and same.

[0069] The first factor, namely the probability that the IBD fragment is inherited from the most recent common ancestor, can be expressed as: ,in,

[0070]

[0071] Where: k is the chromosome number; n A k λ represents the number of IBD segments of length greater than t and originating from a common ancestor among individuals on chromosome k in the disputed sequence. k This represents the average number of shared IBD segments of length greater than t among all individuals with a specific kinship in chromosome k; Let j be the set of IBD segments of length greater than t among the individuals in the dispute on chromosome k that originate from a common ancestor; j is the set The length of an IBD segment; θ k t represents the average length of IBD segments with a length greater than t among all individuals with a specific kinship on chromosome k.

[0072] The second factor, the probability that the IBD fragment originates from the population background, can be expressed as: ,in, ,

[0073] Furthermore, the detected IBD fragments are sorted in ascending order of length, assuming the first n P One line originates from the group's background, while the rest stem from a common ancestor; traversing n... P Calculate the probability from 0 to the total number. The maximum value is taken as the probability under the alternative hypothesis.

[0074] S5. Calculate the likelihood ratio by combining the probabilities of IBD fragment data under the null and alternative hypotheses.

[0075] The likelihood ratio is calculated by combining the probabilities of IBD fragments occurring under both the null and alternative hypotheses, as follows:

[0076] Substitute the probabilities under the null and alternative hypotheses into the likelihood ratio calculation formula; if the likelihood ratio is >1, it is determined that there is a hypothetical kinship; if the likelihood ratio is <1, it is determined that the individuals are unrelated.

[0077] The following compares the systematic evaluation results obtained by the likelihood ratio method based on short tandem repeat sequences obtained by the traditional identification methods for duos, triplets, grandparents and grandchildren, and full siblings provided by the technical specifications of the Ministry of Justice, with the systematic evaluation results obtained by the kinship identification method of the present invention.

[0078] (1) Traditional identification methods for twins, triplets, grandparents and grandchildren and full siblings provided by the Ministry of Justice’s identification technical specifications: As shown in Table 1, the systematic evaluation results of the likelihood ratio method based on short tandem repeat sequences obtained by the identification methods for twins, triplets, grandparents and grandchildren and full siblings provided by the Ministry of Justice’s identification technical specifications show that the systematic effectiveness of the methods in the Ministry of Justice’s identification technical specifications can only reach 66.62% to 99.995%.

[0079] Table 1. Systematic evaluation of the likelihood ratio method based on short tandem repeat sequences.

[0080] Note: In Table 1: Threshold 1 – Threshold for identifying a presumed kinship type; Threshold 2 – Threshold for identifying unrelated individuals; Sensitivity – The proportion of correctly identified kinship individuals; Specificity – The proportion of correctly identified unrelated individuals; False Positive – The proportion of unrelated individuals identified as kinship individuals; False Negative – The proportion of kinship individuals identified as unrelated individuals; Undetermined – The proportion of results between Threshold 1 and Threshold 2 that cannot be definitively concluded; System Efficiency – The probability of providing a definitive conclusion when using a given detection system and corresponding criteria for biological kinship identification.

[0081] Figures 1-3 This indicates the results of kinship and identification of unrelated individuals in the trio. Figures 4-6 The image shows the results of kinship and identification of unrelated individuals in the duo. Figures 7-9 The image shows the kinship test results for grandparent-grandchild relationships and unrelated individuals. Figures 10-12 As shown, the results of kinship and unrelated individual identification for full sibling relationships are presented. Each family lineage includes the identification results for the number of STR loci in three categories. Figures 1-12 The middle red line represents the identification threshold; the left red line represents threshold 1 in Table 1, and the right red line represents threshold 2 in Table 1. Figures 1-12 The arrows pointing left correspond to the specificity in Table 1, and the arrows pointing right correspond to the sensitivity. It can be seen that the identification effect for full siblings is very poor. Table 1 and... Figures 1-12 Correspondingly, the following are... Figures 1-12 The coordinates and curves in the diagram will be explained in detail.

[0082] Figures 1-6 In the graph, the horizontal axis represents the logarithm of the cumulative parentage index (base 10), and the vertical axis represents the corresponding probability density value. Figures 1-6 The curve in the middle is a distribution curve fitted based on the kernel density estimation method. The two red vertical dashed lines represent the thresholds for parentage identification: the left dashed line corresponds to a threshold of -4, and the right dashed line corresponds to a threshold of 4. The red arrow pointing to the left indicates that the identification result is an unrelated individual, and the red arrow pointing to the right indicates that the identification result is a parentage.

[0083] Figures 7-9 In the graph, the horizontal axis represents the logarithmic value of the cumulative grandparent-grandchild relationship index (CGI), base 10; the vertical axis represents the corresponding probability density value. Figures 7-9 The middle curve is a distribution curve fitted based on the kernel density estimation method. The two red vertical dashed lines represent the thresholds for identifying grandparent-grandchild relationships: the left dashed line corresponds to a threshold of -4, and the right dashed line corresponds to a threshold of 4. The red arrow pointing to the left indicates an irrelevant individual, while the red arrow pointing to the right indicates a grandparent-grandchild relationship.

[0084] Figures 10-12In the graph, the horizontal axis represents the logarithm of the Cumulative Sibling Relationship Index (CFSI), base 10; the vertical axis represents the corresponding probability density value. Figures 10-12 The curve in the middle is a distribution curve fitted based on the kernel density estimation method. The two red vertical dashed lines represent the thresholds for identifying full sibling relationships: the left dashed line corresponds to a threshold of -4, and the right dashed line corresponds to a threshold of 4. The red arrow pointing to the left indicates that the identification result is an unrelated individual, and the red arrow pointing to the right indicates that the identification result is a full sibling relationship.

[0085] (2) The kinship identification method of the present invention: Table 2. Systematic evaluation of the method of the present invention in identifying first to fifth degree of kinship.

[0086] Note: Error rate – the proportion of misjudgments made by related or unrelated individuals.

[0087] Table 2 shows the systematic evaluation results of first to fifth degree kinship identification obtained using the invented kinship identification method based on whole-genome data (80 million SNPs), and... Figures 13-20 Correspondingly. Figures 13-20 Specificity is defined as LR greater than 1 (log10LR > 0) for related individuals (corresponding to blue bars) and less than 1 for irrelevant individuals (corresponding to red bars). The following... Figures 13-20 To explain: Figures 13-20 This is a distribution map of LR calculation simulation results for different kinship relationships across the entire genome (approximately 80 million SNPs) obtained based on the kinship identification method of this invention. The results are obtained from the simulation data of different kinship relationships across the entire genome using the kinship identification method of this invention, where each kinship and unrelated individual pair contains 1000 family data sets. According to Table 2 and... Figures 13-20 It can be seen that the present invention has a sensitivity, specificity, and systematic efficacy of 100% for identifying kinship levels 1-4, and a systematic efficacy of 99.85% for level 5. According to Table 1 and... Figures 1-12 It can be seen that while other IBD methods can identify or infer kinship at levels up to six or seven, their systematic efficacy for kinship at levels four and below is not 100%. According to Table 1, and Figures 1-12 It can be seen that the methodological system efficiency in the Ministry of Justice's technical specifications for forensic identification can only reach 66.62% to 99.995%. However, according to Table 2 and... Figures 13-20 It can be seen that the technology of this invention can completely distinguish between siblings and unrelated individuals. This invention exhibits extremely high robustness in identifying genomic data of 800,000 SNPs: as shown in Table 3. Table 3. Systematic assessment of kinship based on 800,000 SNP data.

[0088] Table 3 shows the systematic evaluation results of first to fifth degree kinship identification obtained using the invented kinship identification method based on 800,000 SNP data. Figures 21-28 Correspondingly.

[0089] Figures 21-28 This is a distribution map of LR calculation simulation results based on 800,000 SNPs with different kinship relationships, obtained using the kinship identification method of this invention. Figures 21-28 Even if the number of loci is reduced by 100 times relative to the whole genome, the technology of this invention can still maintain 100% system efficiency. The identification results of 800,000 SNPs with different kinship simulation data obtained by using the kinship identification method of this invention show that the sensitivity, specificity and system efficiency of this technology for the identification of kinship within the fourth degree are still maintained at 100%, and the system efficiency of the fifth degree only drops to 99.52%. Figures 21-28 The distribution map of LR simulation results based on 800,000 SNPs with different kinship relationships not only theoretically verifies the robustness of IBD fragments in kinship identification, but also practically demonstrates the technology's good platform adaptability and cost-effectiveness. Compared to whole-genome sequencing data, medium-density SNP chips have advantages in data acquisition cost, detection cycle, and hardware dependence, making them more suitable for practical applications with high cost and efficiency requirements, such as forensic identification, disaster scene identification, and missing persons database comparison.

[0090] Furthermore, since the SNP chip platform has been widely deployed in multiple national population genome projects, its data structure, quality control processes, and technological maturity have all been standardized, greatly reducing the technical threshold and facilitating its implementation in a wider range of fields. The feasibility of the medium-density SNP scheme also lays the foundation for its promotion under diverse sample conditions. In some practical scenarios, such as historical remains, environmentally exposed samples, or mixed contaminated samples, it is often difficult to obtain high-quality WGS data, while medium-density, uniformly distributed SNP markers are more likely to obtain stable signals. The technology of this invention can still maintain high recognition efficiency under such conditions, demonstrating the method's tolerance for data quality and broad adaptability.

[0091] Example 2 This invention proposes a kinship identification system based on the common ancestor fragment IBD, used to execute the kinship identification method of claim 1. The kinship identification system includes: combining genome-wide IBD fragment data with a likelihood ratio framework to achieve high-precision determination of complex kinship relationships from first to fifth degree. The system has a complete graphical user interface (GUI), is easy to operate, and is suitable for multiple application scenarios such as forensic laboratories, genealogical research, and historical individual identification.

[0092] This kinship identification system includes the following modules: (1) User interface module: Provides users with interactive entry points, including kinship type selection, data import, task execution, result viewing and export.

[0093] The main components of the user interface module include a relation type selection box, a file path selection and display panel, execution buttons (calculate, clear, export), and a real-time log window (displaying analysis progress and prompts). It supports standardized workflow operations and one-click execution; all interface elements are extensible controls, supporting the insertion of new relation types in the future.

[0094] (2) Data reading and format parsing module: Input requirements: A standardized .txt format IBD test result file, which should at least include the sample number, haploid number, start and end physical locations, and genetic length (in cM).

[0095] Processing logic: Automatically detect whether the column format is valid, analyze whether it is a paired individual relationship input, establish a unique identifier for the sample pair and a mapping structure of associated segments, and construct a list of shared segment information for the sample pair to be analyzed.

[0096] Error handling: When the file path is empty or the format is incorrect, a pop-up window will remind the user to recheck the data source.

[0097] (3) IBD Feature Modeling Module: This module models the number and length of IBD fragments, using Poisson and exponential distributions for fitting. The distribution fitting parameters can be estimated from the training sample data or entered manually.

[0098] (4) Likelihood ratio calculation and relationship discrimination module: The discrimination mechanism is as follows: For each IBD segment, the likelihood ratio is calculated under two hypotheses: "assuming a certain kinship" and "unrelated individuals", and the overall judgment is achieved by the cumulative LR value of the 22 autosomes.

[0099] Judgment criterion: Cumulative LR > 1, i.e., log 10 If LR > 0, a hypothetical kinship is determined; if cumulative LR < 1, i.e., log 10 If LR < 0, the individual is considered irrelevant.

[0100] The judgment process is fully automated, built on statistical principles, and eliminates human intervention; it supports batch calculation and parallel evaluation, and is suitable for large-scale relationship comparison and analysis needs.

[0101] (5) Results Output and Visualization Module: After the analysis is completed, a report is automatically generated, displaying the LR value, cumulative LR value and final judgment result of each IBD segment. The analysis results are displayed in a graphical window, and the report can be exported to PDF, TXT or CSV format. The report supports timestamps and batch numbers, and provides visualization tools to display kinship network diagrams, IBD segment distribution diagrams, etc., to facilitate user understanding and analysis.

[0102] This system runs on 64-bit Windows 7 / 10 operating systems. Hardware requirements include 16GB or more of RAM, 20GB or more of hard disk space, and a recommended CPU with at least 4 cores and a recommended CPU clock speed of 2.0GHz or higher. All computations are performed locally, making it compatible with general forensic laboratories and data workstations. For large-scale datasets, multi-core CPUs and sufficient memory resources are recommended to improve computational efficiency. The system design supports parallel computing, enabling efficient processing of large-scale genomic data in a multi-threaded environment.

[0103] Finally, it should be noted that the above content is only used to illustrate the technical solution of the present invention, and is not intended to limit the scope of protection of the present invention. Simple modifications or equivalent substitutions made by those skilled in the art to the technical solution of the present invention do not depart from the essence and scope of the technical solution of the present invention.

Claims

1. A method for kinship identification based on the common ancestor fragment IBD, characterized in that, include: Detect IBD fragments between two disputed individuals, and denote the set of all detected IBD fragments as s, where s = {i1, i2, i3, ..., i n }, where i represents the length of the detected IBD fragment in centimoles, and the set s contains n elements; A likelihood ratio testing framework is constructed. Regarding the existence of kinship, the null hypothesis is that the two disputed individuals are unrelated individuals, and the alternative hypothesis is that there is a specific kinship between the two disputed individuals. This paper models the probability of IBD fragments occurring under the null hypothesis in the likelihood ratio test framework, and calculates the probability of IBD fragments occurring under the null hypothesis. Under the null hypothesis, all fragments in the IBD fragment set come from the population background, denoted as s. P ={i1,i2,i3,…,i np }, n P s represents the number of elements in the IBD fragment set originating from the group background. P =s,n P =n; This paper models the probability of IBD fragments occurring under the alternative hypothesis in the likelihood ratio test framework, and calculates the probability of IBD fragments occurring under the alternative hypothesis. Under the alternative hypothesis, a portion of the IBD fragment set originates from the population background, denoted as s. P Part of it comes from a common ancestor, denoted as s A ;s P and s A Let s be two mutually exclusive subsets of s. A The number of elements in is n A Set set s P The number of elements in is n P And n A +n P =n; The likelihood ratio is calculated by combining the probabilities of IBD fragment data under both the null and alternative hypotheses.

2. The method for identifying kinship according to claim 1, characterized in that, The likelihood ratio test framework is as follows: Where LR represents the likelihood ratio, Pr represents the probability, E represents the IBD fragment data detected between disputed individuals, H1 represents the hypothesis that there is a specific kinship between the two disputed individuals, H1 is the alternative hypothesis; H0 represents the hypothesis that the two disputed individuals are unrelated individuals, H0 is the null hypothesis. Where Pr(E|H1) represents the probability of the IBD segment data appearing under the H1 hypothesis, and Pr(E|H0) represents the probability of the IBD segment data appearing under the H0 hypothesis.

3. The method for identifying kinship according to claim 2, characterized in that, To model the probability of IBD fragment data occurring under the null hypothesis, the following model was established: , Where t is the minimum detection threshold for IBD fragment length; This represents the number of IBD fragments n under the H0 hypothesis, given the minimum detection threshold t. P and length set s P The joint probability; This indicates that, under the H0 hypothesis, given a minimum detection threshold t, the individuals in dispute share n. P The probability of an IBD fragment; This indicates that, under the H0 hypothesis, given a minimum detection threshold t, the individuals involved in the dispute share a set of IBD fragments s. P The probability of.

4. The method for identifying kinship according to claim 3, characterized in that: , Where λ represents the average number of IBD fragments with a length exceeding the minimum detection threshold t shared among all unrelated individuals in the population; θ represents the probability that the IBD fragment length is i given a minimum detection threshold t; θ represents the average length of IBD fragments exceeding the minimum detection threshold t among all unrelated pairs of individuals in the population.

5. The method for identifying kinship according to claim 4, characterized in that, When there is a specific kinship between two disputed individuals, the probability of their IBD fragments appearing is defined as the product of the probability that the IBD fragments originate from the most recent common ancestor and the probability that they originate from the group background.

6. The method for identifying kinship according to claim 5, characterized in that, To model the probability of IBD fragment data occurring under the alternative hypothesis, the following model was established: , This represents the number of IBD fragments n under the H1 hypothesis, given the minimum detection threshold t. A and length set s A The joint probability; This represents the number of IBD fragments n under the H1 hypothesis, given the minimum detection threshold t. p and length set s p The joint probability.

7. The method for identifying kinship according to claim 6, characterized in that, ,in, ,in, , Where: k is the chromosome number; n A k λ represents the number of IBD segments of length greater than t and originating from a common ancestor among individuals on chromosome k in the disputed sequence. k This represents the average number of shared IBD segments of length greater than t among all individuals with a specific kinship in chromosome k; Let j be the set of IBD segments of length greater than t among the individuals in the dispute on chromosome k that originate from a common ancestor; j is the set The length of an IBD segment; θ k t represents the average length of IBD segments with a length greater than t among all individuals with a specific kinship on chromosome k.

8. The method for identifying kinship according to claim 7, characterized in that, The detected IBD fragments are sorted in ascending order of length. Assume the first n... P One line originates from the group's background, while the rest stem from a common ancestor; traversing n... P Calculate the probability from 0 to the total number. The maximum value is taken as the probability under the alternative hypothesis.

9. The method for identifying kinship according to claim 7, characterized in that, The likelihood ratio is calculated by combining the probabilities of IBD fragments occurring under both the null and alternative hypotheses, as follows: Substitute the probabilities under the null and alternative hypotheses into the likelihood ratio calculation formula; if the likelihood ratio is >1, it is determined that there is a hypothetical kinship; if the likelihood ratio is <1, it is determined that the individuals are unrelated.

10. A kinship identification system based on the common ancestor fragment IBD, characterized in that, For performing the kinship identification method according to claims 1-9, the kinship identification system comprises: Data reading and format parsing module: used to receive and parse standardized .txt format IBD test result files; IBD Feature Modeling Module: Models the number and length of IBD fragments separately, using Poisson and exponential distributions for fitting; Likelihood Ratio Calculation and Relationship Determination Module: Calculates the likelihood ratio of each IBD segment under two hypotheses: "existence of a specific kinship" and "unrelated individuals"; and makes an overall determination based on the cumulative LR value of the 22 autosomes: cumulative LR > 1: determined to have a specific kinship; cumulative LR < 1: determined to be an unrelated individual. Results output and visualization module: Automatically generates reports that display the LR value, cumulative LR value and final judgment result for each IBD segment.

Citation Information

Patent Citations

  • Relative relationship identification method and system, computer device and readable storage medium

    CN107609343A

  • Genetic relationship identification method and device for three-individual combination

    CN107832586A

  • SNP (Single Nucleotide Polymorphism)-based genetic relationship identification method

    CN115565604A