Rapid inference method for genetic relationship based on co-progenitor fragment

By encoding the gene data of the individual to be analyzed in segments and determining the data fragment of the co-ancestral gene, the problem of low efficiency in determining the kinship level in the prior art is solved, and fast and efficient kinship inference is achieved.

CN120072046APending Publication Date: 2025-05-30INST OF FORENSIC SCI OF MIN OF PUBLIC SECURITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510129678.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-05
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art takes a long time and is inefficient when determining the level of kinship between individuals to be analyzed.

Method used

By obtaining the genetic data of the individual to be analyzed, performing segment co-ancestral gene data fragments are determined, and a kinship level is determined based on their length or proportion.

Benefits of technology

The rapid and efficient determination of the kinship level is achieved, the steps of homologous chromosome separation are avoided, the processing time is shortened, and the efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120072046A_ABST
    Figure CN120072046A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a genetic relationship rapid inference method based on co-progenitor fragments. The method comprises the following steps: acquiring a gene data set; segmenting the gene data of the to-be-analyzed individual to obtain a plurality of gene data fragments of the to-be-analyzed individual; encoding the gene data fragments of the to-be-analyzed individuals to obtain a first encoding set and a second encoding set of the gene data fragments of the to-be-analyzed individuals; determining a pair of co-progenitor gene data fragments corresponding to the to-be-analyzed individuals according to the first coding set and the second coding set of each gene data fragment in the to-be-analyzed individuals; and according to the co-progenitor gene data fragments corresponding to the to-be-analyzed individuals, determining genetic relationship levels corresponding to the to-be-analyzed individuals. The method is used for achieving the effect of improving the genetic relationship grade determination efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical fields of computational biology and genetics, and particularly to a method for rapidly inferring the genetic relationship based on shared ancestral segments. Background Art

[0002] In genetics, a gene is the basic unit of genetic information, and various traits and functions of an organism can be determined through genes. The genetic information of an organism can be transmitted between parents and offspring through genes. Therefore, the genetic relationship between organisms can be inferred through the genes of the organisms.

[0003] In some techniques, gene fragments with similarity between the individuals to be analyzed are determined by the separation of homologous chromosomes, so as to determine the genetic relationship level between the individuals to be analyzed. The above method of separating homologous chromosomes takes a long time and has low efficiency.

[0004] Therefore, there is an urgent need for a solution that can quickly and efficiently determine the genetic relationship level. Summary of the Invention

[0005] The rapid inference method for genetic relationship based on shared ancestral segments provided by the embodiments of the present application is used to achieve the effect of improving the efficiency of determining the genetic relationship level.

[0006] In a first aspect, an embodiment of the present application provides a rapid inference method for genetic relationship based on shared ancestral segments, including:

[0007] Obtain a gene data set; wherein, the gene data set includes the gene data of each individual to be analyzed in a pair of individuals to be analyzed;

[0008] Segment the gene data of the individuals to be analyzed to obtain multiple gene data segments of the individuals to be analyzed; wherein, the gene data segment includes at least one gene data;

[0009] Encode the gene data segments of the individuals to be analyzed to obtain a first encoding set and a second encoding set of the gene data segments of the individuals to be analyzed; wherein, the first encoding set includes the positions of the gene data with the first encoding result in the gene data segments; the second encoding set includes the positions of the gene data with the second encoding result in the gene data segments;

[0010] Determine the shared ancestral gene data segments corresponding to a pair of individuals to be analyzed according to the first encoding set and the second encoding set of the gene data segments of each individual in the pair of individuals to be analyzed;

[0011] Determine the genetic relationship level corresponding to a pair of individuals to be analyzed according to the shared ancestral gene data segments corresponding to the pair of individuals to be analyzed.

[0012] In a possible implementation, the gene data fragments have fragment sequence numbers; determining the co-ancestral gene data fragments corresponding to a pair of individuals to be analyzed according to the first coding set and the second coding set of each gene data fragment in the pair of individuals to be analyzed includes:

[0013] For each pair of gene data fragments with the same fragment sequence number in a pair of individuals to be analyzed, determine the gene window set corresponding to the pair of gene data fragments with the same fragment sequence number according to the first coding set and the second coding set of the pair of gene data fragments with the same fragment sequence number; wherein, the gene window set includes positions, and the positions in the gene window set represent that the gene data of the pair of gene data fragments with the same fragment sequence number are different at these positions;

[0014] Determine the gene data fragments corresponding to the gene window sets that are empty sets between each pair of non-empty gene window sets, and combine them into the co-ancestral gene data fragments corresponding to the pair of individuals to be analyzed.

[0015] In a possible implementation, for each pair of gene data fragments with the same fragment sequence number in a pair of individuals to be analyzed, determining the gene window set corresponding to the pair of gene data fragments with the same fragment sequence number according to the first coding set and the second coding set of the pair of gene data fragments with the same fragment sequence number includes:

[0016] For each pair of gene data fragments with the same fragment sequence number in a pair of individuals to be analyzed, determine the first coding set of one gene data fragment and the second coding set of the other gene data fragment in the pair of gene data fragments, the intersection of the two is the first coding intersection, and determine the second coding set of one gene data fragment and the first coding set of the other gene data fragment in the pair of gene data fragments, the intersection of the two is the second coding intersection;

[0017] Determine the union of the first coding intersection and the second coding intersection as the gene window set corresponding to the pair of gene data fragments with the same fragment sequence number.

[0018] In a possible implementation, determining the degree of kinship level corresponding to a pair of individuals to be analyzed according to the co-ancestral gene data fragments corresponding to the pair of individuals to be analyzed includes:

[0019] For each individual to be analyzed in a pair of individuals to be analyzed, determine the kinship coefficient of the individual to be analyzed according to the length of the gene data of the individual to be analyzed and the lengths of the co-ancestral gene data fragments;

[0020] If it is determined that the kinship coefficient of the individual to be analyzed is greater than the first threshold, then determine the degree of kinship level of the individual to be analyzed according to the kinship coefficient of the individual to be analyzed;

[0021] Otherwise, determine the kinship level of the individual to be analyzed according to the lengths of the co-ancestral gene data segments.

[0022] In a possible implementation, determining the kinship coefficient of the individual to be analyzed according to the length of the gene data of the individual to be analyzed and the lengths of the co-ancestral gene data segments includes:

[0023] Determine the sum of the lengths of the co-ancestral gene data segments as the total length;

[0024] Determine the ratio of the total length to the length of the gene data of the individual to be analyzed as the kinship coefficient of the individual to be analyzed.

[0025] In a possible implementation, determining the kinship level of the individual to be analyzed according to the kinship coefficient of the individual to be analyzed includes:

[0026] Determine the coefficient range corresponding to the kinship coefficient of the individual to be analyzed according to the first preset mapping relationship and the kinship coefficient of the individual to be analyzed; wherein, the first preset mapping relationship represents the corresponding relationship between the kinship coefficient and the coefficient range; the coefficient range is the value range of the kinship coefficient;

[0027] Determine the kinship level corresponding to the determined coefficient range as the kinship level of the individual to be analyzed according to the second preset mapping relationship and the determined coefficient range; wherein, the second preset mapping relationship is the corresponding relationship between the coefficient range and the kinship level.

[0028] In a possible implementation, determining the kinship level of the individual to be analyzed according to the lengths of the co-ancestral gene data segments includes:

[0029] Determine the sum of the lengths of the co-ancestral gene data segments as the total length;

[0030] Determine the length range corresponding to the total length according to the third preset mapping relationship and the total length; wherein, the third preset mapping relationship represents the corresponding relationship between the total length and the length range; the length range is the value range of the total length of the co-ancestral gene data segments;

[0031] Determine the kinship level corresponding to the determined length range as the kinship level of the individual to be analyzed according to the fourth preset mapping relationship and the determined length range; wherein, the fourth preset mapping relationship is the corresponding relationship between the length range and the kinship level.

[0032] In a possible implementation, before determining the kinship level corresponding to a pair of individuals to be analyzed according to the co-ancestral gene data segments corresponding to the pair of individuals to be analyzed, the method further includes:

[0033] If it is determined that the length of the co-ancestral gene data segment is less than the preset length, then remove the co-ancestral gene data segment; and / or,

[0034] If it is determined that the error rate of the co-ancestral gene data segment is greater than or equal to the preset error rate threshold, then remove the co-ancestral gene data segment; wherein, the error rate indicates the proportion of heterozygous genotypes in the co-ancestral gene data segment.

[0035] In a second aspect, an embodiment of the present application provides a fast inference device for the genetic relationship based on co-ancestral segments, including:

[0036] An acquisition module, configured to acquire a gene data set; wherein, the gene data set includes the gene data of each individual to be analyzed in a pair of individuals to be analyzed;

[0037] A segmentation module, configured to segment the gene data of the individual to be analyzed to obtain a plurality of gene data segments of the individual to be analyzed; wherein, the gene data segment includes at least one gene data;

[0038] An encoding module, configured to encode the gene data segments of the individual to be analyzed to obtain a first encoding set and a second encoding set of the gene data segments of the individual to be analyzed; wherein, the first encoding set includes the positions of the gene data with the first encoding result in the gene data segment; the second encoding set includes the positions of the gene data with the second encoding result in the gene data segment;

[0039] A determination module, configured to determine the co-ancestral gene data segment corresponding to a pair of individuals to be analyzed according to the first encoding set and the second encoding set of the gene data segments in the pair of individuals to be analyzed;

[0040] The determination module is further configured to determine the genetic relationship level corresponding to a pair of individuals to be analyzed according to the co-ancestral gene data segment corresponding to the pair of individuals to be analyzed.

[0041] In a possible implementation manner, the gene data segment has a segment serial number; according to the first encoding set and the second encoding set of the gene data segments in a pair of individuals to be analyzed, to determine the co-ancestral gene data segment corresponding to the pair of individuals to be analyzed, the determination module is configured to:

[0042] For each pair of gene data segments with the same segment serial number in a pair of individuals to be analyzed, according to the first encoding set and the second encoding set of the pair of gene data segments with the same segment serial number, determine the gene window set corresponding to the pair of gene data segments with the same segment serial number; wherein, the gene window set includes positions, and the positions in the gene window set indicate that the gene data of the pair of gene data segments with the same segment serial number are different at this position;

[0043] Determine the gene data segments corresponding to the gene window sets that are empty sets between each pair of gene window sets where the pair is non-empty, and combine them into the co-ancestral gene data segments corresponding to a pair of individuals to be analyzed.

[0044] In a possible implementation, for each pair of gene data segments with the same segment number in a pair of individuals to be analyzed, according to the first coding set and the second coding set of the pair of gene data segments with the same segment number, determine the gene window set corresponding to the pair of gene data segments with the same segment number. The determination module is used for:

[0045] For each pair of gene data segments with the same segment number in a pair of individuals to be analyzed, determine the first coding set of one gene data segment and the second coding set of the other gene data segment in the pair. The intersection between the two is the first coding intersection, and determine the second coding set of one gene data segment and the first coding set of the other gene data segment in the pair. The intersection between the two is the second coding intersection;

[0046] Determine the union of the first coding intersection and the second coding intersection as the gene window set corresponding to the pair of gene data segments with the same segment number.

[0047] In a possible implementation, according to the co-ancestral gene data segments corresponding to a pair of individuals to be analyzed, determine the kinship level corresponding to the pair of individuals to be analyzed. The determination module is used for:

[0048] For each individual to be analyzed in a pair of individuals to be analyzed, determine the kinship coefficient of the individual to be analyzed according to the length of the gene data of the individual to be analyzed and the lengths of the co-ancestral gene data segments;

[0049] If it is determined that the kinship coefficient of the individual to be analyzed is greater than the first threshold, then determine the kinship level of the individual to be analyzed according to the kinship coefficient of the individual to be analyzed;

[0050] Otherwise, determine the kinship level of the individual to be analyzed according to the lengths of the co-ancestral gene data segments.

[0051] In a possible implementation, determine the kinship coefficient of the individual to be analyzed according to the length of the gene data of the individual to be analyzed and the lengths of the co-ancestral gene data segments. The determination module is used for:

[0052] Determine the sum of the lengths of the co-ancestral gene data segments as the total length;

[0053] Determine the ratio of the total length to the length of the gene data of the individual to be analyzed as the kinship coefficient of the individual to be analyzed.

[0054] In a possible implementation, according to the coefficient of kinship of the individual to be analyzed, the kinship level of the individual to be analyzed is determined. The determination module is used for:

[0055] According to the first preset mapping relationship and the coefficient of kinship of the individual to be analyzed, determine the coefficient range corresponding to the coefficient of kinship of the individual to be analyzed; wherein, the first preset mapping relationship represents the corresponding relationship between the coefficient of kinship and the coefficient range; the coefficient range is the value range of the coefficient of kinship;

[0056] According to the second preset mapping relationship and the determined coefficient range, determine the kinship level corresponding to the determined coefficient range as the kinship level of the individual to be analyzed; wherein, the second preset mapping relationship is the corresponding relationship between the coefficient range and the kinship level.

[0057] In a possible implementation, according to the lengths of the co-ancestral gene data segments, the kinship level of the individual to be analyzed is determined. The determination module is used for:

[0058] Determine the sum of the lengths of the co-ancestral gene data segments as the total length;

[0059] According to the third preset mapping relationship and the total length, determine the length range corresponding to the total length; wherein, the third preset mapping relationship represents the corresponding relationship between the total length and the length range; the length range is the value range of the total length of the co-ancestral gene data segments;

[0060] According to the fourth preset mapping relationship and the determined length range, determine the kinship level corresponding to the determined length range as the kinship level of the individual to be analyzed; wherein, the fourth preset mapping relationship is the corresponding relationship between the length range and the kinship level.

[0061] In a possible implementation, before determining the kinship level corresponding to a pair of individuals to be analyzed according to the co-ancestral gene data segments corresponding to the pair of individuals to be analyzed, the determination module is further used for:

[0062] If it is determined that the length of the co-ancestral gene data segment is less than the preset length, then remove the co-ancestral gene data segment; and / or,

[0063] If it is determined that the error rate of the co-ancestral gene data segment is greater than or equal to the preset error rate threshold, then remove the co-ancestral gene data segment; wherein, the error rate indicates the proportion of heterozygous genotypes in the co-ancestral gene data segment.

[0064] In a third aspect, an embodiment of the present application provides an electronic device, including: a memory, a processor;

[0065] The memory stores computer execution instructions;

[0066] The processor executes the computer-executable instructions stored in the memory, such that the processor executes the first aspect and / or various possible implementations of the first aspect as described above.

[0067] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, which are used to implement the first aspect and / or various possible implementations of the first aspect as described above when being executed by a processor.

[0068] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which implements the first aspect and / or various possible implementations of the first aspect as described above when being executed by a processor.

[0069] The fast inference method for genetic relationship based on common ancestor segments provided by the embodiments of the present application obtains the genetic data of a pair of individuals to be analyzed, segments them; encodes each genetic data segment to obtain an encoding set of each genetic data segment; determines the common ancestor genetic data segments for the encoding set of each genetic data segment; and determines the genetic relationship level of the pair of individuals to be analyzed based on the common ancestor genetic data segments. It is possible to determine the common ancestor genetic data segments without the method of homologous chromosome separation, shortening the time consumed for determining the genetic relationship level, thereby improving the efficiency of determining the genetic relationship level. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] The drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0071] Figure 1 Schematic flowchart of the fast inference method for genetic relationship based on common ancestor segments provided by the present application Figure 1 ;

[0072] Figure 2 Schematic flowchart of the fast inference method for genetic relationship based on common ancestor segments provided by the present application Figure 2 ;

[0073] Figure 3 Schematic flowchart of the fast inference method for genetic relationship based on common ancestor segments provided by the present application Figure 3 ;

[0074] Figure 4 Accuracy rate curve for inferring genetic relationship with different numbers of SNP locus data;

[0075] Figure 5 Confidence interval accuracy rate curve for inferring genetic relationship with different numbers of SNP locus data;

[0076] Figure 6 The false negative rate curve for inferring genetic relationships with exemplary SNP locus data of different quantities;

[0077] Figure 7 The accuracy rate curve for inferring genetic relationships with exemplary SNP locus data of different types;

[0078] Figure 8 The confidence interval accuracy rate curve for inferring genetic relationships with exemplary SNP locus data of different types;

[0079] Figure 9 The false negative rate curve for inferring genetic relationships with exemplary SNP locus data of different types;

[0080] Figure 10 The structural schematic diagram of the rapid inference device for genetic relationships based on identical-by-descent segments provided by this application;

[0081] Figure 11 The structural schematic diagram of the electronic device provided by this application.

[0082] Through the above-mentioned accompanying drawings, specific embodiments of this application have been shown, and there will be more detailed descriptions hereinafter. These accompanying drawings and the written description are not intended to limit the scope of the concept of this application in any way, but to illustrate the concept of this application to those skilled in the art by referring to specific embodiments. Detailed Description of the Embodiments

[0083] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numerals in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. On the contrary, they are merely examples of devices and methods consistent with some aspects of this application as detailed in the appended claims.

[0084] First, the terms involved in this application are explained:

[0085] Genetic relationship: It refers to the family relationship established by blood relationship between two or more individuals; it refers to the connection between organisms due to a common ancestor.

[0086] Identical-by-descent (IBD) segment: It refers to the same gene data segment inherited from a common ancestor in the gene data of two or more individuals. The identical-by-descent segment can be used to infer the genetic relationship between two or more individuals.

[0087] Coding set: It refers to the set of positions where the gene data of an individual is located after segmental coding of the gene data of the individual. In the gene data of an organism, there are four types of nucleotides, namely adenine (A), thymine (T), cytosine (C), and guanine (G). Coding can be performed according to the type of nucleotide at the same position of the individual to be analyzed.

[0088] Gene window set: It refers to whether the genotype coding results among the individuals to be analyzed are the same within each gene data segment. Exemplarily, if the gene window set is a non-empty set, it indicates that the genotype coding results among the individuals to be analyzed are different, and the elements in the gene window set represent the positions corresponding to the different genotype coding results among the individuals to be analyzed.

[0089] In genetics, genetic information is transmitted between organisms from parents to offspring through genes. Genes can determine various traits and functions of organisms. Since genetic information is transmitted between organisms with kinship, organisms with kinship have similar traits. Specifically, organisms with kinship have similar gene data.

[0090] In some embodiments, the kinship between individuals is determined by analyzing the gene data between individuals. Specifically, by the way of homologous chromosome separation, the similar gene segments between the individuals to be analyzed are determined. Based on the similar gene segments, the kinship level between the individuals to be analyzed is determined.

[0091] In the above embodiments, by the way of homologous chromosome separation, the time consumed is long and the efficiency of determining similar gene segments is low, thus resulting in low efficiency of determining the kinship level between the individuals to be analyzed.

[0092] The rapid inference method of kinship based on co-ancestral segments provided by this application obtains the gene data of paired individuals to be analyzed and segments them respectively; and codes each gene data segment to obtain the first coding set and the second coding set of each gene data segment; for the first coding set and the second coding set, the co-ancestral gene data segments are determined; according to the co-ancestral gene data segments, the kinship level between the paired individuals to be analyzed is determined. It is not necessary to perform homologous chromosome separation to determine the co-ancestral gene data segments, shortening the time consumed for determining the kinship level, thereby improving the efficiency of determining the kinship level.

[0093] The following uses specific embodiments to elaborate in detail on the technical solution of this application and how the technical solution of this application solves the above technical problems. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this application will be described below in conjunction with the accompanying drawings.

[0094] Figure 1 Schematic flow of the method for quickly inferring genetic relationships based on co-ancestral segments provided by this application Figure 1 , as Figure 1 shown, this method includes:

[0095] Step 101. Obtain a gene data set.

[0096] Among them, the gene data set includes the gene data of each individual to be analyzed in a pair of individuals to be analyzed.

[0097] Exemplarily, to obtain a gene data set, in the gene data set, there are pairs of individuals to be analyzed, and each individual to be analyzed in the pairs of individuals to be analyzed has gene data.

[0098] It should be noted that the gene data mentioned in the embodiments of this application can be single nucleotide polymorphism (SNP) site data. An SNP site refers to a variation of a single nucleotide on the genome, including transitions, transversions, deletions, and insertions, forming genetic markers.

[0099] Optionally, the genetic information of the individuals to be analyzed can be analyzed by means of genome-wide association study (GWAS) to obtain the gene data of the individuals to be analyzed, that is, the SNP site data of the individuals to be analyzed.

[0100] Moreover, the gene data and / or genetic information mentioned in the embodiments of this application are all obtained on the premise of the user's informed consent and authorization, and during the process of processing gene data, strict confidentiality is carried out in accordance with relevant regulations.

[0101] Step 102. Segment the gene data of the individuals to be analyzed to obtain multiple gene data segments of the individuals to be analyzed.

[0102] Among them, a gene data segment includes at least one gene data.

[0103] Exemplarily, based on the pairs of individuals to be analyzed, the gene data of each individual to be analyzed is segmented to obtain multiple gene data segments of the individuals to be analyzed.

[0104] Specifically, each gene data segment formed by segmentation includes at least one gene data.

[0105] For example, taking the case where each gene data segment includes 3 gene data, for a certain individual to be analyzed, the gene data segments can include: segment 1 "C:C, A:A, T:C", segment 2 "G:G, A:A, T:T", segment 3 "C:C, C:C, T:T", etc., until all the gene data of the individual to be analyzed are segmented.

[0106] It should be noted that for each individual to be analyzed in the paired individuals to be analyzed, the same segmentation strategy is adopted. Therefore, the gene data segments between each individual to be analyzed in the paired individuals to be analyzed are corresponding.

[0107] Step 103. Encode the gene data segments of the individual to be analyzed to obtain the first encoding set and the second encoding set of the gene data segments of the individual to be analyzed.

[0108] Among them, the first encoding set includes the positions of the gene data with the first encoding result in the gene data segments; the second encoding set includes the positions of the gene data with the second encoding result in the gene data segments.

[0109] Exemplarily, for the sake of explanation, one of the individuals to be analyzed in the paired individuals to be analyzed is denoted as individual A, and correspondingly, the other individual to be analyzed is denoted as individual B. Taking the encoding of the first gene data segment of individual A and individual B as an example. Exemplarily, there are 3 gene data in the first gene data segments of both individual A and individual B, and the genotype performances of each gene data on different individuals are encoded respectively.

[0110] For example, the gene data of individual A in the first gene data segment are C:C, A:A, C:C; the gene data of individual B in the first gene data segment are G:G, C:C, T:T. Encoding is performed respectively at each corresponding position (Marker) in the first gene data segment. For position 1, the genotype of individual A is C:C and the genotype of individual B is G:G. It can be set that the encoding result of C:C is 1, and correspondingly, the encoding result of G:G is 0; for position 2, the genotype of individual A is A:A and the genotype of individual B is C:C. It can be set that the encoding result of C:C is 1, and correspondingly, the encoding result of A:A is 0; for position 3, the genotype of individual A is C:C and the genotype of individual B is T:T. It can be set that the encoding result of C:C is 1, and correspondingly, the encoding result of T:T is 0. That is, the encoding result of individual A in the first gene data segment is "101", and the encoding result of individual B in the first gene data segment is "010".

[0111] It should be noted that if the genotype of an individual at any position is non - homozygous, that is, heterozygous, the coding result at the corresponding position is empty, which can be exemplarily denoted as "-". Combining the above - mentioned example, the genotypes of individual A in the first gene data segment are C:C, A:A, T:C, then the coding result of individual A in the first gene data segment is "10-". The genotypes of individual B in the first gene data segment are G:G, A:C, T:T, then the coding result of individual B in the first gene data segment is "0-0".

[0112] Based on the coding results, the coding sets can be determined. For example, on the first gene data segment, the coding result of individual A is "10-", and the coding result of individual B is "0-0". The position where the coding result of individual A in the first gene data segment is 0 is position 2. Then, in the coding set used to represent the position where the coding result of individual A in the first gene data segment is 0, there is an element "2", and this element is used to indicate that in the first gene data segment, the genotype coding result at position 2 is 0, that is, the first coding set of individual A in the first gene data segment. Exemplarily, the first coding set can be denoted as The position where the coding result of individual A in the first gene data segment is 1 is position 1. Then, in the coding set used to represent the position where the coding result of individual A in the first gene data segment is 1, there is an element "1", and this element is used to indicate that in the first gene data segment, the genotype coding result at position 1 is 1, that is, the second coding set of individual A in the first gene data segment. Exemplarily, the second coding set can be denoted as

[0113] Similarly, the coding set used to represent the position where the coding result of individual B in the first gene data segment is 0 can be denoted as That is, the first coding set of individual B in the first gene data segment; the coding set used to represent the position where the coding result of individual B in the first gene data segment is 1 can be denoted as That is, the second coding set of individual B in the first gene data segment, which is an empty set.

[0114] Through the above - mentioned coding method of examples, coding is performed for each pair of individuals to be analyzed on each gene data segment, and the first coding set and the second coding set of each pair of individuals to be analyzed on each gene data segment are respectively determined.

[0115] Step 104. Determine the co - ancestral gene data segments corresponding to a pair of individuals to be analyzed according to the first coding set and the second coding set of each gene data segment in a pair of individuals to be analyzed.

[0116] Exemplarily, according to the foregoing examples, the positions of the gene data with the first coding result in the gene data fragment are included in the first coding set, and the positions of the gene data with the second coding result in the gene data fragment are included in the second coding set.

[0117] By respectively comparing the first coding set and the second coding set of each gene data fragment in a pair of individuals to be analyzed, one or more co-ancestral gene data fragments in the gene data fragments of the pair of individuals to be analyzed can be determined.

[0118] It should be noted that the co-ancestral gene data fragment can also be abbreviated as the co-ancestral fragment. That is, the co-ancestral gene data fragments in this article can all be abbreviated as the co-ancestral fragment.

[0119] Step 105. Determine the kinship level corresponding to a pair of individuals to be analyzed according to the co-ancestral gene data fragments corresponding to the pair of individuals to be analyzed.

[0120] Exemplarily, according to the determined co-ancestral gene data fragments, the kinship level corresponding to the pair of individuals to be analyzed can be determined.

[0121] In one example, the kinship level can be judged according to the length of the co-ancestral gene data fragment.

[0122] Exemplarily, the longer the length of the co-ancestral gene data fragment, the closer the kinship level corresponding to the pair of individuals to be analyzed represents the kinship between the pair of individuals to be analyzed.

[0123] In another example, the kinship level can be judged according to the ratio of the length of the co-ancestral gene data fragment to the length of the total gene data fragment.

[0124] Exemplarily, the larger the ratio of the length of the co-ancestral gene data fragment to the length of the total gene data fragment, the closer the kinship level corresponding to the pair of individuals to be analyzed represents the kinship between the pair of individuals to be analyzed.

[0125] It should be noted that through the determined kinship level, the inferred kinship between the individuals to be analyzed can be obtained.

[0126] The rapid inference method for kinship based on co-ancestral segments provided by this application obtains the gene data of a pair of individuals to be analyzed, segments them respectively, encodes each gene data segment to obtain the first coding set and the second coding set of each gene data segment, determines the co-ancestral gene data segments of a pair of individuals to be analyzed for these two coding sets, and determines the kinship level between a pair of individuals to be analyzed according to the co-ancestral gene data segments. It is not necessary to separate homologous chromosomes for each individual to be analyzed to determine the co-ancestral gene data segments, which shortens the time-consuming for determining the kinship level, can rapidly infer the kinship of individuals to be analyzed, and thus improves the efficiency of determining the kinship level. Further, due to the improvement in the inference speed of kinship, the rapid inference method for kinship based on co-ancestral segments provided by this application can be applied to the inference of kinship in a large-scale population.

[0127] Combined with the foregoing embodiments, it can be seen that in the rapid inference method for kinship based on co-ancestral segments provided by this application, the co-ancestral gene data segments are determined by the first coding set and the second coding set on each gene data segment of a pair of individuals to be analyzed. On the basis of the foregoing embodiments, this embodiment specifically explains how to determine the co-ancestral gene data segments according to the first coding set and the second coding set on the gene data segments.

[0128] In one example, the gene data segment has a segment serial number. The segment serial number of the gene data segment is used to identify the gene data segment on an individual to be analyzed. It can be understood that the gene data segments corresponding to a pair of individuals to be analyzed have the same segment serial number.

[0129] Figure 2 Schematic flow of the rapid inference method for kinship based on co-ancestral segments provided by this application Figure 2 , such as Figure 2 shown, the co-ancestral gene data segments corresponding to a pair of individuals to be analyzed can be determined through the following steps:

[0130] Step 201. For each pair of gene data segments with the same segment serial number among a pair of individuals to be analyzed, according to the first coding set and the second coding set of the pair of gene data segments with the same segment serial number, determine the gene window set corresponding to the pair of gene data segments with the same segment serial number.

[0131] Among them, the gene window set includes positions, and the positions in the gene window set represent that the gene data of a pair of gene data segments with the same segment serial number are different at these positions.

[0132] Exemplarily, in a pair of individuals to be analyzed, the corresponding gene data segments have the same segment number. The gene data segments with the same segment number are regarded as a pair of gene data segments. Each gene data segment in a pair of gene data segments has a first coding set and a second coding set. According to the first coding set, the second coding set of the first gene data segment in a pair of gene data segments, and the first coding set, the second coding set of the second gene data segment, the gene window set of this pair of gene data segments can be determined.

[0133] In the gene window set of this pair of gene data segments, it includes one or more positions. The positions in the gene window set indicate that in this pair of gene data segments, the gene data at this position is different.

[0134] Specifically, the gene window set corresponding to a pair of gene data segments with the same segment number can be determined through the following steps:

[0135] Step 2011. For each pair of gene data segments with the same segment number in a pair of individuals to be analyzed, determine the first coding set of one gene data segment and the second coding set of the other gene data segment in this pair of gene data segments. The intersection between the two is the first coding intersection, and determine the second coding set of one gene data segment and the first coding set of the other gene data segment in this pair of gene data segments. The intersection between the two is the second coding intersection.

[0136] Step 2012. Determine the union of the first coding intersection and the second coding intersection, which is the gene window set corresponding to a pair of gene data segments with the same segment number.

[0137] Exemplarily, for individual A and individual B in a pair of individuals to be analyzed, for a pair of gene data segments with a segment number of 1, that is, a pair of first gene data segments, they both have a first coding set and a second coding set respectively.

[0138] First, determine the intersection of the first coding set of individual A in the first gene data segment and the second coding set of individual B in the first gene data segment, which is the first intersection; and determine the intersection of the second coding set of individual A in the first gene data segment and the first coding set of individual B in the first gene data segment, which is the second intersection.

[0139] Then determine the union of the first intersection and the second intersection, which is the gene window set corresponding to a pair of individuals to be analyzed in the first gene data segment.

[0140] Exemplarily, the gene window set can be determined through the following formula:

[0141]

[0142] Wherein, W i represents the gene window set corresponding to the i-th pair of gene data segments, represents the first coding set of individual A for the i-th pair of gene data segments, represents the second coding set of individual A for the i-th pair of gene data segments, represents the first coding set of individual B for the i-th pair of gene data segments, represents the second coding set of individual B for the i-th pair of gene data segments.

[0143] It can be understood that through the method provided by the above example, the gene window set corresponding to each pair of gene data segments in a pair of individuals to be analyzed can be calculated.

[0144] In the above example, through the first coding set and the second coding set of each pair of gene data segments, the gene window set of this pair of gene data segments can be determined. The gene window set can indicate whether the coding results of the gene data on this pair of gene data segments are the same. It lays a foundation for subsequently determining the co-ancestral gene data segments based on the gene window set.

[0145] Step 202. Determine the gene data segments corresponding to the gene window sets that are empty sets between each pair of non-empty gene window sets, and combine them into the co-ancestral gene data segments corresponding to a pair of individuals to be analyzed.

[0146] Exemplarily, the gene window set can be an empty set or a non-empty set. If it is determined that there are empty gene window sets among the non-empty gene window sets, then the gene data segments corresponding to the one or more empty gene window sets are combined into the co-ancestral gene data segments corresponding to this pair of individuals to be analyzed.

[0147] For example, for multiple gene window sets W 1 、W 2 、W 3 、W 4 、W 5 , if it is determined that both W 1 and W 5 are non-empty sets, that is, there are elements in the gene window set W 1 and the gene window set W 5 ; and W 2 、W 3 、W 4 are all empty sets, that is, there are no elements in the x gene window x sets W 2 、W 3 、W 4 , then the co-ancestral gene data segments between this pair of individuals to be analyzed are determined to be W 2 、W3 , W 4 The gene data segments corresponding to them, namely the second gene data segment, the third gene data segment, and the fourth gene data segment, together constitute the co-ancestral gene data segment of the individual to be analyzed.

[0148] For another example, for multiple gene window sets W 1 , W 2 , W 3 , W 4 , W 5 , W 6 , W 7 , W 8 , W 9 , W 10 , if it is determined that W 1 , W 5 , W 8 , W 10 are all non-empty sets, and W 2 , W 3 , W 4 , W 6 , W 7 , W 9 are all empty sets, then the gene data segments corresponding to W 2 , W 3 , W 4 , and the gene data segments corresponding to W 6 , W 7 , and the gene data segment corresponding to W 9 are combined to obtain the co-ancestral gene data segment of the individual to be analyzed.

[0149] In the above embodiments, according to the two coding sets of each pair of gene data segments of an individual to be analyzed, the gene window set corresponding to each pair of gene data segments is determined; then, according to the gene data segments corresponding to the gene window sets that are non-empty sets and empty sets among the gene window sets, the co-ancestral gene data segment of the individual to be analyzed is determined. It can quickly and efficiently determine the co-ancestral gene data segment between a pair of individuals to be analyzed, improving the efficiency of determining the co-ancestral gene data segment. It lays a foundation for subsequently determining the kinship level between a pair of individuals to be analyzed based on the co-ancestral gene data segment.

[0150] Combined with the foregoing embodiments, it can be seen that in the method for quickly inferring kinship based on co-ancestral segments provided in this application, the kinship level between a pair of individuals to be analyzed can be determined through the co-ancestral gene data segment. On the basis of the foregoing embodiments, this embodiment specifically explains how to determine the kinship level according to the co-ancestral gene data segment.

[0151] Figure 3Schematic flowchart of the method for quickly inferring the genetic relationship based on the co-ancestral segments provided by this application Figure 3 , as Figure 3 shown, the genetic relationship level corresponding to a pair of individuals to be analyzed can be determined through the following steps:

[0152] Step 301. For each individual to be analyzed in a pair of individuals to be analyzed, determine the genetic relationship coefficient of the individual to be analyzed according to the length of the genetic data of the individual to be analyzed and the lengths of the co-ancestral genetic data segments.

[0153] Exemplarily, for each individual to be analyzed in a pair of individuals to be analyzed, the genetic relationship coefficient of the individual to be analyzed can be determined according to the length of the genetic data of the individual to be analyzed and the lengths of the co-ancestral genetic data segments of this pair of individuals to be analyzed.

[0154] Specifically, the genetic relationship coefficient of the individual to be analyzed can be determined through the following steps:

[0155] Step 3011. Determine the sum of the lengths of the co-ancestral genetic data segments as the total length.

[0156] Step 3012. Determine the ratio of the total length to the length of the genetic data of the individual to be analyzed as the genetic relationship coefficient of the individual to be analyzed.

[0157] Exemplarily, the co-ancestral genetic data segments corresponding to a pair of individuals to be analyzed can be one or more, and each co-ancestral genetic data segment is composed of a combination of one or more genetic data segments. For each individual to be analyzed in a pair of individuals to be analyzed, determine the sum of the lengths of the co-ancestral gene segments as the total length.

[0158] Furthermore, determine the ratio of the total length to the length of the genetic data of the individual to be analyzed as the genetic relationship coefficient of the individual to be analyzed. Among them, the genetic relationship coefficient is the probability that two alleles randomly selected from a pair of individuals to be analyzed come from the same ancestral allele, and its value is the ratio of the total length to the length of the genetic data of the individual to be analyzed. The genetic relationship coefficient is usually between 0 and 1, where 0 indicates that there is no genetic relationship between a pair of individuals to be analyzed, and 1 indicates that a pair of genetic relationships have exactly the same genetic data.

[0159] In the above example, the ratio between the sum of the lengths of the co-ancestral genetic data segments and the length of the genetic data of the individual to be analyzed is used as the genetic relationship coefficient. Subsequently, the genetic relationship level between a pair of individuals to be analyzed can be further determined through the genetic relationship coefficient.

[0160] Step 302. If it is determined that the kinship coefficient of the individual to be analyzed is greater than the first threshold, determine the kinship level of the individual to be analyzed according to the kinship coefficient of the individual to be analyzed; otherwise, determine the kinship level of the individual to be analyzed according to the lengths of the co-ancestral gene data segments.

[0161] Exemplarily, the determined kinship coefficient based on the foregoing example is compared with the first threshold.

[0162] In one example, if the kinship coefficient is greater than the first threshold, the kinship level of the pair of individuals to be analyzed is determined by the kinship coefficient of the pair of individuals to be analyzed.

[0163] Among them, the first threshold can be exemplarily set to 1 / (2 17 / 2 ). It should be noted that the specific value of the first threshold can also be set to any real number greater than 0 and less than 1. The foregoing first threshold is only for exemplary illustration, and the specific value of the first threshold is not limited in this example.

[0164] Specifically, according to the first preset mapping relationship and the kinship coefficient of the individual to be analyzed, determine the coefficient range corresponding to the kinship coefficient of the individual to be analyzed.

[0165] Among them, the first preset mapping relationship represents the corresponding relationship between the kinship coefficient and the coefficient range; the coefficient range is the value range of the kinship coefficient.

[0166] According to the second preset mapping relationship and the determined coefficient range, determine the kinship level corresponding to the determined coefficient range as the kinship level of the individual to be analyzed.

[0167] Among them, the second preset mapping relationship is the corresponding relationship between the coefficient range and the kinship level.

[0168] Exemplarily, the kinship level can be divided into multiple levels. For example, the kinship level can include: first-degree kinship (1st), second-degree kinship (2nd), third-degree kinship (3rd), fourth-degree kinship (4th), fifth-degree kinship (5th), sixth-degree kinship (6th), seventh-degree kinship (7th).

[0169] Furthermore, for each kinship level, a corresponding coefficient range of the kinship coefficient can be set. Based on the relationship between the value of the kinship coefficient and the coefficient range, determine the coefficient range; and then based on the determined coefficient range, determine the corresponding kinship level.

[0170] For example, based on Table 1, the kinship of the individual to be analyzed can be determined by the kinship coefficient.

[0171] Table 1 Example Table of Coefficient Ranges of Kinship Corresponding to Kinship Levels

[0172]

[0173] With reference to Table 1 for explanation, according to the kinship coefficient determined by the foregoing example, and according to the first preset mapping relationship, the coefficient range corresponding to the kinship coefficient of the individual to be analyzed is determined. Among them, the first preset mapping relationship can be set as follows: if the value of the kinship coefficient is within a certain coefficient range, then this coefficient range is determined as the coefficient range of the kinship coefficient of the individual to be analyzed.

[0174] For example, if the calculated value of the kinship coefficient of a pair of individuals to be analyzed is 1 / (2 16 / 2 ), according to the coefficient range in the example in Table 1, it can be known that the value of this kinship coefficient is greater than 1 / (2 17 / 2 ), and less than 1 / (2 15 / 2 ), then the coefficient range of this kinship coefficient is (1 / (2 17 / 2 ), 1 / (2 15 / 2 )).

[0175] Furthermore, according to the coefficient range determined by the foregoing example, the kinship level corresponding to the coefficient range of the individual to be analyzed is determined according to the second preset mapping relationship. And the determined kinship level is used as the kinship level between this pair of individuals to be analyzed.

[0176] Continuing with the foregoing example, according to the kinship levels in the example in Table 1, it can be known that the kinship level corresponding to the coefficient range (1 / (2 17 / 2 ), 1 / (2 15 / 2 )) is the seventh-level kinship (7th). Then the kinship level between this pair of individuals to be analyzed is the seventh-level kinship.

[0177] It should be noted that for each pair of individuals to be analyzed, the kinship is mutual. For example, a pair of individuals to be analyzed includes individual A and individual B. For a pair of individuals with a seventh-level kinship, the kinship level of individual A to individual B is the seventh-level kinship. Correspondingly, the kinship level of individual B to individual A is also the seventh-level kinship.

[0178] In the above example, when the value of the kinship coefficient is greater than the first threshold, the coefficient range to which the kinship coefficient belongs is determined through the kinship coefficient, and then the kinship level of the individual to be analyzed is determined according to the kinship level corresponding to the coefficient range. It can subdivide the kinship level and improve the fineness of kinship level determination.

[0179] In another example, if the coefficient of kinship is less than or equal to the first threshold, the kinship level of the pair of individuals to be analyzed is determined by the lengths of the respective identical-by-descent gene data segments of the pair of individuals to be analyzed.

[0180] Specifically, the sum of the lengths of the respective identical-by-descent gene data segments is determined as the total length.

[0181] According to the third preset mapping relationship and the total length, the length range corresponding to the total length is determined.

[0182] Among them, the third preset mapping relationship represents the corresponding relationship between the total length and the length range; the length range is the value range of the total length of the identical-by-descent gene data segments.

[0183] According to the fourth preset mapping relationship and the determined length range, the kinship level corresponding to the determined length range is determined as the kinship level of the individual to be analyzed.

[0184] Among them, the fourth preset mapping relationship is the corresponding relationship between the length range and the kinship level.

[0185] Exemplarily, in combination with the foregoing example, the kinship level can be further divided into levels. For example, the kinship level can include: first-degree kinship (1st), second-degree kinship (2nd), third-degree kinship (3rd), fourth-degree kinship (4th), fifth-degree kinship (5th), sixth-degree kinship (6th), seventh-degree kinship (7th), eighth-degree kinship (8th), ninth-degree kinship (9th), and unrelated kinship (UN).

[0186] First, the lengths of the respective identical-by-descent gene data segments of the pair of individuals to be analyzed are summed to obtain the total length. Among them, the unit of the length of the identical-by-descent gene data segment can be centimorgan (cM).

[0187] Furthermore, for each kinship level, a corresponding length range of the total length can be set. Based on the relationship between the value of the total length and the length range, the length range is determined; and then based on the determined length range, the corresponding kinship level is determined.

[0188] For example, based on Table 2, the kinship of the individual to be analyzed can be determined by the total length.

[0189] Table 2 Example table of the length range of the total length corresponding to the kinship level

[0190] Degree of kinship Length range of the total length 1st [2200,+∞) 2nd [1300,2200) 3rd [650,1300) 4th [340,650) 5th [200,340) 6th [90,200) 7th [60,90) 8th [30,60) 9th [5,30) UN [0,5)

[0191] With reference to Table 2 for explanation, the sum of the lengths of each co-ancestral gene data segment determined according to the foregoing example, that is, the total length, determines the length range corresponding to the total length of the individual to be analyzed according to the third preset mapping relationship. The third preset mapping relationship can be set as follows: if the value of the total length falls within a certain length range, then determine this length range as the length range of the total length of the individual to be analyzed.

[0192] For example, if the calculated total length of a pair of individuals to be analyzed is 580, according to the length range in the example in Table 2, the value of this total length is greater than 340 and less than 650, then the length range of this total length is [340, 650).

[0193] Furthermore, according to the length range determined according to the foregoing example, determine the kinship level corresponding to the length range of the individual to be analyzed according to the fourth preset mapping relationship. And take the determined kinship level as the kinship level between this pair of individuals to be analyzed.

[0194] Continuing with the foregoing example, according to the kinship level in the example in Table 2, the kinship level corresponding to the length range of [340, 650) is the fourth-degree kinship (4th). Then the kinship level between this pair of individuals to be analyzed is the fourth-degree kinship.

[0195] In the above example, when the value of the kinship coefficient is less than or equal to the first threshold, the length range to which the total length belongs is determined through the total length of the co-ancestral gene data segments, and then the kinship level of the individual to be analyzed is determined according to the kinship level corresponding to the length range. It can further subdivide the kinship level on the basis of the foregoing example. And for kinship levels above the seventh-degree kinship, due to the decrease in genetic similarity between the individuals to be analyzed, the accuracy of the kinship coefficient decreases. At this time, the total length of the co-ancestral gene data segments is directly used to judge the kinship, and the kinship level of the individuals to be analyzed above the seventh-degree kinship can be accurately determined.

[0196] In the above embodiment, first, according to each individual to be analyzed in a pair of individuals to be analyzed, the kinship coefficient of the individual to be analyzed is determined. When the kinship coefficient is greater than the first threshold, the coefficient range is determined through the kinship coefficient, and then the corresponding kinship level is determined; when the kinship coefficient is less than or equal to the first threshold, the total length is determined by summing the lengths of each co-ancestral gene data segment, the length range is determined, and then the corresponding kinship level is determined. It can improve the fineness of determining the kinship and the accuracy of determining the kinship level of the individuals to be analyzed above the seventh-degree kinship level.

[0197] As can be seen from the foregoing examples, the determination of the kinship level is based on the kinship coefficient, and the calculation of the kinship coefficient is based on the co-ancestral gene data segments of a pair of individuals to be analyzed that have been determined. Therefore, before calculating the kinship coefficient, the co-ancestral gene data segments can be screened. Based on any of the foregoing embodiments, this embodiment specifically explains how to screen the co-ancestral gene data segments.

[0198] In one example, before determining the kinship level corresponding to a pair of individuals to be analyzed based on the co-ancestral gene data segments corresponding to the pair of individuals to be analyzed, the method further includes:

[0199] If it is determined that the length of the co-ancestral gene data segment is less than the preset length, then remove the co-ancestral gene data segment; and / or, if it is determined that the error rate of the co-ancestral gene data segment is greater than or equal to the preset error rate threshold, then remove the co-ancestral gene data segment.

[0200] Among them, the error rate indicates the proportion of heterozygous genotypes in the co-ancestral gene data segment.

[0201] Exemplarily, before determining the kinship level corresponding to a pair of individuals to be analyzed, the determined co-ancestral gene data segments can be screened. The screening strategy can include, but is not limited to: screening based on the length of the co-ancestral gene data segment, and / or screening based on the error rate of the co-ancestral gene data segment.

[0202] In one example, if it is determined that the length of the co-ancestral gene data segment is less than the preset length, then remove the co-ancestral gene data segment.

[0203] Exemplarily, co-ancestral gene data segments with too small a length of the co-ancestral gene data segment are excluded. Among them, different preset lengths can be set respectively. For example, the preset length can be set to any of the following values: 0 cM, 2 cM, 3 cM, 6 cM, 7 cM, 8 cM, 9 cM, 12 cM, 15 cM, 20 cM, etc. Optionally, the preset length can be set to 8 cM.

[0204] In another example, if it is determined that the error rate of the co-ancestral gene data segment is greater than or equal to the preset error rate threshold, then remove the co-ancestral gene data segment.

[0205] Exemplarily, co-ancestral gene data segments with too high an error rate in the co-ancestral gene data segment are excluded. Among them, different error rates can be set respectively. For example, the error rate can be set to any of the following values: 0, 0.002, 0.004, 0.005, 0.01. Optionally, the error rate can be set to 0.002.

[0206] It should be noted that in the foregoing embodiments, in the process of determining the co-ancestral gene data segments through the gene window set, when the gene window set is a non-empty set, that is, the gene window set includes positions, which indicates that at these positions of the pair of gene data segments, the coding results of the gene data between the individuals to be analyzed are different, that is, the gene data is different. And when the gene window set is an empty set, it indicates that on the pair of gene data segments, the coding results of the gene data between the individuals to be analyzed are the same.

[0207] However, for heterozygous genotypes, the coding result is "-". In the process of calculating the gene window set, after the calculations of taking intersections and unions, the heterozygous genotypes are not considered in the co-ancestral gene data segments.

[0208] Therefore, for the co-ancestral gene data segments with heterozygous genotypes, the paired individuals to be analyzed cannot be matched at the gene positions of the heterozygous genotypes, that is, it is regarded as an error.

[0209] Count the number of errors that occur in the determined co-ancestral gene data segments to obtain a count value; and determine the ratio of the count value to the total genotype quantity value in the co-ancestral gene data segments as the error rate.

[0210] It should be noted that the screening strategies for the co-ancestral gene data segments shown in the above two examples can be implemented separately or in combination. And when implemented in combination, the implementation order of the screening strategies shown in the above two examples is not limited.

[0211] In the foregoing embodiments, before calculating the kinship coefficient, screen the determined co-ancestral gene data segments, and remove the co-ancestral gene data segments with too high error rate and / or too small length. Thereby, the accuracy of the co-ancestral gene data segments can be improved, and the accuracy of determining the kinship level for a pair of individuals to be analyzed can be improved indirectly.

[0212] Optionally, to verify the feasibility and accuracy of the fast inference method for kinship based on co-ancestral segments provided by the present application, perform accuracy tests on the determination of the kinship levels of pairs of individuals to be analyzed through different strategies.

[0213] In one example, obtain the gene data of different numbers of individuals to be analyzed respectively. Specifically, obtain the SNP locus data of different numbers of individuals to be analyzed respectively, and determine the kinship levels between the individuals to be analyzed based on the fast inference method for kinship based on co-ancestral segments provided by any of the foregoing embodiments and / or combinations of the foregoing embodiments. And calculate the accuracy of determining the kinship levels of the individuals to be analyzed for different amounts of SNP locus data respectively.

[0214] Exemplarily, the number of SNP sites can be selected as one or more of the following numbers: 650,000, 550,000, 450,000, 350,000, 250,000, 150,000, 100,000, and 50,000. The numbers of SNP sites are respectively selected as 250,000, 150,000, 100,000, and 50,000 to determine the kinship level, and the accuracy rate, confidence interval accuracy rate, and false negative rate of each kinship level determination are calculated.

[0215] Figure 4 It is a graph showing the accuracy rate of inferring kinship for exemplary SNP site data with different numbers. As Figure 4 shown, the abscissa represents different kinship levels, and the ordinate represents the accuracy rate of determining the kinship level in this application. Among them, the accuracy rate refers to: among the total individuals to be analyzed with a known kinship level of a certain kinship level, the proportion of individuals to be analyzed whose kinship level determined by this application is the known kinship level among the total individuals to be analyzed.

[0216] Figure 5 It is a graph showing the confidence interval accuracy rate of inferring kinship for exemplary SNP site data with different numbers. As Figure 5 shown, the abscissa represents different kinship levels, and the ordinate represents the confidence interval accuracy rate of determining the kinship level in this application. Among them, the confidence interval accuracy rate refers to: among the total individuals to be analyzed with a known kinship level of a certain kinship level, the proportion of individuals to be analyzed whose kinship level determined by this application is the known kinship level or the kinship level adjacent to the known kinship level among the total individuals to be analyzed.

[0217] Figure 6 It is a graph showing the false negative rate of inferring kinship for exemplary SNP site data with different numbers. As Figure 6 shown, the abscissa represents different kinship levels, and the ordinate represents the false negative rate of determining the kinship level in this application. Among them, the false negative rate refers to: among the total individuals to be analyzed with a known kinship level of a certain kinship level, the proportion of individuals to be analyzed whose kinship level determined by this application is unrelated (UN) among the total individuals to be analyzed.

[0218] It should be noted that experiments have shown that when the number of SNP sites is more than 250,000, further increasing the number of SNP sites does not significantly improve the accuracy rate of kinship level determination. Therefore, Figures 4 to 6 it is only shown as an exemplary experimental result and is not construed as a limitation on the selection of the number of SNP sites.

[0219] In another example, gene data of individuals to be analyzed of different types are obtained respectively. Specifically, SNP locus data of individuals to be analyzed of different types are obtained respectively, and the kinship level between the individuals to be analyzed is determined based on the rapid inference method of kinship based on co-ancestral segments provided by any one of the foregoing embodiments and / or a combination of the foregoing embodiments. And the influence of SNP locus data of different types on the accuracy rate of kinship level determination of the individuals to be analyzed is calculated respectively.

[0220] Exemplarily, one or more of the following types of SNP loci can be selected: type A, type B, type C, type D, type E, type F. Among them, type D is the SNP locus shared by type A and type B; type E is the SNP locus shared by type A and type C; type F is the SNP locus shared by type A, type B, and type C.

[0221] The types of SNP loci selected are type D, type E, and type F respectively for kinship level determination, and the accuracy rate, confidence interval accuracy rate, and false negative rate of each kinship level determination are calculated. Figure 7 It is a graph showing the accuracy rate of inferring kinship for SNP locus data of different types exemplarily. Figure 8 It is a graph showing the confidence interval accuracy rate of inferring kinship for SNP locus data of different types exemplarily. Figure 9 It is a graph showing the false negative rate of inferring kinship for SNP locus data of different types exemplarily.

[0222] Based on the above two examples, it can be known the influence of different numbers of SNP loci and different types of SNP loci on the accuracy rate of the kinship determined by the rapid inference method of kinship based on co-ancestral segments provided by the present application, providing accuracy data and credibility data for subsequent work.

[0223] The rapid inference method for genetic relationship based on co-ancestral segments provided by the embodiments of the present application obtains the genetic data of pairs of individuals to be analyzed, segments the genetic data of each individual to be analyzed in a pair of individuals to be analyzed, and encodes each genetic data separately, so as to obtain the first encoding set and the second encoding set of each genetic data segment; for a pair of genetic data segments with the same segment number in a pair of individuals to be analyzed, determines the genetic window set of the pair of genetic data segments; determines the co-ancestral genetic data segments of a pair of individuals to be analyzed according to the genetic window set. Avoid using the method of homologous chromosome separation, shorten the time consumed to obtain the co-ancestral genetic data segments of the individuals to be analyzed, thereby improving the efficiency of obtaining the co-ancestral genetic data segments of the individuals to be analyzed, and improving the efficiency of determining the genetic relationship level between the individuals to be analyzed. Based on the improvement of the inference speed of the genetic relationship, the rapid inference method for genetic relationship based on co-ancestral segments provided by the present application can be applied to the inference of genetic relationship in a large-scale population. In addition, the co-ancestral genetic data segments are screened, the genetic relationship coefficient is determined based on the screened co-ancestral genetic data segments, and the genetic relationship level between the pair of individuals to be analyzed is determined based on the genetic relationship coefficient or the lengths of the co-ancestral genetic data segments. It can improve the fineness and accuracy of the determined genetic relationship level.

[0224] Figure 10 is a schematic structural diagram of the rapid inference device for genetic relationship based on co-ancestral segments provided by the present application, as Figure 10 shown, the rapid inference device 100 for genetic relationship based on co-ancestral segments provided by this embodiment includes:

[0225] An acquisition module 1001, configured to acquire a genetic data set; wherein, the genetic data set includes the genetic data of each individual to be analyzed in a pair of individuals to be analyzed;

[0226] A segmentation module 1002, configured to segment the genetic data of the individual to be analyzed to obtain multiple genetic data segments of the individual to be analyzed; wherein, the genetic data segment includes at least one genetic data;

[0227] An encoding module 1003, configured to encode the genetic data segments of the individual to be analyzed to obtain a first encoding set and a second encoding set of the genetic data segments of the individual to be analyzed; wherein, the first encoding set includes the positions of the genetic data with the first encoding result in the genetic data segment; the second encoding set includes the positions of the genetic data with the second encoding result in the genetic data segment;

[0228] A determination module 1004, configured to determine the co-ancestral genetic data segments corresponding to a pair of individuals to be analyzed according to the first encoding set and the second encoding set of the genetic data segments of the pair of individuals to be analyzed;

[0229] The determination module 1004 is further configured to determine the kinship level corresponding to a pair of individuals to be analyzed according to the co-ancestral gene data segments corresponding to the pair of individuals to be analyzed.

[0230] In a possible implementation manner, the gene data segment has a segment serial number; according to the first coding set and the second coding set of each gene data segment in a pair of individuals to be analyzed, the co-ancestral gene data segment corresponding to the pair of individuals to be analyzed is determined, and the determination module 1004 is configured to:

[0231] For each pair of gene data segments with the same segment serial number in a pair of individuals to be analyzed, according to the first coding set and the second coding set of the pair of gene data segments with the same segment serial number, determine the gene window set corresponding to the pair of gene data segments with the same segment serial number; wherein, the gene window set includes positions, and the positions in the gene window set indicate that the gene data of the pair of gene data segments with the same segment serial number at this position is different;

[0232] Determine the gene data segments corresponding to the gene window sets that are empty sets between each pair of non-empty gene window sets, and combine them into the co-ancestral gene data segments corresponding to the pair of individuals to be analyzed.

[0233] In a possible implementation manner, for each pair of gene data segments with the same segment serial number in a pair of individuals to be analyzed, according to the first coding set and the second coding set of the pair of gene data segments with the same segment serial number, determine the gene window set corresponding to the pair of gene data segments with the same segment serial number, and the determination module 1004 is configured to:

[0234] For each pair of gene data segments with the same segment serial number in a pair of individuals to be analyzed, determine the first coding set of one gene data segment and the second coding set of the other gene data segment in the pair of gene data segments, the intersection of the two is the first coding intersection, and determine the second coding set of one gene data segment and the first coding set of the other gene data segment in the pair of gene data segments, the intersection of the two is the second coding intersection;

[0235] Determine the union of the first coding intersection and the second coding intersection as the gene window set corresponding to the pair of gene data segments with the same segment serial number.

[0236] In a possible implementation manner, according to the co-ancestral gene data segments corresponding to the pair of individuals to be analyzed, determine the kinship level corresponding to the pair of individuals to be analyzed, and the determination module 1004 is configured to:

[0237] For each individual to be analyzed in a pair of individuals to be analyzed, determine the coefficient of kinship of the individual to be analyzed according to the length of the genetic data of the individual to be analyzed and the lengths of the respective co-ancestral gene data segments;

[0238] If it is determined that the coefficient of kinship of the individual to be analyzed is greater than the first threshold, determine the kinship level of the individual to be analyzed according to the coefficient of kinship of the individual to be analyzed;

[0239] Otherwise, determine the kinship level of the individual to be analyzed according to the lengths of the respective co-ancestral gene data segments.

[0240] In a possible implementation manner, to determine the coefficient of kinship of the individual to be analyzed according to the length of the genetic data of the individual to be analyzed and the lengths of the respective co-ancestral gene data segments, the determination module 1004 is used for:

[0241] Determine the sum of the lengths of the respective co-ancestral gene data segments as the total length;

[0242] Determine the ratio of the total length to the length of the genetic data of the individual to be analyzed as the coefficient of kinship of the individual to be analyzed.

[0243] In a possible implementation manner, to determine the kinship level of the individual to be analyzed according to the coefficient of kinship of the individual to be analyzed, the determination module 1004 is used for:

[0244] According to the first preset mapping relationship and the coefficient of kinship of the individual to be analyzed, determine the coefficient range corresponding to the coefficient of kinship of the individual to be analyzed; wherein, the first preset mapping relationship represents the corresponding relationship between the coefficient of kinship and the coefficient range; the coefficient range is the value range of the coefficient of kinship;

[0245] According to the second preset mapping relationship and the determined coefficient range, determine the kinship level corresponding to the determined coefficient range as the kinship level of the individual to be analyzed; wherein, the second preset mapping relationship is the corresponding relationship between the coefficient range and the kinship level.

[0246] In a possible implementation manner, to determine the kinship level of the individual to be analyzed according to the lengths of the respective co-ancestral gene data segments, the determination module 1004 is used for:

[0247] Determine the sum of the lengths of the respective co-ancestral gene data segments as the total length;

[0248] According to the third preset mapping relationship and the total length, determine the length range corresponding to the total length; wherein, the third preset mapping relationship represents the corresponding relationship between the total length and the length range; the length range is the value range of the total length of the co-ancestral gene data segments;

[0249] Determine the kinship level corresponding to the determined length range according to the fourth preset mapping relationship as the kinship level of the individual to be analyzed; wherein, the fourth preset mapping relationship is the corresponding relationship between the length range and the kinship level.

[0250] In a possible implementation manner, before determining the kinship level corresponding to a pair of individuals to be analyzed according to the co-ancestral gene data segments corresponding to the pair of individuals to be analyzed, the determining module 1004 is further configured to:

[0251] If it is determined that the length of the co-ancestral gene data segment is less than the preset length, then remove the co-ancestral gene data segment; and / or,

[0252] If it is determined that the error rate of the co-ancestral gene data segment is greater than or equal to the preset error rate threshold, then remove the co-ancestral gene data segment; wherein, the error rate indicates the proportion of heterozygous genotypes in the co-ancestral gene data segment.

[0253] The fast inference device for kinship based on co-ancestral segments provided in this embodiment can execute the method provided in the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here in this embodiment.

[0254] Figure 11 It is a schematic structural diagram of an electronic device provided in this application. As Figure 11 shown, the electronic device 110 provided in this embodiment includes: at least one processor 1101 and a memory 1102. Optionally, the electronic device 110 further includes a communication component 1103. Among them, the processor 1101, the memory 1102, and the communication component 1103 are connected through a bus 1104.

[0255] In a specific implementation process, at least one processor 1101 executes the computer execution instructions stored in the memory 1102, so that at least one processor 1101 executes the above method.

[0256] The specific implementation process of the processor 1101 can refer to the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here in this embodiment.

[0257] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU for short), or other general-purpose processors, digital signal processors (DSP for short), application specific integrated circuits (ASIC for short), etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly implemented by the execution of the hardware processor, or implemented by the combination of hardware and software modules in the processor.

[0258] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (Non-volatile Memory, NVM), such as at least one disk memory.

[0259] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, the bus in the drawings of this application is not limited to only one bus or one type of bus.

[0260] This application also provides a computer program product, including a computer program, which implements the above method when executed by a processor.

[0261] This application also provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the processor executes the computer-executable instructions, the above method is implemented.

[0262] The above-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk or an optical disc. The readable storage medium can be any available medium accessible by a general-purpose or special-purpose computer.

[0263] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an Application Specific Integrated Circuits (ASIC). Of course, the processor and the readable storage medium can also exist as discrete components in a device.

[0264] The division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Additionally, the couplings or direct couplings or communication connections shown or discussed between each other can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.

[0265] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0266] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0267] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0268] Those of ordinary skill in the art will understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disks, or optical disks that can store program code.

[0269] Finally, it should be noted that: after considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present invention. The present invention is intended to cover any variations, uses, or adaptations of the present invention, which follow the general principles of the present invention and include known common knowledge or conventional technical means in the technical field not disclosed in the present invention. It is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.

Claims

1. A method for rapid inference of kinship based on common ancestral fragments, characterized in that: include: Acquire a gene data set; wherein the gene data set includes gene data of each individual to be analyzed in a pair of individuals to be analyzed; Segmenting the gene data of the individual to be analyzed to obtain a plurality of gene data segments of the individual to be analyzed; wherein the gene data segments include at least one gene data; Encoding the genetic data segments of the individual to be analyzed to obtain a first encoding set and a second encoding set of the genetic data segments of the individual to be analyzed; wherein the first encoding set includes the position of the genetic data whose encoding result is the first encoding result in the genetic data segments; and the second encoding set includes the position of the genetic data whose encoding result is the second encoding result in the genetic data segments; Determine, according to the first coding set and the second coding set of each gene data segment in the pair of individuals to be analyzed, a common ancestor gene data segment corresponding to the pair of individuals to be analyzed; The kinship levels corresponding to the pair of individuals to be analyzed are determined according to the common ancestor gene data segments corresponding to the pair of individuals to be analyzed.

2. The method according to claim 1, characterized in that The gene data segments have segment serial numbers; determining the common ancestor gene data segments corresponding to the pair of individuals to be analyzed according to the first coding set and the second coding set of each gene data segment in the pair of individuals to be analyzed comprises: For each pair of gene data segments with the same segment number in the pair of individuals to be analyzed, a gene window set corresponding to the pair of gene data segments with the same segment number is determined according to a first coding set and a second coding set of the pair of gene data segments with the same segment number; wherein the gene window set includes positions, and the positions in the gene window set represent that the gene data of the pair of gene data segments with the same segment number at the positions are different; The gene data segments corresponding to the gene window sets that are empty sets between each pair of gene window sets that are non-empty sets are determined, and combined into the common ancestor gene data segments corresponding to the pair of individuals to be analyzed.

3. The method according to claim 2, characterized in that For each pair of genetic data segments with the same segment sequence number in the pair of individuals to be analyzed, determining a genetic window set corresponding to the pair of genetic data segments with the same segment sequence number according to a first coding set and a second coding set of the pair of genetic data segments with the same segment sequence number, including: For each pair of genetic data segments with the same segment sequence number in the pair of individuals to be analyzed, determine a first coding set of one genetic data segment in the pair of genetic data segments and a second coding set of the other genetic data segment, the intersection of the two being a first coding intersection, and determine a second coding set of one genetic data segment in the pair of genetic data segments and a first coding set of the other genetic data segment, the intersection of the two being a second coding intersection; The union of the first encoding intersection and the second encoding intersection is determined to be a gene window set corresponding to a pair of gene data segments with the same segment sequence number.

4. The method according to claim 1, characterized in that: Determining the kinship level corresponding to the pair of individuals to be analyzed according to the common ancestor gene data segments corresponding to the pair of individuals to be analyzed includes: For each individual to be analyzed in the pair of individuals to be analyzed, determining the kinship coefficient of the individual to be analyzed according to the length of the gene data of the individual to be analyzed and the length of each common ancestor gene data fragment; If it is determined that the kinship coefficient of the individual to be analyzed is greater than a first threshold, determining the kinship level of the individual to be analyzed according to the kinship coefficient of the individual to be analyzed; Otherwise, the kinship level of the individuals to be analyzed is determined according to the length of each of the common ancestor gene data segments.

5. The method according to claim 4, characterized in that Determining the kinship coefficient of the individual to be analyzed according to the length of the gene data of the individual to be analyzed and the length of each of the common ancestor gene data fragments includes: Determine the sum of the lengths of the common ancestor gene data fragments as the total length; The ratio of the total length to the length of the gene data of the individual to be analyzed is determined as the kinship coefficient of the individual to be analyzed.

6. The method according to claim 4, characterized in that Determining the kinship level of the individual to be analyzed according to the kinship coefficient of the individual to be analyzed includes: Determining a coefficient range corresponding to the kinship coefficient of the individual to be analyzed according to a first preset mapping relationship and the kinship coefficient of the individual to be analyzed; wherein the first preset mapping relationship represents a corresponding relationship between the kinship coefficient and the coefficient range; and the coefficient range is a value range of the kinship coefficient; According to the second preset mapping relationship and the determined coefficient range, the kinship level corresponding to the determined coefficient range is determined as the kinship level of the individual to be analyzed; wherein the second preset mapping relationship is the correspondence between the coefficient range and the kinship level.

7. The method according to claim 4, characterized in that Determining the level of kinship of the individuals to be analyzed according to the lengths of the common ancestor gene data segments includes: Determine the sum of the lengths of the common ancestor gene data fragments as the total length; According to a third preset mapping relationship and the total length, determining a length range corresponding to the total length; wherein the third preset mapping relationship represents a corresponding relationship between the total length and the length range; the length range is a value range of the total length of the common ancestor gene data segment; According to the fourth preset mapping relationship and the determined length range, the kinship level corresponding to the determined length range is determined as the kinship level of the individual to be analyzed; wherein the fourth preset mapping relationship is the correspondence between the length range and the kinship level.

8. The method according to any one of claims 1 to 7, characterized in that Before determining the kinship levels corresponding to the pair of individuals to be analyzed based on the common ancestor gene data segments corresponding to the pair of individuals to be analyzed, the method further includes: If it is determined that the length of the common ancestor gene data segment is less than a preset length, then the common ancestor gene data segment is removed; and / or, If it is determined that the error rate of the common ancestral gene data segment is greater than or equal to a preset error rate threshold, the common ancestral gene data segment is removed; wherein the error rate indicates the proportion of heterozygous genotypes in the common ancestral gene data segment.

9. A rapid inference device for kinship based on common ancestral fragments, characterized in that: include: An acquisition module, used to acquire a gene data set; wherein the gene data set includes gene data of each individual to be analyzed in a pair of individuals to be analyzed; A segmentation module, used to segment the gene data of the individual to be analyzed to obtain a plurality of gene data segments of the individual to be analyzed; wherein the gene data segment includes at least one gene data; An encoding module, used to encode the genetic data segments of the individual to be analyzed to obtain a first encoding set and a second encoding set of the genetic data segments of the individual to be analyzed; wherein the first encoding set includes the position of the genetic data whose encoding result is the first encoding result in the genetic data segment; and the second encoding set includes the position of the genetic data whose encoding result is the second encoding result in the genetic data segment; A determination module, configured to determine a common ancestor gene data segment corresponding to a pair of individuals to be analyzed according to a first coding set and a second coding set of each gene data segment in the pair of individuals to be analyzed; The determination module is further used to determine the kinship levels corresponding to the pair of individuals to be analyzed based on the common ancestor gene data segments corresponding to the pair of individuals to be analyzed.

10. An electronic device, characterized in that: include: Memory, processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the method according to any one of claims 1 to 8.