Method for predicting genetic relationship within three levels among multiple individuals based on deep learning

The two-individual, mother-child, and full-sibling models established through deep learning algorithms solve the problem of insufficient accuracy of traditional methods in judging kinship, realize efficient kinship prediction in public security cases, and are suitable for a variety of detection systems.

CN120656535APending Publication Date: 2025-09-16WUHU PUBLIC SECURITY BUREAU
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510619454.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing technologies make it difficult to accurately determine the kinship relationship between two or more individuals, especially when public security organs are handling major criminal cases, anti-trafficking and family search, and disaster events. Traditional methods cannot effectively distinguish between full siblings, half siblings, and more distant relatives, and have high requirements for experimental data quality and analysis accuracy.

Method used

A deep learning algorithm is used to establish three models: two-individual, mother-and-son, and full-sibling models. Through random simulation of family generation and data preprocessing, autosomal STR genetic markers are used to predict kinship, including the two-individual model, the mother-and-son model, and the full-sibling model, which are suitable for various detection systems.

Benefits of technology

The accuracy of kinship prediction has been improved, and it can accurately predict first-, second-, and third-degree kinship and unrelated individuals in any detection system. It is suitable for criminal case investigation, anti-trafficking and family search, and identity recognition in disaster events.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120656535A_ABST
    Figure CN120656535A_ABST
Patent Text Reader

Abstract

A method for predicting genetic relationships within three levels among multiple individuals based on deep learning comprises the steps that S1, a random simulation family is generated, simulated random individuals and simulated random families of various detection systems are generated through a random simulation method, and the random families comprise father children, full sibs, half sibs and the like; the kit is singly used or used in combination to form a detection system of 15-74 gene loci; s2, data preprocessing: (1) a two-individual model, (2) a parent-mother-child-individual model, and (3) a full-sib-individual model; 60% of data of each model is randomly selected as a training set, 10% of data is randomly selected as a verification set, and 30% of data is randomly selected as a test set; s3, modeling and evaluation are carried out, a deep neural network algorithm is used, and models are all multi-feature-input single-output multi-classification models; and comparing the predicted value with the real genetic relationship of the corresponding sample group, and calculating the percentage of the genetic relationship consistent group as the accuracy rate of the specific model prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for kinship identification, in particular to an analysis method for kinship identification based on autosomal STR and deep learning prediction. Background Art

[0002] Autosomal STRs are currently one of the most important DNA genetic markers used for kinship determination. The Inductively Transmitted Tolerance (ITO) and Inductively Transmitted Bases (IBS) methods are the primary analytical methods for autosomal STR kinship determination. The "Technical Specifications for Identifying Full Sibling Relationships in Biology" use the ITO and IBS methods to determine full sibship between two individuals. The "Technical Specifications for Identifying Half Sibling Relationships in Biology" use the ITO method to determine half-sibship between two pairs, half-sibship between three pairs, and half-sibship between four pairs. However, in practical applications, public security agencies often need to determine the specific level of kinship between two or more individuals, providing key suspect clues for investigation and solving crimes.

[0003] However, in cases handled by public security agencies, such as investigating the relatives of suspects in major criminal cases, locating relatives in anti-trafficking operations, identifying missing persons, and identifying the source of corpses in disasters, it is often necessary to accurately determine the kinship between two or more individuals. To fully utilize the massive autosomal STR data already available in public security agency DNA databases and anti-trafficking databases, algorithms for determining kinship based on autosomal STR typing have high practical application value. Traditional methods can only calculate IBS and ITO values ​​based on autosomal STR typing of two or more individuals, using standardized thresholds to determine whether they are full siblings or half siblings. Lu Huiling et al. reported that by calculating the full sibling index (FSI) and half sibling index (HSI), threshold methods and discriminant functions can be used to distinguish between full siblings and unrelated individuals, half siblings and unrelated individuals, and full siblings and half siblings. Liu Jing et al. reported the construction of an algorithm for predicting distant kinship based on high-density SNP genetic markers using the shared ancestor (IBD) fragment algorithm. However, there is still no report on the algorithm for determining the close relationship between two or more individuals such as full siblings, half siblings, and first cousins ​​based on autosomal STR genetic markers.

[0004] CN202010423850.8 discloses a set of reagents and applications for the composite amplification detection of 17 human autosomal STR and 28 Y chromosome STR loci. The reagent set disclosed in the present invention consists of 86 single-stranded DNAs shown in sequence 1 86 in the sequence table, and the single-stranded DNAs with even sequence numbers are labeled with a fluorescent substance. All amplified fragments of the reagent set are less than 580bp and can be used for individual identification, kinship retrieval, analysis, and family investigation. The requirements for instruments and equipment are not high, and it can be used in cases involving mixed, constant, trace, and degraded samples, with a high individual recognition rate.

[0005] CN 202411318193.5 discloses a kit for simultaneously detecting STR loci and SNP sites and a method for using the kit. The kit includes amplification primers for amplifying the following 132 STR loci and 306 SNP sites, and probes for identifying the following 132 STR loci and 306 SNP sites.

[0006] FSI (Full Sibling Index) and HSI (Half Sibling Index) are statistical indicators used to assess the sibling relationship between two individuals in kinship testing. Their calculation is based on the probability of allele matching of genetic markers. The following are the core points of both:

[0007] FSI (Full Sibling Index): Used to determine the probability that two individuals are full siblings (same father and same mother). When FSI ≥ 19, full sibling relationships are supported.

[0008] HSI (Half Sibling Index): Used to determine the probability that two individuals are half-siblings (having only the same father or mother). When HSI ≥ 19, a half-sibling relationship is supported.

[0009] Unified algorithm basis: Both indices are based on the ITO method, which is calculated by analyzing allele frequencies and relatedness coefficients (r)46:

[0010] o FSI: When both individuals are heterozygous, the calculation formula involves adding 1 to the inverse of the state-identical allele frequency; if both individuals are homozygous, a different formula is used. HSI: The calculation logic is similar, but the kinship coefficient (r) is different (r = 0.5 for full sibs and r = 0.25 for half sibs).46

[0011] Full-sib identification: An FSI ≥ 19 supports full-sib relationships, with 96.4% accuracy in distinguishing unrelated individuals. This standard applies to cases where both parents are missing and uses the ITO method to calculate the full-sib relationship index based on typing results at 15 STR loci.

[0012] Half-sibling identification: When HSI ≥ 19, half-sibling relationships are supported with an accuracy rate of 85.3%. Also based on STR locus typing, it is suitable for scenarios where only partial kinship needs to be determined (such as single-parent tracing).

[0013] Sibling type differentiation: Full siblings and half siblings can be differentiated by comparing the ratio of FSI to HSI (e.g., FSI ≥ 1), with an accuracy rate of 87.5%.

[0014] Standard Basis: The calculation of FSI and HSI complies with the "Technical Specification for Identification of Full Sibling Relationships in Biology" (SF / T 0117-2021) and the "Technical Specification for Identification of Half Sibling Relationships in Biology" (SF / T 0131-2023). The ITO method is used to calculate the index and set thresholds. The standards clarify the inspection procedures, parameter calculations, and the requirements for issuing identification opinions.

[0015] Application scenarios: Commonly used for full-sibling or half-sibling relationship identification in the absence of both parents. It is not suitable for more distant relatives (such as cousins).

[0016] Genetic information dilution: As generations of kinship increase (e.g., grandparents and grandchildren, great-grandparents and grandchildren), the probability of matching genetic markers decreases and the accuracy decreases.

[0017] Effect of heterozygosity: Calculations require distinguishing whether an individual is heterozygous, which places high demands on the quality of experimental data and analytical accuracy.

[0018] In summary, FSI and HSI provide a scientific basis for sibling relationship identification through statistical analysis of genetic markers, but their application needs to be combined with technical specifications and the rigor of experimental data.

[0019] The ITO method is a kinship index calculation method based on genetic markers (such as STR loci). It is primarily used to assess sibling relationships (e.g., full siblings, half siblings) or other related relationships (e.g., uncle-nephew, grandparent-grandchild). Its core is to calculate relationship indices (e.g., FSI, HSI) by analyzing allele frequencies and relatedness coefficients (r).

[0020] Core formula: The calculation of different kinship indices shares a common factor, that is, the inverse of the state consistency allele frequency plus 1. The specific formula is derived by combining the kinship coefficient (such as full sibling r = 0.5, half sibling r = 0.25).

[0021] The typical application steps of the ITO method are as follows: Genotyping: using a standardized STR detection system (such as PowerPlex TM 16) Obtain allele typing results for two individuals at multiple loci (e.g., 15 STR loci).

[0022] Matching statistics: Statistics of allele matching for each pair of loci, divided into three categories:

[0023] All different (x0): no shared alleles;

[0024] Half identical (x1): shares one allele;

[0025] Identical (x2): Shares two alleles.

[0026] This paper is the first to use a deep learning algorithm to study the problem of autosomal STR kinship analysis and establish three models: two-individual, mother-offspring-individual, and full-sibling-individual. These models cover the most common kinship prediction application scenarios with high accuracy and broad practical application value. Summary of the Invention

[0027] The present invention aims to provide a method for determining close kinship relationships between two or more individuals, such as full siblings, half-siblings, and first cousins, based on autosomal STR genetic markers. Three deep learning kinship prediction models based on first-generation STR test data are proposed: a two-individual model, a mother-child model, and a full-sibling model. Predictable kinship relationships include first-degree, second-degree, third-degree, and unrelated individuals. In any testing system, the two-individual model can accurately predict the kinship between two individuals; the mother-child model can accurately predict the kinship between the mother-child and the suspected father's close relatives; and the full-sibling model can accurately predict the kinship between two full siblings and the suspected individual.

[0028] The technical solution of the present invention is a method for predicting the kinship within the third degree among multiple individuals based on deep learning, comprising the following steps:

[0029] Step S1: Random simulated pedigree generation. Random simulation methods are used to generate simulated random individuals and simulated random pedigrees for various detection systems. The random pedigrees include father and son, full siblings, half siblings, and first cousins. The detection systems include: Identifiler Plus (ABI), VeriFiler Plus (ABI), PowerPlex 21 (Promega), AGCU EX30 (Zhongde Midland), AGCU EX38 (Zhongde Midland), AGCU 21+1 (Zhongde Midland), AGCU21HS (Zhongde Midland), and AGCU 21+1FS (Zhongde Midland). These kits, used alone or in combination, can form a detection system covering 15 to 74 loci.

[0030] Step S2: Data preprocessing. Based on the application scenario, it is divided into the following three models:

[0031] (1) Two-body model: predict the relationship between two individuals based on autosomal STR typing; the training data is the simulated father-son, full sibling, half sibling, first cousin, and unrelated individual typing groups; based on the autosomal typing of the two individuals, the cumulative parentage index (CPI) of the two individuals is used. duo ) formula to calculate the CPI of two individuals duo , calculate the full sibling ITO of two individuals according to the ITO formula (ITO HS )、Half-sibling ITO(ITO FS ), the first generation of ITO (ITO 1C ).

[0032] (2) Mother-child-individual model: predict the relationship between the biological mother and child and the suspected individual based on the autosomal STR typing. The training data are simulated biological mother and child + biological father, biological mother and child + biological father's full siblings, biological mother and child + biological father's half siblings, biological mother and child + biological father's first cousins, biological mother and child + unrelated individual typing groups. According to the autosomal typing of the biological mother and child and the suspected individual, according to the triplet cumulative paternity index (CPI trio ) formula to calculate the CPI of the biological mother and child and the suspected individual trio The IBS of the parent-child-individual was calculated according to the formula of the cumulative modified consistency score (CIBS) of the parent-child-individual (see Table 3), and the ITO of the child and the suspected individual was calculated according to the ITO formula (ITO HS )、Half-sibling ITO(ITO FS ), the first generation of ITO (ITO 1C ).

[0033] (3) Full-sibling-individual model: predicts the relationship between two known full siblings and a suspected individual based on autosomal STR typing. The training data are simulated typing groups of two full siblings + full siblings, two full siblings + half siblings, two full siblings + first cousins, and two full siblings + unrelated individuals.

[0034] The data for each model was divided into three independent datasets: 60% of the data was randomly selected as the training set for model training and parameter optimization; 10% of the data was randomly selected as the validation set for adjusting the model's hyperparameters, such as learning rate and training time; and 30% of the data was randomly selected as the test set for testing the accuracy of the deep learning model in assessing kinship.

[0035] In step S3 (modeling and evaluation), the three models built using the independently developed deep neural network algorithm are all multi-classification models with multiple feature inputs and a single output. By inputting the training set and the validation set into the improved deep learning model at the same time, a series of complex operations such as feature extraction and reverse transfer parameter adjustment are performed iteratively, and the hyperparameters of the model are adjusted using the validation set to obtain the best deep learning model. The test set that has never participated in the model construction is used to test the model. The typing of the test set of each model is calculated with the same preprocessing, and the obtained parameters are input into the corresponding model to obtain the predicted value. The predicted value is compared with the actual kinship of the corresponding sample group, and the percentage of groups with consistent kinship is calculated as the accuracy of the specific kinship prediction of the specific model.

[0036] In step S1, a random simulation family generation algorithm is designed, specifically as follows:

[0037] (1) Kinship classification

[0038] The kinship relationships and corresponding kinship coefficients to be determined in the present invention are shown in Table 1.

[0039] Table 1 Kinship and kinship coefficient

[0040]

[0041] (2) Random simulation of family generation

[0042] Randomly simulate the autosomal STR locus typing of a certain testing system for N families. Each family has a kinship group consisting of father and son, siblings, half-siblings, first cousins, and second cousins. N kinship groups are simulated for each kinship group. The specific steps are as follows:

[0043] (1) Random Individual: At a single locus, two alleles are randomly generated based on the allele frequencies of the test kit to form the genotype of a random individual. Similarly, all autosomal alleles of the test system are randomly generated to form the complete STR profile of a random individual. This design simulates the genotype of a random Chinese Han population using a test system.

[0044] (2) Random child: At a single locus, one allele from the father and one allele from the mother are randomly selected to form the two alleles of the child, with a probability of 50%. Similarly, the sampling operation is performed at all autosomal loci in the test system to form the complete STR profile of a random child.

[0045] (3) Pedigree of N group family: Family diagram see Figure 2, including four generations, seven random individuals, and seven pairs of "father-mother-son" relationships, totaling 14 individuals. This pedigree includes five different relatedness relationships: ① Father-son (PO): I1 and II2; ② Sibling (FS): II2 and II4; ③ Second-degree relatedness (D2): III1 and III2; and ④ Third-degree relatedness (D3): III2 and III3. By simulating this pedigree N times, all STR loci for these N families are generated. Figure 2 simulated pedigree chart;

[0046] (3) Sample sources of each detection system

[0047] Random simulations were used to generate simulated random individuals and families (including father-son, full siblings, half-sibs, and first cousins) for various testing systems (see Table 1). The testing systems included eight kits: Identifiler Plus (ABI), VeriFiler Plus (ABI), PowerPlex 21 (Promega), AGCU EX30 (Zhongde Midland), AGCU EX38 (Zhongde Midland), AGCU 21+1 (Zhongde Midland), AGCU 21HS (Zhongde Midland), and AGCU21+1FS (Zhongde Midland). These kits, used individually or in combination, formed testing systems covering 15 to 74 loci. Allele probabilities were based on gene frequencies in the Chinese Han population.

[0048] The deep learning training process uses the Adam optimizer to train the input parameters, and the mean squared error (MSE) is used as the loss function. The deep neural network is set to three layers, with 10 neural nodes in each layer 1-2. The labeled parameters are used as the training target, and the number of training cycles is set to 500. By simultaneously inputting the training set and the validation set into the improved deep learning model, a series of complex operations such as feature extraction and backpropagation are continuously iterated to increase parameters. The validation set is used to adjust the model's hyperparameters to obtain the optimal deep learning model.

[0049] Three deep learning-based kinship prediction models based on first-generation STR test data are available: the two-individual model, the mother-son model, and the full-sib model. Predicted relatedness includes first-degree, second-degree, third-degree, and unrelated individuals. Within any testing system, the two-individual model accurately predicts the relatedness between two individuals; the mother-son model accurately predicts the relatedness between the mother and son and the suspected father's close relatives; and the full-sib model accurately predicts the relatedness between two full siblings and the suspected individual.

[0050] This paper establishes three deep learning kinship prediction models based on first-generation STR test data: a two-individual model, a mother-child model, and a full-sib model. These predict relatedness includes first-degree, second-degree, third-degree, and unrelated individuals. In any test system, the two-individual model accurately predicts the relatedness between two individuals; the mother-child model accurately predicts the relatedness between the mother and child and the suspected father's close relatives; and the full-sib model accurately predicts the relatedness between two full siblings and the suspected individual.

[0051] Beneficial effect: The present invention uses a deep learning algorithm for the first time to study the problem of autosomal STR kinship detection and judgment, and establishes three models: two individuals, mother-child-individual, and full sibling-individual, covering the most common kinship prediction application scenarios, with high accuracy and wide practical application value. The present invention is a method for predicting kinship between multiple individuals in the field of forensic DNA, which performs AI prediction on whether multiple individuals belong to the first-level (single parent and full sibling), second-level, third-level or unrelated individuals. The invention is based on STR detection and includes three models: a two-individual model, a mother-child model, and a full sibling model. Among them, the two-individual model can predict the kinship between two individuals, the mother-child model can predict the kinship between the mother and child and the close relatives of the suspected father, and the full sibling model can predict the kinship between two full siblings and the suspected individual. The above three models can be used for prediction using any test kit or a test kit combined with a detection system. The invention can play an important role in cases such as criminal suspect kinship screening, anti-trafficking and family search, missing persons identification, and disaster corpse source identification. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 This is a simplified modeling diagram of the kinship prediction model of the present invention

[0053] Figure 2 simulated pedigree chart;

[0054] Figure 3 Schematic diagrams of three models. DETAILED DESCRIPTION

[0055] As shown in the figure, in step S1, a random simulation family generation algorithm is designed, specifically as follows:

[0056] (1) Kinship Classification The kinship relationships and corresponding kinship coefficients to be determined in the present invention are shown in Table 1.

[0057] Table 1 Kinship and kinship coefficient

[0058] Kinship coefficient Kinship example

[0059] First degree kinship (D1) 1 / 2 single parent (PO, i.e. father and son and mother and son), full siblings (FS, i.e. brothers and sisters with the same father and mother)

[0060] Second degree kinship (D2): 1 / 4 half siblings (i.e. half brothers and sisters), grandparents, uncles and nephews

[0061] Third degree kinship (D3): 1 / 8th generation cousins, great-grandsons

[0062] Unrelated Entity (UN) 0

[0063] (2) Random simulation of family generation

[0064] Randomly simulate the autosomal STR locus typing of a certain testing system for N families. Each family has a kinship group consisting of father and son, siblings, half-siblings, first cousins, and second cousins. N kinship groups are simulated for each kinship group. The specific steps are as follows:

[0065] (1) Random Individual: At a single locus, two alleles are randomly generated based on the allele frequencies of the test kit to form the genotype of a random individual. Similarly, all autosomal alleles of the test system are randomly generated to form the complete STR profile of a random individual. This design simulates the genotype of a random Chinese Han population using a test system.

[0066] (2) Random child: At a single locus, one allele from the father and one allele from the mother are randomly selected to form the two alleles of the child, with a probability of 50%. Similarly, the sampling operation is performed at all autosomal loci in the test system to form the complete STR profile of a random child.

[0067] (3) Pedigree of N group family: Family diagram see Figure 2 , including four generations, seven random individuals, and seven pairs of "father-mother-son" relationships, totaling 14 individuals. This pedigree includes five different relatedness relationships: ① Father-son (PO): I1 and II2; ② Sibling (FS): II2 and II4; ③ Second-degree relatedness (D2): III1 and III2; and ④ Third-degree relatedness (D3): III2 and III3. By simulating this pedigree N times, all STR loci for these N families are generated.

[0068] Data preprocessing: The training data of the present invention are all derived from random simulated families. Before training, the STR typing of each group needs to be preprocessed and converted into probability values ​​through calculation.

[0069] (3) Sample sources of each detection system

[0070] Random simulation was used to generate simulated random individuals and families (including fathers and sons, full siblings, half siblings, and first cousins) for various testing systems (see Table 1). The testing systems included eight kits: Identifiler Plus (ABI), VeriFiler Plus (ABI), PowerPlex 21 (Promega), AGCU EX30 (Zhongde Midland), AGCU EX38 (Zhongde Midland), AGCU 21+1 (Zhongde Midland), AGCU 21HS (Zhongde Midland), and AGCU21+1FS (Zhongde Midland), as shown in Table 2. These kits were used individually or in combination to create testing systems covering 15 to 74 loci. Allele probabilities were based on gene frequencies in the Chinese Han population.

[0071] Table 2 Detection system and number of loci

[0072]

[0073] In step S2, the data preprocessing algorithm is as follows:

[0074] Data preprocessing

[0075] The training data of the present invention are all derived from random simulated families. Before training, the STR typing of each group needs to be preprocessed and converted into probability values ​​through calculation.

[0076] (1) Two-body model: Based on the autosomal typing of the two individuals, the cumulative parentage index (CPI) of the two individuals is used. duo ) formula to calculate the CPI of two individuals duo , calculate the full sibling ITO of two individuals according to the ITO formula (ITO HS )、Half-sibling ITO(ITO FS ), the first generation of ITO (ITO 1C ), a total of four parameters. The number of loci in the detection system serves as the fifth parameter. The kinship groups include: PO, FS, D2, D3, and UN, labeled 0 to 4, respectively, serving as the sixth parameter (marker parameter). A total of six parameters serve as variable inputs to the model.

[0077] Predict the relationship between two individuals based on autosomal STR typing. The training data consists of simulated typing groups of father and son, full siblings, half siblings, first cousins, and unrelated individuals, with 10,000 individuals in each group. The detection system includes all 10 groups listed in Table 1. The total number of typing groups is 500,000.

[0078] (2) Parent-child-individual model: Based on the autosomal typing of the parent-child and suspected individual, the cumulative parentage index (CPI) of the triplet is used to determine the parent-child relationship. trio) formula to calculate the CPI of the biological mother and child and the suspected individual trio The IBS of the parent-child-individual was calculated according to the formula of the cumulative modified consistency score (CIBS) of the parent-child-individual (see Table 3), and the ITO of the child and the suspected individual was calculated according to the ITO formula (ITO HS )、Half-sibling ITO(ITO FS ), the first generation of ITO (ITO 1C ), a total of 5 parameters. The number of loci in the detection system serves as the sixth parameter. The kinship group includes: biological father (BF), D1, D2, D3, and UN, labeled 0 to 4, serving as the seventh parameter (marker parameter). A total of 7 parameters serve as variable inputs to the model. The kinship between biological mother and child and suspected individuals is predicted based on autosomal STR typing. The training data consists of simulated typing groups of biological mother and child, biological mother and child and biological father's full siblings, biological mother and child and biological father's half siblings, biological mother and child and biological father's first cousins, and biological mother and child and unrelated individuals, with 10,000 individuals in each group, for a total of 10 detection systems. The total number of typing groups is 500,000.

[0079] (3) Full-sibling-individual model: In a certain detection system, the full-sibling ITO (ITO) of full-sibling A and suspected individual is calculated based on the autosomal typing of full-sibling A and suspected individual. HS1 )、Half-sibling ITO(ITO FS1 ), the first generation of ITO (ITO 1C1 ), and full sibling B and full sibling ITO of the suspected individual (ITO HS2 )、Half-sibling ITO(ITO FS2 ), the first generation of ITO (ITO 1C2 ), for a total of 6 parameters. The number of loci in the detection system serves as the 7th parameter. The relatedness groups include: D1, D2, D3, and UN, labeled 0 to 3, respectively, serving as the 8th parameter (marker parameter). A total of 8 parameters serve as variable inputs to the model. The relatedness between two known full siblings and a suspected individual is predicted based on autosomal STR typing. The training data consists of simulated typing groups of 2 full siblings + full siblings, 2 full siblings + half siblings, 2 full siblings + first cousins, and 2 full siblings + unrelated individuals, with 10,000 individuals in each group, for a total of 10 detection systems. The total number of typing groups is 400,000.

[0080] The data for each model was divided into three independent datasets: 60% of the data was randomly selected as the training set for model training and parameter optimization; 10% of the data was randomly selected as the validation set for adjusting the model's hyperparameters, including learning rate and training time; and 30% of the data was randomly selected as the test set for testing the accuracy of the deep learning model in assessing kinship.

[0081] Model Implementation

[0082] The deep learning training process uses the Adam optimizer to train the input parameters, and the mean squared error (MSE) is used as the loss function. The deep neural network is set to 3 layers, with 10 neural nodes in each layer 1-2. The labeled parameters are used as the training target, and the number of training cycles is set to 500. By simultaneously inputting the training set and the validation set into the improved deep learning model, feature extraction and parameter adjustment are continuously iterated through backpropagation.

[0083] Model and dataset

[0084] This paper uses a self-developed deep neural network algorithm and JavaScript language. The three models built are all multi-classification models with multiple feature inputs and a single output. According to the application scenario, they are divided into the following three models:

[0085] (1) Two body models;

[0086] (2) the parent-child-individual model;

[0087] (3) Full-sibling-individual model;

[0088] After calculation, the validation set is used to adjust the model's hyperparameters to obtain the optimal deep learning model.

[0089] (1) Maternal Suspected Father's Rights Index (PI) trio )algorithm

[0090] The paternity index (CPI) is the confidence probability of multiple loci in paternity testing. trio ) refers to the ratio of the probability (X) that the genotypes of a hypothetical father and biological mother and child at a single locus are consistent with the parental relationship to the probability (Y) that the genotypes of a random father and biological mother and child are consistent with the parental relationship. Suppose the allele frequencies of alleles A, B, C, D, E, and F at a given locus are a, b, c, d, e, and f.

[0091] <1> PI that conforms to genetic laws trio Formula, see Table 3

[0092] Table 3 PIs that conform to genetic laws trio formula

[0093]

[0094]

[0095] Cumulative paternity index (CPI) is the confidence probability of multiple loci in paternity testing. Let n be the number of loci, PI i is the PI of the i-th locus. CPI formula:

[0096]

[0097] The above formula covers the confidence probability of all types of paternity tests, including cases that conform to genetic laws and mutation cases.

[0098] <2> PI that conforms to genetic laws trio formula

[0099] The mutation rate (μ) refers to the rate at which mutations occur in each generation of cells. The formula for μ is:

[0100] μ = number of mutant alleles / number of meiotic divisions

[0101] Let A and B be the alleles of the parent and offspring respectively, f represents the father, and m represents the mother. Let μ f(A→B) It means that if the father's allele A mutates to B, μ m(A→B) Indicates that the mother's allele A mutated to B. μ f(A→B) and μ m(A→B) The calculation of is based on the empirical decreasing model. Assuming the number of mutation steps is λ and the ratio of paternal mutation to maternal mutation is α, the formula is:

[0102] μ f(A→A+λ) =μ f(A→A-λ) =(μ / 2)×10 1-λ

[0103] μ m(A→A+λ) =μ m(A→A-λ) =(μ / 2α)×10 1-λ

[0104] Assumptions: ① There is only one mutation in a case; ② Mutations are not considered when they are heritable; ③ Mutations are only considered when they can provide allele changes for children; ④ Mutations are not considered in random individuals.

[0105] In practical applications, μ=0.002 and α=3.5.

[0106] In paternity testing under mutation conditions, usually only some loci are mutated. Therefore, the situation should be judged first according to the genetic laws, and the PE and PI should be calculated at each locus using the corresponding formulas. Finally, the cumulative probability should be calculated using the CPE and CPI formulas.

[0107] <3> Single mutation PI trio Formula, see Table 4

[0108] Table 4 PI of single mutation trio formula

[0109]

[0110]

[0111] <4> Secondary mutation of PI trio Formula, see Table 5:

[0112] Table 5 PI of secondary mutation trio formula

[0113]

[0114] (2) Parentage Index (PI) duo )algorithm

[0115] The parentage index (PI) FC ) refers to the ratio of the probability (X) that the genotypes of the hypothetical father and biological mother and child at a single locus are consistent with the parental relationship to the probability (Y) that the genotypes of the random father and biological mother and child are consistent with the parental relationship.

[0116] <1> PI that conforms to genetic laws duo Formula, see Table 6

[0117] Table 6 PIs that conform to genetic laws duo formula

[0118]

[0119] <2> Mutated PI duo Formula, see Table 7

[0120] Table 7 Mutated PI duo formula

[0121]

[0122]

[0123] (3) ITO formula, see Table 8

[0124] Table 8 ITO formula

[0125]

[0126] (4) Modified IBS formula for mother and child

[0127] First, the suspected paternal gene is inferred based on the genotypes of the mother and child, and then compared with the suspected individual. If the father and the suspected individual do not share the same gene, the IBS value is assigned to 0, the same as the traditional IBS method. If the father and the suspected individual share gene A (frequency a), the IBS value is assigned to 1-a. Because the higher the frequency of the shared gene, the greater the probability of coincidence and the lower the likelihood of hereditary inheritance, the lower the value of this shared gene, and the lower the IBS value (the lowest being 0). Conversely, the lower the frequency of the shared gene, the higher the IBS value (the highest being 1). The modified IBS formula for the mother, child, and individual is shown in Table 9.

[0128] Table 9 Modified IBS formula for mother and child

[0129]

[0130] Cumulative Consistency of Improved Status Score (CIBS) formula:

[0131]

[0132] (5) Model data input

[0133] The three models built in this invention are all multi-classification models with multiple feature inputs and a single output. The data input methods of the three models are:

[0134] (1) Two-body model: Based on the autosomal typing of two individuals, according to CPI duo Formula to calculate the CPI of two individuals duo , calculate the full sibling ITO of two individuals according to the ITO formula (ITO HS )、Half-sibling ITO(ITO FS ), the first generation of ITO (ITO 1C ), a total of four parameters. The number of loci in the detection system serves as the fifth parameter. The kinship groups include: PO, FS, D2, D3, and UN, labeled 0 to 4, respectively, serving as the sixth parameter (marker parameter). A total of six parameters serve as variable inputs to the model.

[0135] (2) Parent-child-individual model: Based on the autosomal typing of the parent-child and suspected individual, the cumulative parentage index (CPI) of the triplet is used to determine the parent-child relationship. trio ) formula to calculate the CPI of the biological mother and child and the suspected individual trio The IBS of the parent-child-individual was calculated according to the formula of the cumulative modified consistency score (CIBS) of the parent-child-individual (see Table 3), and the ITO of the child and the suspected individual was calculated according to the ITO formula (ITO HS )、Half-sibling ITO(ITO FS ), the first generation of ITO (ITO 1C), for a total of five parameters. The number of loci in the detection system serves as the sixth parameter. The kinship groups include the father (BF), D1, D2, D3, and UN, labeled 0 to 4, serving as the seventh parameter (marker parameter). A total of seven parameters serve as variable inputs to the model.

[0136] (3) Full-sibling-individual model: In a certain detection system, the full-sibling ITO (ITO) of full-sibling A and suspected individual is calculated based on the autosomal typing of full-sibling A and suspected individual. HS1 )、Half-sibling ITO(ITO FS1 ), the first generation of ITO (ITO 1C1 ), and full sibling B and full sibling ITO of the suspected individual (ITO HS2 )、Half-sibling ITO(ITO FS2 ), the first generation of ITO (ITO 1C2 ), for a total of 6 parameters. The number of loci in the detection system serves as the 7th parameter. The kinship groups D1, D2, D3, and UN, labeled 0 to 3, serve as the 8th parameter (marker parameter). A total of 8 parameters serve as variable inputs to the model.

[0137] The data for each model was divided into three independent datasets: 60% of the data was randomly selected as the training set for model training and parameter optimization; 10% of the data was randomly selected as the validation set for adjusting the model's hyperparameters, such as learning rate and training time; and 30% of the data was randomly selected as the test set for testing the accuracy of the deep learning model in assessing kinship.

[0138] In step S3, the modeling and evaluation algorithm is as follows:

[0139] (1) Modeling method

[0140] Deep learning utilizes a proprietary deep neural network algorithm developed in JavaScript. The training process uses the Adam optimizer to optimize input parameters, with the mean squared error (MSE) as the loss function. The deep neural network is configured with three layers, with 10 nodes in each of layers 1-2. Model training targets the labeled parameters, with a training cycle of 1500. By simultaneously feeding the training and validation sets into the improved deep learning model, a series of complex operations, including feature extraction and parameter adjustment, are iterated. The validation set is used to adjust the model's hyperparameters to achieve the optimal deep learning model.

[0141] (2) Evaluation of model performance

[0142] The models were tested using a test set that had never been used in model construction. The typing of the test set for each model was calculated using the same preprocessing method, and the resulting parameters were input into the corresponding model to obtain predicted values. The predicted values ​​were then compared with the actual kinship relationships of the corresponding sample groups. Groups with consistent kinship relationships were referred to as the matching groups. Each model's kinship group was validated against 10,000 groups, and the percentage of matching groups was calculated as the accuracy of the specific kinship prediction for that particular model. This was used to evaluate the performance of the three kinship prediction models developed by the present invention.

[0143] The prediction accuracy (AC) of the three models was tested using the test set that did not participate in training. A total of 30 groups were tested, see Tables 10-12.

[0144] Table 10 Prediction accuracy of the two-individual model

[0145]

[0146] Table 11 Prediction accuracy of the parent-child-individual model

[0147]

[0148]

[0149] Table 12 Prediction accuracy of full sibling-individual model

[0150]

[0151] The above description is merely a specific embodiment of the present invention and does not limit the scope of the invention. Therefore, the substitution of equivalent components, or equivalent changes and modifications made within the scope of protection of the present invention, shall still fall within the scope of the present invention. In addition, the technical features of the present invention may be freely combined with each other, with each other's technical solutions, and with each other's technical solutions.

Claims

1. A method for predicting kinship within three degrees among multiple individuals based on deep learning, characterized in that: The following steps are involved: Step S1: random simulated pedigree generation, using random simulation to generate simulated random individuals and simulated random pedigrees for various detection systems, wherein the random families include father and son, full siblings, half siblings, and first cousins; the detection systems include: Identifiler Plus, VeriFiler Plus, PowerPlex 21, AGCU EX30, AGCU EX38, AGCU 21+1, AGCU 21HS, and AGCU21+1FS kits, which can be used alone or in combination to form a detection system covering 15 to 74 loci; Step S2: Data preprocessing. Based on the application scenario, it is divided into the following three models: (1) Two-body model: predict the relationship between two individuals based on autosomal STR typing; the training data is simulated father-son, full siblings, half siblings, first cousins, and unrelated individual typing groups; based on the autosomal typing of the two individuals, the cumulative paternity index (CPI) of the two-body model is used. duo Formula to calculate the CPI of two individuals duo , calculate the full sibling ITO, half sibling ITO, and first cousin ITO of two individuals according to the ITO formula; (2) Mother-child-individual model: predict the relationship between the biological mother and child and the suspected individual based on the autosomal STR typing; the training data are simulated biological mother and child + biological father, biological mother and child + biological father's full siblings, biological mother and child + biological father's half siblings, biological mother and child + biological father's first cousins, biological mother and child + unrelated individual typing groups; based on the autosomal typing of the biological mother and child and the suspected individual, the cumulative paternity index (CPI) of the triplet is used. trio Formula to calculate the CPI of biological mother and child and suspected individuals trio The IBS of the parent-child relationship was calculated based on the CIBS formula (Table 3). The ITO formula was used to calculate the full sibling ITO, half sibling ITO, and first cousin ITO of the child and the suspected individual. (3) Full-sibling-individual model: predicts the relationship between two known full siblings and a suspected individual based on autosomal STR typing; the training data are simulated typing groups of 2 full siblings + full siblings, 2 full siblings + half siblings, 2 full siblings + first cousins, and 2 full siblings + unrelated individuals; The data for each model was divided into three independent datasets: 60% of the data was randomly selected as the training set for model training and parameter optimization; 10% of the data was randomly selected as the validation set for adjusting the model's hyperparameters: learning rate and training time; and 30% of the data was randomly selected as the test set for testing the accuracy of the deep learning model in assessing kinship. Step S3 is modeling and evaluation, using a deep neural network algorithm. The three models built are all multi-classification models with multiple feature inputs and a single output. By simultaneously inputting the training set and the validation set into the improved deep learning model, feature extraction and reverse transfer parameter adjustment operations are performed iteratively, and the hyperparameters of the model are adjusted using the validation set to obtain the best deep learning model. The test set that has never participated in model construction is used to test the model. The typing of the test set of each model is calculated with the same preprocessing, and the obtained parameters are input into the corresponding model to obtain the predicted value. The predicted value is compared with the actual kinship of the corresponding sample group, and the percentage of groups with consistent kinship is calculated as the accuracy of the specific kinship prediction of the specific model.

2. The method for predicting the relationship within the third degree among multiple individuals based on deep learning according to claim 1, characterized in that: In step S1, a random simulation family generation algorithm is designed, specifically as follows: (1) Kinship classification The kinship relationships to be determined and the corresponding kinship coefficients are shown in Table 1; Table 1 Kinship and kinship coefficient (2) Random simulation of family generation Randomly simulate the autosomal STR locus typing of a certain testing system in N families; each family has a kinship group consisting of father and son, siblings, half-siblings, first cousins, and second cousins, and N kinship groups are simulated for each relationship group; the specific steps are as follows: (1) Random individual: At a single locus, two alleles are randomly generated based on the allele frequencies of the test kit to form the genotype of a random individual; similarly, all autosomal alleles of the test system are randomly generated to form the complete STR typing of a random individual; this design simulates the genotype of a random Chinese Han population with a test system; (2) Random child: At a single locus, one allele from the father and one allele from the mother are randomly selected to form the two alleles of the child, with a probability of 50%; similarly, the sampling operation is performed at all autosomal loci in the test system to form a complete STR profile of a random child; (3) N groups of family pedigrees: The family includes 4 generations, 7 random individuals, 7 pairs of "father-mother-son" relationships, and a total of 14 people; the pedigree contains 5 different kinship relationships: ① father-son (PO): Ⅰ1 and Ⅱ2; ② siblings (FS): Ⅱ2 and Ⅱ4; ③ second-level kinship relationship (D2): Ⅲ1 and Ⅲ2; ④ third-level kinship relationship (D3): Ⅲ2 and Ⅲ3; simulate the above pedigrees, repeat N times, and simulate all STR loci typing of N families. (3) Sample sources of each detection system Random simulation was used to generate simulated random individuals and pedigrees for various test systems, including father and son, full siblings, half siblings, and first cousins. Test systems are shown in Table 2 , used individually or in combination to form test systems with 15 to 74 loci. Allele probabilities were based on gene frequencies in the Han Chinese population. Table 2 Detection system and number of loci 3. The method for predicting the relationship within the third degree among multiple individuals based on deep learning according to claim 1, characterized in that: In step S2, the data preprocessing algorithm is as follows: (1) Maternal Suspected Father's Rights Index (PI) trio )algorithm The paternity index (CPI) is the confidence probability of multiple loci in paternity testing; the paternity index (PI) of the mother and the father is the same. trio ) refers to the ratio of the probability (X) that the genotypes of a hypothetical father and biological mother and child at a single locus are consistent with the parental relationship to the probability (Y) that the genotypes of a random father and biological mother and child are consistent with the parental relationship; suppose the allele frequencies of alleles A, B, C, D, E, and F at a certain locus are a, b, c, d, e, and f; <1> PI that conforms to genetic laws trio Formula, see Table 3 Table 3 PIs that conform to genetic laws trio formula Cumulative paternity index (CPI) is the confidence probability of multiple loci in paternity testing; let n be the number of loci, PI i is the PI of the i-th locus; CPI formula: The above formula covers the confidence probability of all types of paternity tests, including cases that conform to genetic laws and mutation cases; <2> PI that conforms to genetic laws trio formula The mutation rate (μ) refers to the rate at which mutations occur in each generation of cells; the μ formula is: μ = number of mutant alleles / number of meiotic divisions Let A and B be the alleles of the parent and offspring respectively, f represents the father, and m represents the mother; let μ f(A→B) It means that if the father's allele A mutates to B, μ m(A→B) Indicates that the mother's allele A mutated to B; μ f(A→B) and μ m(A→B) The calculation is based on the empirical decreasing model; let the number of mutation steps be λ, and the ratio of paternal mutation to maternal mutation be α, then the formula is: m f(A→A+λ) =μ f(A→A-λ) =(μ / 2)×10 1-λ m m(A→A+λ) =μ m(A→A-λ) =(μ / 2a)×10 1-λ Assumptions: ① There is only one mutation in a case; ② Mutations are not considered if they are heritable; ③ Mutations are only considered if they can provide allele changes for children; ④ Random individuals are not considered mutations; In practical applications, μ = 0.002, α = 3.5; In the case of mutation, paternity testing usually only has mutations in some loci. Therefore, the situation should be determined first according to the genetic law, and the PE and PI should be calculated at each locus using the corresponding formula. Finally, the cumulative probability is calculated using the CPE and CPI formulas. <3> Single mutation PI trio Formula, see Table 4 Table 4 PI of single mutation trio formula <4> Secondary mutation of PI trio Formula, see Table 5: Table 5 PI of secondary mutation trio formula (2) Parentage Index (PI) duo )algorithm The parentage index (PI) FC ) refers to the ratio of the probability (X) that the genotypes of a hypothetical father and biological mother and child at a single locus are consistent with the parental relationship to the probability (Y) that the genotypes of a random father and biological mother and child are consistent with the parental relationship; <1> PI that conforms to genetic laws duo Formula, see Table 6 Table 6 PIs that conform to genetic laws duo formula <2> Mutated PI duo Formula, see Table 7 Table 7 Mutated PI duo formula (3) ITO formula, see Table 8 Table 8 ITO formula (4) Modified IBS formula for mother and child First, the suspected father's genes are inferred based on the genotypes of the mother and child, and then compared with the suspected individual. If the father and the suspected individual do not share the same genes, the IBS value is assigned to 0, the same as the traditional IBS method. If the father and the suspected individual share the same gene A (frequency a), the IBS value is assigned to 1-a. Because the higher the frequency of the same gene, the greater the probability of accidental similarity and the lower the possibility of hereditary similarity, the lower the value of this similarity, and the lower the IBS value should be (the lowest is 0). Conversely, the lower the frequency of the same gene, the higher the IBS value should be (the highest is 1). The modified IBS formula for the mother, child, and individual is shown in Table 9. Table 9 Modified IBS formula for mother and child Cumulative Consistency of Improved Status Score (CIBS) formula: (5) Model data input All three models are multi-classification models with multiple feature inputs and a single output. The data input methods for the three models are: (1) Two-body model: Based on the autosomal typing of two individuals, according to CPI duo Formula to calculate the CPI of two individuals duo , calculate the full sibling ITO of two individuals according to the ITO formula (ITO HS )、Half-sibling ITO(ITO FS ), the first generation of ITO (ITO 1C ), a total of 4 parameters; The number of loci in the detection system is used as the fifth parameter; the kinship groups include PO, FS, D2, D3, and UN, which are marked as 0 to 4, respectively, as the sixth parameter; a total of 6 parameters are used as variable inputs of the model; (2) Parent-child-individual model: Based on the autosomal typing of the parent-child and suspected individual, the cumulative parentage index (CPI) of the triplet is used to determine the parent-child relationship. trio ) formula to calculate the CPI of the biological mother and child and the suspected individual trio The IBS of the parent-child-individual was calculated according to the formula of the cumulative modified consistency score (CIBS) of the parent-child-individual (see Table 3), and the ITO of the child and the suspected individual was calculated according to the ITO formula (ITO HS )、Half-sibling ITO(ITO FS ), the first generation of ITO (ITO 1C ), a total of 5 parameters; the number of loci in the detection system is the sixth parameter; the kinship group includes: biological father (BF), D1, D2, D3 and UN, which are marked as 0 to 4 respectively, as the seventh parameter; a total of 7 parameters are used as variable inputs of the model; (3) Full-sibling-individual model: In a certain detection system, the full-sibling ITO (ITO) of full-sibling A and suspected individual is calculated based on the autosomal typing of full-sibling A and suspected individual. HS1 )、Half-sibling ITO(ITO FS1 ), the first generation of ITO (ITO 1C1 ), and full sibling B and full sibling ITO of the suspected individual (ITO HS2 )、Half-sibling ITO(ITO FS2 ), the first generation of ITO (ITO 1C2 ), a total of 6 parameters; the number of loci in the detection system is the 7th parameter; the kinship groups include: D1, D2, D3 and UN, marked as 0 to 3 respectively, as the 8th parameter; a total of 8 parameters are used as variable inputs of the model; The data for each model was divided into three independent datasets: 60% of the data was randomly selected as the training set for model training and parameter optimization; 10% of the data was randomly selected as the validation set for adjusting the model's hyperparameters, such as learning rate and training time; and 30% of the data was randomly selected as the test set for testing the accuracy of the deep learning model in assessing kinship.

4. The method for predicting the relationship within three levels among multiple individuals based on deep learning according to claim 1, characterized in that: In step S3, the modeling and evaluation algorithm is as follows: (1) Modeling method Deep learning uses a deep neural network algorithm and JavaScript. The training process uses the Adam optimizer to train the input parameters, and the mean squared error (MSE) loss function. The deep neural network is set to 3 layers, with 10 neural nodes in each layer 1-2. The labeled parameters were used as the target for model training, and the number of training cycles was set to 1500. The improved deep learning model was fed with both the training and validation sets, and a series of complex operations such as feature extraction and parameter adjustment were iteratively performed. The validation set was used to adjust the model's hyperparameters to obtain the optimal deep learning model. (2) Evaluation of model performance A test set that has never participated in model construction is used to test the model; the typing of the test set of each model is calculated with the same preprocessing, and the obtained parameters are input into the corresponding model to obtain the predicted value, which is compared with the actual kinship of the corresponding sample group, where the group with consistent kinship is called the matching group; each kinship group of each model is verified in 10,000 groups, and the percentage of the matching group is calculated as the accuracy of the specific kinship prediction of the specific model; this is used to evaluate the performance of the three kinship prediction models established by the present invention; and the prediction accuracy of the three models is tested using a test set that has not participated in training and learning.

Citation Information

Patent Citations

  • Complete set of reagent and application for complex amplification detection of 17 autosomal STR and 28 Y chromosome STR loci in human

    CN111621499A

  • Kit for simultaneously detecting STR (short tandem repeat) loci and SNP (single nucleotide polymorphism) loci and use method of kit

    CN118910286A