Hierarchical population-wide genome-wide kinship inference method and system

By using whole-genome principal component analysis and EigenGWAS regression, non-ancestor information markers were screened and statistically corrected, solving the problem of distinguishing kinship and population structure in stratified populations and improving the accuracy of kinship inference.

CN122290702APending Publication Date: 2026-06-26ZHEJIANG CHINESE MEDICAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG CHINESE MEDICAL UNIVERSITY
Filing Date
2026-03-19
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing technologies cannot effectively distinguish between kinship and population structure when processing genotype data of stratified populations, leading to incorrect inferences about kinship. This is especially true when it is necessary to assume population homogeneity, as existing methods require prior information or cluster analysis and have poor adaptability.

Method used

Principal component analysis was performed using whole-genome information to screen for non-ancestor information markers. EigenGWAS regression analysis was used to determine the significance p-value, and ancestor information markers were removed. Statistical significance was tested using theoretical expectation and asymptotic variance. Kinship coefficients were calculated and corrected.

Benefits of technology

It significantly improves the accuracy of kinship estimation in large-scale stratified populations, effectively controls the computational bias caused by group stratification, and improves the accuracy of kinship inference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122290702A_ABST
    Figure CN122290702A_ABST
Patent Text Reader

Abstract

This invention relates to the field of computational biology, specifically to a method and system for inferring kinship relationships in stratified populations across the entire genome. The method involves performing principal component analysis on the genotype data of the entire genome to determine c principal components for analysis, performing regression analysis to obtain the significance p-value of each molecular marker on each principal component, identifying molecular markers with significance p-values ​​exceeding a given threshold as ancestry information markers, and removing ancestry information markers to obtain non-ancestry information markers. Based on the non-ancestry information markers and the standardized genotype codes of all individuals on the non-ancestry information markers, the method calculates estimated kinship coefficients between all pairs of individuals. Statistical significance tests are performed on the estimated kinship coefficients based on theoretical expectation and asymptotic variance to determine whether a significant kinship exists between each pair of individuals. This invention is applicable to kinship inference in large-scale stratified populations and effectively controls the calculation bias caused by population stratification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computational biology, specifically to a method and system for inferring genome-wide kinship. Background Technology

[0002] Kinship inference based on genotype data has wide applications in population genetics, genetic association and linkage studies, pedigree studies, and forensic medicine. With the increasing prevalence of large-scale population cohorts, the calculation and quality control of kinship within and between cohorts have become increasingly important. Natural population samples are typically drawn from groups with different ancestry. However, distinguishing kinship and population structure using genotype data in heterogeneous samples or samples with population stratification is very challenging because both measure genetic similarity by measuring the degree of shared alleles.

[0003] Most existing methods for estimating kinship assume strong homogeneity in population structure and cannot cope with genetic bias caused by population stratification. This may lead to incorrect inferences about the kinship of individuals with mixed ancestry. On the other hand, estimation methods that take into account population structure often require prior ancestral information or cluster analysis to recalculate allele frequencies, which makes them poorly compatible with kinship algorithms. Summary of the Invention

[0004] To address the above problems, the purpose of this invention is to provide a method for inferring genome-wide kinship in stratified populations; Another objective of this invention is to provide a stratified population whole-genome kinship inference system.

[0005] A method for inferring genome-wide kinship in stratified populations includes the following steps: Step 1: Perform principal component analysis on the genotype data of the whole genome. Based on the results of the principal component analysis, determine c principal components to be analyzed. Perform regression analysis on c principal components and m molecular markers respectively to obtain the significance p-value of each molecular marker on each principal component. Molecular markers with significance p-values ​​exceeding a given threshold are identified as ancestry information markers. Remove ancestry information markers to obtain non-ancestry information markers. Step 2: Based on the non-ancestor information markers and the standardized genotype codes of all individuals on the non-ancestor information markers, calculate the estimated kinship coefficients between all individuals; Step 3: Perform a statistical significance test on the estimated kinship coefficients based on theoretical expectation and asymptotic variance to determine whether there is a significant kinship relationship between each pair of individuals.

[0006] The method for inferring kinship of whole genome in stratified populations described in this invention includes the principal component analysis results in step 1, which includes each principal component and the proportion of genetic variation explained by each principal component. The number of principal components to be analyzed, c, is determined based on the inflection point of the proportion of genetic variation explained by each principal component.

[0007] The stratified population genome-wide kinship inference method of this invention, in step 1, performs regression analysis with c principal components as dependent variables and m molecular markers as independent variables. The regression analysis formula is as follows:

[0008] in, Let be the projection vector of the k-th principal component. Represents the mean vector. Represents the regression coefficient of the j-th molecular marker. Indicates the first j Genotype coding vectors obtained by standardizing molecular markers j The value range is 1, 2, ..., m , Represents the residual vector; Through the estimated value of regression coefficients and variance Construct a chi-square distribution with 1 degree of freedom. The significance p-value of each molecular marker on each principal component was obtained. The given threshold is equal to 0.05.

[0009] The stratified population genome-wide kinship inference method of this invention, in step 2, calculates the estimated kinship coefficients between all individuals based on the following formula:

[0010] in, This indicates the number of non-ancestry information tags after quality control. and They represent the first i The individual and the first j The individual l Standardized genotype encoding of molecular markers.

[0011] The theoretical expectation in step 3 of the stratified population whole-genome kinship inference method described in this invention is: ; The asymptotic variance is:

[0012] in, Indicates the number of valid molecular markers. ,in Molecular markers and molecular markers The square of the Pearson correlation coefficient between them.

[0013] The stratified population genome-wide kinship inference method of this invention, in step 3, calculates a statistical index based on the cumulative distribution function of the standard normal distribution. Value, of which , The cumulative distribution function represents the standard normal distribution. ; when p The value is greater than a given significance threshold. When it indicates that the individuals have no significant kinship, p The value is less than a given significance threshold. This indicates a significant kinship relationship between individuals. A stratified population genome-wide kinship inference system includes, The ancestry information marker screening module performs principal component analysis on the genotype data of the whole genome. Based on the results of the principal component analysis, it determines c principal components to be analyzed. Regression analysis is performed on c principal components and m molecular markers respectively to obtain the significance p-value of each molecular marker on each principal component. Molecular markers with significance p-values ​​exceeding a given threshold are identified as ancestry information markers. Ancestry information markers are removed to obtain non-ancestry information markers. The kinship calculation module calculates the estimated kinship coefficients between all individuals based on the non-ancestor information markers and the standardized genotype codes of all individuals on the non-ancestor information markers; The statistical test and result output module performs a statistical significance test on the estimated kinship coefficient based on theoretical expectation and asymptotic variance to determine whether there is a significant kinship relationship between each pair of individuals.

[0014] The method for inferring kinship of whole genome in stratified populations described in this invention includes the results of principal component analysis, which includes each principal component and the proportion of genetic variation explained by each principal component. The number of principal components to be analyzed, c, is determined based on the inflection point of the proportion of genetic variation explained by each principal component. The formula for the regression analysis is:

[0015] in, Let be the projection vector of the k-th principal component. Represents the mean vector. Represents the regression coefficient of the j-th molecular marker. Indicates the first jGenotype coding vectors obtained by standardizing molecular markers j The value range is 1, 2, ..., m , Represents the residual vector; Through the estimated value of regression coefficients and variance Construct a chi-square distribution with 1 degree of freedom. The significance p-value of each molecular marker on each principal component was obtained. The given threshold is equal to 0.05.

[0016] The stratified population whole-genome kinship inference method of the present invention includes a kinship calculation module that calculates the estimated kinship coefficients between all pairs of individuals based on the following formula:

[0017] in, This indicates the number of non-ancestry information tags after quality control. and They represent the first i The individual and the first j The individual l Standardized genotype encoding of molecular markers.

[0018] The theoretical expectation of the stratified population whole-genome kinship inference method described in this invention is: ; The asymptotic variance is:

[0019] in, Indicates the number of valid molecular markers. ,in Molecular markers and molecular markers The square of the Pearson correlation coefficient between them; Calculate a statistical index based on the cumulative distribution function of the standard normal distribution. Value, of which , The cumulative distribution function represents the standard normal distribution. ; when p The value is greater than a given significance threshold. When it indicates that the individuals have no significant kinship, p The value is less than a given significance threshold. This indicates that individuals have a significant kinship relationship.

[0020] Beneficial effects: This invention utilizes whole-genome information for principal component analysis to determine the number of principal components to be selected. c The method assesses the ancestry information level of markers, removes ancestry information markers, and achieves ancestry information marker (AIM) quality control. This significantly improves the accuracy of existing tools in estimating kinship when analyzing large-scale stratified populations. It is applicable to kinship inference in large-scale stratified populations and effectively controls the kinship calculation bias caused by group stratification. Attached Figure Description

[0021] Figure 1 This is a flowchart of the stratified population whole-genome kinship inference method of the present invention; Figure 2 This is a schematic diagram of the principle of the stratified population whole genome kinship inference system of the present invention; Figure 3 The results of kinship calculations using KING and deepKin are obtained before and after quality control of ancestry information marking in the simulation data of this invention. Detailed Implementation

[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0023] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other.

[0024] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, but this is not intended to limit the scope of the invention.

[0025] Reference Figures 1 to 3 A method for inferring genome-wide kinship in stratified populations, comprising the following steps: Step 1: Perform principal component analysis on the genotype data of the whole genome. Based on the results of the principal component analysis, determine the c principal components to be analyzed. Perform regression analysis on the c principal components and m molecular markers respectively to obtain the significance p-value of each molecular marker on each principal component. Molecular markers with significance p-values ​​exceeding a given threshold are identified as ancestry information markers. Remove ancestry information markers to obtain non-ancestry information markers. Step 2: Based on non-ancestor information markers and the standardized genotype codes of all individuals on non-ancestor information markers, calculate the estimated kinship coefficients between all individuals; Step 3: Based on the theoretical expectation and asymptotic variance, perform a statistical significance test on the estimated kinship coefficient to determine whether there is a significant kinship relationship between each pair of individuals.

[0026] This invention first utilizes whole-genome information to perform principal component analysis and obtains the proportion of variation explained by each principal component. Based on the proportion of variation explained by each principal component, the number of principal components to be selected is determined according to the number of principal components at the time of the inflection point. c This invention assesses the ancestral information level of markers, removes ancestral information markers, and implements ancestral information marker (AIM) quality control, significantly improving the accuracy of existing tools in estimating kinship when analyzing large-scale stratified populations. This invention is applicable to kinship inference in large-scale stratified populations, effectively controlling the bias in kinship calculation caused by group stratification. The proposed ancestral information marker quality control method is applicable to multiple kinship calculation software programs that require assumptions of group homogeneity, demonstrating strong generalizability.

[0027] The stratified population whole-genome kinship inference method of the present invention includes, in step 1, the principal component analysis results of each principal component and the proportion of genetic variation explained by each principal component. The number of principal components to be analyzed, *c*, is determined based on the inflection point of the proportion of genetic variation explained by each principal component. Step 1 utilizes whole-genome information to perform principal component analysis and obtains the proportion of variation explained by each principal component. Based on the magnitude of the proportion of variation explained by each principal component, the number of selected principal components, *c*, is determined according to the number of principal components at the time of the inflection point (e.g., *c*). ).

[0028] The method for inferring genome-wide kinship in stratified populations of this invention involves performing EigenGWAS regression analysis in step 1, with c principal components as dependent variables and m molecular markers as independent variables (see Chen, GB et al. (2016) EigenGWAS: Finding loci under selection through genome-wide associations studies of eigenvectors in structured populations. Heredity, 117, 51–61). The formula for EigenGWAS regression analysis is:

[0029] in, Let be the projection vector of the k-th principal component. Represents the mean vector. Represents the regression coefficient of the j-th molecular marker. Indicates the first j Genotype coding vectors obtained by standardizing molecular markersj The value range is 1, 2, ..., m , Represents the residual vector; Through the estimated value of regression coefficients and variance Construct a chi-square distribution with 1 degree of freedom. The significance p-value of each molecular marker on each principal component was obtained. The given threshold for the significance p-value is 0.05. Molecular markers with a significance p-value exceeding the given threshold are identified as ancestry information markers and are not used for subsequent kinship estimation; they should be removed.

[0030] EigenGWAS was used to assess the ancestral information level of the markers. Regression analysis was performed with each principal component as the dependent variable and each molecular marker as the independent variable to obtain the significance of the regression. p Value, significance p The value represents the strength of the association between the molecular marker and the population structure represented by the principal component. By using a statistical model to identify and remove noisy data that mainly reflects the differences in ancestry of the population rather than the kinship between individuals, a "corrected" genotype dataset that more accurately reflects kinship can be provided for subsequent calculations.

[0031] Kinship is generally represented by Greek letters. This indicates that the theoretical expected value is 1, 0.5, and 0.25, which are integer powers of 0.5, corresponding to kinship relationships such as identical twins, father-son relationships, and grandfather-grandson relationships, respectively. When population stratification exists, kinship within subgroups is often overestimated, while kinship between subgroups is often underestimated. Therefore, it is necessary to screen ancestry information markers before performing genome-wide kinship estimation, retaining only non-ancestry information markers for kinship calculation. This invention uses ancestry-quality-controlled markers for genome-wide kinship calculation, and is applicable to existing kinship calculation tools KING and deepKin.

[0032] Taking deepKin as an example, step 2 calculates the estimated kinship coefficients between all pairs of individuals based on the following formula:

[0033] in, This indicates the number of non-ancestry information tags after quality control. and They represent the first i The individual and the first j The individual l Standardized genotype encoding of molecular markers.

[0034] Using the screening results from step 1 as input ensures that the labels used to calculate kinship are free from the influence of population structure, which directly solves the problem of estimation bias caused by population stratification.

[0035] The theoretical expectation in step 3 of the stratified population whole-genome kinship inference method of this invention is: ; The asymptotic variance is:

[0036] in, Indicates the number of valid molecular markers. ,in Molecular markers and molecular markers The square of the Pearson correlation coefficient between them.

[0037] In the stratified population genome-wide kinship inference method of the present invention, step 3 calculates a statistical index p-value based on the cumulative distribution function of the standard normal distribution, wherein... , The cumulative distribution function represents the standard normal distribution. ; when p The value is greater than a given significance threshold. When it indicates that the individuals have no significant kinship, p The value is less than a given significance threshold. This indicates that individuals have a significant kinship relationship.

[0038] Based on the theoretical distribution of kinship estimates, a statistical testing model is constructed to make a significant inference on the kinship of each pair of individuals, thus distinguishing true kinship from random noise. By introducing an effective number of molecular markers, the influence of linkage disequilibrium between genomic markers on variance estimation is corrected, thereby making the statistical test based on the normal approximation more accurate. The final output is no longer a simple estimate, but a reliable conclusion with statistical significance.

[0039] The null hypothesis is that the pair of individuals are unrelated, and the estimated kinship of this pair of individuals follows a normal distribution. After adjusting to a standard normal distribution, the Z-statistic is constructed as follows: and obtain the corresponding p Value , This represents the cumulative distribution function of the standard normal distribution. When p The value is greater than a given significance threshold. If the null hypothesis is accepted, that is, there is no significant kinship between the individuals; otherwise, if the null hypothesis is accepted, the null hypothesis is accepted. p The value is less than a given significance threshold. The null hypothesis was rejected, which stated that the individuals were significantly related.

[0040] In step 3, the asymptotic variance depends on the normal distribution and the normality is easily affected by low-frequency variation. The frequency of the minor allele of the molecular marker is taken to be above 0.05.

[0041] Reference Figure 2 A stratified population whole-genome kinship inference system, comprising, Ancestry information marker screening module 1 performs principal component analysis on the genotype data of the whole genome. Based on the results of the principal component analysis, it determines the c principal components to be analyzed. Regression analysis is performed on the c principal components and m molecular markers respectively to obtain the significance p-value of each molecular marker on each principal component. Molecular markers with significance p-values ​​exceeding a given threshold are identified as ancestry information markers. Ancestry information markers are removed to obtain non-ancestry information markers. Module 2 for calculating kinship relationships calculates the estimated kinship coefficients between all individuals based on non-ancestral information markers and the standardized genotype codes of all individuals on non-ancestral information markers. Module 3, Statistical Testing and Result Output, performs statistical significance testing on the estimated kinship coefficients based on theoretical expectation and asymptotic variance to determine whether there is a significant kinship relationship between each pair of individuals.

[0042] A specific embodiment: This invention simulates two subpopulations, each with 500 individuals, with F between the subpopulations. ST Individuals with a threshold of 0.1 and no phylogenetic relationship within or between subpopulations were included. A total of 10,000 molecular markers were simulated, with minor allele frequencies randomly selected from a uniform distribution U(0.05, 0.5) and measured using Lewontin LD (…). This describes the degree of linkage disequilibrium between molecular markers. Randomly selected from a uniform distribution U(0.5, 0.8). AIM quality control uses the first PC (processed sample), with a screening threshold of 0.05, corrected to 5 × 10⁻⁵ by Bonferroni. 6 The kinship calculations were performed before and after AIM quality control, using both KING and deepKin kinship methods, for both within and between subgroups. The simulation was repeated 10 times, and the mean and standard deviation of the kinship estimates were calculated. A Student's t-test was used to assess the significance of the means before and after AIM quality control.

[0043] The results show that, Figure 3As shown in Table 1, before AIM quality control, both the KING and deepKin methods showed significant biases in estimating inter-group kinship, and the deepKin method also showed a significant bias in estimating intra-group kinship. However, after AIM quality control, the biases in estimating inter-group kinship using both the KING and deepKin methods were significantly reduced, with the mean values ​​increasing from -0.1871 to -0.0120, respectively. p <0.0001) and increased from –0.0705 to –0.0053 ( p <0.0001), the bias in estimating kinship within a population using deepKin was also significantly reduced, with the mean decreasing from 0.0707 to 0.0053. p <0.0001). It is evident that this method, by implementing quality control on AIM markers, can significantly improve the accuracy of estimating individual kinship among subgroups, especially when the target population exhibits a group structure.

[0044] Table 1. Mean and significance test results of ancestry information markers before and after quality control using KING and deepKin kinship (SD represents standard deviation) in simulated data.

[0045] The description and accompanying drawings provide typical embodiments of specific structures for specific implementations. Other modifications are possible based on the spirit of the invention. While the above-described invention presents preferred embodiments, these are not intended to be limiting.

[0046] For those skilled in the art, various changes and modifications will undoubtedly be apparent after reading the above description. Therefore, the appended claims should be construed as covering all changes and modifications that encompass the true intent and scope of the invention. Any and all equivalent scope and content within the scope of the claims should be considered to remain within the intent and scope of the invention.

Claims

1. A hierarchical population-wide genome-wide kinship inference method, characterized in that, Includes the following steps: Step 1: Perform principal component analysis on the genotype data of the whole genome. Based on the results of the principal component analysis, determine c principal components to be analyzed. Perform regression analysis on c principal components and m molecular markers respectively to obtain the significance p-value of each molecular marker on each principal component. Molecular markers with significance p-values ​​exceeding a given threshold are identified as ancestry information markers. Remove ancestry information markers to obtain non-ancestry information markers. Step 2: Based on the non-ancestor information markers and the standardized genotype codes of all individuals on the non-ancestor information markers, calculate the estimated kinship coefficients between all individuals; Step 3: Perform a statistical significance test on the estimated kinship coefficients based on theoretical expectation and asymptotic variance to determine whether there is a significant kinship relationship between each pair of individuals.

2. The method for inferring genome-wide kinship in stratified populations according to claim 1, characterized in that, The results of the principal component analysis described in step 1 include each principal component and the proportion of genetic variation explained by each principal component. The number of principal components to be analyzed, c, is determined based on the inflection point of the proportion of genetic variation explained by each principal component.

3. The method for inferring genome-wide kinship in stratified populations according to claim 1, characterized in that, In step 1, regression analysis is performed with c principal components as dependent variables and m molecular markers as independent variables. The regression analysis formula is as follows: ; in, Let be the projection vector of the k-th principal component. Represents the mean vector. Represents the regression coefficient of the j-th molecular marker. Indicates the first j Genotype coding vectors obtained by standardizing molecular markers j The value range is 1, 2, ..., m , Represents the residual vector; Through the estimated value of regression coefficients and variance Construct a chi-square distribution with 1 degree of freedom. The significance p-value of each molecular marker on each principal component was obtained. The given threshold is equal to 0.

05.

4. The method for inferring genome-wide kinship in stratified populations according to claim 1, characterized in that, In step 2, the estimated kinship coefficients between all individuals are calculated based on the following formula: ; in, This indicates the number of non-ancestry information tags after quality control. and They represent the first i The individual and the first j The individual l Standardized genotype encoding of molecular markers.

5. The method for inferring genome-wide kinship in stratified populations according to claim 1, characterized in that, The theoretical expectation mentioned in step 3 is: ; The asymptotic variance is: ; in, Indicates the number of valid molecular markers. ,in Molecular markers and molecular markers The square of the Pearson correlation coefficient between them.

6. The method for inferring genome-wide kinship in stratified populations according to claim 1, characterized in that, In step 3, a statistical index is calculated based on the cumulative distribution function of the standard normal distribution. Value, of which , The cumulative distribution function represents the standard normal distribution. ; when p The value is greater than a given significance threshold. When it indicates that the individuals have no significant kinship, p The value is less than a given significance threshold. This indicates that the individuals have a significant kinship relationship.

7. A stratified population whole-genome kinship inference system, characterized in that, include, The ancestry information marker screening module performs principal component analysis on the genotype data of the whole genome. Based on the results of the principal component analysis, it determines c principal components to be analyzed. Regression analysis is performed on c principal components and m molecular markers respectively to obtain the significance p-value of each molecular marker on each principal component. Molecular markers with significance p-values ​​exceeding a given threshold are identified as ancestry information markers. Ancestry information markers are removed to obtain non-ancestry information markers. The kinship calculation module calculates the estimated kinship coefficients between all individuals based on the non-ancestor information markers and the standardized genotype codes of all individuals on the non-ancestor information markers; The statistical test and result output module performs a statistical significance test on the estimated kinship coefficient based on theoretical expectation and asymptotic variance to determine whether there is a significant kinship relationship between each pair of individuals.

8. The stratified population whole-genome kinship inference system according to claim 7, characterized in that, The results of the principal component analysis include each principal component, the proportion of genetic variation explained by each principal component, and the number of principal components c to be analyzed based on the inflection point of the proportion of genetic variation explained by each principal component. The formula for the regression analysis is: ; in, Let be the projection vector of the k-th principal component. Represents the mean vector. Represents the regression coefficient of the j-th molecular marker. Indicates the first j Genotype coding vectors obtained by standardizing molecular markers j The value range is 1, 2, ..., m , Represents the residual vector; Through the estimated value of regression coefficients and variance Construct a chi-square distribution with 1 degree of freedom. The significance p-value of each molecular marker on each principal component was obtained. The given threshold is equal to 0.

05.

9. The stratified population whole-genome kinship inference system according to claim 7, characterized in that, The kinship calculation module calculates the estimated kinship coefficients between all pairs of individuals based on the following formula: ; in, This indicates the number of non-ancestry information tags after quality control. and They represent the first i The individual and the first j The individual l Standardized genotype encoding of molecular markers.

10. The stratified population whole-genome kinship inference system according to claim 7, characterized in that, The theoretical expectation is: ; The asymptotic variance is: ; in, Indicates the number of valid molecular markers. ,in Molecular markers and molecular markers The square of the Pearson correlation coefficient between them; Calculate a statistical index based on the cumulative distribution function of the standard normal distribution. Value, of which , The cumulative distribution function represents the standard normal distribution. ; when p The value is greater than a given significance threshold. When it indicates that the individuals have no significant kinship, p The value is less than a given significance threshold. This indicates that the individuals have a significant kinship relationship.