A method for constructing a genetic risk prediction model integrating functional annotation information and its application

By calculating the specific heritability of each tissue and integrating functional annotation information, a genetic risk prediction model was constructed, which solved the problem of inaccurate genetic risk prediction in existing technologies and achieved higher prediction accuracy and comprehensive evaluation.

CN119229963BActive Publication Date: 2025-09-23HUAZHONG UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411332006.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-24
Publication Date
2025-09-23
Estimated Expiration
2044-09-24

AI Technical Summary

Technical Problem

In the existing technology, the genetic risk prediction method based on the integration of genetic-related traits and functional annotation information is not very accurate, resulting in inaccurate genetic risk prediction.

Method used

By calculating the heritability of the specific annotations of each cell type corresponding to each tissue of the sample to be tested, a random effects model is used to estimate the joint effect size of SNPs in the tissue, and the optimal linear unbiased estimate is obtained based on the covariance matrix. A genetic risk prediction model is constructed, and the functional annotation information of different tissues is integrated to improve prediction accuracy.

Benefits of technology

It improves the accuracy of genetic risk prediction, quantifies the relative contribution of different tissues to the overall genetic risk of a disease or trait, provides a comprehensive measurement standard for individual diseases or traits, and constructs a robust genetic risk prediction model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119229963B_ABST
    Figure CN119229963B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of genetic risk prediction technology and discloses a method for constructing a genetic risk prediction model integrating functional annotation information and its application, comprising: calculating the heritability of the specific annotations of each cell type corresponding to each tissue of the sample to be tested based on the marginal chi-square statistic of each SNP, and then obtaining the heritability of the j-th SNP in the current tissue f and estimating the joint effect size b of M SNPs in tissue f. f , and based on b f The covariance between the phenotype vector y of the n samples to be tested is b f The optimal linear unbiased estimate of the genetic risk prediction model is then obtained. Furthermore, the genetic risk scores corresponding to each tissue are integrated, and the functional annotation information of each tissue is incorporated into the genetic risk prediction. A genetic risk prediction method integrating functional annotation information is also provided. The present invention can improve the accuracy of genetic risk prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of genetic risk prediction, and more specifically, relates to a method for constructing a genetic risk prediction model integrating functional annotation information and its application. Background Art

[0002] Genetic risk score (PRS), as an important method for disease risk prediction in systems epidemiology, can be used for population risk prediction, optimization of screening programs and precise prevention.

[0003] PRS involves screening for SNPs (single nucleotide polymorphisms) and assessing the true effects of the screened SNPs. Some existing PRS methods integrate genetically correlated traits and functional annotation information to further screen and estimate effects. For example, by directly incorporating SNP functional annotations into the prior distribution of effect size, SNPs can be screened and their true effects estimated based on the prior distribution. However, this approach directly calculates the heritability of all SNPs, resulting in low accuracy in the final genetic risk prediction. Summary of the Invention

[0004] In response to the above-mentioned defects or improvement needs of the prior art, the present invention provides a method and application for constructing a genetic risk prediction model that integrates functional annotation information, the purpose of which is to improve the accuracy of genetic risk prediction.

[0005] To achieve the above objectives, according to a first aspect of the present invention, a method for constructing a genetic risk prediction model is provided, comprising:

[0006] S1, marginal chi-square statistic based on the j-th SNP Calculate the heritability of the specific annotations of each cell type corresponding to each tissue of the sample to be tested, and then obtain the heritability of the j-th SNP in the current tissue f Where j is 1, 2, ... M, and M represents the total number of SNPs in each tissue of the sample to be tested; Indicates the cell type corresponding to tissue f The heritability of the specific annotation, c takes 1, 2, ... C, C represents the total number of cell types corresponding to tissue f;

[0007] S2, using the random effects model y = Xb f +ε f Estimate the joint effect size b of M SNPs in tissue f f ; Where y is the n-dimensional phenotype vector of the sample to be tested, X is the M*n-dimensional genotype matrix, ε f is the n-dimensional residual vector of tissue f, and n is the total number of samples to be tested;

[0008] S3, based on bf The covariance between y and b f The optimal linear unbiased estimate of Among them, D f and Respectively represent b f and ε f The covariance matrix of σ 2 Pick u is a constant;

[0009] S4. Constructing a genetic risk prediction model PRS f It represents the genetic risk score corresponding to tissue f, and the high or low genetic risk score is used to reflect the high or low genetic risk.

[0010] Furthermore, after obtaining the genetic risk scores corresponding to each tissue, it also includes:

[0011] The genetic risk scores corresponding to each tissue are integrated, and the integrated genetic risk score is used as the overall genetic risk score.

[0012] Furthermore, the genetic risk scores corresponding to the tissues are integrated using an equal weight method, a Bayesian model averaging method, or a least absolute shrinkage and selection operator to obtain the integrated genetic risk score.

[0013] Furthermore, in S1, according to the marginal effect value β of the j-th SNP j Get the marginal chi-square statistic of the j-th SNP Among them, the marginal effect value β of the j-th SNP j Calculated by the following formula:

[0014] y=X j β j +ε

[0015] Where, X j is the n-dimensional genotype vector of the j-th SNP; ε is the n-dimensional residual vector, independent and identically distributed with variance The normal distribution of ε f The covariance matrix of

[0016] Furthermore, in S1, according to the marginal effect value β of the j-th SNP j Get the marginal chi-square statistic of the j-th SNP Among them, the marginal effect value β of the j-th SNP j It is estimated by the following formula:

[0017]

[0018] in, Represents β j estimated value of; Represents X j The transpose of X j is the n-dimensional genotype vector of the j-th SNP.

[0019] Further, take Then the b f The optimal linear unbiased estimate of can be simplified to:

[0020]

[0021] Among them, I M represents the M-order unit matrix; R represents the correlation matrix between SNPs obtained from the external data set; It is composed of M The matrix formed by , j takes 1, 2, …M.

[0022] Furthermore, in S1, the marginal chi-square statistic based on the j-th SNP is The heritability of the specific annotations of each cell type corresponding to each tissue in the sample to be tested is calculated by the following formula:

[0023]

[0024] in, represents the marginal chi-square statistic of the j-th SNP The expected value of l(j,c,f) represents the cth cell type in tissue f LD score; Indicates the cell type corresponding to tissue f The heritability of the specific annotation.

[0025] According to a second aspect of the present invention, a method for predicting genetic risk by integrating functional annotation information is provided, comprising:

[0026] Genetic risk prediction is performed using a genetic risk prediction model constructed using the genetic risk prediction model construction method integrating functional annotation information described in any one of the first aspects.

[0027] According to a third aspect of the present invention, there is provided an electronic device comprising a computer-readable storage medium and a processor;

[0028] The computer-readable storage medium is used to store executable instructions;

[0029] The processor is used to read the executable instructions stored in the computer-readable storage medium to execute the method for constructing a genetic risk prediction model integrating functional annotation information as described in any one of the first aspects, and / or to execute the method for predicting genetic risk integrating functional annotation information as described in the second aspect.

[0030] According to the fourth aspect of the present invention, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, it implements the method for constructing a genetic risk prediction model integrating functional annotation information as described in any one of the first aspects, and / or implements the genetic risk prediction method integrating functional annotation information as described in the second aspect.

[0031] In general, the above technical solutions conceived by the present invention can achieve the following beneficial effects:

[0032] (1) The method of the present invention is based on the marginal chi-square statistic of each SNP Calculate the heritability of the specific annotations of each cell type corresponding to each tissue of the sample to be tested, and then obtain the heritability of the j-th SNP in the current tissue f The random effects model is used to estimate the joint effect size b of M SNPs in tissue f f , and based on b f The covariance between the phenotype vector y of the n samples to be tested is b f The optimal linear unbiased estimate of Then we get the genetic risk prediction model In the process of constructing the genetic risk prediction model for each tissue, the heritability of each SNP is used. Rather than the heritability of the entire SNP, the accuracy of genetic risk prediction can be improved when genetic risk prediction is performed using a genetic risk prediction model constructed based on the heritability of each SNP.

[0033] In addition, the present invention calculates the heritability of the specific annotations of each cell type corresponding to each tissue of the sample to be tested, incorporates the SNP functional annotation information of different tissues into the genetic risk prediction, and constructs a genetic risk prediction model corresponding to each tissue. Compared with the existing single functional annotation information model, the method of the present invention can integrate the functional annotation information of different tissues for prediction, thereby improving the accuracy of genetic risk prediction.

[0034] (2) Furthermore, the method of the present invention integrates the genetic risk scores corresponding to each tissue and uses the genetic risk score obtained after integrating the functional annotation information as the genetic risk score of the overall sample to be tested. This overall genetic risk score takes into account the functional annotation information of each tissue, further improving the accuracy of genetic risk prediction. Moreover, this overall genetic risk score integrates the genetic risk scores of each tissue and can further quantify the relative contribution of different tissues to the overall genetic risk of a disease or trait.

[0035] (3) As a preference, the genetic variance (heritability) of a single SNP ) is small, and the phenotypic vector is usually normalized before analysis. Use the M-order identity matrix I M Approximately, that is You can b f The optimal linear unbiased estimate of Further simplification can improve computational efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 Schematic diagram of the method for constructing a genetic risk prediction model in an embodiment of the present invention. DETAILED DESCRIPTION

[0037] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only intended to illustrate the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below may be combined with each other as long as they do not conflict with each other.

[0038] Example 1

[0039] like Figure 1 As shown, an embodiment of the present invention provides a method for constructing a genetic risk prediction model integrating functional annotation information, which mainly includes:

[0040] S1, marginal chi-square statistic based on the j-th SNP Calculate the heritability of the specific annotations of each cell type corresponding to each tissue of the sample to be tested, and then obtain the heritability of the j-th SNP in the current tissue f Wherein, j is 1, 2, ..., M, and M represents the total number of SNPs in each tissue of the sample to be tested. In the embodiment of the present invention, the sample to be tested is a human body; Indicates the cell type corresponding to tissue f The heritability of the specific annotation, c takes 1, 2, ... C, and C represents the total number of cell types corresponding to tissue f.

[0041] S2, using the random effects model y = Xb f +ε f Estimate the joint effect size b of M SNPs in tissue f f ; Where y represents the n-dimensional phenotype vector of the sample to be tested, X is the M*n-dimensional genotype matrix, which is composed of the n-dimensional genotype vectors of M SNPs, ε f Represents the n-dimensional residual vector of tissue f, representing the difference between y and phenotype prediction value Xb f The difference between them, n is the total number of samples to be tested.

[0042] S3, based on b f The covariance between y and b f The optimal linear unbiased estimate of Among them, D f Indicates b f The covariance matrix of σ 2 Pick and X T represents the transpose of X, Represents ε f The covariance matrix of Determine that I is the identity matrix, and its order is the same as The number of h f It represents the heritability corresponding to tissue f, which is calculated as follows: l j is the linkage disequilibrium score of the j-th SNP, and n represents the number of samples to be tested.

[0043] S4. Constructing a genetic risk prediction model PRS f It represents the genetic risk score corresponding to tissue f, and the high or low score is used to reflect the high or low genetic risk.

[0044] Specifically, in S1, according to the marginal effect value β of the j-th SNP j Get the marginal chi-square statistic for the j-th SNP Among them, the marginal effect value β of the j-th SNP j By fitting the n-dimensional phenotype vector y of the sample to be tested and the n-dimensional genotype vector X of the j-th SNP j The relationship between them is obtained.

[0045] As a preferred embodiment, the standard linear model in GWAS is used to fit the n-dimensional phenotype vector y of the sample to be tested and the n-dimensional genotype vector X of the j-th SNP. j The relationship between:

[0046] y=X j β j +ε

[0047] Among them, ε is the n-dimensional residual vector, which represents the difference between the n-dimensional phenotype vector y and the phenotype prediction value X. j β j The residuals between are independent and identically distributed with variance Normal distribution; where ε f The covariance matrix of for:

[0048] Typically, both the phenotype vector and the genotype vector are standardized to follow a standard normal distribution with a mean of 0 and a variance of 1. Therefore, the marginal effect value β of the j-th SNP is j It can also be estimated by the following formula:

[0049]

[0050] in, Represents β j The estimated value of ; generally the number of samples to be tested n is relatively large.

[0051] Preferably, the sample tissues to be tested in the embodiment of the present invention include adrenal gland and pancreas, central nervous system, cardiovascular, connective tissue and bone, gastrointestinal tract, hematopoietic, kidney, liver, skeletal muscle and other tissues, a total of 10 tissues, and the C cell types corresponding to tissue f are

[0052] As a preference, the present invention uses a stratified LD score regression (S-LDSC) method based on the marginal chi-square statistic of the j-th SNP. Calculate the heritability of the specific annotations for each cell type corresponding to each tissue in the sample to be tested:

[0053]

[0054] in, represents the marginal chi-square statistic for the j-th SNP The expected value of l(j,c,f) represents the cth cell type in tissue f LD score, that is, the linkage disequilibrium score between M SNPs; Indicates the cell type corresponding to tissue f The heritability of the specific annotation, that is, the cell type corresponding to tissue f In other embodiments, the restricted maximum likelihood estimation (REML) and other algorithms can also be used to calculate

[0055] As a preferred implementation, in S3, b f The covariance between y and y is:

[0056]

[0057] Among them, Cov represents the covariance operator, D f for:

[0058]

[0059] Furthermore, in S3, due to the genetic variance (heritability) of a single SNP ) is small, and the phenotypic vector is usually normalized before analysis. Use the M-order identity matrix I M Approximately, that is According to the Woodbury matrix, b f The optimal linear unbiased estimate of Simplified to:

[0060]

[0061] Where R represents the correlation matrix between SNPs, which is obtained from external datasets, such as 1000 Genomes (1000G) or UK Biobank; Because of M The matrix formed by , j takes 1, 2, ... M. In this way, the computational efficiency can be improved.

[0062] Furthermore, the genetic risk prediction model construction method of the present invention also includes:

[0063] The genetic risk scores corresponding to each tissue are integrated, and the integrated genetic risk score is used as the genetic risk score of the overall sample to be tested.

[0064] Preferably, the ensemble includes Equal Weight (EW), Bayesian Model Averaging (BMA), or Least Absolute Shrinkage and Selection Operator (LASSO) models. Each strategy for adjusting PRS hyperparameters provides a method for assessing the relative contribution of each cell type to the overall genetic risk of a disease or trait.

[0065] The method of the present invention is based on the marginal chi-square statistic of each SNP Calculate the heritability of the specific annotations of each cell type corresponding to each tissue of the sample to be tested, and then obtain the heritability of the j-th SNP in the current tissue f The random effects model is used to estimate the joint effect size b of M SNPs in tissue f. f , and based on b f The covariance between the phenotype vector y of the n samples to be tested is b f The optimal linear unbiased estimate of Then we get the genetic risk prediction model In the process of constructing the genetic risk prediction model for each tissue, the heritability of each SNP is used. Rather than the heritability of the entire SNP, the accuracy of genetic risk prediction can be improved when genetic risk prediction is performed using a genetic risk prediction model constructed based on the heritability of each SNP.

[0066] As a further design of the present invention, the method of the present invention incorporates the SNP functional annotation information of different tissues into the genetic risk prediction by calculating the heritability of the specific annotations of each cell type corresponding to each tissue of the sample to be tested, so as to construct a genetic risk prediction model corresponding to each tissue. The genetic risk score corresponding to each tissue is integrated, and the genetic risk score obtained after the integration is used as the genetic risk score of the overall sample to be tested. The overall genetic risk score takes into account the functional annotation information of each tissue, further improving the accuracy of the genetic risk prediction. Moreover, the overall genetic risk score integrates the genetic risk scores of each tissue, which can further quantify the relative contribution of different tissues to the overall genetic risk of the disease or trait.

[0067] In summary, the present invention integrates PRSs from functional annotation information sets across multiple tissues to provide a comprehensive measure of the overall genetic risk of an individual disease or trait. The methods of the present invention facilitate the construction of a broadly applicable and robust PRS, enabling effective genetic prediction of phenotypes.

[0068] Example 2

[0069] An embodiment of the present invention provides a genetic risk prediction method integrating functional annotation information, comprising:

[0070] The genetic risk prediction model constructed by the genetic risk prediction model construction method integrating functional annotation information in Example 1 is used to obtain the genetic risk score corresponding to each tissue. The high or low score is used to reflect the high or low genetic risk, thereby completing the genetic risk prediction.

[0071] Furthermore, it also includes integrating the genetic risk scores corresponding to each tissue and using the integrated genetic risk score as the genetic risk score of the overall sample to be tested.

[0072] For the related technical solutions, please refer to the corresponding description in Example 1 and will not be repeated here.

[0073] It should be noted that the above method is completely implemented by computer.

[0074] Example 3

[0075] An embodiment of the present invention provides an electronic device, including a computer-readable storage medium and a processor;

[0076] Computer-readable storage media for storing executable instructions;

[0077] The processor is configured to read executable instructions stored in a computer-readable storage medium to execute the method for constructing a genetic risk prediction model integrating functional annotation information in Example 1, and / or to execute the method for predicting genetic risk integrating functional annotation information in Example 2. For related technical solutions, please refer to the description of the corresponding embodiments and will not be repeated here.

[0078] Example 4

[0079] Embodiments of the present invention provide a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the method for constructing a genetic risk prediction model integrating functional annotation information as described in Example 1 and / or the method for genetic risk prediction integrating functional annotation information as described in Example 2 are implemented. For related technical solutions, please refer to the description of the corresponding embodiments and will not be repeated here.

[0080] It will be easily understood by those skilled in the art that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for constructing a genetic risk prediction model integrating functional annotation information, characterized in that: include: S1, marginal chi-square statistic based on the j-th SNP Calculate the heritability of the specific annotations of each cell type corresponding to each tissue of the sample to be tested, and then obtain the heritability of the j-th SNP in the current tissue f Where j is 1, 2, ... M, and M represents the total number of SNPs in each tissue of the sample to be tested; Indicates the cell type corresponding to tissue f The heritability of the specific annotation, c takes 1, 2, ... C, C represents the total number of cell types corresponding to tissue f; S2, using the random effects model y = Xb f +ε f Estimate the joint effect size b of M SNPs in tissue f f ; Where y is the n-dimensional phenotype vector of the sample to be tested, X is the M*n-dimensional genotype matrix, ε f is the n-dimensional residual vector of tissue f, and n is the total number of samples to be tested; S3, based on b f The covariance between y and b f The optimal linear unbiased estimate of Among them, D f and Respectively represent b f and ε f The covariance matrix of σ 2 Pick u is a constant; S4. Constructing a genetic risk prediction model PRS f It represents the genetic risk score corresponding to tissue f, and the high or low genetic risk score is used to reflect the high or low genetic risk.

2. The method for constructing a genetic risk prediction model according to claim 1, wherein After obtaining the genetic risk scores corresponding to each tissue, it also includes: The genetic risk scores corresponding to each tissue are integrated, and the integrated genetic risk score is used as the overall genetic risk score.

3. The method for constructing a genetic risk prediction model according to claim 2, wherein: The genetic risk scores corresponding to the tissues are integrated using an equal weight method, a Bayesian model averaging method, or a least absolute shrinkage and selection operator to obtain the integrated genetic risk score.

4. The method for constructing a genetic risk prediction model according to claim 1 or 2, wherein: In S1, according to the marginal effect value β of the j-th SNP j Get the marginal chi-square statistic of the j-th SNP Among them, the marginal effect value β of the j-th SNP j Calculated by the following formula: y=X j b j +e Where, X j is the n-dimensional genotype vector of the j-th SNP; ε is the n-dimensional residual vector, independent and identically distributed with variance The normal distribution of ε f The covariance matrix of 5. The method for constructing a genetic risk prediction model according to claim 1 or 2, wherein: In S1, according to the marginal effect value β of the j-th SNP j Get the marginal chi-square statistic of the j-th SNP Among them, the marginal effect value β of the j-th SNP j It is estimated by the following formula: in, Represents β j estimated value of; Represents X j The transpose of X j is the n-dimensional genotype vector of the j-th SNP.

6. The method for constructing a genetic risk prediction model according to claim 5, wherein: Pick Then the b f The optimal linear unbiased estimate of can be simplified to: Among them, I M represents the M-order unit matrix; R represents the correlation matrix between SNPs obtained from the external data set; It is composed of M The matrix formed by , j takes 1, 2, …M.

7. The method for constructing a genetic risk prediction model according to claim 1 or 2, wherein: In S1, the marginal chi-square statistic based on the j-th SNP The heritability of the specific annotations of each cell type corresponding to each tissue in the sample to be tested is calculated by the following formula: in, represents the marginal chi-square statistic of the j-th SNP expected value; represents the cth cell type in tissue f LD score; Indicates the cell type corresponding to tissue f The heritability of the specific annotation.

8. A genetic risk prediction method integrating functional annotation information, characterized in that: include: Genetic risk prediction is performed using a genetic risk prediction model constructed using the genetic risk prediction model construction method integrating functional annotation information described in any one of claims 1-7.

9. An electronic device, characterized in that: comprising a computer-readable storage medium and a processor; The computer-readable storage medium is used to store executable instructions; The processor is used to read the executable instructions stored in the computer-readable storage medium to execute the method for constructing a genetic risk prediction model integrating functional annotation information as described in any one of claims 1-7, and / or to execute the method for predicting genetic risk integrating functional annotation information as described in claim 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by the processor, it implements the method for constructing a genetic risk prediction model integrating functional annotation information as described in any one of claims 1 to 7, and / or implements the method for predicting genetic risk integrating functional annotation information as described in claim 8.

Citation Information

Patent Citations

  • Genomic genetic hierarchical conjoint analysis method and system

    CN114496076A

  • Multi-gene genetic risk score calculation method and system based on tissue specific regulatory network atlas

    CN117409860A