A multi-gene genetic risk score calculation method and system based on a tissue-specific regulatory network map

CN117409860BActive Publication Date: 2026-09-18ACAD OF MATHEMATICS & SYSTEMS SCIENCE - CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311080327.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-25
Publication Date
2026-09-18
Estimated Expiration
2043-08-25

AI Technical Summary

Technical Problem

[0007](1)目前的多基因风险评分没有考虑蕴含在多组学数据中的环境因素,复杂疾病是由多基因遗传、表观遗传和环境等多种因素引起的非线性和动态的多细胞类型功能紊乱

Benefits of technology

[0033] This invention discloses a method for predicting PRS risk scores based on the integration of multi-omics data using regulatory networks and utilizing specific regulatory networks of relevant tissue and cell types. Its beneficial effects include:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117409860B_ABST
    Figure CN117409860B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of multi-gene genetic risk score calculation method and system based on tissue-specific regulatory network atlas.The steps of the method include: for a given phenotype or disease, obtaining the summary statistics data of its GWAS and the phenotype and genotype data of population;Using the data obtained, a PRS score prediction model based on tissue-specific regulatory network atlas is constructed, and it is trained;The genotype of the individual to be tested is input into the PRS score prediction model after training, and its PRS score is predicted.The present application proposes to integrate the regulatory network specific to multiple tissues of human body for PRS risk prediction, by constructing human tissue-specific regulatory network atlas, the screening of risk loci is realized, the re-estimation of risk locus effect value and cell environment combination, and the effects of genetic variation in multiple tissues are integrated, which provides a new model for genetic risk prediction of complex diseases.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of medical technology and information technology, and specifically relates to a method and system for calculating multi-gene genetic risk scores based on tissue-specific regulatory network maps. Background Technology

[0002] Precision medicine is a novel strategy for disease prevention and treatment that takes into account individual differences in genes, environment, and lifestyle. It aims to accurately classify and diagnose diseases, providing patients with personalized and more targeted prevention and treatment measures. Large-scale population cohort studies are a crucial foundation for precision medicine research, providing the best evidence-based medicine for precision medicine practice. For example, the UK Biobank measured the genotypes of 500,000 people and matched thousands of phenotypes. Similar data resources include 23andMe, deCODE Genetics, the Estonian Biobank, Biobank Japan, China Kadoori Biobank, FinnGen in Finland, Lifelines in the Netherlands, the Million Veteran Program, and the All of US research program. It is expected that large-scale biobanks will continue to expand to more diverse populations and capture more allele spectra through whole-genome sequencing, further statistically identifying human genetic variations associated with complex traits and diseases. These genetic variations can be integrated with environmental risk factors for disease risk prediction, enabling personalized precision prevention—a significant research hotspot in precision medicine research and practice.

[0003] Statistical analysis of genome-wide association studies (GWAS) data for specific phenotypes or diseases can typically identify thousands of genetic variation sites in the genome associated with complex diseases. These sites mostly fall within non-coding regions of the genome, suggesting that these genetic variations may have different effects in different tissues, developmental stages, and even different individuals by regulating proximal or distal gene expression. Moreover, each site may only carry a small risk of disease. A feasible strategy is to mathematically weight and accumulate disease risks to quantify risk categories, which can then be used to provide early warning of complex diseases based on the polygenic risk score (PRS). This method has seen accelerated development in recent years, and in 2018 it was selected as one of MIT Technology Review's "Top 10 Breakthrough Technologies." PubMed publishes 200 scientific papers on polygenic risk scores annually. From a clinical application perspective, polygenic risk scores can achieve better stratification and intervention. For patients with traditional intermediate- or high clinical risk, polygenic risk scores can significantly improve the ability to re-stratify traditional clinical risks, providing solutions for clinical decision-making and intervention selection. For example, a team from Fuwai Hospital of the Chinese Academy of Medical Sciences published a paper in the top cardiology research journal, the European Heart Journal, which successfully established my country's first multi-gene risk scoring model for coronary heart disease. The study integrated genomic data from a total of 267,465 East Asian individuals from 12 cohorts, including those from China, Japan, South Korea, and Singapore. It identified 540 genetic variants that influence coronary heart disease and major risk factors in Chinese and East Asian populations, and constructed the first multi-gene risk scoring model for coronary heart disease suitable for the Chinese population. Based on this, cardiovascular health management pathways and programs can be proposed for different genetic and clinical risk groups (especially young people who have not yet developed clinical risk factors).

[0004] The concept of polygenic genetic scoring technology can be traced back to 2001, and its ideas were first used in plant breeding and animal breeding. From a methodological perspective, a research team at the University of Michigan systematically compared and summarized 46 current construction methods (Ma Y, Zhou X. Genetic prediction of complex traits with polygenic scores: a statistical review. Trends Genet, 2021.). According to the framework of statistical methods, they can be mainly divided into four categories: (1) based on the complete Bayesian method, using Markov chain Monte Carlo method for model fitting, such as BayesR, BSLMM, etc.; (2) based on the empirical Bayesian method, using site linkage and site function information to optimize the model, such as LDpred, AnnoPred, etc.; (3) based on the frequency theory method, such as MultiBLUP, DBSLMM, etc.; (4) based on penalized regression, using iterative algorithms for parameter estimation, such as lassosum, CTPR, etc. These methods have been applied in many studies. The basic idea behind polygenic genetic scoring is to sum the effects of risk loci as an indicator to predict phenotype and disease risk. For example, the PRS prediction in Plink2 software is based on this idea. Most existing improvements to this classic method are achieved by re-estimating the effect sizes of risk loci. The effect sizes of risk loci can be improved through prior distributions; for example, the SDPR method re-estimates the effect sizes of risk loci by assuming a Beta distribution of their effects. Improvements can also be achieved by introducing individual-level genotype and phenotype data; for example, SNPNet applies LASSO sparsity constraints to the effect sizes of risk loci when regressing phenotypes. Recently, some methods have considered specific functional regions composed of multi-omics data. For example, the LDpred-funct method considers multiple functional regions of the genome and linkage disequilibrium between effect loci, re-estimates the effect sizes of effect loci, groups effect loci, integrates risk loci within each group, and finally weights and sums the risk scores from multiple groups to obtain the final PRS score.

[0005] Existing polygenic risk scoring techniques have improved PRS risk score prediction to some extent, but many challenges remain. The first challenge is identifying risk loci with causal effects. Since most genetic variations occur in non-coding regions, their related distant regulation needs to be considered to identify causal risk loci. Secondly, the biological functions and effects of risk loci operate in specific times and spaces, and their effect sizes vary with different cellular environments; therefore, the cellular environments of multiple tissues in the human body need to be considered. Finally, the development of complex polygenic phenotypes and diseases is the result of the combined effects of multiple cell types, tissues, and organs; integrating the influence of genetic variations across multiple cell types, tissues, and organs is necessary for accurate prediction.

[0006] The main disadvantages of existing technologies are as follows:

[0007] (1) Current polygenic risk scores do not consider environmental factors contained in multi-omics data. Complex diseases are nonlinear and dynamic multicellular functional disorders caused by multiple factors such as polygenic inheritance, epigenetics, and environment. Most existing polygenic risk scores for disease occurrence and development are based on genetic information, lack environmental and interaction factors, and are limited to linear and static perspectives, thus failing to form an effective early warning.

[0008] (2) Current polygenic risk scores do not take into account the tissue specificity of the biological function of genetic variations, that is, they do not take into account that genetic variations only play a role in specific spatiotemporal conditions, affecting phenotypes and causing diseases.

[0009] (3) Current polygenic risk scores do not take into account the relationship between genetic variations, i.e., multiple genetic variations work together to produce utility.

[0010] (4) Current polygenic risk scores lack interpretability, that is, they cannot establish the regulatory network of genetic variations to specific tissues and the pathways that affect phenotypes or diseases. Summary of the Invention

[0011] To overcome the aforementioned difficulties and challenges, this invention proposes integrating multi-tissue-specific regulatory networks in the human body for PRS risk prediction. By constructing a human tissue-specific regulatory network map, it achieves the screening of risk sites, the re-estimation of the effect values ​​of risk sites in conjunction with the cellular environment, and integrates the effects of genetic variations in multiple tissues, providing a new model for predicting PRS risk scores.

[0012] The technical solution adopted in this invention is as follows:

[0013] A method for calculating polygenic genetic risk scores based on tissue-specific regulatory network maps includes the following steps:

[0014] For a given phenotype or disease, obtain the summary statistics of its genome-wide association analysis (GWAS) as well as the phenotypic and genotypic data of the population;

[0015] Using the acquired data, a PRS score prediction model based on the tissue-specific regulatory network map was constructed and trained.

[0016] The PRS score prediction model is trained by inputting the genotype of the individual to be tested, and its PRS score is predicted.

[0017] Furthermore, for a given phenotype or disease, obtaining its GWAS summary statistics and population phenotype and genotype data includes:

[0018] The acquired phenotypic and genotypic data are numericalized, noise-removed, outlier removed, and normalized.

[0019] For GWAS of phenotypes or diseases, aggregated statistical data are obtained from public databases, including information on chromosome name, SNP location, SNP(rs)-identifier, effect allele frequency (MAF), SNP effect size, standard deviation, and p-value.

[0020] Furthermore, the construction of the PRS score prediction model based on the tissue-specific regulatory network map includes:

[0021] For the disease or phenotype to be studied, the relevant tissue or cell types are identified through publicly available GWAS data;

[0022] In tissue-specific regulatory networks, genetic variations near each regulatory element and its own regulatory information are integrated, and the effects of genetic variations are summarized into the genetic effects of the regulatory elements.

[0023] Using regulatory elements as basic units, tissue-specific regulatory elements are integrated to achieve tissue-specific genetic effects;

[0024] Integrate the genetic effects of multiple related tissues on the integrated phenotype, calculate the PRS score, and assess the phenotype or disease.

[0025] Furthermore, a two-step strategy is employed to train the parameters of the PRS score prediction model. First, the weights ω of the regulatory elements for each relevant tissue or cell type are trained. ik Then train the weights γ for all M tissue or cell types. i .

[0026] Furthermore, the hyperparameter α of the PRS score prediction model i The following steps are used to determine:

[0027] Select undetermined values ​​for the hyperparameters;

[0028] For each undetermined value of a hyperparameter, predict the PRS in the calibrated population, then calculate the accuracy of the predicted PRS compared to the true PRS, and select the hyperparameter with the highest accuracy.

[0029] A multi-gene genetic risk scoring system based on a multi-tissue regulatory network map, comprising:

[0030] The data acquisition module is used to acquire the summary statistics of GWAS and the phenotypic and genotypic data of the population for a given phenotype or disease.

[0031] The mathematical modeling module is used to construct and train a PRS score prediction model based on the organization-specific regulatory network map using the acquired data.

[0032] The PRS score prediction module is used to predict the PRS score of an individual by inputting their genotype into a pre-trained PRS score prediction model.

[0033] This invention discloses a method for predicting PRS risk scores based on the integration of multi-omics data using regulatory networks and utilizing specific regulatory networks of relevant tissue and cell types. Its beneficial effects include:

[0034] 1. It can more effectively screen for genetic variation risk sites;

[0035] 2. It can more accurately predict disease status using PRS risk scores;

[0036] 3. It can output disease-related tissue or cell types, as well as regulatory elements and genes related to genetic variations, thereby improving the interpretability of risk warnings. Attached Figure Description

[0037] Figure 1 This is a schematic diagram of a multi-gene risk scoring system based on a tissue-specific regulatory network map, as presented in this invention.

[0038] Figure 2 This is a graph showing the performance of this invention in predicting PRS scores based on the height phenotype and its comparison with other methods.

[0039] Figure 3 This is a graph showing the performance of this invention in predicting PRS risk in diabetes and its comparison with other methods. Detailed Implementation

[0040] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to specific embodiments and accompanying drawings.

[0041] The technical problem this invention aims to solve is to provide a method for assessing the genetic risk of complex human diseases. This method utilizes individual genomic data obtained through high-throughput sequencing technology, identifies thousands of disease phenotype-related loci in the genome based on large-scale cohort genome-wide association analysis, integrates human tissue-specific regulatory network maps to design a polygenic genetic risk score, and mathematically weights and accumulates the disease risk of these loci, thereby quantifying the risk of an individual's associated disease phenotype. Challenges to overcome include: the small effect of a single locus is insufficient to predict disease phenotype; existing polygenic genetic risk scoring methods ignore genetically relevant information encoded in the epigenome; lack of cross-ethnic generality; and inability to detect causally related genetic variations. To overcome these technical bottlenecks, this invention proposes a novel method for calculating polygenic genetic risk scores based on human tissue-specific regulatory network maps. This method involves fully automated biochemical analysis of individual genomic data, designing models to identify tissues or cell types associated with disease phenotypes, integrating the effects of individual genetic variations and the regulatory strength of the regulatory network within the regulatory networks of disease-related tissues, and defining a polygenic genetic risk score. Based on this new polygenic genetic risk score, a complete workflow and computational system for predicting individual disease phenotype risk are established.

[0042] To overcome the shortcomings of the prior art, the present invention proposes the following... Figure 1 The system shown integrates multi-omics data to calculate PRS risk scores. This system comprises three modules:

[0043] (1) Data Acquisition Module

[0044] For a given phenotype or disease, the relevant data that needs to be collected consists of two parts:

[0045] (1.1) Summary statistics of GWAS for this phenotype in a large cohort population.

[0046] (1.2) Genotypic data of a population consisting of several individuals, along with the corresponding phenotypic measurements or disease status. The data of this population is divided into three groups: a training group (including phenotypic and genotypic data of 80% of the population, used for training parameters in the model), a calibration group (including phenotypic and genotypic data of 10% of the population, used for selecting hyperparameters of the model), and a test group (including phenotypic and genotypic data of 10% of the population, used for model validation and comparison with other methods).

[0047] (2) Mathematical Modeling Module

[0048] This module models the relationship between genotype and phenotype based on tissue-specific gene regulatory networks, selects the correct tissue or cell type, integrates SNPs (Single Nucleotide Polymorphisms) using regulatory elements, obtains multi-tissue effects from these integrated regulatory elements, and integrates multi-tissue effect scores to achieve more accurate PRS risk score prediction. This module consists of the following five steps:

[0049] (2.1) For the disease or phenotype to be studied, the relevant tissue or cell type is identified by using the SpecVar software (Feng, Z. et al. (2022) 'Heritability enrichment in context-specific regulatory networks improves phenotype-relevant tissue identification', eLife, 11. doi:10.7554 / elife.82535.) through publicly available GWAS data.

[0050] (2.2) In tissue-specific regulatory networks, genetic variations near each regulatory element and its own regulatory information are integrated, and the effects of genetic variations are summarized into the genetic effects of the regulatory elements.

[0051] (2.3) Using regulatory elements as the basic unit, tissue-specific regulatory elements are integrated to form tissue-specific genetic effects.

[0052] (2.4) Integrate the genetic effects of multiple related tissues of the phenotype, calculate the PRS score, and evaluate the phenotype or disease.

[0053] (2.5) Model Hyperparameter Determination. To determine several hyperparameters in the model, different hyperparameters were tested using a calibration group, and the optimal hyperparameters were selected to construct the model. This process involved the following three steps:

[0054] (2.5.1) Predict PRS scores for each person in the calibration group and calculate the accuracy within the calibration group population.

[0055] (2.5.2) For each hyperparameter, select a set of undetermined values ​​and perform step (2.5.1) to calculate the accuracy of each undetermined value in the calibration group.

[0056] (2.5.3) Select the optimal hyperparameters.

[0057] (3) Model Testing Module

[0058] The calibrated model is tested using samples from the test group to objectively determine the model's predictive accuracy for individual phenotypes or diseases, while also analyzing and diagnosing the distribution of health assessment scores at the population level.

[0059] The PRS risk score prediction model proposed in this invention is divided into three modules, which are described in detail below for data acquisition, model construction and solution, and model testing.

[0060] (1) Data Acquisition

[0061] For a selected phenotype or disease, a population group is first collected. The phenotype or disease status of interest is obtained through technical means, questionnaires, etc., and the phenotype is simultaneously obtained using SNP microarray technology or whole-genome sequencing technology. Low-quality samples are first removed, and the phenotypic and genotypic data are then subjected to numericalization, noise reduction, outlier removal, and normalization.

[0062] For GWAS of this phenotype or disease, summary statistics can be obtained from public databases (such as the GWAS Catalog). These statistics include the results obtained after GWAS, such as information on chromosome name, SNP location, SNP(rs)-identifier, effect genotype frequency (MAF), SNP effect size (odds ratio / beta), standard deviation, and p-value.

[0063] (2.1) Model Construction

[0064] The first step involves using the aggregated statistical data from GWAS as input. Existing SpecVar software is then used to identify M tissue or cell types associated with this phenotype or disease, denoted as C1, C2, ..., C6. M For each relevant tissue or cell type C i Using multi-omics data and PECA2 software (Duren, Z. et al. (2020) 'Time Course Regulatory Analysis based on paired expression and chromatin accessibility data', Genome Research, 30(4), pp.622-634. doi: 10.1101 / gr.257063.119.), active tissue-specific regulatory networks were obtained. These networks are directed graphs with transcription factors (TF), regulatory elements (RE), and target genes (TG) as nodes and their regulatory relationships as edges. TF binds to the regulatory element RE and exerts a regulatory effect on nearby TG. Therefore, each regulatory network contains N... iA collection of control elements And there is the regulatory intensity s of the k-th regulatory element RE in the i-th organization. ik (Output from PECA2). For the k-th RE in the i-th tissue, the set of related SNPs V that are adjacent to the k-th RE can be obtained by calculating the genomic distance. ik V ik The genotype of the j-th SNP is g ikj ∈{0, 1, 2}, the effect size β in the GWAS summary statistics of this SNP. ikj Chained imbalance fractions l ikj and the distance d between the SNP and the RE ikj Regulatory elements refer to specific, active 200-2000 bp segments on DNA that regulate the expression of downstream genes.

[0065] The second step involves taking each regulatory element of each relevant tissue or cell type as the core, integrating its own regulatory strength and the effect value of nearby SNPs to obtain the effect score of the k-th RE in the i-th tissue:

[0066]

[0067] The third step is to integrate the weighted sum of all regulatory elements in the i-th tissue or cell type to obtain the effect score of that tissue or cell type (where the weight ω). ik (Requires training to obtain):

[0068]

[0069] Where, N i This represents the total number of regulatory elements in the i-th organization.

[0070] The fourth step is to calculate the effect score R for this tissue or cell type. i Perform Z-score normalization, and denote μ as R i The average value, σ is R i Standard deviation, Z(R) i The effect score is the normalized value.

[0071]

[0072] The PRS risk score is obtained by integrating the weighted sum of the effect scores of M related tissue or cell types (where the weight γ). i (Requires training to obtain):

[0073]

[0074] In summary, this invention ultimately constructs a model for predicting PRS scores based on tissue-specific regulatory network maps, as follows:

[0075]

[0076] Among them, the organization-specific regulatory network refers to the regulatory network of a certain organization output by PECA2 software, and the regulatory network map is formed by summing up multiple organizations.

[0077] The following steps require using the collected genotype and phenotypic data of the population as input to train the weights ω. ik and γ i .

[0078] (2.2) Model Training

[0079] A two-step strategy is employed to train the model parameters: first, the weights ω of the regulatory elements are trained for each relevant tissue or cell type. ik Then train the weights γ for all M tissue or cell types. i .

[0080] Using the phenotype y and genotype g of each person in the training set as input, the first step is to base the C of genotype g on each tissue or cell type. i The regulating element RE ik Calculate its effect fraction R ik Then, the weights ω are estimated using the Least Absolute Value Selection and Shrinkage Operator (LASSO) model. ik :

[0081]

[0082] in, ω ik Let α represent the weight of the k-th RE in the i-th organization to be trained in step 3 of (2.1), where n represents the number of people in the training set, and α represents the weight of the RE. i This represents the hyperparameter corresponding to the i-th organization.

[0083] When the weight ω of each regulatory element in each tissue or cell type ik Once identified, the effect score for each tissue or cell type is calculated using a weighted sum of the regulatory elements. Then, the weight γ for each tissue or cell type is estimated using the following linear model. i :

[0084]

[0085] Among them, γ = (γ1, γ2,…, γ M ), γ iThis represents the weight of the i-th organization to be trained in step 4 of (2.1).

[0086] (2.3) Determination of hyperparameters

[0087] During the training process of the model, there is a set of sparse hyperparameters. First, the undetermined hyperparameter values ​​are selected as follows:

[0088] α i ∈{5e -7 6e -7 , ..., 9e -7 ,1e -6 ,2e -6 , ..., 9e -6 ,1e -5 ,2e -5 , ..., 9e -5 ,1e -4 ,2e -4 , ..., 9e -4}

[0089] Then, for each undetermined value of hyperparameter, the PRS is predicted in the calibration group, and the accuracy of the predicted PRS compared to the true PRS is calculated. Specifically, assuming the calibration group has M individuals, their phenotypic predictions are {y′1, y′2, ..., y′...} M}, its actual phenotype is {y1, y2, ..., y}. M For continuous phenotypes, the following PCC score is used to measure accuracy:

[0090]

[0091] in, This represents the average of the actual phenotypic values ​​in the calibration group. This represents the average predicted phenotypic value of the calibration group.

[0092] For the 0-1 phenotype, the accuracy ACC score is used as a measure of accuracy:

[0093]

[0094] Where # indicates the number of predicted phenotypic values ​​that equal the actual phenotypic values.

[0095] Finally, select the hyperparameter with the highest accuracy.

[0096] (3) Model Testing

[0097] After the hyperparameters are determined, the genotypes of the test group population are used as input to predict their PRS risk scores, which are compared with the true phenotypes. The Pearson correlation coefficient (PCC) score or ACC score in (2.3) is calculated to measure the accuracy, and the prediction accuracy is compared with Plink2, SDPR, and LDpref-funct.

[0098] This invention proposes a novel framework for integrating multi-omics data to calculate polygenic genetic scores, with key features including:

[0099] (1) Methods for identifying the tissue or cell type most relevant to the phenotype.

[0100] (2) A method of integrating the effects of multiple genetic variations into a regulatory element.

[0101] (3) A model for predicting polygenic genetic scores using regulatory networks of multiple related tissue or cell types.

[0102] (4) By scoring each tissue or cell type and each regulatory unit, a system can be established to identify the tissue or cell type in which genetic variations play a role, as well as the regulatory elements and genes therein, providing a reference for personalized medicine.

[0103] This invention has been experimentally verified. Taking height as an example, in terms of continuous phenotype, phenotypes and genotypes of 10,000 individuals were collected for training and testing, and PCC scores were calculated, such as... Figure 2 As shown, compared with Plink2, SDPR, and LDpref-funct, the model of this invention achieves higher accuracy than other methods. Regarding disease phenotypes, taking diabetes as an example, training and testing were performed using a dataset of 10,000 individuals, and ACC scores were calculated, such as... Figure 3 As shown, the accuracy of the model in this invention is also higher than that of other methods.

[0104] Another embodiment of the present invention provides a multi-gene genetic risk scoring system (or device) based on a multi-tissue regulatory network map, comprising:

[0105] The data acquisition module is used to acquire the summary statistics of GWAS and the phenotypic and genotypic data of the population for a given phenotype or disease.

[0106] The mathematical modeling module is used to construct and train a PRS score prediction model based on the organization-specific regulatory network map using the acquired data.

[0107] The PRS score prediction module is used to predict the PRS score of an individual by inputting their genotype into a pre-trained PRS score prediction model.

[0108] For the specific implementation process of each module, please refer to the description of the method of the present invention above.

[0109] Another embodiment of the present invention provides a computer device (computer, server, smartphone, etc.) including a memory and a processor, the memory storing a computer program configured to be executed by the processor, the computer program including instructions for performing the steps of the method of the present invention.

[0110] Another embodiment of the present invention provides a computer-readable storage medium (such as ROM / RAM, disk, optical disk) storing a computer program that, when executed by a computer, implements the various steps of the method of the present invention.

[0111] It should be understood that the application of this invention is not limited to the examples described above. Those skilled in the art can make improvements or modifications based on the above description, or adjust and select relevant parameters. All such improvements and modifications, as well as parameter adjustments and selections, should fall within the protection scope of the appended claims.

Claims

1. A method for calculating polygenic genetic risk scores based on tissue-specific regulatory network maps, characterized in that, Includes the following steps: For a given phenotype or disease, obtain the summary statistics of its genome-wide association study (GWAS) and the phenotypic and genotypic data of the population. Using the acquired data, a PRS score prediction model based on the tissue-specific regulatory network map was constructed and trained. The genotype of the individual to be tested is input into the trained PRS score prediction model to predict their PRS score. The PRS score prediction model is constructed using the following steps: The first step involves using the aggregated statistical data from GWAS as input and employing SpecVar software to identify M tissue or cell types associated with the phenotype or disease; for each tissue or cell type... Tissue-specific regulatory networks were obtained using PECA2 software with multi-omics data; each regulatory network contains... A collection of control elements And there is a first The first in the organization The control strength of each control element RE ; for the first The first in the organization The RE is obtained by calculating the distance on the genome to the first RE. The set of SNPs that are neighbors of RE ,and The Middle The genotype of each SNP is The effect size in the GWAS summary statistics of this SNP Chained imbalance fractions and the SNP and the first Distance between REs ; The second step involves taking each regulatory element of each relevant tissue or cell type as the core, integrating its own regulatory strength and the effect values ​​of nearby SNPs, to obtain the... The first in the organization Effect score of each RE : The third step is to integrate the first... The effect score for a given tissue or cell type is obtained by weighting all regulatory elements in that tissue or cell type. : Among them, weight Obtained through training; Fourth step, for the first Effect score for each tissue or cell type Z-score normalization was performed to obtain the normalized effect score. The PRS risk score is obtained by integrating the weighted sum of the effect scores of M related tissue or cell types. Among them, weight Obtained through training; Finally, the PRS score prediction model based on the tissue-specific regulatory network map is constructed as follows:

2. The method according to claim 1, characterized in that, For a given phenotype or disease, obtaining its GWAS summary statistics and population phenotype and genotype data includes: The acquired phenotypic and genotypic data are numericalized, noise-removed, outlier removed, and normalized. For GWAS of phenotypes or diseases, aggregated statistical data are obtained from public databases, including information on chromosome name, SNP location, SNP(rs)-identifier, effector allele frequency (MAF), SNP effect size, standard deviation, and p-value.

3. The method according to claim 1, characterized in that, The construction of the PRS score prediction model based on the tissue-specific regulatory network map includes: For the disease or phenotype to be studied, the relevant tissue or cell types are identified through publicly available GWAS data; In tissue-specific regulatory networks, genetic variations near each regulatory element and its own regulatory information are integrated, and the effects of genetic variations are summarized into the genetic effects of the regulatory elements. Using regulatory elements as basic units, tissue-specific regulatory elements are integrated to achieve tissue-specific genetic effects; Integrate the genetic effects of multiple related tissues on the integrated phenotype, calculate the PRS score, and assess the phenotype or disease.

4. The method according to claim 1, characterized in that, The parameters of the PRS score prediction model are trained using a two-step strategy. First, the weights of the regulatory elements are trained for each relevant tissue or cell type. Retrain all Weight of each tissue or cell type The training process of the PRS score prediction model includes: Using the phenotypes of each person in the training set and genotype As input, based on genotype For each tissue or cell type Control element Calculate its effect score The LASSO model, which uses the minimum absolute value selection and shrinkage operator, is used to estimate the weights. : in, This represents the number of people in the training set. Indicates the relationship with the first Hyperparameters corresponding to each organization; The effect score for each tissue or cell type was obtained by weighted summation of the regulatory elements. Then, a linear model is used to estimate the weights for each tissue or cell type. :

5. The method according to claim 1, characterized in that, The hyperparameters of the PRS score prediction model The following steps are used to determine: Select undetermined values ​​for the hyperparameters; For each undetermined value of a hyperparameter, predict the PRS in the calibrated population, then calculate the accuracy of the predicted PRS compared to the true PRS, and select the hyperparameter with the highest accuracy.

6. The method according to claim 5, characterized in that, The accuracy of the calculated predicted PRS compared to the actual PRS includes: Assuming the calibration group has M individuals, their phenotypic predictions are as follows: Its actual phenotype is ; For continuous phenotypes, the following PCC score is used to measure accuracy: in, This represents the average of the actual phenotypic values ​​in the calibration group. This represents the average predicted phenotypic value of the calibration group. For the 0-1 phenotype, the accuracy ACC score is used as a measure of accuracy: in, This indicates the number of predicted phenotypic values ​​that equal the actual phenotypic values.

7. A multi-gene genetic risk scoring system based on a multi-tissue regulatory network map, characterized in that, The system comprising performing the method of any one of claims 1 to 6, wherein the system includes: The data acquisition module is used to acquire the summary statistics of GWAS and the phenotypic and genotypic data of the population for a given phenotype or disease. The mathematical modeling module is used to construct and train a PRS score prediction model based on the organization-specific regulatory network map using the acquired data. The PRS score prediction module is used to predict the PRS score of an individual by inputting their genotype into a pre-trained PRS score prediction model.

8. A computer device, characterized in that, It includes a memory and a processor, the memory storing a computer program configured to be executed by the processor, the computer program including instructions for performing the method of any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a computer, implements the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Method for predicting survival prognosis of bladder cancer patient based on lncRNA optimization model

    CN115762792A

  • Systems, devices and media for de novo prediction of regulatory mutations based on single cell sequencing

    CN116486913A