Disease risk prediction and nutrient recommendation device based on personal genome
By constructing the optimal PRS model and personalized nutrient recommendations that are suitable for East Asian populations, combined with the human protein-compound-food relationship network, the problem of insufficient accuracy in the prediction of disease risk in East Asian populations is solved, and the effectiveness of personalized nutrition intervention is achieved.
Patent Information
- Application Number
- CN202510335593.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art has insufficient accuracy in the prediction of disease risk in East Asian populations, lack of effective nutritional intervention measures, and fails to fully combine traditional wisdom with food activity factors, resulting in users being aware of the risks but no effective intervention plan.
The optimal PRS model is constructed to adapt to East Asian populations, identify high-risk genes based on high-risk sites, and use the GeneRank algorithm to construct personalized nutrient recommendations, and combine human protein-compound-food relationship network to recommend personalized nutrients.
It has improved the accuracy and effectiveness of individual health management and disease prevention in East Asian populations, provided practical and effective nutritional intervention plans, and filled the gap in the existing technology.
Smart Images

Figure CN120299626A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of health and food, and particularly relates to a device for disease risk prediction and nutrient recommendation based on an individual's genome. Background Art
[0002] With the progress of genomics and bioinformatics, health management based on an individual's genetic data (such as disease prediction, drug response assessment services provided by BGI, 23andMe, etc.) has become the core direction of precision medicine. However, existing technologies mostly focus on disease risk early warning and genetic trait analysis. Although they cover diverse information such as ancestry and exercise genes, there are gaps in the field of nutritional intervention: especially for the East Asian population, due to data scarcity and insufficient model pertinence, the accuracy of existing technologies in disease risk prediction is limited. In addition, existing technologies have not fully combined the traditional wisdom of "medicine and food having the same origin" and "the superior doctor prevents diseases". There are also many active factors in foods and nutrients that have regulatory and interaction effects on genes, thereby achieving the effects of preventing diseases and improving physical conditions. For example, red dates nourish qi and blood, and Chinese yams strengthen the spleen and stomach, converting genomic data into personalized dietary therapy plans. Previous health reports based on an individual's genome would make users aware of their own disease risks but lack effective intervention measures, falling into the dilemma of "knowing the risks but having no countermeasures". Summary of the Invention
[0003] The present invention is to make up for the deficiencies existing in the above-mentioned prior art, and provides a device for disease risk prediction and nutrient recommendation based on an individual's genome. The present invention constructs an optimal PRS model adapted to the East Asian population, identifies high-risk diseases of an individual based on the optimal PRS model, screens high-risk loci, annotates and obtains potential high-risk genes based on the high-risk loci, and accurately recommends personalized nutrients using the information of the potential high-risk genes, thereby improving the accuracy and effectiveness of individual health management and disease prevention for the East Asian population.
[0004] In order to achieve the above-mentioned invention purpose, the present invention adopts the following technical solutions:
[0005] A device for disease risk prediction and nutrient recommendation based on an individual's genome, comprising:
[0006] A quality control module for personal genome data of a sample to be tested: used to perform quality control on the personal genome data of the sample to be tested;
[0007] An optimal PRS model construction module: used to obtain basic data and target data, construct an optimal PRS model for each disease type, calculate the PRS score corresponding to the disease type of the sample to be tested and the samples in the target data, and further calculate the Z score and risk ratio corresponding to the disease type of the sample to be tested. The disease types with the risk ratio of the sample to be tested greater than the risk ratio threshold are high-risk diseases;
[0008] Identification and annotation module: used to identify and gene-annotate single nucleotide polymorphism sites associated with high-risk diseases;
[0009] Personalized nutrient recommendation module: based on high-risk genes, obtain high-risk human proteins, construct pairs of human protein-associated human proteins, human protein-compound pairs, and food-compound pairs, obtain associated human proteins associated with high-risk human proteins based on the pairs of human protein-associated human proteins, each human protein, associated human protein, compound, and food is used as a node in the association network, use the GeneRank recommendation algorithm to score the nodes, and select the top set number of foods with the highest scores as recommended nutrients.
[0010] Quality control of the personal genomic data of the test sample as described above includes:
[0011] Unify the human reference genome version on which the genomic data of the test sample is based to the hg38 version;
[0012] Impute the genotypes of the missing genotype data of the test sample;
[0013] Delete low-quality sites in the genotype data.
[0014] Obtaining the basic data and target data as described above includes:
[0015] Retrieve the target disease in the genome-wide association study catalog database, screen and download the research data with East Asian populations as samples as the basic data related to the disease type;
[0016] The basic data includes the following parameters: the race, disease type, sample size, reference genome version, and single nucleotide polymorphism sites of the sample, and also includes the effect allele, effect allele frequency, effect value, and its significance level associated with the single nucleotide polymorphism sites;
[0017] Download the research data with East Asian populations from the UK Biobank and record it as the target data;
[0018] The target data includes multiple samples, and each sample includes: the genotype of the sample, sample ID, age, gender, disease status, disease type, and reference genome version;
[0019] Verify the consistency of the race, disease type, and reference genome version of the basic data and the target data;
[0020] Perform quality control on the basic data;
[0021] Perform quality control on the target data.
[0022] The quality control of the basic data as described above includes:
[0023] Using linkage disequilibrium score regression to estimate the heritability of the basic data. If the heritability > 0.05, the basic data is considered qualified; otherwise, it is unqualified.
[0024] Identifying the effect alleles in the basic data;
[0025] Removing single nucleotide polymorphism (SNP) sites in the basic data with a minor allele frequency < 0.01 and an information quality score < 0.8;
[0026] The quality control of the target data as described above includes:
[0027] Excluding samples with inconsistent reported gender and physiological gender and a genotype missing rate > 0.01;
[0028] Removing SNP sites with a minor allele frequency (MAF) < 0.01 and a p-value of the Hardy-Weinberg equilibrium test < 1e-6 from the target data.
[0029] Constructing the optimal PRS model for each disease category as described above includes the following steps:
[0030] In the basic data, use the SNP clumping method of linkage disequilibrium in PLINK software for screening, and select SNP sites with a P-value not greater than the P-value threshold of 1 in the linkage disequilibrium block as significant SNP sites;
[0031] In the basic data, screen out SNP sites with each P-value not greater than the currently traversed P-value threshold through the SNP clumping method of linkage disequilibrium, and intersect the SNP sites according to the currently traversed P-value threshold with the significant SNP sites to obtain the set of SNP sites under the currently traversed P-value threshold. Further obtain the sets of SNP sites under each P-value threshold and construct the PRS model:
[0032]
[0033] where PRS ik is the PRS score for the k-th disease of the i-th sample, M i is the total number of SNP sites in the set of SNP sites under the P-value threshold, β ijk is the effect value of the j-th SNP site of the i-th sample in the target data for the k-th disease in the basic data, G ij is the genotype of the i-th sample in the target data at the j-th SNP site;
[0034] Using the disease status of the selected disease type for each sample in the target data as the dependent variable, and the PRS score corresponding to the same selected disease type for each sample in the target data as the independent variable, calculate the coefficient of determination R corresponding to the PRS score obtained by the PRS model corresponding to the selected disease type at different P-value thresholds through regression analysis 2 , select the P-value threshold corresponding to the maximum coefficient of determination as the optimal P-value threshold for this disease type, and select the PRS model corresponding to the optimal P-value threshold as the optimal PRS model for the corresponding disease type.
[0035] Calculating the Z-score and risk ratio corresponding to the disease type of the sample to be tested as described above includes:[[]]
[0036] Calculate the PRS scores of the same disease type for all samples in the target data, further calculate the PRS score corresponding to the disease type of the sample to be tested, calculate the position of the PRS score of the disease type of the sample to be tested among the PRS scores of the same disease type for all samples in the target data, and obtain the Z-score corresponding to the disease type of the sample to be tested;
[0037] Use the Z-score corresponding to the disease type of the sample to be tested to calculate through the cumulative distribution function of the normal distribution, and obtain the risk ratio corresponding to each disease type of the sample to be tested.
[0038] Identifying and gene annotating the single nucleotide polymorphism sites associated with high-risk diseases as described above includes:[[]]
[0039] Obtain the single nucleotide polymorphism sites associated with the optimal PRS model corresponding to the high-risk disease as the screened single nucleotide polymorphism sites, find the single nucleotide polymorphism sites identical to the screened single nucleotide polymorphism sites in the basic data as the matching single nucleotide polymorphism sites, obtain the effect value and effect allele frequency corresponding to the matching single nucleotide polymorphism sites, and perform the following effect value screening and conversion rules on the matching single nucleotide polymorphism sites in the basic data:[[]]
[0040] If the effect allele frequency is greater than or equal to 0.5 and the effect value is greater than or equal to 0, or the effect allele frequency is less than 0.5 and the effect value is less than 0, then the corresponding effect value is set to empty and no subsequent analysis is performed;
[0041] If the effect allele frequency is less than 0.5 and the effect value is greater than or equal to 0, then keep the effect value unchanged.
[0042] If the effect allele frequency is greater than or equal to 0.5 and the effect value is less than 0, then use the absolute value of the effect value as the new effect value;
[0043] The basic data after effect value screening and conversion is denoted as the new basic data;
[0044] Sort the corresponding matched single nucleotide polymorphism sites in the new basic data from largest to smallest effect value from front to back, and denote it as sorted data;
[0045] Select the top 100 single nucleotide polymorphism sites with the largest effect values from the sorted data;
[0046] Perform high-risk gene annotation on the above top 100 single nucleotide polymorphism sites.
[0047] Constructing the relationship pairs of human protein - associated human protein, human protein - compound, and food - compound as described above includes:
[0048] Download the food - compound relationship pairs and corresponding scoring information from the FOODB database; search and download the relationship pairs of human protein - associated human protein and corresponding scoring information in the STITCH database; search and download the relationship pairs of human protein - compound and corresponding scoring information in the STRING database, and select the relationship pairs with a scoring information greater than 400;
[0049] Convert the compounds collected from the FOODB database into CID format, and convert the high-risk genes into a protein format that matches the STITCH database through the DAVID database to obtain high-risk human proteins.
[0050] Perform node scoring using the GeneRank recommendation algorithm as described above, and select the top set number of foods with the highest scores as recommended nutrients, including:
[0051] Based on the relationship pairs of human protein - associated human protein, obtain the associated human proteins associated with the high-risk human proteins. Based on the relationship pairs of human protein - associated human protein, human protein - compound, and food - compound, construct an association network, and each human protein, associated human protein, compound, and food serves as a node of the association network. Set the initial score of the node corresponding to the annotated high-risk gene in the association network to 1, and set the initial score of other nodes to 0.
[0052] Let the weight d of the associated nodes associated with the node corresponding to the high-risk gene in the association network be 0.85, and the weight d of the node corresponding to the high-risk gene be 0.15, and set the convergence threshold to 10 -5 Using the constructed association network above, adopt the GeneRank recommendation algorithm, spread the nodes with initial scores in the association network, and calculate the scores of all nodes on the association network until the scores of all nodes on the association network converge, obtain the scores of each food and sort them, and select the top several foods with the highest scores as recommended nutrients.
[0053] A computer device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, it implements the personal genome data quality control module, the optimal PRS model construction module, the identification and annotation module, and the personalized nutrient recommendation module of the above-mentioned device.
[0054] The present invention has the following beneficial effects compared with the prior art:
[0055] 1. The method of the present invention uses a variety of polygenic risk score (PRS) algorithms, combines their advantages, and constructs an optimal PRS model suitable for the East Asian population.
[0056] 2. In the field of personalized nutrient recommendation, the method of the present invention deeply explores the traditional concept of "diet therapy", closely associates genomic information with food nutrition. By identifying high-risk diseases of an individual through individual genome, screening high-risk genes related to high-risk diseases, and using the pairs of human protein - associated human proteins, the pairs of human protein - compounds, and the pairs of food - compounds, deeply analyzes the mutual relationship between food and genes, constructs a large-scale heterogeneous knowledge graph association network, and based on this association network, uses the GeneRank algorithm to accurately recommend suitable nutrients for each individual, filling the blank in this aspect of the prior art and providing an effective nutritional intervention plan. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 It is a schematic structural diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] To facilitate the understanding and implementation of the present invention by those of ordinary skill in the art, the present invention will be further described in detail below in conjunction with embodiments. It should be understood that the embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0059] In this embodiment, a method for disease risk prediction and nutrient recommendation based on personal genome starts to analyze from the obtained user personal genome data. Specifically, as Figure 1 shown, it is carried out according to the following steps:
[0060] Step 1: Perform quality control on the personal genome data of the test sample.
[0061] Step 1 as described above includes the following steps:
[0062] Step 1.1. Check whether the human reference genome version on which the genomic data uploaded for the sample to be tested is based is the hg38 version. If the versions are inconsistent, use the LiftOver tool to convert the genotype data to the hg38 version to ensure data compatibility;
[0063] Step 1.2. Detect whether there are any missing values in the genotype data uploaded for the sample to be tested. If there are missing values, use the reference panel data of the 1000 Genomes Project and perform genotype imputation through the beagle tool.
[0064] Step 1.3. Delete the low-quality loci with a genotype quality score < 30 in the genotype data.
[0065] Step 2. Obtain the basic data and target data of the East Asian population, construct the optimal PRS model for each disease, calculate the PRS scores corresponding to the disease types of the sample to be tested and the samples in the target data, and further calculate the Z scores corresponding to the disease types of the sample to be tested. Calculate using the cumulative distribution function of the normal distribution with the Z scores corresponding to the disease types of the sample to be tested to obtain the risk ratios corresponding to each disease type of the sample to be tested. The disease types with the risk ratio of the sample to be tested greater than the risk ratio threshold are high-risk diseases.
[0066] Step 2 as described above includes the following steps:
[0067] Step 2.1. Search for the target disease in the Genome-Wide Association Study Catalog (GWAS Catalog) database, screen and download the research data with the East Asian population as the sample as the basic data related to the disease type, base data.
[0068] The basic data includes the following parameters: the race of the sample (limited to the East Asian population), disease type, sample size, reference genome version, single nucleotide polymorphism (SNP), and also includes the effect allele (EA), effect allele frequency (EAF), effect value (beta value) and its significance level (P value) associated with the single nucleotide polymorphism;
[0069] Step 2.2. Download the research data with the East Asian population from the UK Biobank database, denoted as target data.
[0070] The target data includes multiple samples, and each sample includes: the genotype of the sample (encoded as 0, 1, 2, representing the copy number of the risk allele respectively), sample ID, age, gender, disease status, disease type, and reference genome version;
[0071] Step 2.3: Conduct a preliminary check on the downloaded basic data and target data, verify the consistency of the race, disease type, and reference genome version of the basic data and target data to ensure data applicability.
[0072] Step 2.4: Perform quality control on the basic data (base data).
[0073] Step 2.4.1: Use Linkage Disequilibrium Score Regression (LDSC) to estimate the heritability of the basic data. If the heritability > 0.05, the basic data is considered qualified; otherwise, the basic data is unqualified.
[0074] Step 2.4.2: Confirm the effect allele (EA) in the basic data (base data).
[0075] Step 2.4.3: Remove single nucleotide polymorphism (SNP) sites in the basic data with a minor allele frequency (MAF) < 0.01 and an information quality score (INFO, referring to the reliability score of SNP genotype inference) < 0.8.
[0076] Step 2.5: Perform quality control on the target data (target data).
[0077] Step 2.5.1: Exclude samples with inconsistent reported gender and physiological gender and a genotype missing rate > 0.01.
[0078] Step 2.5.2: Remove single nucleotide polymorphism (SNP) sites with a minor allele frequency (MAF) < 0.01 and a Hardy - Weinberg equilibrium test p - value < 1e - 6 from the target data.
[0079] Step 2.6: Construct and evaluate the PRS model.
[0080] Step 2.6.1: Based on the basic data after quality control in Step 2.4, use the SNP clumping method (clumping process) of linkage disequilibrium (LD) in PLINK software for screening. Select single nucleotide polymorphism sites with a P-value (significance level) not greater than the P-value threshold of 1 in the linkage disequilibrium (LD) block as significant single nucleotide polymorphism sites for analysis, reducing the correlation between single nucleotide polymorphism sites;
[0081] The specific commands are as follows:
[0082] plink --bfile tragetdata.QC --clump-p1 1 --clump-r2 0.1 --clump-kb 250\
[0083] --clump basedata.QC --clump-snp-field SNP --clump-field P\
[0084] --out targetdata
[0085] In this command:
[0086] --bfile specifies the prefix of the input binary file.
[0087] --clump-p1 sets the significance threshold, which is set to 1 here, indicating no significance filtering is considered.
[0088] --clump-r2 sets the r threshold of linkage disequilibrium LD 2 threshold, and SNPs with r 2 less than 0.1 are retained.
[0089] --clump-kb sets the physical distance threshold, and SNPs with a distance greater than 250kb are retained.
[0090] --clump specifies the file containing SNPs and P-values.
[0091] --clump-snp-field and --clump-field specify the fields of SNP identification and P-value respectively.
[0092] --out specifies the prefix of the output file.
[0093] Step 2.6.2: Based on the basic data after quality control in Step 2.4, use the SNP aggregation method of linkage disequilibrium to screen out single nucleotide polymorphism sites with P-values not greater than the currently traversed P-value threshold, and intersect the single nucleotide polymorphism sites corresponding to the currently traversed P-value threshold with the significant single nucleotide polymorphism sites in Step 2.6.1 to obtain the set of single nucleotide polymorphism sites under the currently traversed P-value threshold, and further obtain the sets of single nucleotide polymorphism sites under each P-value threshold for subsequent model construction and calculation. Calculate the PRS score based on the effect values of the basic data and the genotypes of the target data. The PRS score calculated by the following formula is used as the PRS model:
[0094]
[0095] where PRS ik is the PRS score of the k-th disease for the i-th sample, M i is the total number of single nucleotide polymorphism sites in the set of single nucleotide polymorphism sites under the P-value threshold (the total number of single nucleotide polymorphism sites corresponding to different P-value thresholds in the previous steps is different), β ijk is the effect value of the j-th single nucleotide polymorphism site of the i-th sample in the target data for the k-th disease in the basic data (from the basic data), G ij is the genotype of the i-th sample in the target data at the j-th single nucleotide polymorphism site (from the target data, encoded as 0, 1, 2);
[0096] Step 2.6.3: Evaluate the predictive ability of the PRS model and determine the optimal P-value threshold. Specifically, use the disease status of the selected disease type of each sample in the target data as the dependent variable, and the PRS scores corresponding to the same selected disease type of each sample in the target data as the independent variable to construct a regression model (for example, logistic regression for binary disease phenotypes or linear regression for continuous phenotypes). Calculate the proportion of the phenotypic variation explained by the PRS scores calculated by the PRS model through regression analysis, that is, calculate the coefficient of determination R 2 corresponding to the PRS scores obtained by the PRS model for the selected disease type under different P-value thresholds. For different P-value thresholds selected in the previous steps, the calculated coefficient of determination R 2 for the selected disease type is also different. Traverse the P-value threshold, calculate the coefficient of determination corresponding to the same disease type, and select the P-value threshold corresponding to the maximum coefficient of determination as the optimal P-value threshold for this disease type. Select the PRS score of the k-th disease corresponding to the optimal P-value threshold as the optimal PRS model for the k-th disease.
[0097] Step 2.7: Sample disease risk prediction.
[0098] Using the optimal PRS model of the k-th disease determined in step 2.6, calculate the PRS scores of all samples in the target data for the same disease category (55 in this embodiment), further calculate the PRS score corresponding to the disease category of the sample to be tested, calculate the position of the PRS score of the disease category of the sample to be tested among the PRS scores of all samples in the target data for the same disease category, and obtain the Z-score (Z-score, which refers to the standardized score, indicating the deviation degree of an individual's PRS from the population mean) corresponding to the disease category of the sample to be tested;
[0099] Use the Z-score corresponding to the disease category of the sample to be tested to calculate through the cumulative distribution function (CDF) of the normal distribution to obtain the risk ratio corresponding to each disease category of the sample to be tested.
[0100] If the risk ratio of the disease category of the sample to be tested is greater than the risk ratio threshold (90%), then the disease category corresponding to the sample to be tested is a high-risk disease, and subsequent steps 3 and 4 are executed.
[0101] Specifically, when implementing, taking a male sample as an example for disease risk prediction analysis, the prediction results of some of its diseases are as follows: the risk ratio of gastric cancer is 35.06%, higher than 35% of the samples in the target data; the risk ratio of prostate cancer is 71.48%, higher than 70% of the samples in the target data; the risk ratio of pulmonary fibrosis is 93.09%, significantly higher than 90% of the samples in the target data; the risk ratio of type 2 diabetes is 60.48%, higher than 60% of the samples in the target data set. Among them, the risk ratio of pulmonary fibrosis of this male sample is the highest, belonging to an individual high-risk disease.
[0102] Step 3: Identification and gene annotation of single nucleotide polymorphism sites (SNP sites) of high-risk diseases.
[0103] Step 3 as described above includes the following steps:
[0104] Step 3.1: Identify the single nucleotide polymorphism sites related to the high-risk diseases of the sample to be tested, that is, obtain the single nucleotide polymorphism sites associated with the optimal PRS model corresponding to the high-risk diseases as the screening single nucleotide polymorphism sites, find the same single nucleotide polymorphism sites as the screening single nucleotide polymorphism sites in the basic data as the matching single nucleotide polymorphism sites, obtain the effect value (beta value) and the effect allele frequency (EAF) corresponding to the matching single nucleotide polymorphism sites, and execute the following effect value (beta value) screening and conversion rules for the matching single nucleotide polymorphism sites in the basic data:
[0105] If the effect allele frequency (EAF) is greater than or equal to 0.5 and the effect value (beta value) is greater than or equal to 0, or the effect allele frequency (EAF) is less than 0.5 and the effect value (beta value) is less than 0, then the corresponding effect value (beta value) is set to null NA and no subsequent analysis is performed.
[0106] If the effect allele frequency (EAF) is less than 0.5 and the effect value (beta value) is greater than or equal to 0, then the effect value (beta value) remains unchanged.
[0107] If the effect allele frequency (EAF) is greater than or equal to 0.5 and the effect value (beta value) is less than 0, then the absolute value of the effect value (beta value) is used as the new effect value (beta value).
[0108] The basic data after filtering and conversion of the effect value (beta value) is recorded as the new basic data trans data.
[0109] In specific implementation, taking the pulmonary fibrosis disease as an example, the data conversion information of some SNP loci in step 3.1 is shown in Table 1.
[0110] Table 1 Genetic information table of single nucleotide polymorphism loci related to some risks of pulmonary fibrosis
[0111]
[0112] Step 3.2: For the new basic data trans data processed in step 3.1, sort the corresponding matching single nucleotide polymorphism loci in the new basic data trans data in descending order according to the beta value, and record it as the sorted data sorted data;
[0113] In specific implementation, the sorting result of step 3.2 is shown in Table 2.
[0114] Table 2 Sorting result table of single nucleotide polymorphism loci related to some risks of pulmonary fibrosis
[0115] Single nucleotide polymorphism site Converted beta value rs28733018 1.40 rs12099011 0.87 rs12593207 0.66 rs2928267 0.64 rs11043769 0.60 rs2280594 0.57
[0116] Step 3.3: From the sorted data sorted data processed in step 3.2, screen out the first 100 single nucleotide polymorphism loci with the beta value;
[0117] Step 3.4: Use the gprofiler R package to annotate the top 100 single nucleotide polymorphism (SNP) sites in Step 3.2 for high-risk genes. During the annotation process, only the SNP sites carrying the effect allele (EA) are annotated, and the SNP sites that cannot be annotated by the gprofiler R package are discarded to ensure the accuracy and relevance of the annotation results;
[0118] Specifically, the gene annotation results of Step 3.4 are shown in Table 3, where the entries with the gene column marked as " / " indicate that no genes were annotated.
[0119] Table 3 Gene annotation results corresponding to SNP sites for some risks of pulmonary fibrosis
[0120] Single nucleotide polymorphism site High-risk gene annotation rs28733018 LINC01892 rs12099011 / rs12593207 CHRNB4 rs2928267 / rs11043769 / rs2280594 RORA
[0121] Step 4: Based on the high-risk gene annotation in Step 3.4 and combined with the food database, construct a personalized nutrient recommendation model based on the individual genome.
[0122] Step 4 as described above includes the following steps:
[0123] Step 4.1: Download the food-compound relationship pairs and the corresponding scoring information for these relationship pairs from the FOODB database; search and download the human protein-associated human protein relationship pairs and the corresponding scoring information for these relationship pairs from the STITCH database; search and download the human protein-compound relationship pairs and the corresponding scoring information for these relationship pairs from the STRING database, and select the relationship pairs with a scoring information greater than 400;
[0124] The FooDB (The Food Database) is the world's largest and most comprehensive database of food ingredients, chemistry, and biology. It is a freely accessible electronic database that provides detailed information on the chemical components in food and is applicable to fields such as nutrition, food science, diet planning, and education.
[0125] The STITCH (Search Tool for Interactions of Chemicals) database is a database for exploring known and predicted interactions between chemicals and proteins. It provides a comprehensive compound-protein interaction network by integrating multiple data sources (including experimental data, prediction models, and text mining results).
[0126] The STRING database (Search Tool for the Retrieval of Interacting Genes / Proteins) is a comprehensive bioinformatics database aimed at providing information on gene and protein interactions. It integrates multiple data sources, including experimentally verified, literature-reported, computationally predicted, and interaction information obtained from other databases.
[0127] Step 4.2: Convert the compounds collected from the FOODB database into the CID format to make it correspond to the format of the compounds in the STRING database. Convert the high-risk genes annotated in Step 3.4 into the protein format matching the STITCH database through the DAVID database to obtain high-risk human proteins.
[0128] Step 4.3: Based on the relationship pairs of human protein-associated human proteins, obtain the associated human proteins associated with the high-risk human proteins. Construct an association network based on the relationship pairs of human protein-associated human proteins, human protein-compound relationship pairs, and food-compound relationship pairs. Each human protein, associated human protein, compound, and food serves as a node in the association network. Set the initial score of the node corresponding to the high-risk gene annotated in Step 3.4 to 1, and set the initial scores of other nodes to 0.
[0129] Step 4.4: Let the weight d of the associated nodes associated with the node corresponding to the high-risk gene in the association network be 0.85, and the weight d of the node corresponding to the high-risk gene be 0.15. Set the convergence threshold to 10 -5 , using the constructed association network above, adopt the GeneRank recommendation algorithm. By spreading the nodes with initial scores in the association network and calculating the scores of all nodes on the association network until the scores of all nodes on the association network converge, obtain the scores of each food and sort them. Select the top several foods with the highest scores, which can accurately recommend the most suitable nutrients for each test sample corresponding to the high-risk gene.
[0130] In specific implementation, step 4.4 obtains the following partial food recommendation results and their corresponding scores through the Generank recommendation algorithm: Buckwheat - 3.03, Broccoli - 3.02, Spinach - 2.99, Carrot - 2.96, Celery - 2.93, Watermelon - 2.93, Mangos - 2.92, Kale - 2.90, Honey - 2.88, and Brussels_sprouts - 2.88. Based on the above scores, it is recommended that users increase the intake of the following foods in their daily diet: Buckwheat, Broccoli, Spinach, Carrot, Celery, Watermelon, Mangos, Kale, Honey, and Brussels_sprouts.
[0131] An automated process developed using Python integrates the user's genomic analysis, annotation information, and personalized nutrient recommendations to generate a comprehensive and easy-to-understand user report.
[0132] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods.
[0133] Embodiment 2:
[0134] A device for disease risk prediction and nutrient recommendation based on an individual's genome, comprising:
[0135] A quality control module for personal genomic data of the sample to be tested: used to implement step 1 of the above Embodiment 1, that is, to perform quality control on the personal genomic data of the sample to be tested;
[0136] An optimal PRS model construction module: used to implement step 2 of the above Embodiment 1, that is, to obtain basic data and target data, construct an optimal PRS model for each disease type, calculate the PRS scores corresponding to the disease types of the sample to be tested and the samples in the target data, and further calculate the Z scores and risk ratios corresponding to the disease types of the sample to be tested. The disease types with the risk ratio of the sample to be tested greater than the risk ratio threshold are high-risk diseases;
[0137] An identification and annotation module: used to implement step 3 of the above Embodiment 1, that is, to identify and gene-annotate single nucleotide polymorphism sites associated with high-risk diseases;
[0138] Personalized Nutrient Recommendation Module: It is used to implement step 4 of the above-mentioned Embodiment 1, that is, to obtain high-risk human proteins based on high-risk genes, construct relationship pairs of human protein-associated human proteins, human protein-compound relationship pairs, and food-compound relationship pairs, obtain the associated human proteins associated with high-risk human proteins based on the relationship pairs of human protein-associated human proteins. Each human protein, associated human protein, compound, and food is used as a node of the association network, and the GeneRank recommendation algorithm is used to score the nodes, and the top set number of foods with the highest scores are selected as recommended nutrients.
[0139] Embodiment 3:
[0140] In this embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above-mentioned method embodiments are implemented.
[0141] Embodiment 4:
[0142] In this embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0143] Embodiment 5:
[0144] In this embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0145] It should be noted that the specific embodiments described in the present invention are only illustrative of the spirit of the present invention. Those skilled in the art of the present invention can make various modifications or supplements to the described specific embodiments or use similar ways to replace them, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.
Claims
1. A device for disease risk prediction and nutrient recommendation based on an individual's genome, characterized in that Including: Personal genomic data quality control module for the sample to be tested: used to perform quality control on the personal genomic data of the sample to be tested; Optimal PRS model construction module: used to obtain basic data and target data, construct the optimal PRS model for each disease type, calculate the PRS scores corresponding to the disease types of the sample to be tested and the samples in the target data, and further calculate the Z scores and risk ratios corresponding to the disease types of the sample to be tested. The disease types with the risk ratio of the sample to be tested greater than the risk ratio threshold are high-risk diseases; Identification and annotation module: used to identify and gene annotate single nucleotide polymorphism sites associated with high-risk diseases; Personalized nutrient recommendation module: obtain high-risk human proteins based on high-risk genes, construct pairs of related human proteins, pairs of human protein-compound relationships, and pairs of food-compound relationships. Based on the pairs of related human proteins, obtain the related human proteins associated with high-risk human proteins. Each human protein, related human protein, compound, and food is used as a node in the association network. Use the GeneRank recommendation algorithm to score the nodes, and select the top set number of foods with the highest scores as the recommended nutrients.
2. The device for disease risk prediction and nutrient recommendation based on an individual genome according to claim 1, wherein The quality control of the personal genomic data of the sample to be tested includes: Unifying the human reference genome version on which the genomic data of the sample to be tested is based to the hg38 version; Performing genotype imputation on the missing genotype data of the sample to be tested; Deleting low-quality sites in the genotype data.
3. The device for predicting disease risk and recommending nutrients based on an individual's genome according to claim 1, characterized in that, The obtaining of the basic data and the target data includes: Searching for the target disease in the genome-wide association study catalog database, screening and downloading the research data with East Asian populations as samples as the basic data related to the disease type; The basic data includes the following parameters: the race, disease type, sample size, reference genome version, and single nucleotide polymorphism sites of the sample, and also includes the effect allele, effect allele frequency, effect value, and its significance level associated with the single nucleotide polymorphism sites; Downloading the research data with East Asian populations from the UK Biobank as the target data; The target data includes multiple samples, and each sample includes: the genotype of the sample, sample ID, age, gender, disease status, disease type, and reference genome version; Verifying the consistency of the race, disease type, and reference genome version of the basic data and the target data; Performing quality control on the basic data; Performing quality control on the target data.
4. The device for disease risk prediction and nutrient recommendation based on an individual's genome according to claim 3, wherein The quality control of the basic data includes: Using linkage disequilibrium score regression to estimate the heritability of the basic data. If the heritability > 0.05, the basic data is considered qualified, otherwise it is unqualified; Confirming the effect alleles in the basic data; Removing single nucleotide polymorphism sites with a minor allele frequency less than 0.01 and an information quality score less than 0.8 in the basic data; The quality control of the target data includes: Excluding samples with inconsistent reported gender and physiological gender and a genotype missing rate > 0.01; Remove single nucleotide polymorphism (SNP) sites with minor allele frequency (MAF) < 0.01 and Hardy-Weinberg equilibrium test p-value < 1e-6 from the target data.
5. The device for predicting disease risk and recommending nutrients based on an individual's genome according to claim 3, characterized in that, The construction of the optimal polygenic risk score (PRS) model for each disease category includes the following steps: In the basic data, use the SNP clumping method of linkage disequilibrium in PLINK software for screening, and select the SNP sites with P-values not greater than the P-value threshold of 1 in the linkage disequilibrium block as significant SNP sites. In the basic data, screen out SNP sites with each P-value not greater than the currently traversed P-value threshold through the SNP clumping method of linkage disequilibrium, and intersect the SNP sites according to the currently traversed P-value threshold with the significant SNP sites to obtain the set of SNP sites under the currently traversed P-value threshold, and further obtain the sets of SNP sites under each P-value threshold to construct the PRS model. Among them, PRS ik is the PRS score of the k-th disease for the i-th sample, M i is the total number of single nucleotide polymorphism sites in the set of single nucleotide polymorphism sites under the P-value threshold, β ijk is the effect value of the j-th single nucleotide polymorphism site of the i-th sample of the target data for the k-th disease in the basic data, G ij is the genotype of the i-th sample of the target data at the j-th single nucleotide polymorphism site; Using the disease status of the selected disease type of each sample in the target data as the dependent variable and the PRS score corresponding to the same selected disease type of each sample in the target data as the independent variable, calculate the coefficient of determination R corresponding to the PRS score obtained by the PRS model corresponding to the selected disease type at different P-value thresholds through regression analysis 2 , select the P-value threshold corresponding to the maximum coefficient of determination as the optimal P-value threshold for this disease type, and select the PRS model corresponding to the optimal P-value threshold as the optimal PRS model for the corresponding disease type.
6. The device for disease risk prediction and nutrient recommendation based on an individual genome according to claim 5, wherein The calculation of the Z-score and risk ratio corresponding to the disease category of the test sample includes: Calculate the PRS scores of all samples in the target data for the same disease category, further calculate the PRS score corresponding to the disease category of the test sample, calculate the position of the PRS score of the disease category of the test sample among the PRS scores of all samples in the target data for the same disease category, and obtain the Z-score corresponding to the disease category of the test sample. Use the Z-score corresponding to the disease category of the test sample to calculate through the cumulative distribution function of the normal distribution to obtain the risk ratio corresponding to each disease category of the test sample.
7. The device for predicting disease risk and recommending nutrients based on an individual's genome according to claim 1, wherein The identification and gene annotation of SNP sites associated with high-risk diseases include: Obtain the SNP sites associated with the optimal PRS model corresponding to the high-risk disease as the screening SNP sites, find the SNP sites identical to the screening SNP sites in the basic data as the matching SNP sites, obtain the effect value and effect allele frequency corresponding to the matching SNP sites, and perform the following effect value screening and conversion rules on the matching SNP sites in the basic data: If the effect allele frequency is greater than or equal to 0.5 and the effect value is greater than or equal to 0, or the effect allele frequency is less than 0.5 and the effect value is less than 0, then the corresponding effect value is set to null and no further analysis is performed. If the effect allele frequency is less than 0.5 and the effect value is greater than or equal to 0, then keep the effect value unchanged. If the effect allele frequency is greater than or equal to 0.5 and the effect value is less than 0, then take the absolute value of the effect value as the new effect value. The basic data after effect value screening and conversion is denoted as the new basic data. Sort the corresponding matching SNP sites in the new basic data in descending order of the effect value from front to back, denoted as the sorted data. Select the top 100 SNP sites with the largest effect values from the sorted data. Perform high-risk gene annotation on the above top 100 SNP sites.
8. The device for predicting disease risks and recommending nutrients based on an individual's genome according to claim 7, wherein The construction of human protein-related human protein relationship pairs, human protein-compound relationship pairs, and food-compound relationship pairs comprises: Download food-compound relationship pairs and corresponding scoring information from the FOODB database; search and download human protein-associated human protein relationship pairs and corresponding scoring information from the STITCH database; search and download human protein-compound relationship pairs and corresponding scoring information from the STRING database, and select relationship pairs with scoring information greater than 400; The compounds collected in the FOODB database were converted into CID format, and the high-risk genes were converted into a protein format that matched the STITCH database via the DAVID database to obtain high-risk human proteins.
9. The device for predicting disease risk and recommending nutrients based on an individual's genome according to claim 8, wherein The GeneRank recommendation algorithm is used to score nodes, and the top-scoring foods are selected as recommended nutrients, including: Based on the human protein-associated human protein relationship pairs, the associated human proteins associated with the high-risk human proteins are obtained. Based on the human protein-associated human protein relationship pairs, human protein-compound relationship pairs, and food-compound relationship pairs, the association network is constructed. Each human protein, associated human protein, compound, and food are used as nodes of the association network. The initial score of the node corresponding to the annotated high-risk gene in the association network is set to 1, and the initial score of other nodes is set to 0. Let the weight d of the associated nodes associated with the nodes corresponding to the high-risk genes in the association network be 0.85, and the weight d of the nodes corresponding to the high-risk genes be 0.
15. The convergence threshold is set to 10 -5 , using the constructed association network above, adopting the GeneRank recommendation algorithm, by spreading the nodes with the initial scores in the association network and calculating the scores of all nodes on the association network until the scores of all nodes on the association network converge, obtaining the scores of each food and sorting them, and selecting the top several foods with the highest scores as the recommended nutrients.
10. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the personal genome data quality control module of the test sample, the optimal PRS model construction module, the identification and annotation module, and the personalized nutrient recommendation module of the device according to any one of claims 1 to 9 are implemented.