Adenovirus susceptibility risk assessment model and biomarkers
By assessing adenovirus susceptibility through genome-wide association analysis and multi-gene risk prediction models, the problem of inaccurate adenovirus susceptibility assessment was solved, enabling individualized risk assessment and preventive measures, and reducing the side effects of adenovirus vector vaccines.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGZHOU KINGMED TRANSFORMATIVE MEDICINE INST CO LTD
- Filing Date
- 2022-12-27
- Publication Date
- 2026-06-30
AI Technical Summary
Current technology cannot effectively determine an individual's susceptibility to adenovirus, leading to inaccurate risk assessment of adenovirus infection and making it difficult to predict the side effects of adenovirus vector vaccines.
Adenovirus susceptibility loci and susceptibility effect values were obtained through genome-wide association analysis (GWAS), a multi-gene risk prediction model for adenovirus susceptibility was established, an SNP locus set was used to assess individual adenovirus susceptibility risk, and risk assessment was performed in conjunction with biomarkers.
It improves the accuracy of adenovirus susceptibility assessment, can distinguish between adenovirus-susceptible populations and control groups, provides individualized preventive measures, and reduces the risk of side effects from adenovirus vector vaccines.
Smart Images

Figure CN116052757B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of gene analysis technology, and in particular to an adenovirus susceptibility risk assessment model and biomarkers. Background Technology
[0002] Adenovirus infection can cause infections of the respiratory tract, eyes, gastrointestinal tract, urinary tract, and nervous system. Symptoms of adenovirus infection are generally mild, but due to individual genetic differences, some people are asymptomatic while others experience serious symptoms and consequences. For example, some patients may experience a severe dry cough similar to whooping cough; some adenovirus infections can cause acute otitis media, and when affecting the lower respiratory tract, can lead to bronchiolitis and pneumonia, although uncommon, these can become severe in infants; furthermore, some adenovirus infections can cause keratoconjunctivitis, a more serious type of infection affecting both corneas and conjunctiva of both eyes; additionally, some adenovirus infections can sometimes cause meningitis and encephalitis.
[0003] These symptoms are determined by doctors based on their experience and the phenotype of the disease to determine whether it is caused by adenovirus infection. Further methods, such as qPCR testing for adenovirus DNA or antigen testing for the presence of adenovirus antigens and antibodies, are used to confirm whether it is an adenovirus infection.
[0004] However, while these methods can determine whether an infection is caused by adenovirus and what type of adenovirus it is, they cannot determine whether an individual is susceptible to adenovirus infection.
[0005] Furthermore, many inactivated vaccines use adenovirus as a vector, and the side effects experienced by some individuals after receiving adenovirus-vectored vaccines are likely due to susceptibility to the adenovirus itself. However, it is currently impossible to determine whether the side effects of adenovirus vector vaccines are caused by human genetic background. Therefore, adenovirus susceptibility can provide some reference value for understanding the side effects of adenovirus vector vaccines. Summary of the Invention
[0006] Therefore, it is necessary to address the aforementioned problem of the lack of methods for determining adenovirus susceptibility by providing a method for establishing an adenovirus susceptibility risk assessment model. The adenovirus susceptibility risk assessment model obtained using this method can score the adenovirus susceptibility risk of the person being assessed, thereby assessing the person's risk of adenovirus infection and enabling more appropriate preventive measures to be taken.
[0007] A method for establishing an adenovirus susceptibility risk assessment model includes the following steps:
[0008] Susceptibility loci and susceptibility effect value analysis: Whole genome SNP locus data of adenovirus-infected and non-infected individuals were obtained and analyzed using the GWAS method to obtain susceptibility loci for adenovirus infection and corresponding susceptibility effect values.
[0009] Establishment of a polygenic risk prediction model for adenovirus susceptibility: The adenovirus susceptibility polygenic risk score (PRS) was calculated using the following formula:
[0010] PRS=β1x1+β2x2+…+β k x k +…+β n x n
[0011] In the above formula, β is the susceptibility effect value corresponding to the susceptibility site obtained by GWAS analysis, x is the number of risky allele mutations, n is the total number of SNPs included in the PRS analysis model, and k is each SNP included in the PRS analysis model.
[0012] The aforementioned method for establishing an adenovirus susceptibility risk assessment model involves analyzing variation data obtained from whole-genome microarray analysis of adenovirus-infected patients and comparing this data with GWAS analysis of a control group (non-infected group) to identify susceptibility loci in the population. These loci can provide a preliminary assessment of an individual's likelihood of adenovirus infection. However, relying on only one or a few loci is insufficient to determine susceptibility, its degree, or probability. This invention further utilizes a multi-gene risk prediction model for adenovirus susceptibility, employing a comprehensive assessment based on a set of SNP loci to improve assessment accuracy.
[0013] Understandably, increasing the sample size can improve the accuracy of susceptibility loci and corresponding susceptibility effect values in the analysis steps. These susceptibility loci and corresponding susceptibility effect values can be archived for future use. In subsequent adenovirus susceptibility multi-gene risk prediction models, the above data can be directly called to obtain the PRS value of a specific individual. The higher the PRS value, the higher the individual's susceptibility to adenovirus.
[0014] In one embodiment, x is determined by the following method:
[0015] If the gene locus is a risk-free allele, then x = 0;
[0016] If the gene locus is a heterozygous risk gene, then x = 1;
[0017] If the gene locus is a homozygous risk gene, then x = 2.
[0018] In one embodiment, the susceptible site and susceptible effect value analysis step is performed according to the following method:
[0019] 1) DNA extraction: Genomic DNA was extracted from adenovirus-infected and non-infected individuals to obtain samples for testing;
[0020] 2) Mutation detection: The above-mentioned samples were tested using a whole-genome SNP chip to obtain mutation site data;
[0021] 3) Quality control: Take the above mutation site data, delete duplicate quality control mutations and abnormal chromosome-related CNV sites to obtain a reference mutation set, and filter the samples with a site deletion rate ≥0.03 to obtain the mutation set;
[0022] 4) Genotype filling: The above-mentioned quality-controlled data are first processed for haplotype analysis using the East Asian population database, and then the whole genome data is filled using the Chinese population database and the East Asian population database as references, and then merged into whole genome data.
[0023] 5) GWAS analysis: Perform GWAS analysis on the above data to obtain adenovirus susceptibility sites and corresponding susceptibility effect values for later use.
[0024] The aforementioned East Asian population database, Chinese population database (WBBC database of Westlake University), and East Asian population database are all existing publicly available databases that can be populated online using the website https: / / imputationserver.westlake.edu.cn / stat.html. Those skilled in the art can also choose different databases for population based on the available data.
[0025] In one embodiment, in step 3), the sites with a mutated genotype deletion rate greater than 2%, Hardy-Weinberg equilibrium < 0.00001, and minimum allele frequency less than 1% are also filtered.
[0026] In one embodiment, in step 4) genotype filling, the filled data is subjected to quality control screening using MAF>0.01 and R2>0.8 conditions, and then merged into whole genome data.
[0027] In one embodiment, in step 5) GWAS analysis, GWAS analysis is performed using the SPA method.
[0028] The present invention also discloses the adenovirus susceptibility risk assessment model obtained by the above-described method.
[0029] This model can score the adenovirus susceptibility risk of individuals being assessed, thereby evaluating their risk of adenovirus infection and enabling more appropriate preventative measures to be taken.
[0030] This invention also discloses a method for assessing adenovirus susceptibility risk for non-diagnostic and therapeutic purposes, comprising the following steps:
[0031] Detection: Take biological samples of the object to be evaluated and obtain data on its whole genome mutation sites;
[0032] Analysis: Substituting the above data into the adenovirus susceptibility risk assessment model, the adenovirus susceptibility polygenic risk score (PRS) is calculated.
[0033] The present invention also discloses a terminal device, including a processor and a storage medium communicatively connected to the processor, the storage medium being used to store multiple instructions; the processor being used to call the instructions in the storage medium to execute the steps of establishing the above-described adenovirus susceptibility risk assessment model.
[0034] It is understandable that the aforementioned processors, storage media, and other devices, with conventional equipment settings, only need to be able to execute the steps of the method for establishing the adenovirus susceptibility risk assessment model.
[0035] This invention also discloses a biomarker for assessing adenovirus susceptibility, selected from: rs2260051, rs6990526, rs73603631, rs140305398, rs3135044, rs2736171, rs10215532, rs7969580, rs11609803, rs849744, rs849743, rs4767789, rs11833097, rs4382429, rs849736, rs7803717, rs702481, rs12044011, rs73165889, rs6772702, rs2057424, rs205 7423, rs10158840, rs9290399, rs13315738, rs13315773, rs6424776, rs28548537, rs3923815, rs4328802, rs60663273, rs113855439, rs7005987, rs 12401416, rs9290400, rs73165885, rs2736172, rs7975974, rs7979652, rs2260000, rs7980351, rs4407416, rs9784390, rs112052539, rs9831063, rs 9831327, rs9818414, rs73165892, rs73165893, rs176917, rs849739, rs9836242, rs12667701, rs9816216, rs12636032, rs7627847, rs9848167, rs19 86917, rs9851225, rs35871527, rs13290620, rs2260050, rs7642107, rs176913, rs72663326, rs112066938, rs7420912, rs2351144, rs11900438, rs1 1900638, rs34528777, rs35512426, rs10173843, rs10200441, rs2948097, rs10804267, rs11897511, rs13387832, rs57049829, rs176914, rs7316587 0, rs437499, rs10038005, rs4920928, rs4920860, rs6738903, rs12995490, rs1534752, rs7586890, rs78667444, rs10037986, rs451488, rs6777429,rs7645350,rs7355958,rs9817616,rs9874923,rs892666,rs12516894,rs550438,rs2933608,rs55639451,rs148102884,rs9872371,rs9854758,rs9854918,rs75797766,rs9845497,rs4858767,rs7564233,rs13018514,rs1527756,rs2929394,rs10162372,rs10162541,rs2933607,rs75877424,rs10202061,rs10193997,rs3993756,rs11860053,rs13420750,rs10198409,rs7597534,rs148371058,rs8050478,rs10527169,rs34080124,rs4388421,rs10190385,rs79962582,rs11887446,rs16839701,rs13020355,rs12622351,rs12622374,rs12622403,rs73818865,rs73818866,rs9575485,rs9575486,rs8019339,rs35316644,rs13422010,rs7583113,rs72808892,rs16839638,rs3130545,rs1197,rs9860918,rs9840899,rs8014484,rs12636078,rs72663325,rs2964662,rs2933617,rs2933616,rs55946451,rs4920925,rs7728276,rs4841399,rs200198541,rs381542,rs10133167,rs34148965,rs1953419,rs7141326,rs1959146,rs1959147,rs12638247,rs13267246,rs13259091,rs142841935,rs7592399,rs7604707,rs748815,rs4819925,rs73442564,rs79242093,rs57097666,rs77589922,rs80280231,rs73165898,rs6032869,rs1402071,rs17015335,rs4012781,rs10687441,rs7016825, rs74918533, rs1722879, rs10040214, rs7150435, rs12619668, rs71084407, rs12694437, rs10166901, rs75328290, rs11332290 5. rs117367684, rs79575740, rs4451367, rs7565549, rs17015360, rs72808888, rs10496206, rs72808889, rs55819663, rs72808890, rs7280 8891, rs59189165, rs17015366, rs72808894, rs17015370, rs7607586, rs5748768, rs140839175, rs76129222, rs77749053, rs113154037, rs 33952221, rs1234192, rs142242611, rs36110121, rs2365375, rs147534986, rs77668443, rs72808902, rs11295813, rs77913079, rs3752247 80, rs369843708, rs376152889, rs4819926, rs201781347, rs2964659, rs2933614, rs4920926, rs148122239, rs57676824, rs10466424, rs57 978018, rs10220303, rs10220304, rs59538561, rs2411819, rs2411818, rs33920097, rs7701326, rs2964653, rs10043155, rs2365361, rs355 At least one gene locus from the following: rs17140312, rs2964658, rs76013707, rs17140330, rs2964655, rs4920927, rs7711109, rs7728633, rs2933593, rs16839720, rs2308575, rs73284149, rs2898817, rs10146179, rs10162506, rs743814, rs10958480, rs12155578, rs743813, rs7616764, rs70983797.
[0036] The above-mentioned sites are all adenovirus susceptibility sites with p values lower than 5e-5 obtained from GWAS analysis. These susceptibility sites can be used as biomarkers to further develop products for analyzing and evaluating adenovirus susceptibility.
[0037] Compared with the prior art, the present invention has the following beneficial effects:
[0038] The present invention provides a method for establishing an adenovirus susceptibility risk assessment model. First, a genome-wide association analysis (GWAS) method is used to obtain the susceptibility sites for adenovirus infection in the population. Then, an adenovirus susceptibility multi-gene risk prediction model is used to comprehensively assess the risk of the individual to adenovirus infection using a set of SNP sites, thereby assessing the risk and adopting more appropriate preventive measures.
[0039] Furthermore, the adenovirus susceptibility risk assessment model established in this invention, when validated with sample groups from different data sources, can distinguish between adenovirus-susceptible populations and control groups, demonstrating that the model of this invention has superior discrimination efficacy. Attached Figure Description
[0040] Figure 1 This is a QQ plot of the GWAS results in Example 1;
[0041] Figure 2 The Manhattan plot of the GWAS results in Example 1;
[0042] Figure 3 The PRS score distribution results for adenovirus-infected and non-infected samples of the optimal performance model in Example 1;
[0043] Figure 4 The PRS-predicted ROC curve of the optimal performance model in Example 1;
[0044] Figure 5 The PRS score distribution results for adenovirus-infected and non-infected samples in the optimized model of Example 1;
[0045] Figure 6 The PRS-predicted ROC curve of the optimized model in Example 1;
[0046] Figure 7 The PRS score distribution results for adenovirus-infected and non-infected samples of the optimal performance model in Example 2;
[0047] Figure 8 The PRS-predicted ROC curve of the optimal performance model in Example 2;
[0048] Figure 9 The PRS score distribution results for adenovirus-infected and non-infected samples in the optimized model of Example 2;
[0049] Figure 10 The ROC curve for PRS prediction of the optimized model in Example 2. Detailed Implementation
[0050] To facilitate understanding of the present invention, a more complete description will be given below with reference to the accompanying drawings. Preferred embodiments of the invention are shown in the drawings. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a thorough and complete understanding of the disclosure of the invention.
[0051] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention.
[0052] Unless otherwise specified, all reagents used in the following examples are commercially available; and all methods used in the following examples are conventional methods unless otherwise specified.
[0053] Example 1
[0054] An adenovirus susceptibility risk assessment model was established using the following method.
[0055] 1. Analysis of susceptible sites and susceptible effect values.
[0056] The first step of this invention is to identify susceptibility sites for adenovirus and calculate their effect size in adenovirus susceptibility. We conducted a GWAS analysis using 832 adenovirus-infected individuals and 431 non-adenovirus-infected individuals as controls to obtain susceptibility sites for adenovirus infection and their corresponding susceptibility effect sizes.
[0057] Specific steps:
[0058] (1) DNA extraction:
[0059] All adenovirus-infected samples underwent DNA extraction using standard methods. Genomic DNA was extracted from the samples using the QIAamp DNA Extraction Kit (QIAamp DNA Mini Kit) according to the kit's instructions. DNA concentration and A260 / A280 and A260 / A230 ratios were measured using a Nanodrop 2000. The extracted DNA samples were stored at -20°C for later use.
[0060] (2) Mutation detection using whole-genome microarrays:
[0061] DNA from adenovirus-infected samples and control groups was detected using a whole-genome SNP array. The experiment employed Illumina's Infinium Chinese Genotyping Array-24v1.0 BeadChip. All qualified samples were analyzed according to the procedures outlined in the Illumina Infinium HTS Assay Reference Guide. The main process was as follows: DNA was denatured with alkali and incubated overnight at a constant temperature for whole-genome DNA amplification; after amplification, the DNA was fragmented by enzyme digestion, followed by precipitation and resuspending; after resuspending, array hybridization, single-base extension, and staining were performed, and the array was scanned using an iScan scanning system to read the data after vacuum drying. After obtaining the raw data, preliminary data analysis was performed using Genomestudio software to calculate the sample call rate, generate a report, and obtain mutation site data.
[0062] (3) Sample and mutation quality control:
[0063] After deleting duplicate quality control mutations, non-chromosomal (such as XY, MT and related non-chromosomal sites in quality control) and related CNV sites, the above mutation sites yielded 681,640 genome-wide sites, which will serve as the reference mutation set for our subsequent studies. The mutation information of each individual will be used as a reference to construct the VCF data of the sample.
[0064] Since the chip is based on the human reference genome version hg19 / GRCh37, this genome version will be used as the reference version for subsequent data analysis.
[0065] We calculated the deletion rate for autosomes (chr1-22). Mutations that did not meet quality control or could not be detected in the microarray were marked with "- / -". Therefore, we calculated the deletion site ratio of all autosomes to the total number of autosomes detected, and obtained the site deletion rate for each sample.
[0066] After constructing the mutation data VCF for each sample, samples with a missing rate of less than 0.03 were selected for subsequent analysis to ensure the quality of the analysis.
[0067] In addition, sites with a genotype deletion rate greater than 2%, Hardy-Weinberg equilibrium < 0.00001, and minimum allele frequency less than 1% were filtered to obtain a high-quality mutation set.
[0068] Specifically, in this embodiment, preliminary quality control was first performed. Among them, 62,576 sites had a deletion rate greater than 0.02, another 24,161 sites did not meet Hardy-Weinberg equilibrium (p<0.00001), and 143,048 sites had a minimum allele frequency of less than 0.01. After filtering out low-quality sites, 443,805 high-quality sites remained, which were then used for subsequent whole-genome filling.
[0069] (4) Genome genotyping:
[0070] Since the microarray data analyzed above contained only about 700,000 mutations, less than 500,000 remained after quality control. This represents only one ten-thousandth of the 3 billion base pairs in the entire genome, leaving large areas blank. To avoid problems caused by missing data, we performed whole-genome genotyping.
[0071] After genotype imputation, the SNP density increases significantly. If there are no SNPs at positions associated with the phenotype, there is no significance before imputation, but significance may appear after imputation.
[0072] In this embodiment, we used Westlake University's Chinese population imputation website (https: / / imputationserver.westlake.edu.cn / stat.html) to imputate whole-genome data. To improve accuracy, we performed haplotype processing using the East Asian population (EAS), and then used Westlake University's Chinese population data (WBBC) and the East Asian population data from the 1000 Genomes Project (1000G EAS) as reference panels. This reference panel contains 9,968 haplotypes and 40,249,755 mutations, which can effectively imput our microarray mutation data into whole-genome data.
[0073] The imputation data were subjected to quality control screening using conditions such as MAF>0.01 and R2>0.8 (Estimated Imputation Accuracy) to obtain 4,315,312 high-quality loci. Finally, these were merged into whole-genome data for subsequent analysis.
[0074] (5) GWAS analysis:
[0075] Based on the above steps, we performed GWAS analysis on 4,315,312 high-quality loci. Specifically, the analysis was conducted using REGENIE v3.2.2 software, employing the SPA (Saddlepoint approximation) method, which corrects the p-values of the GWAS using the saddlepoint approximation method.
[0076] In general, traditional methods assume that T asymptotically follows a Gaussian distribution. However, when the proportion of case-control (i.e., adenovirus infection group - control group) is unbalanced and the tested variant has a low MAC (Minorallele count), the T statistic will deviate significantly from the Gaussian distribution. This invention uses the SPA method to adjust the test, which can obtain a more accurate P value.
[0077] Through GWAS analysis, we obtained 274 adenovirus susceptibility sites with p-values below 5e-5 and their corresponding susceptibility effect values beta (see Table 1 below).
[0078] Table 1. Adenovirus susceptibility sites and their corresponding susceptibility effect values beta (β)
[0079]
[0080]
[0081]
[0082] Note: Under the SNP locus information field, the first digit indicates the chromosome where the SNP is located, and the digits after the colon indicate the position of the locus on that chromosome (based on the human reference genome version hg19 / GRCh37). The last letter indicates the mutation requiring susceptibility assessment. To the left of the rightmost colon are the alleles identical to those in the reference genome, and to the right are the base information of the allele mutation whose susceptibility risk needs to be assessed (i.e., the allele whose effect value needs to be evaluated). For example, 6:31591918:A:T indicates that the SNP is located at base position 31591918 on chromosome 6, A is the allele identical to that in the reference genome, and T is the mutated allele base, requiring susceptibility effect assessment.
[0083] Table 1 above shows the relevant significant sites and their corresponding functions and genes. The most significant sites are: rs2260051 (7.86e-08), which is an intron mutation in the PRRC2A gene; rs6990526 (3.02e-7), which is a mutation in the enhancer regulatory region; and rs73603631 (3.29e-7), which is a mutation in the binding region of the CTCF transcription factor, which also belongs to the regulatory region mutation, etc. (see Table 1 for details).
[0084] Using the qqman package in R (v4.2.2) to generate QQ plots and Manhattan plots for quality control inspection and result display, such as... Figure 1-2 As shown. Figure 1 The QQ graph of the GWAS results. Figure 2 The Manhattan plot is the result of GWAS.
[0085] The results show that the observed and expected values of the p-values are mostly on the diagonal (lambda GC = 1.036), indicating that the analysis results are reasonable.
[0086] Through the above steps, we obtained the susceptibility sites for adenovirus infection based on the p-value using GWAS analysis, and further obtained the susceptibility effect value of each site in adenovirus infection.
[0087] 2. Establishment of a polygenic risk prediction model for adenovirus susceptibility
[0088] The polygenic risk score (PRS) can predict an individual's risk of developing a complex genetic disease or their susceptibility to a specific pathogen based on large-scale population genetic background data. Therefore, we established a polygenic risk prediction model for adenovirus susceptibility using a population that did not overlap with the population used to calculate GWAS. We also established corresponding PRS models using 133 individuals with adenovirus infection and 63 control groups without adenovirus infection.
[0089] PRS is calculated using the following formula:
[0090] PRS=β1x1+β2x2+…+β k x k +…+β n x n
[0091] β: The effect size of the mutation as assessed by GWAS;
[0092] X: The number of risky allele mutations (0 for risk-free alleles, 1 for heterozygous alleles, and 2 for homozygous alleles);
[0093] N: The total number of SNPs included in the PRS model;
[0094] k: represents each mutation that is included.
[0095] Among them, the β value and k value can be obtained from the GWAS results, namely the effect alleles and their beta values; the x value can be obtained from the mutation detection in the test population; and the n value represents the total number of SNPs included in the PRS model. The selection of which SNPs to analyze needs to be obtained by establishing a specific model and then calculating it.
[0096] The following example illustrates the PRS calculation above.
[0097] Table 2. Examples of PRS Model Analysis
[0098]
[0099] Figure 2 For the PRS calculations of three different subjects, we can see that the effect alleles and effect value beta can be obtained from GWAS data; the gene locus set is the set of all SNPs used in the PRS model analysis, which includes 8 SNPs in Table 2, i.e., n=8. The mutation sets of the three subjects are shown above, and x represents the number of effect alleles, which is determined according to each person's genotype. Finally, the PRS score for each person is obtained by weighting the effect alleles and effect values in the gene set of the SNPs.
[0100] Different SNP sets can be used to calculate different PRS values, and the best-performing (best performance in evaluation) SNP sets have their own specific SNP sets.
[0101] In this embodiment, we selected 196 samples that differed from the GWAS baseline data, including 133 adenovirus-infected samples and 63 control samples without adenovirus infection. The C+T method was used to calculate PRS, which involves SNP clumping based on the total SNP linkage in the baseline data (73,723 clumped SNPs in this invention) and p-value selection based on the GWAS results. Using PRSice 2 with different p-values as thresholds, 9,855 different SNP sets were obtained based on the SNP clumping results. The optimal model was selected using the confidence level R², and model5962 was chosen as the optimal PRS model for this invention. This model includes 44,817 clumped SNP sets with p-values below 0.29855, making it the optimal performance model.
[0102] The optimal performance PRS model can effectively distinguish between adenovirus-susceptible populations and control groups, such as... Figure 3 As shown, Figure 3 The distribution of PRS scores for adenovirus-infected and non-infected samples in model 5962 is shown; additionally, its predictive power is as follows: Figure 4 As shown, Figure 4 The ROC curve predicted for the PRS model of model5962 has an AUC of 0.7112 (95% CI: 0.6373-0.7851), which demonstrates the predictive power of the model.
[0103] However, considering the large number of SNPs used in this model, it may lead to high costs in subsequent applications. We further optimized the SNP set of the model, obtaining an SNP set of 1993 SNPs, the details of which are shown in the table below.
[0104] Table 3. Optimization of SNP set
[0105]
[0106]
[0107]
[0108]
[0109]
[0110]
[0111]
[0112]
[0113]
[0114]
[0115]
[0116]
[0117]
[0118]
[0119]
[0120]
[0121] Note: Under the SNP locus information field, the first number indicates the chromosome on which the SNP is located, the number after the colon indicates the position of the locus on that chromosome (based on the human reference genome version hg19 / GRCh37), the last letter indicates the mutated base for susceptibility assessment, the allele bases to the left of the colon are the same as those in the reference genome, and the allele base information to the right of the colon is the allele base information for which susceptibility risk assessment is required (i.e., the allele whose effect value needs to be calculated).
[0122] Using the optimized SNP set model to calculate the 196 samples that differed from the GWAS baseline data, the optimized SNP set model could also effectively distinguish between adenovirus-susceptible and control groups. Figure 5 As shown, Figure 5 To optimize the PRS score distribution of adenovirus-infected and non-infected samples in the model; additionally, its predictive performance is as follows: Figure 6 As shown, Figure 6The ROC curve predicted by the optimized PRS model has an AUC of 0.6788 (95% CI: 0.5981-0.7596), indicating that the optimized model also has the corresponding predictive function.
[0123] Example 2
[0124] Because PRS calculation involves parameter optimization, overfitting is a potential issue. To address this, we validated the two PRS models (optimal performance model and optimized model) in Example 1 using GWAS data and 132 samples (90 adenovirus-infected individuals and 42 non-adenovirus-infected individuals) outside of the PRS modeling dataset. We performed discriminative validation by calculating each individual's PRS score.
[0125] The results are as follows Figure 7-10 As shown, where, Figure 7-8 The figures show the PRS score distributions for adenovirus-infected and non-infected samples, respectively, and the ROC curves predicted by the PRS model for the optimal performance model. Figure 9-10 The figures show the PRS score distributions for adenovirus-infected and non-infected samples in the optimized model, and the ROC curves predicted by the PRS model, respectively.
[0126] The results above show that the PRS scores calculated by the optimal power model and the optimized model can distinguish between the adenovirus-susceptible population and the control population in the validation set. Their AUC predictive power is 0.7056 (95% CI: 0.6109-0.8002) and AUC = 0.7058 (95% CI: 0.6165-0.7951), respectively, indicating that the model of the present invention can be applied to practical applications to assess adenovirus susceptibility in populations.
[0127] Furthermore, the optimized model achieves excellent predictive performance with a smaller SNP set, making it more suitable for a wide range of applications and offering advantages such as low cost and high accuracy.
[0128] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0129] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.
Claims
1. A method for establishing an adenovirus susceptibility risk assessment model, characterized in that, Includes the following steps: Susceptibility site and susceptibility effect value analysis: conducted according to the following methods: 1) DNA extraction: Genomic DNA was extracted from adenovirus-infected and non-infected individuals to obtain samples for testing; 2) Mutation detection: The above-mentioned samples were tested using a whole-genome SNP chip to obtain mutation site data; 3) Quality control: Take the above mutation site data, delete duplicate quality control mutations and abnormal chromosome-related CNV sites to obtain a reference mutation set. Filter samples with a site deletion rate ≥0.03, and also filter sites with a genotype deletion rate greater than 2%, Hardy-Weinberg equilibrium <0.00001, and minimum allele frequency less than 1% to obtain the mutation set. 4) Genotype imputation: The data that has undergone quality control processing above is first processed for haplotype in the East Asian population database, and then the whole genome data is imputed using the Chinese population database and the East Asian population database as references. The imputed data is then subjected to quality control screening with MAF>0.01 and R2>0.8 conditions, and then merged into whole genome data. 5) GWAS analysis: Perform GWAS analysis on the above data to obtain adenovirus susceptibility sites and corresponding susceptibility effect values for later use; Establishment of a polygenic risk prediction model for adenovirus susceptibility: The adenovirus susceptibility polygenic risk score (PRS) was calculated using the following formula: PRS=β 1 x 1 +β 2 x 2 +…+β k x k +…+β n x n In the above formula, β is the susceptibility effect value corresponding to the susceptibility site obtained by GWAS analysis, x is the number of risky allele mutations, n is the total number of SNPs included in the PRS analysis model, and k is the SNP included in the PRS analysis model. The x is determined by the following method: If the gene locus is a risk-free allele, then x=0; If the gene locus is a heterozygous risk gene, then x=1; If the gene locus is a homozygous risk gene, then x=2.
2. The method for establishing an adenovirus susceptibility risk assessment model according to claim 1, characterized in that, In step 5) of the GWAS analysis, the SPA method is used to perform the GWAS analysis.
3. The adenovirus susceptibility risk assessment model obtained by the method described in any one of claims 1-2.
4. A method for assessing adenovirus susceptibility risk for non-diagnostic and therapeutic purposes, characterized in that, Includes the following steps: Detection: Take biological samples of the object to be evaluated and obtain data on its whole genome mutation sites; Analysis: Substituting the above data into the adenovirus susceptibility risk assessment model described in claim 3, the adenovirus susceptibility polygenic risk score (PRS) is calculated.
5. A terminal device, characterized in that, The method includes a processor and a storage medium communicatively connected to the processor, the storage medium being used to store multiple instructions; the processor is used to invoke the instructions in the storage medium to execute the steps of establishing the adenovirus susceptibility risk assessment model according to any one of claims 1-2.
Citation Information
Patent Citations
Method for constructing disease classification model based on multi-gene risk scoring
CN113066586A