Arteriosclerosis risk assessment model construction method
By constructing an arteriosclerosis risk assessment model based on the Chinese population, and utilizing indicators such as SNP sites and brachial-ankle pulse wave velocity, the problem of low accuracy in predicting arteriosclerosis risk in existing technologies has been solved, enabling accurate risk assessment of the Chinese population and early identification of high-risk individuals.
Patent Information
- Application Number
- CN202411294780.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-14
- Publication Date
- 2026-03-17
AI Technical Summary
The lack of a precise arteriosclerosis risk assessment model applicable to the Chinese population in the current technology results in low accuracy in predicting arteriosclerosis risk, making it difficult to identify high-risk individuals early and develop personalized prevention strategies.
A risk assessment model for arteriosclerosis was constructed. Based on genetic and background data of the Chinese population, association analysis was performed. Using SNP loci and association significance values, combined with indicators such as brachial-ankle pulse wave conduction velocity, a linear mixture model was constructed. SNP loci below the significance threshold were screened, and multiple general linear models were constructed. The accuracy of the model was evaluated through a validation set.
It achieves accurate risk assessment of arteriosclerosis in the Chinese population, with an AUC of 0.693, which can effectively identify high-risk individuals and reduce the incidence of arteriosclerosis and its complications.
Smart Images

Figure CN121674549A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biotechnology, specifically to a method for constructing an arteriosclerosis risk assessment model, an arteriosclerosis risk assessment method, arteriosclerosis biomarkers, and the use of reagents for detecting biomarkers in the preparation of kits for predicting, preventing, and / or treating arteriosclerosis. Background Technology
[0002] Arteriosclerosis, especially atherosclerosis, is a serious vascular disease and one of the leading causes of premature death worldwide. Arteriosclerosis not only affects the structure and function of arterial walls, leading to thickening, hardening, loss of elasticity, and narrowing of the lumen, but it is also often accompanied by the accumulation of lipids and complex carbohydrates in the arterial intima, forming plaques and triggering a series of cardiovascular events. Currently, brachial-ankle pulse wave velocity (baPWV), as a non-invasive and convenient indicator, has gradually been recognized by the medical community as an important means of assessing arteriosclerosis. baPWV measures the conduction time of the pulse wave between the brachial and ankle arteries, reflecting changes in arterial stiffness and elasticity. Its advantage lies in its relatively simple measurement method, requiring only a blood pressure cuff to be applied to the limbs, making it particularly suitable for large-scale epidemiological studies.
[0003] Current research on atherosclerosis genes is limited. The main technique used is the arterial stiffness index (ASI) for genome-wide association studies; however, no predictive models based on ASI for atherosclerosis risk have been developed. Furthermore, differences in physiological characteristics, lifestyles, and genetic factors among different populations make it difficult for relevant risk assessment models to accurately reflect specific situations, resulting in low accuracy in predicting atherosclerosis risk. Specifically, models capable of accurately predicting atherosclerosis risk for the Chinese population are currently lacking.
[0004] Therefore, there is an urgent need to provide a method for constructing an arteriosclerosis risk assessment model suitable for the Chinese population by combining brachial-ankle pulse wave velocity. This is of great significance for early identification of high-risk individuals, development of personalized prevention strategies, and reduction of the incidence of arteriosclerosis and its complications. Summary of the Invention
[0005] The present invention aims to at least partially solve one of the technical problems in the related art.
[0006] Therefore, an embodiment of the first aspect of the present invention proposes a method for constructing an arteriosclerosis risk assessment model, comprising: performing association analysis based on population genetic data and background data to obtain association significance values corresponding to SNP loci, wherein the genetic data includes the SNP loci and the genotype of the SNP loci; constructing the arteriosclerosis risk assessment model based on the SNP loci and the association significance values corresponding to the SNP loci, wherein the arteriosclerosis risk assessment model includes one or more SNP loci as shown in Table 1. The arteriosclerosis risk assessment model obtained by the construction method according to the embodiment of the present invention considers the genetic background, lifestyle, environmental factors, etc., unique to the Chinese population, and the AUC of the model obtained by it reaches 0.693, which can accurately and effectively assess the arteriosclerosis risk of the individual to be analyzed.
[0007] In some embodiments, the association analysis based on population genetic data and background data to obtain the association significance value corresponding to the SNP locus includes: constructing a linear mixture model based on population genetic data and background data, wherein the genetic data includes the SNP locus and the genotype of the SNP locus; and performing a genome-wide association analysis based on the linear mixture model to obtain the association significance value corresponding to the SNP locus.
[0008] In some embodiments, the genetic data further includes a population genetic relationship matrix; and the background data includes at least one selected from the group consisting of brachial-ankle pulse wave velocity and arteriosclerosis index, and optionally at least one selected from the group consisting of sex, age, systolic blood pressure, diastolic blood pressure, total cholesterol, low-density lipoprotein, and population structure clustering.
[0009] In some embodiments, the formula for the linear mixture model is:
[0010] Y = β i *X i +β j *Q j +Kinship
[0011] The dependent variable Y in the linear mixed model is at least one of the following groups: brachial-ankle pulse wave velocity and arteriosclerosis index; the independent variable X... i Q represents the genotype of the SNP locus. j Kinship is a population genetic relationship matrix, which selects at least one of the following groups: sex, age, systolic blood pressure, diastolic blood pressure, total cholesterol, low-density lipoprotein, and population structure clustering.
[0012] In some embodiments, the dependent variable Y is the brachial-ankle pulse wave conduction velocity.
[0013] In some embodiments, the group is a population of 1037 Chinese individuals.
[0014] In some embodiments, constructing the arteriosclerosis risk assessment model based on the SNP site and the association significance value corresponding to the SNP site includes: constructing multiple general linear models based on the SNP site and the association significance value corresponding to the SNP site; and evaluating the multiple general linear models to obtain the arteriosclerosis risk assessment model.
[0015] In some embodiments, constructing multiple general linear models based on the SNP sites and the association significance values corresponding to the SNP sites includes: using 1×10 -4 5×10 -5 2×10 -5 1×10 -5 1×10 -6 1×10 -7 5×10 -8 and 1×10 -8 Using each of these threshold values as a significance threshold, multiple groups of SNP sites with significance values lower than the threshold values are included in the plurality of general linear models, wherein the formula of the general linear model is: N is the total number of significant SNPs selected based on the association significance threshold; i is the i-th SNP; j is the j-th individual; S i G represents the effect size of the i-th SNP; ij M represents the number of SNPs carrying the i-th effect carried by the j-th individual; j The total number of effect SNPs carried by the j-th individual.
[0016] In some embodiments, evaluating the plurality of general linear models to obtain the arteriosclerosis risk assessment model includes: obtaining model evaluation indicators based on the plurality of general linear models in a validation set; and obtaining the arteriosclerosis risk assessment model based on the model evaluation indicators.
[0017] In some embodiments, the model evaluation metrics include rho value, AUC, and R0. 2 Adjusted R 2 And one or more of the F1 scores.
[0018] An embodiment of the second aspect of the present invention provides a method for assessing the risk of arteriosclerosis, comprising: obtaining genetic data from a sample; inputting the genetic data into an arteriosclerosis risk assessment model obtained according to the method of any embodiment of the first aspect to obtain an arteriosclerosis risk assessment score; and indicating that the sample has an arteriosclerosis risk based on the arteriosclerosis risk assessment score being higher than 2.19727.
[0019] A third aspect of the present invention provides a biomarker for arteriosclerosis, comprising one or more genes selected from the group consisting of CD163, SDC2, UBE2U, CD163L1, BMPR2, ZNF514, VPS54, SEC63, ECI1, RNPS1, SRCIN1, NGEF, ZNF708, SPATA5L1, DNMT3B, GALNT8, SLC30A4, MRPS5, BRICD5, SRP9, CCL24, CADPS, POR, KCNA6, CACNG8, PPM1F, BLOC1S6, WDR61, C4orf19, and IRX6.
[0020] In some embodiments, the biomarkers include one or more SNP sites as shown in Table 1, wherein the SNP sites are located in the one or more genes.
[0021] Embodiments of the fourth aspect of the present invention provide for the use of reagents for detecting biomarkers as described in any embodiment of the third aspect in the preparation of kits for predicting, preventing and / or treating arteriosclerosis.
[0022] A fifth aspect of the present invention provides a computer program for assessing the risk of atherosclerosis in an individual to be analyzed, comprising a list of instructions, which, when executed on an electronic computer, are provided with genetic data obtained from a sample of the individual to be analyzed to perform the following steps: inputting the genetic data into an atherosclerosis risk assessment model obtained according to the method of any embodiment of the first aspect to obtain an atherosclerosis risk assessment score; comparing the atherosclerosis risk assessment score with a first value; and indicating that the individual has a risk of atherosclerosis when the atherosclerosis risk assessment score is greater than the first value, wherein the first value is 2.19727.
[0023] The embodiments of the present invention achieve the following beneficial effects:
[0024] The method for constructing the arteriosclerosis risk assessment model provided in this invention takes into account the unique genetic background, lifestyle, and environmental factors of the Chinese population. The model obtained achieves an AUC of 0.693, enabling it to accurately and effectively assess the arteriosclerosis risk of the individual being analyzed. This is of great significance for early identification of high-risk individuals, development of personalized prevention strategies, and reduction of the incidence of arteriosclerosis and its complications. Attached Figure Description
[0025] Figure 1 The Manhattan plot obtained through association analysis in the construction method of the arteriosclerosis risk assessment model in this embodiment of the invention shows the SNP sites and their corresponding association significance values.
[0026] Figure 2 The following diagram illustrates the evaluation results of different models obtained by constructing the arteriosclerosis risk assessment model according to an embodiment of the present invention.
[0027] Figure 3 The evaluation results of the optimal model obtained by constructing the arteriosclerosis risk assessment model according to the embodiment of the present invention are shown.
[0028] Figure 4 The evaluation results of the optimal model obtained by the method of constructing the arteriosclerosis risk assessment model according to the embodiments of the present invention are shown on the validation set. Detailed Implementation
[0029] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.
[0030] This invention is based on the inventor's discoveries and understanding of the following facts and problems:
[0031] Arteriosclerosis, especially atherosclerosis, is a serious vascular disease and one of the leading causes of premature death worldwide. Arteriosclerosis not only affects the structure and function of the arterial walls, leading to thickening, hardening, loss of elasticity, and narrowing of the lumen, but it is also often accompanied by the accumulation of lipids and complex carbohydrates in the arterial intima, forming plaques, which in turn trigger a series of cardiovascular events.
[0032] Currently, various indicators can be used to detect and assess arteriosclerosis, such as baPWV and ASI. The advantage of baPWV lies in its relatively simple measurement method, requiring only a blood pressure cuff to be applied to the limbs, making it particularly suitable for large-scale population-based epidemiological studies. However, there are currently few PRS studies on arteriosclerosis genes. Related technologies mainly use ASI for genome-wide association analysis, but no predictive models based on it for arteriosclerosis risk have been provided. Furthermore, related risk assessment models struggle to accurately reflect specific situations, resulting in low accuracy in predicting arteriosclerosis risk. And specifically for the Chinese population, a model capable of accurately predicting arteriosclerosis risk is still lacking.
[0033] In this paper, the term "poly-genetic risk score" (PRS) refers to a score based on variations at multiple genetic loci and their associated weights. The PRS is constructed based on the effect size of each risk allele or effect allele and typically follows this form: Individual PRS is equal to the individual marker genotype SNP in n genetic variants or nucleotide polymorphisms.i The weighted sum is estimated using regression analysis. The PRS (Prognostic Risk Score) sums the numerical values of the relationship between multiple genetic variations and phenotypes, and is a weighted linear combination of alleles on the genome associated with disease phenotypes. For complex traits, many genetic loci typically have a small impact on the phenotype; in such cases, a single variation is insufficient to assess an individual's risk for a particular complex trait. Therefore, a polygenic risk score can utilize multiple genetic variations to comprehensively assess an individual's disease risk. It is understood that "PRS" in this application can also specifically refer to the "arteriosclerosis risk assessment score" in this application.
[0034] In this paper, the term "polymorphism" refers to genetic polymorphism, which is used to describe the diversity of the genome of a species (such as humans), and is essentially the inter-individual variation in a DNA sequence unique to an individual. In other words, genetic polymorphism is the occurrence of multiple discrete allelic states within the same population. Polymorphism involves one of two or more variants of a particular DNA sequence. The most common type of polymorphism involves variation in a single nucleotide, known as single nucleotide polymorphism (SNP).
[0035] As used herein, the term "variant" or "genetic variant" refers to a specific region of the genome that differs from a reference genome. Depending on the type of alteration, the term "genetic variant" can refer to (but is not limited to) a single nucleotide variant (SNV) or an SNP. As used herein, the term "SNV" or "SNP" refers to a variant that has a single nucleotide substitution in its DNA sequence. Traditionally, an SNP is an SNV present in a population at some perceptible level (e.g., more than 1% of the population).
[0036] Based on the aforementioned practical needs, an embodiment of the first aspect of the present invention proposes a method for constructing an arteriosclerosis risk assessment model, comprising: performing association analysis based on population genetic data and background data to obtain association significance values corresponding to SNP loci, wherein the genetic data includes the SNP loci and the genotype of the SNP loci; and constructing the arteriosclerosis risk assessment model based on the SNP loci and the association significance values corresponding to the SNP loci, wherein the arteriosclerosis risk assessment model includes one or more SNP loci as shown in Table 1. The arteriosclerosis risk assessment model obtained by the construction method according to the embodiment of the present invention considers the genetic background, lifestyle, and environmental factors unique to the Chinese population, and the AUC of the model obtained reaches 0.693, which can accurately and effectively assess the arteriosclerosis risk of the individual to be analyzed.
[0037] Table 1
[0038]
[0039]
[0040]
[0041]
[0042]
[0043]
[0044]
[0045]
[0046]
[0047]
[0048]
[0049]
[0050]
[0051]
[0052] *Table 1 lists the reference genotype and variant genotype in the population in columns A1-Reference Genotype and A2-Variant Genotype, respectively. Specific variants include common single base variations (e.g., T→G), base insertions (e.g., C→CTT), and base loss (e.g., CAGAG→C).
[0053] *In the rs number column of Table 1, for SNPs that do not yet have rs numbers, a combination of chromosome and location (e.g., chr17:31826378) is used for naming.
[0054] In some embodiments, the association analysis based on population genetic data and background data to obtain the association significance value corresponding to the SNP locus includes: constructing a linear mixture model based on population genetic data and background data, wherein the genetic data includes the SNP locus and the genotype of the SNP locus; and performing a genome-wide association analysis based on the linear mixture model to obtain the association significance value corresponding to the SNP locus.
[0055] What is understandable is... Figure 1As shown in the figure, the Manhattan plot obtained through association analysis in the method for constructing the arteriosclerosis risk assessment model of this embodiment of the invention illustrates the SNP sites and their corresponding association significance values. The association significance value corresponding to the SNP site is the P-value, representing the correlation between the SNP and arteriosclerosis.
[0056] In some embodiments, the genetic data further includes a population genetic relationship matrix; and the background data includes at least one selected from the group consisting of brachial-ankle pulse wave velocity and arteriosclerosis index, and optionally at least one selected from the group consisting of sex, age, systolic blood pressure, diastolic blood pressure, total cholesterol, low-density lipoprotein, and population structure clustering.
[0057] In some embodiments, the formula for the linear mixture model is:
[0058] Y = β i *X i +β j *Q j +Kinship
[0059] The dependent variable Y in the linear mixed model is at least one of the following groups: brachial-ankle pulse wave velocity and arteriosclerosis index; the independent variable X... i Q represents the genotype of the SNP locus. j Kinship is a population genetic relationship matrix, which selects at least one of the following groups: sex, age, systolic blood pressure, diastolic blood pressure, total cholesterol, low-density lipoprotein, and population structure clustering.
[0060] In some embodiments, the dependent variable Y is the brachial-ankle pulse wave conduction velocity.
[0061] In some embodiments, the group is a group of 1037 Chinese individuals.
[0062] In some embodiments, constructing the arteriosclerosis risk assessment model based on the SNP site and the association significance value corresponding to the SNP site includes: constructing multiple general linear models based on the SNP site and the association significance value corresponding to the SNP site; and evaluating the multiple general linear models to obtain the arteriosclerosis risk assessment model.
[0063] In some embodiments, constructing multiple general linear models based on the SNP sites and the association significance values corresponding to the SNP sites includes: using 1×10 -4 5×10 -5 2×10 -5 1×10 -5 1×10 -6 1×10 -75×10 -8 and 1×10 -8 Using each of these threshold values as a significance threshold, multiple groups of SNP sites with significance values lower than the threshold values are included in the plurality of general linear models, wherein the formula of the general linear model is: The PRS j S is the PRS value of the j-th individual; N is the total number of significant SNPs selected according to the association significance threshold, which is 607 in this invention; i is the i-th SNP; j is the j-th individual; S i G represents the effect size of the i-th SNP; ij M represents the number of SNPs carrying the i-th effect from the j-th individual; j This represents the total number of effector SNPs carried by the j-th individual. It should be noted that this effect value is the effect of significant SNPs selected based on the association significance threshold on the brachial-ankle pulse wave velocity. In some embodiments, the effect values are shown in Table 1. Gij refers to the number of i-th effector SNPs carried by the j-th individual, where the effector SNP indicates whether the genotype of the i-th SNP is an A2 variant genotype. Based on the DNA double helix, the number of A2 variant genotypes can be 0, 1, or 2. Furthermore, it is understood that due to genetic differences among individuals and sequencing quality, the above SNPs may be missing values (NA), therefore M... j This represents the total number of non-missing SNPs actually carried by the j-th individual. Specifically, it is represented by 1 × 10⁻⁶. -4 5×10 -5 2×10 -5 1×10 -5 5×10 -6 1×10 -6 5×10 -7 1×10 -7 and 5×10 -8 Using each SNP site as a threshold for association significance, multiple groups of SNP sites with association significance values lower than the threshold are included in the multiple general linear models, and their corresponding model evaluation results are as follows. Figure 2 As shown. With threshold P from 5 × 10... -8 Up to 1×10 -4 As the threshold gradually increases, when it equals 2 × 10 -5At this threshold, the model's AUC, F1 Score, rho, and corrected R2 are all at high levels, and compared to other thresholds with high AUC, F1 Score, rho, and corrected R2, the number of SNPs included is the fewest. This threshold is the optimal threshold for balancing model accuracy and the number of included SNPs, and the corresponding model is the optimal model. Specific parameters are shown in Table 1.
[0064] In some embodiments, evaluating the plurality of general linear models to obtain the arteriosclerosis risk assessment model includes: obtaining model evaluation indicators based on the plurality of general linear models in a validation set; and obtaining the arteriosclerosis risk assessment model based on the model evaluation indicators.
[0065] In some embodiments, the model evaluation metrics include rho value, AUC, and R0. 2 Adjusted R 2 And one or more of the F1 scores.
[0066] like Figure 3 As shown, the evaluation results of the optimal model obtained by the construction method of the arteriosclerosis risk assessment model in this embodiment of the invention show the specificity-sensitivity curve of the optimal model of the present invention, with an area under the curve of 0.693, which indicates that it has excellent accuracy in assessing arteriosclerosis risk.
[0067] An embodiment of the second aspect of the present invention provides a method for assessing the risk of arteriosclerosis, comprising: obtaining genetic data from a sample; inputting the genetic data into an arteriosclerosis risk assessment model obtained according to the method of any embodiment of the first aspect to obtain an arteriosclerosis risk assessment score; and indicating that the sample has an arteriosclerosis risk based on the arteriosclerosis risk assessment score being higher than 2.19727.
[0068] A third aspect of the present invention provides a biomarker for arteriosclerosis, comprising one or more genes selected from the group consisting of CD163, SDC2, UBE2U, CD163L1, BMPR2, ZNF514, VPS54, SEC63, ECI1, RNPS1, SRCIN1, NGEF, ZNF708, SPATA5L1, DNMT3B, GALNT8, SLC30A4, MRPS5, BRICD5, SRP9, CCL24, CADPS, POR, KCNA6, CACNG8, PPM1F, BLOC1S6, WDR61, C4orf19, and IRX6.
[0069] In some embodiments, the biomarkers include one or more SNP sites as shown in Table 1, wherein the SNP sites are located in the one or more genes. It is understood that the inventors, through comparative statistical analysis of the location information of the SNP sites, found that all 607 SNP sites are enriched in the aforementioned 30 genes. This indicates that the aforementioned 30 genes are significantly associated with atherosclerosis and can serve as potential biomarkers for atherosclerosis-related research and applications.
[0070] Embodiments of the fourth aspect of the present invention provide for the use of reagents for detecting biomarkers as described in any embodiment of the third aspect in the preparation of kits for predicting, preventing and / or treating arteriosclerosis.
[0071] A fifth aspect of the present invention provides a computer program for assessing the risk of atherosclerosis in an individual to be analyzed, comprising a list of instructions, which, when executed on an electronic computer, are provided with genetic data obtained from a sample of the individual to be analyzed to perform the following steps: inputting the genetic data into an atherosclerosis risk assessment model obtained according to the method of any embodiment of the first aspect to obtain an atherosclerosis risk assessment score; comparing the atherosclerosis risk assessment score with a first value; and indicating that the individual has a risk of atherosclerosis when the atherosclerosis risk assessment score is greater than the first value, wherein the first value is 2.19727.
[0072] Unless otherwise specified, the experimental methods in the following embodiments are conventional methods, performed in accordance with the techniques or conditions described in the literature in this field or in accordance with the product instructions.
[0073] The acquisition, storage, and application of user personal information involved in the technical solutions of this invention comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0074] Example
[0075] This invention, based on genetic data from the Chinese population, identified gene loci related to the arteriosclerosis marker baPWV and constructed an arteriosclerosis risk assessment model. The sample consisted of 1037 participants from Central China, with a male-to-female ratio of 5:2. Participants were divided into an arteriosclerosis group and a normal group based on PWV ≥ 1400 cm / s, with a ratio of approximately 5.5:4.5. This invention was approved by the BGI Bioethics Review Committee (No.: BGI-IRB 20144-T4). All eligible participants submitted informed consent forms before joining this study.
[0076] 1.1 Model Construction
[0077] (1) The raw whole genome data (30x) obtained from whole genome sequencing was filtered, aligned, sorted, deduplicated, locally realigned, base quality recalibrated, and variant detected to obtain individual base variation information (MegaBOLT BQSRv2.3.7WGS integrated platform --type full --runtype WGS). The individual base variation information was merged using the GatherVcfs function of GATK version 4.1.8.1, and the VariantFiltration and VariantRecalibrator functions of GATK version 4.1.8.1 and the ApplyVQSR function were used for variant quality control filtering to obtain genomic variation information.
[0078] (2) Using the parameters --missing, --check-sex, --freq, --maf, --hardy, --het, --remove, and --extract in Plink 1.9, the obtained variation information was sequentially processed through SNP missing value filtering, individual missing value filtering, sex difference test, minor allele test, Hardy-Weinberg balance test, heterozygosity test, population structure and kinship test to complete GWAS data quality control. Principal component analysis was then performed on the quality-controlled data, and the portion that could explain the first 80%-90% of the sample was selected as covariates for association analysis.
[0079] (3) Using SNP variation information as independent variables and the arteriosclerosis index baPWV as the dependent variable, with sex, age, and principal components as covariates for adjustment, and the genetic relationship matrix as a fixed effect, a mixed linear model was constructed using GCTA fastGWA for association analysis to obtain SNPs related to the arteriosclerosis index baPWV. Furthermore, LD linkage disequilibrium analysis was used to further screen for SNPs significantly associated with arteriosclerosis.
[0080] (4) Use plink to construct a general linear model of baPWV-related SNPs.
[0081] 5×10 -8 1×10 -8 1×10 -7 1×10 -6 5×10 -5 2×10 -5 1×10 -5 1×10 -4 To construct multiple PRS models, the threshold for P is included in SNPs. The specific formula is as follows:
[0082]
[0083] Among them, PRSj is the PRS value of the j-th individual; N is the total number of SNPs included in the PRS model, which is 607 in this invention; i is the i-th SNP; j is the j-th individual; S i G represents the effect weight of the i-th SNP; ij M represents the number of SNPs carrying the i-th effect in the j-th individual; j The total number of non-deletion effect SNPs carried by the j-th individual.
[0084] The results of the association analysis are as follows Figure 1 As shown. Figure 1 The Manhattan plot obtained through association analysis in the method for constructing the arteriosclerosis risk assessment model in this embodiment of the invention shows SNP sites and their corresponding association significance values. The association significance value corresponding to the SNP site is the P-value, representing the correlation between the SNP and arteriosclerosis.
[0085] Specifically, 20 P < 5 × 10 were obtained. -8 Significantly independent sites, 397 with P < 1 × 10⁻⁶. -5 607 suggestive sites, 2 × 10 -5 The risk SNPs were identified, and multiple PRS models were constructed based on them.
[0086] 1.2 Model Evaluation
[0087] For the multiple PRS models obtained in Example 1.1, the rho value, AUC, and adjusted R are used. 2 The F1 score is used as a model evaluation indicator for comprehensive evaluation.
[0088] Model evaluation results for multiple general linear models are as follows: Figure 2 As shown. With threshold P from 5 × 10... -8 Up to 1×10 -4 As the threshold gradually increases, when it equals 2 × 10 -5 At this point, the model's AUC, F1 score, rho, and adjusted R2 were all at high levels, and compared to other thresholds where AUC, F1 score, rho, and adjusted R2 were at high levels, the number of SNPs included was the fewest. This threshold is the relatively optimal threshold for balancing model accuracy and the number of SNPs included. Ultimately, P = 2 × 10⁻⁶ was used as the threshold. -5 The optimal PRS model was obtained by setting the threshold, and the included SNP sites and their corresponding effect values are shown in Table 1.
[0089] The evaluation results of the optimal model obtained by the method for constructing the arteriosclerosis risk assessment model according to the embodiments of the present invention are as follows: Figure 3As shown, it illustrates the specificity-sensitivity curve of the optimal model of the present invention, with an area under the curve of 0.693, representing its excellent accuracy in assessing the risk of arteriosclerosis.
[0090] In this optimal model, after calculating the PRS score by weighting the SNPs, a PRS score higher than 2.19727 indicates a risk of arteriosclerosis.
[0091] Verification Example
[0092] The accuracy of the present invention was verified using an optimal model in a population of 916 cases as validation cases.
[0093] The evaluation results of the optimal model obtained by the method for constructing the arteriosclerosis risk assessment model according to the embodiments of the present invention on the validation set are as follows: Figure 4 As shown, the specificity-sensitivity curve of the model in the validation example is displayed. The area under the curve is 0.689, which is not significantly different from the area under the curve of 0.693 in the training set, indicating that it has stable accuracy and applicability in assessing the risk of arteriosclerosis.
[0094] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0095] In this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," "linking," and "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection, an electrical connection, or a connection that allows communication between them; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components, unless otherwise explicitly limited. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0096] In this invention, the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to a specific feature, structure, material, or characteristic described in connection with that embodiment or example, which is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0097] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A method of constructing an arteriosclerosis risk assessment model, characterized by, The method comprises: performing association analysis on population-based genetic data and background data to obtain association significance values corresponding to SNP sites, wherein the genetic data comprises the SNP sites and genotypes of the SNP sites; constructing the arteriosclerosis risk assessment model based on the SNP sites and the association significance values corresponding to the SNP sites, wherein the arteriosclerosis risk assessment model comprises one or more of the SNP sites shown in Table 1.
2. The construction method according to claim 1, characterized in that, The association analysis on population-based genetic data and background data to obtain association significance values corresponding to SNP sites comprises: constructing a linear mixed model based on population-based genetic data and background data, wherein the genetic data comprises SNP sites and genotypes of the SNP sites; performing whole genome association analysis based on the linear mixed model to obtain association significance values corresponding to the SNP sites.
3. The construction method of claim 2, wherein, wherein the genetic data further comprises a population genetic relationship matrix, and wherein the background data comprises at least one selected from the group consisting of brachial-ankle pulse wave velocity, arteriosclerosis indicators, and optionally at least one selected from the group consisting of gender, age, systolic pressure, diastolic pressure, total cholesterol, low-density lipoprotein, and population structure clustering.
4. The construction method according to claim 3, characterized in that, wherein the formula of the linear mixed model is: Y = β i X i + β j Q j + Kinship wherein the dependent variable Y of the linear mixed model is at least one of the group consisting of brachial ankle pulse wave conduction velocity, arteriosclerosis index, the independent variable X i is the genotype of the SNP site, Q j is at least one selected from the group consisting of gender, age, systolic pressure, diastolic pressure, total cholesterol, low density lipoprotein, population structure cluster, Kinship is a population genetic relationship matrix, Preferably, the dependent variable Y is brachial-ankle pulse wave velocity.
5. The construction method of claim 3, wherein, The construction of the arteriosclerosis risk assessment model based on the SNP sites and the association significance values corresponding to the SNP sites comprises: constructing a plurality of general linear models based on the SNP sites and the association significance values corresponding to the SNP sites; evaluating the general linear models to obtain the arteriosclerosis risk assessment model.
6. The construction method of claim 5, wherein, The construction of a plurality of general linear models based on the SNP sites and the association significance values corresponding to the SNP sites comprises: 1 x 10 -4 , 5 x 10 -5 , 2 x 10 -5 , 1 x 10 -5 , 1 x 10 -6 , 1 x 10 -7 , 5 x 10 -8 , and 1 x 10 -8 , respectively, as the association significance value threshold, and a plurality of groups of SNP sites corresponding to the SNP sites with association significance values lower than the threshold are respectively included in the plurality of general linear models, wherein the formula of the general linear model is wherein the PRS j is the PRS value of the jth individual; N is the total number of significant SNPs filtered according to the association significance value threshold; i is the ith SNP; j is the jth individual; S i is the effect value of the ith SNP; G ij is the number of the ith effect SNP carried by the jth individual; M j is the total number of effect SNPs carried by the jth individual.
7. The construction method according to claim 6, characterized in that, The evaluation of the plurality of general linear models to obtain the arteriosclerosis risk assessment model comprises: obtaining model evaluation indicators based on the plurality of general linear models in a validation set; obtaining the arteriosclerosis risk assessment model based on the model evaluation indicators, Preferably, the model evaluation metrics include one or more of rho value, AUC, R 2 , adjusted R 2 , and F1 score.
8. A method of assessing the risk of arteriosclerosis, characterized by, The method comprises: obtaining genetic data from a sample; inputting the genetic data into the arteriosclerosis risk assessment model obtained according to the method of any one of claims 1 to 7 to obtain an arteriosclerosis risk assessment score; and prompting that the sample has a risk of arteriosclerosis based on that the arteriosclerosis risk assessment score is higher than 2.19727.
9. An arterial stiffness biomarker, characterized by, comprises one or more genes selected from the group consisting of CD163, SDC2, UBE2U, CD163L1, BMPR2, ZNF514, VPS54, SEC63, ECI1, RNPS1, SRCIN1, NGEF, ZNF708, SPATA5L1, DNMT3B, GALNT8, SLC30A4, MRPS5, BRICD5, SRP9, CCL24, CADPS, POR, KCNA6, CACNG8, PPM1F, BLOC1S6, WDR61, C4orf19, IRX6.
10. The biomarker of claim 9, wherein, The biomarker comprises one or more of the SNP loci as shown in Table 1, wherein the SNP loci are located in the one or more genes.
11. Use of a reagent for detecting the biomarker according to claim 9 or 10 in the manufacture of a kit for predicting, preventing and / or treating arteriosclerosis.
12. A computer program for assessing the risk of arteriosclerosis in an individual to be analyzed, comprising a list of instructions, when executed on an electronic computer, provided with genetic data obtained in a sample of the individual to be analyzed, to perform the following steps: inputting the genetic data into an arteriosclerosis risk assessment model obtained according to the method of any one of claims 1 to 7, to obtain an arteriosclerosis risk assessment score; comparing the arteriosclerosis risk assessment score to a first value; and when the arteriosclerosis risk assessment score is greater than the first value, prompting that the individual is at risk of arteriosclerosis, wherein the first value is 2.19727.