Risk prediction model for ischemic stroke and intracerebral hemorrhage based on novel lipidomic predictors and applications thereof
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- FUDAN UNIVERSITY
- Filing Date
- 2026-04-21
- Publication Date
- 2026-08-07
AI Technical Summary
[0005](1)预测效能有限,难以支撑精准一级预防:现有模型普遍存在预测精度不足的问题,其区分度(如AUC值)多处于中等水平
[0056]本发明涉及基于新型脂质组学预测因子的缺血性卒中和脑出血风险预测模型及其应用。所述模型具有预测效能高、特异性强的特点,能够更有效地识别缺血性卒中与脑出血的高危人群,为临床早期干预提供可靠依据,在卒中精准预防、临床风险评估及疾病负担控制等方面具有良好的应用前景。
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention relates to the field of biotechnology, specifically to a risk prediction model for ischemic stroke and cerebral hemorrhage based on novel lipidomics predictive factors and its application. Background Technology
[0002] Existing research has established various predictive models for stroke risk, addressing the numerous influencing factors. A series of studies based on European and American populations have constructed stroke risk prediction models based on the modified Framingham Stroke Risk Score published in *Circulation*, using factors such as age, sex, systolic blood pressure, use of antihypertensive medications, diabetes, smoking, and prior cardiovascular disease to predict stroke risk. Correspondingly, several studies have established stroke risk prediction models for the Chinese population. For example, the stroke prediction model for the Chinese population published in *Stroke* incorporates variables such as age, treated and untreated systolic blood pressure, current smoking status, diabetes, total cholesterol, high-density lipoprotein cholesterol, and waist circumference.
[0004] However, despite the development of various stroke risk prediction models in existing population epidemiological studies, these publicly available models still have significant limitations in clinical application and precision prevention, specifically:
[0005] (1) Limited predictive efficacy, making it difficult to support precise primary prevention: Existing models generally suffer from insufficient predictive accuracy, and their discrimination (such as AUC value) is mostly at a medium level. This results in limited efficacy of the models in identifying high-risk groups and risk stratification, making it difficult to provide sufficient reliable decision-making basis for clinical practice, and thus making it difficult to achieve precise prevention resource allocation and early intervention for specific high-risk individuals.
[0006] (2) Limited predictive targets and lack of specific risk assessment for cerebral hemorrhage: Although ischemic stroke and cerebral hemorrhage are the two most common types of stroke, most existing models are constructed only for the overall risk of stroke or the risk of ischemic stroke, while systematically neglecting the prediction of cerebral hemorrhage. Given that the proportion of stroke patients in my country is higher than the global average, and that its disability and mortality rates are extremely high, developing predictive tools that can specifically identify high-risk individuals for cerebral hemorrhage is of great significance for reducing the overall burden of stroke in my country.
[0007] (3) Limited predictive factors and insufficient integration of emerging biomarkers: The predictive factors relied upon by existing models are mostly limited to traditional clinical and biochemical indicators. For example, at the lipid level, only traditional predictive factors such as total cholesterol, high-density lipoprotein cholesterol, and low-density lipoprotein cholesterol are included, which are insufficient in characterizing the pathophysiological changes of the disease and cannot fully capture the risk of stroke. It is particularly noteworthy that the above-mentioned traditional lipid detection measures the concentration of lipid "total" or "major categories", but these lipid categories are actually composed of a variety of specific lipid molecules; while emerging detection technologies such as lipidomics can accurately analyze the specific structural information of lipid molecules - such as carbon chain length, number of double bonds, etc. - thereby revealing lipid molecule categories with stronger predictive value that are masked by traditional indicators. However, existing risk assessment models have not systematically evaluated and integrated these novel lipid biomarkers with structural features, resulting in the omission of a large amount of key pathophysiological information, which seriously limits the further improvement of the model's predictive performance and its ability to accurately stratify risk.
[0008] Emerging lipidomics technologies offer significant advantages in constructing stroke risk prediction models. Lipid abnormalities are a core driver of stroke pathophysiology and one of the most easily intervened and clinically valuable variable risk factors. However, traditional clinical lipid detection methods can only roughly measure the levels of fewer than ten lipid subclasses, limiting their ability to capture the rich diversity of lipid subclasses and their complex structures, all of which significantly impact stroke risk. Conversely, lipidomics technologies can perform high-throughput analysis of thousands of different lipid subclasses, obtaining their specific molecular structures and systematically evaluating their predictive ability for stroke risk, thus constructing more efficient stroke risk prediction models.
[0009] In summary, the prevention and treatment of stroke remains a serious challenge, and existing risk prediction models fall short in terms of predictive efficacy, target specificity, and biomarker advancement. Therefore, there is an urgent need in this field for a risk prediction model that integrates novel lipidomics predictive factors, capable of conducting specific risk assessments for ischemic stroke and cerebral hemorrhage respectively. This would allow for more accurate identification of high-risk populations for different stroke types, providing a more advanced and reliable technological means for implementing precise and effective early interventions. Summary of the Invention
[0010] In view of this, the technical problem to be solved by the present invention is to provide a risk prediction model that integrates novel lipidomics predictive factors. This model has the characteristics of high predictive efficacy and strong specificity, and can more effectively identify high-risk groups for ischemic stroke and cerebral hemorrhage, providing a reliable basis for early clinical intervention and possessing significant clinical and social value.
[0011] A first aspect of the present invention provides a diagnostic factor composition for cerebrovascular diseases, comprising at least one of sphingosine-1-phosphate, butyrylcarnitine, ethanolamine acetal phospholipid, phosphatidylethanolamine, phosphatidylinositol, sphingomyelin, 13-hydroxy-9Z,11E-octadecadienoic acid, 9,10-dihydroxy-12Z-octadecadienoic acid, eicostrienoic carnitine, triacylglycerol, diacylglycerol, phosphatidylcholine, or lysophosphatidylcholine.
[0012] The composition can more effectively identify high-risk individuals for ischemic stroke and cerebral hemorrhage. Compared to traditional diagnostic factors, this composition exhibits significant advantages in sensitivity, specificity, and predictive stability, which helps to achieve early risk warning and provides a reliable molecular basis for clinical stratification management and personalized intervention.
[0013] In some embodiments, the cerebrovascular-related diseases include ischemic stroke and cerebral hemorrhage.
[0014] In some embodiments, the sphingosine-1-phosphate comprises sphingosine-1-phosphate (d18:2), the ethanolamine acetal phospholipid comprises ethanolamine acetal phospholipid (P-18:0 / 22:5), the phosphatidylethanolamine comprises phosphatidylethanolamine (18:2 / 18:2), the phosphatidylinositol comprises phosphatidylinositol (18:2-18:2), the sphingomyelin comprises sphingomyelin (d44:4), the triacylglycerol comprises triacylglycerol (58:3)-fatty acyl (22:1), the diacylglycerol comprises diacylglycerol (20:4 / 18:2), the phosphatidylcholine comprises phosphatidylcholine (18:2 / 18:2), and the lysophosphatidylcholine comprises lysophosphatidylcholine (24:0)-sn2.
[0015] In some specific embodiments, the ischemic stroke diagnostic factor composition includes sphingosine-1-phosphate (d18:2), butyryl carnitine, ethanolamine acetal phospholipid (P-18:0 / 22:5), phosphatidylethanolamine (18:2 / 18:2), phosphatidylinositol (18:2-18:2), sphingomyelin (d44:4), 13-hydroxy-9Z,11E-octadecadienoic acid, and 9,10-dihydroxy-12Z-octadecadienoic acid.
[0016] In other specific embodiments, the diagnostic factor composition for cerebral hemorrhage includes eicosatrienoic carnitine, phosphatidylinositol (18:2-18:2), triacylglycerol (58:3)-fatty acyl (22:1), diacylglycerol (20:4 / 18:2), phosphatidylcholine (18:2 / 18:2), lysophosphatidylcholine (24:0)-sn2, 13-hydroxy-9Z,11E-octadecadienoic acid, and 9,10-dihydroxy-12Z-octadecadienoic acid.
[0017] A second aspect of the invention provides a model for predicting the risk of ischemic stroke, the model being:
[0018] Ischemic stroke lipid risk score = α + β1*Sphingosine-1-phosphate (d18:2) + β2*Butyryl-carnitine + β3*PE (P-18:0 / 22:5) + β4*PE (18:2 / 18:2) + β5*PI (18:2_18:2) + β6*SM (d44:4) + β7*13-HODE + β8*9,10-diHOME;
[0019] Wherein, Sphingosine-1-phosphate (d18:2) is sphingosine-1-phosphate (d18:2), Butyryl-carnitine is butyrylcarnitine, PE (P-18:0 / 22:5) is ethanolamine acetal phospholipid (P-18:0 / 22:5), PE (18:2 / 18:2) is phosphatidylethanolamine (18:2 / 18:2), PI (18:2_18:2) is phosphatidylinositol (18:2_18:2), SM (d44:4) is sphingomyelin (d44:4), 13-HODE is 13-hydroxy-9Z,11E-octadecadienoic acid, and 9,10-diHOME is 9,10-dihydroxy-12Z-octadecadienoic acid;
[0020] α, β1, β2, β3, β4, β5, β6, β7, and β8 are model coefficients, which are obtained through logistic regression.
[0021] In some embodiments, α is -0.15 to -0.01, β1 is 0.5 to 1.5, β2 is -1.0 to -0.01, β3 is -1.0 to -0.01, β4 is -1.5 to -0.5, β5 is 0.8 to 1.8, β6 is -1.0 to -0.1, β7 is 0.1 to 1.5, and β8 is -1.2 to -0.1. For example, α is -0.15, -0.10, -0.08, -0.05, or -0.01; β1 is 0.5, 0.8, 1.0, 1.2, or 1.5; β2 is -1.0, -0.8, -0.5, -0.2, or -0.01; β3 is -1.0, -0.8, -0.5, -0.3, -0.1, or -0.01; and β4 is -1.5 or -1.2. β5 is 0.8, 1.0, 1.3, 1.5 or 1.8; β6 is -1.0, -0.8, -0.6, -0.4, -0.2 or -0.1; β7 is 0.1, 0.3, 0.5, 0.8, 1.0, 1.2 or 1.5; β8 is -1.2, -1.0, -0.8, -0.5, -0.3 or -0.1.
[0022] In some specific embodiments, α is -0.089, β1 is 0.944, β2 is -0.399, β3 is -0.359, β4 is -0.984, β5 is 1.311, β6 is -0.549, β7 is 0.883, and β8 is -0.69. The model coefficients are solved through logistic regression. Compared to other coefficient combinations, this gives the model superior predictive efficacy and high specificity, enabling more accurate and stable identification of high-risk individuals for ischemic stroke, and facilitating early risk stratification and targeted intervention.
[0023] In some embodiments, the prediction threshold for ischemic stroke risk is 0.33. When the ischemic stroke lipid risk score exceeds this threshold, the risk of ischemic stroke is considered high. Experimental verification shows that the AUC value of this model for predicting ischemic stroke risk is 0.90, the sensitivity is 0.87, and the specificity is 0.73. Under the same experimental conditions, the AUC of traditional predictive factors for predicting ischemic stroke risk is only 0.64, the sensitivity is 0.52, and the specificity is 0.71. The comparative results indicate that the prediction model provided by this invention has significantly better overall predictive performance than traditional predictive factors and can more efficiently identify high-risk individuals for ischemic stroke.
[0024] Based on the relative risk stratification results indicated by this risk score, patients at high risk of ischemic stroke can be given priority for intensive treatment and / or more intensive disease monitoring, thereby providing an objective basis for clinical decision-making.
[0025] A third aspect of the present invention provides a model for predicting the risk of cerebral hemorrhage, said model being:
[0026] Cerebral hemorrhage lipid risk score = α + β1*Eicosatrienoyl-carnitine + β2*PI (18:2_18:2) + β3*TAG (58:3)-FA22:1 + β4*DAG (20:4 / 18:2) + β5*PC (18:2 / 18:2) + β6*LPC(24:0)-sn2 + β7*13-HODE + β8*9,10-diHOME;
[0027] Wherein, Eicosatrienoyl-carnitine is eicosatrienoylcarnitine, PI (18:2_18:2) is phosphatidylinositol (18:2_18:2), TAG (58:3)-FA (22:1) is triacylglycerol (58:3)-fatty acyl (22:1), DAG (20:4 / 18:2) is diacylglycerol (20:4 / 18:2), PC (18:2 / 18:2) is phosphatidylcholine (18:2 / 18:2), LPC (24:0)-sn2 is lysophosphatidylcholine (24:0)-sn2, 13-HODE is 13-hydroxy-9Z,11E-octadecadienoic acid, and 9,10-diHOME is 9,10-dihydroxy-12Z-octadecadienoic acid;
[0028] α, β1, β2, β3, β4, β5, β6, β7, and β8 are model coefficients, which are obtained through logistic regression.
[0029] In some embodiments, α is -5 to -1, β1 is 0.1 to 1.0, β2 is 0.8 to 1.8, β3 is 0.1 to 1.0, β4 is -1.0 to -0.1, β5 is -1.0 to -0.1, β6 is -1.0 to -0.1, β7 is 0.1 to 1.0, and β8 is -1.0 to -0.1. For example, α is -5, -4, -3, -2, or -1; β1 is 0.1, 0.3, 0.5, 0.8, or 1.0; β2 is 0.8, 1.0, 1.3, 1.5, or 1.8; β3 is 0.1, 0.3, 0.5, 0.8, or 1.0; β4 is -1.0, -0.8, -0.6, -0.3, or -0.1; β5 is -1.0, -0.8, -0.5, -0.3, or -0.1; β6 is -1.0, -0.8, -0.5, -0.3, or -0.1; β7 is 0.1, 0.2, 0.5, 0.8, or 1.0; and β8 is -1.0, -0.8, -0.5, -0.3, or -0.1.
[0030] In some specific embodiments, α is -2.009, β1 is 0.581, β2 is 1.363, β3 is 0.433, β4 is -0.529, β5 is -0.705, β6 is -0.394, β7 is 0.681, and β8 is -0.631. The model coefficients are determined through logistic regression, which, compared to other coefficient combinations, significantly improves the model's predictive efficacy and specificity, thereby more accurately and stably identifying high-risk individuals for cerebral hemorrhage and facilitating early risk stratification and targeted intervention.
[0031] In some embodiments, the prediction threshold for the risk of cerebral hemorrhage is 0.07. When the lipid risk score for cerebral hemorrhage exceeds this threshold, the risk of cerebral hemorrhage is considered high. Experimental verification shows that the AUC value of this model for predicting the risk of cerebral hemorrhage is 0.87, the sensitivity is 0.84, and the specificity is 0.72. Under the same experimental conditions, the AUC of traditional predictive factors for predicting the risk of ischemic stroke is only 0.64, the sensitivity is 0.55, and the specificity is 0.71. The comparative results indicate that the prediction model provided by this invention has significantly better overall predictive performance than traditional predictive factors and can more efficiently identify high-risk individuals for cerebral hemorrhage.
[0032] Based on the relative risk stratification results indicated by this risk score, patients at high risk of cerebral hemorrhage can be given priority for intensive treatment and / or more intensive monitoring, thus providing an objective basis for clinical decision-making.
[0033] A fourth aspect of the present invention provides a method for constructing a model for predicting ischemic stroke and / or cerebral hemorrhage, comprising:
[0034] Lipidome data of plasma samples are obtained, and the lipidome data is preprocessed.
[0035] The preprocessed lipidome data were screened using forward selection and optimal subset analysis to obtain the optimal lipid subset associated with ischemic stroke and / or cerebral hemorrhage.
[0036] Based on the optimal lipid subset, a prediction model is constructed using logistic regression.
[0037] In some embodiments, the acquisition of the lipidomics data includes detecting lipids in a plasma sample using chromatography, mass spectrometry, or a combination thereof.
[0038] In some specific embodiments, the lipidomics data are acquired using a high-coverage lipidomics detection method based on an ultra-high performance liquid chromatography-tandem mass spectrometry (UHPLC-MS / MS) system.
[0039] In some embodiments, experimental bias and batch effects need to be corrected during the detection process. This correction step can effectively eliminate systematic errors introduced by reagents, instruments, and operation time, improve the accuracy and comparability of data, and thus ensure that the lipid diagnostic factors subsequently screened have higher reliability.
[0040] In some embodiments, the data preprocessing includes at least one of the following: lipid feature data processing, missing value processing, batch effect correction, variant feature removal, or standardization. This preprocessing can significantly improve the reliability and comparability of the data, help to screen out more stable lipid diagnostic factors, and thus improve the reproducibility and applicability of the constructed prediction model in different sample sets and experimental environments.
[0041] In some embodiments, the screening criteria for the optimal lipid subset include:
[0042] Lipid pairs with a Pearson correlation coefficient greater than 0.8 were excluded;
[0043] Using the forward selection method, lipids with high contribution increments were selected by ranking them according to the area under the ROC curve. At this time, the model AUC exceeded 0.9.
[0044] The optimal subset analysis method is used to compare the Bayesian information criterion values of different lipid subsets and select the lipid subset with the smallest criterion value.
[0045] During the screening process, lipid pairs with a Pearson correlation coefficient greater than 0.8 are prioritized to effectively reduce the impact of multicollinearity among variables. Lipids with high average absolute correlation to the remaining lipids are further eliminated to improve model stability. Forward selection gradually identifies the lipid molecules most critical to improving the model's discriminative performance, ensuring that the selected features have clear predictive value and improving the model's interpretability and effectiveness. The Bayesian information criterion introduces a penalty term related to the number of parameters while evaluating the model's goodness of fit, thereby automatically balancing model complexity and fitting ability while avoiding overfitting, ultimately selecting the optimal lipid subset with strong generalization ability and statistical robustness.
[0046] In some embodiments, the forward selection method includes: starting from an empty feature set, adding a single feature that optimizes the model's AUC in each iteration to obtain the lipid molecule that most significantly improves model performance. The model AUC is consistently above 0.90, and adding any single lipid molecule does not increase the model AUC by more than 0.002.
[0047] In some specific embodiments, the forward selection method includes: starting from an empty feature set, adding a single feature that optimizes the model's AUC in each iteration to obtain the top 15 lipid molecules that have the most significant effect on improving model performance.
[0048] In some embodiments, the lipid data is input into a logistic regression model as a continuous variable for modeling and analysis.
[0049] The fifth aspect of the invention provides the use of the previously described diagnostic factor composition for cerebrovascular diseases, the previously described model for predicting the risk of ischemic stroke, the previously described model for predicting the risk of cerebral hemorrhage, or the model constructed by the previously described method, in the preparation of products for predicting the risk of cerebrovascular diseases, or in the preparation of products for diagnosing cerebrovascular diseases.
[0050] The sixth aspect of the present invention provides a computer program product including computer-readable instructions that, when the computer-readable instructions are executed on an electronic device, cause the electronic device to perform the prediction function of the model as described above, output the lipid risk score of ischemic stroke and / or lipid risk score of cerebral hemorrhage, or perform the model construction method as described above.
[0051] A seventh aspect of the present invention provides an electronic device comprising at least one processor and a memory connected to said processor, wherein:
[0052] The memory is used to store computer programs;
[0053] The processor is used to execute the computer program to enable the electronic device to realize the predictive function of the model as described above, output the lipid risk score of ischemic stroke and / or lipid risk score of cerebral hemorrhage, or implement the model construction method as described above.
[0054] An eighth aspect of the present invention provides a computer storage medium carrying one or more computer programs that, when executed by an electronic device, enable the electronic device to perform the predictive function of the model as described above, output the lipid risk score for ischemic stroke and / or the lipid risk score for cerebral hemorrhage, or implement the model construction method as described above.
[0055] A ninth aspect of the present invention provides a product for diagnosing cerebrovascular-related diseases, comprising the cerebrovascular-related disease diagnostic factor composition as described above and pharmaceutically acceptable excipients.
[0056] This invention relates to a risk prediction model for ischemic stroke and intracerebral hemorrhage based on novel lipidomics predictive factors and its application. The model features high predictive efficacy and specificity, enabling more effective identification of high-risk populations for ischemic stroke and intracerebral hemorrhage, providing a reliable basis for early clinical intervention, and showing promising application prospects in precision stroke prevention, clinical risk assessment, and disease burden control. Attached Figure Description
[0057] Figure 1A flowchart for model building and validation;
[0058] Figure 2 AUC curves for lipidome predictors of ischemic stroke (IS);
[0059] Figure 3 AUC curves for lipidomics predictors of intracerebral hemorrhage (ICH);
[0060] Figure 4 AUC curves for predicting ischemic stroke (IS) using traditional predictors;
[0061] Figure 5 AUC curves for predicting intracerebral hemorrhage (ICH) using traditional predictors;
[0062] Figure 6 AUC curves for predicting ischemic stroke (IS) using a combination of traditional predictors and lipidome predictors;
[0063] Figure 7 AUC curves for predicting intracerebral hemorrhage (ICH) using a combination of traditional predictors and lipidomics predictors.
[0064] Figure 8 This is a schematic diagram of the composition structure of an electronic device provided by the present invention. Detailed Implementation
[0065] This invention provides a risk prediction model for ischemic stroke and cerebral hemorrhage based on novel lipidomics predictive factors and its application. Those skilled in the art can refer to the content of this document and appropriately modify the process parameters to achieve the model. It should be particularly noted that all similar substitutions and modifications are obvious to those skilled in the art and are considered to be included in this invention. The methods and applications of this invention have been described through preferred embodiments. Those skilled in the art can obviously modify or appropriately change and combine the methods and applications described herein without departing from the content, spirit, and scope of this invention to implement and apply the technology of this invention.
[0066] Unless otherwise defined in this invention, the scientific and technical terms associated with this invention shall have the meanings understood by one of ordinary skill in the art.
[0067] The terms “comprising,” “including,” and “having” are used interchangeably to indicate the inclusiveness of a scheme, meaning that the scheme may contain elements other than those listed. It should also be understood that the use of “comprising,” “including,” and “having” herein also provides for schemes “consisting of…”.
[0068] When used herein, the term “and / or” includes the meaning of “and,” “or,” and “all or any other combination of elements linked by the term.”
[0069] The term "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or multiple items.
[0070] The terms "traditional predictors," "traditional diagnostic factors," or "traditional predictive factors" refer to age, sex, smoking status, BMI, systolic blood pressure, total cholesterol, high-density lipoprotein cholesterol, use of antihypertensive drugs, use of lipid-lowering drugs, and history of diabetes. This information primarily comes from the highly authoritative cardiovascular disease risk prediction model developed by China-PAR (Predicting the 10-Year Risks of Atherosclerotic Cardiovascular Disease in Chinese Population: The China-PAR Project (Prediction for ASCVD Risk in China) - PubMed).
[0071] The Bayesian Information Criterion (BIC) evaluates the trade-off between the goodness of fit and complexity of a statistical model. By introducing a penalty term related to the sample size and the number of parameters, it selects the optimal model while avoiding overfitting.
[0072] It should be understood that in the various embodiments of this application, the order of the above processes does not imply the order of execution. Some or all steps may be executed in parallel or sequentially. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0073] The numerical ranges and parameters involved in this invention have been presented as precisely as possible in the specific embodiments. However, any numerical value inevitably contains standard deviations due to individual test methods. Therefore, unless otherwise expressly stated, it should be understood that all numerical ranges or specific data used in this disclosure may have a reasonable deviation within a certain range, such as ±10%, ±5%, ±1%, or ±0.5%.
[0074] This patent application pertains to the construction of a risk prediction model for ischemic stroke and cerebral hemorrhage based on novel lipidomics predictive molecules. The core content of this invention is as follows:
[0075] (1) Detection of novel lipidome predictors: Ultra-high performance liquid chromatography-tandem mass spectrometry is used for lipidome detection, which has a higher coverage of lipidome than existing technologies and can detect more novel lipid molecules that have not been previously studied. Compared with existing lipid predictors, the types of lipids used in this patent have been expanded to 37 subclasses, and the structural information (such as carbon chain length, number of double bonds, etc.) of 1067 lipid molecules has been accurately resolved.
[0076] (2) Standardized screening of the optimal predictive molecule combination: The detected novel lipid molecules were screened. The forward selection method and the best subset analysis method were used to screen out the lipid molecule combination with the best predictive ability for ischemic stroke and cerebral hemorrhage, respectively.
[0077] (3) Evaluation of the predictive ability of emerging lipid biomarkers: Based on the screened lipid molecules, lipid risk scores were calculated for ischemic stroke and cerebral hemorrhage, and the predictive ability of the models was evaluated and compared with traditional predictive models to determine whether the new lipid biomarker combination has achieved an improvement in predictive ability compared with traditional models.
[0078] (4) Novel lipidome predictors can more effectively predict the risk of ischemic stroke and cerebral hemorrhage: This patent uses novel lipid predictors with specific molecular structures. By comparing the predictive ability of novel lipid biomarker scores with traditional models, and by searching relevant literature and patents, it was found that the predictive ability of the lipidome molecular combination provided by this patent is significantly better than that of traditional predictive models, and there have been no similar reports before. This patent creatively proposes a new and more effective predictor.
[0079] This invention represents the largest and most comprehensive study to date in the field of ischemic stroke and cerebral hemorrhage risk prediction, incorporating the most lipid molecules and achieving the highest lipidome coverage. Furthermore, it optimizes the results through extensive comparison of the performance of predictive factors. The study screened predictive factors from 1,067 lipid molecules across 37 lipid subclasses to ensure optimal combinations. These lipids showed the most significant improvement in model AUC and the lowest BIC, outperforming over 1,000 other lipid molecules.
[0080] The test materials used in this invention are all common commercial products and can be purchased on the market.
[0081] The present invention will be further illustrated below with reference to the embodiments:
[0082] Example 1: Detection of Novel Lipomic Biomarkers
[0083] Lipid extraction from plasma samples was performed using a combination of Sarafian and Matyash protocols, followed by high-coverage lipidomics detection using an ultra-high performance liquid chromatography-tandem mass spectrometry (UHPLC-MS / MS) system. Samples were analyzed in a randomized order, and the case-control grouping information was blinded to minimize bias. Quality control samples were inserted every 6–8 tests to monitor and correct for batch effects.
[0084] Lipid feature data were processed using SCIEX OS (V 1.7) software, and peak extraction, verification, and false positive correction were performed using R packages. Lipid features with more than 20% missing values were removed, and the remaining missing values were imputed using a random forest algorithm. Batch effects were corrected using the batch ratio method based on quality control samples, and features with a coefficient of variation >30% in the quality control samples were excluded. For ease of subsequent analysis, lipid levels were standardized using a rank-inverse normal transformation.
[0085] The resulting lipid dataset covers 1,067 specific lipid molecules across 37 subclasses, demonstrating a higher coverage of the lipidome than other related studies and patents, and enabling the quantification of novel lipidome biomarkers.
[0086] Table 1. Subclasses covered by the lipid dataset
[0087]
[0088] Example 2: Standardized Predictor Screening
[0089] The study screened predictive factors among 1,067 lipid molecules, encompassing 37 lipid subclasses, to ensure optimal combination of predictive factors.
[0090] To address multicollinearity, lipid pairs with a Pearson correlation coefficient > 0.8 were screened, and those with a high mean absolute correlation to the remaining lipids were removed. The population dataset (total sample size 2,539: 1,103 ischemic stroke patients, 159 cerebral hemorrhage patients, and 1,277 healthy controls) was then randomly divided into a training set (75%) and a validation set (25%). In the training set, lipids significantly associated with ischemic stroke or cerebral hemorrhage were screened using forward selection. Starting with an empty feature set, each iteration added a single feature that optimized the model's AUC, resulting in the top 15 lipid molecules that most significantly improved model performance (i.e., the top 15 lipids that contributed the most to the increase in the area under the curve (AUC) for ischemic stroke and cerebral hemorrhage). At this point, the AUC of the ischemic stroke and cerebral hemorrhage prediction models both reached above 0.90, and the addition of any single lipid molecule did not increase the model's AUC by more than 0.002.
[0091] For ischemic stroke, the top 15 lipids with the largest increase in area under the ROC curve were Sphingosine-1-phosphate (d16:1), PE (P-18:0 / 22:5), 13-HODE, 9,10-diHOME, PI (18:2 / 18:2), PE (18:2 / 18:2), SM (d44:4), Sphingosine-1-phosphate (d18:2), Butyryl-carnitine, Cer-NS (d18:1 / 18:0), DAG (18:2 / 18:2), MAG (18:2), LPC (20:1)-sn2, Eicosadienoyl-carnitine, and Pentadecanoic acid.
[0092] For cerebral hemorrhage, the top 15 lipids with the largest increase in area under the ROC curve were Eicosatrienoyl-carnitine, PE (P-20:0 / 20:4), PI (18:2 / 18:2), PC (18:2 / 18:2), TAG (58:3)-FA22:1, 13-HODE, 9,10-diHOME, Ganglioside GM3 (d18:1 / 24:1), DAG (20:4 / 18:2), Sphingosine-1-phosphate (d16:1), LPC (24:0)-sn2, PS (40:2), PS (38:6), PE (20:2 / 18:2), and PE (20:2 / 20:4).
[0093] Next, an optimal subset analysis was performed on these 15 lipid molecules. By exhaustively exploring all possible combinations of predictive factors, the Bayesian Information Criterion (BIC) of the model was directly compared to find the combination that minimizes the model's BIC. This global search method determined the optimal feature subset, forming the final predictive model. The results showed that the final optimal lipid subset for ischemic stroke and cerebral hemorrhage included eight lipid molecules, respectively.
[0094] For ischemic stroke, the optimal subset of lipids obtained from screening included PE (P-18:0 / 22:5), 13-HODE, 9,10-diHOME, PI (18:2 / 18:2), PE (18:2 / 18:2), SM (d44:4), Sphingosine-1-phosphate (d18:2), and Butyryl-carnitine.
[0095] For cerebral hemorrhage, the optimal subset of lipids obtained from screening included Eicosatrienoyl-carnitine, PI (18:2 / 18:2), PC (18:2 / 18:2), TAG (58:3)-FA22:1, 13-HODE, 9,10-diHOME, DAG (20:4 / 18:2), and LPC (24:0)-sn2.
[0096] Table 2 BIC results for ischemic stroke
[0097]
[0098] Table 3. BIC results of cerebral hemorrhage
[0099]
[0100] Example 3: Evaluation of the predictive ability of novel lipid biomarkers
[0101] In the independent validation set (25%) obtained by random sampling, lipid risk scores for ischemic stroke and cerebral hemorrhage were constructed based on the two optimal lipid subsets mentioned above, with weights derived from the coefficients in the logistic regression equation.
[0102] Compare the predictive performance of three types of models for the risk of ischemic stroke and cerebral hemorrhage: (1) lipid risk score alone; (2) traditional predictors alone (age, sex, smoking status, body mass index, systolic blood pressure, total cholesterol, high-density lipoprotein cholesterol, use of antihypertensive drugs, use of lipid-lowering drugs and history of diabetes); (3) lipid risk score and traditional predictors in combination.
[0103] The results showed that the lipid risk score significantly outperformed traditional predictive factors in predicting ischemic stroke risk, achieving an AUC of 0.90 (95% CI: 0.90–0.90) and an AUC of 0.87 (95% CI: 0.87–0.87) in predicting intracerebral hemorrhage risk. Traditional predictive factors had an AUC of 0.64 (95% CI: 0.63, 0.65) for ischemic stroke and 0.64 (95% CI: 0.61, 0.66) for intracerebral hemorrhage. Adding traditional predictive factors to the lipid risk score model did not significantly improve the predictive ability of ischemic stroke (AUC: 0.90–0.90) but slightly improved it (AUC: 0.90 (95% CI: 0.87–0.91) for intracerebral hemorrhage.
[0104] After thorough research, the lipid predictor discovered in this patent has not been reported in previous studies, and its predictive ability exceeds that of existing traditional predictive models. The lipid risk scores constructed using eight lipids for ischemic stroke and cerebral hemorrhage are as follows:
[0105] Lipid risk score for ischemic stroke = -0.089 + Sphingosine-1-phosphate (d18:2) ×0.944 + Butyryl-carnitine × (-0.399) + PE (P-18:0 / 22:5) × (-0.359) + PE(18:2 / 18:2) × (-0.984) + PI (18:2_18:2) × 1.311 + SM (d44:4) × (-0.549) + 13-HODE × 0.883 + 9,10-diHOME × (-0.69);
[0106] Cerebral hemorrhage lipid risk score = -2.009 + Eicosatrienoyl-carnitine × 0.581 + PI(18:2_18:2) × 1.363 + TAG (58:3)-FA22:1 × 0.433 + DAG (20:4 / 18:2) × (-0.529) + PC (18:2 / 18:2) × (-0.705) + LPC (24:0)-sn2 × (-0.394) + 13-HODE ×0.681 + 9,10-diHOME × (-0.631).
[0107] Note: Sphingosine-1-phosphate (S1P); Butyryl-carnitine; Alkenyl-phosphatidylethanolamine (PE-P); Phosphatidylethanolamine (PE); Phosphatidylinositol (PI); Sphingomyelin (SM); 13-Hydroxy-9Z,11E-octadecadienoic acid (13-HODE); 9,10-Dihydroxy-12Z-octadecenoic acid 9,10-diHOME); Eicosatrienoyl-carnitine; Triacylglycerol (TAG); Fatty acyl (FA); Diacylglycerol (DAG); Phosphatidylcholine (PC); Lysophosphatidylcholine (LPC).
[0108] For stroke, the model threshold was 0.33, with a sensitivity of 0.87 and a specificity of 0.73; for cerebral hemorrhage, the threshold was 0.07, with a sensitivity of 0.84 and a specificity of 0.72. Values exceeding these thresholds were considered to indicate a higher risk of developing the disease.
[0109] The threshold for predicting ischemic stroke risk using traditional predictive factors was 0.49, the sensitivity was 0.52, and the specificity was 0.71. The threshold for predicting ischemic stroke risk using a combination of lipid risk score and traditional predictive factors was 0.30, the sensitivity was 0.88, and the specificity was 0.71.
[0110] The threshold for predicting the risk of cerebral hemorrhage using traditional predictive factors was 0.13, the sensitivity was 0.55, and the specificity was 0.71. The threshold for predicting the risk of cerebral hemorrhage using a combination of lipid risk score and traditional predictive factors was 0.12, the sensitivity was 0.84, and the specificity was 0.83.
[0111] This invention provides a risk prediction model for ischemic stroke and cerebral hemorrhage based on novel lipidomics predictive factors and its applications. The model features high predictive efficacy and specificity, enabling more effective identification of high-risk populations for ischemic stroke and cerebral hemorrhage, providing a reliable basis for early clinical intervention, and showing promising application prospects in precision stroke prevention, clinical risk assessment, and disease burden control.
[0112] The above are merely preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A diagnostic factor composition for cerebrovascular-related diseases, characterized in that, Including at least one of sphingosine-1-phosphate, butyrylcarnitine, ethanolamine acetal phospholipid, phosphatidylethanolamine, phosphatidylinositol, sphingomyelin, 13-hydroxy-9Z,11E-octadecadienoic acid, 9,10-dihydroxy-12Z-octadecadienoic acid, eicostrienoic carnitine, triacylglycerol, diacylglycerol, phosphatidylcholine, or lysophosphatidylcholine; The cerebrovascular-related diseases include ischemic stroke and cerebral hemorrhage.
2. The diagnostic factor composition for cerebrovascular diseases according to claim 1, characterized in that, The sphingosine-1-phosphate includes sphingosine-1-phosphate (d18:2), the ethanolamine acetal phospholipid includes ethanolamine acetal phospholipid (P-18:0 / 22:5), the phosphatidylethanolamine includes phosphatidylethanolamine (18:2 / 18:2), the phosphatidylinositol includes phosphatidylinositol (18:2-18:2), the sphingomyelin includes sphingomyelin (d44:4), the triacylglycerol includes triacylglycerol (58:3)-fatty acyl (22:1), the diacylglycerol includes diacylglycerol (20:4 / 18:2), the phosphatidylcholine includes phosphatidylcholine (18:2 / 18:2), and the lysophosphatidylcholine includes lysophosphatidylcholine (24:0)-sn2.
3. A model for predicting the risk of ischemic stroke, characterized in that, The model is as follows: Ischemic stroke lipid risk score = α + β1*Sphingosine-1-phosphate (d18:2) + β2*Butyryl-carnitine + β3*PE (P-18:0 / 22:5) + β4*PE (18:2 / 18:2) + β5*PI (18:2_18:2) + β6*SM (d44:4) + β7*13-HODE + β8*9,10-diHOME; Wherein, Sphingosine-1-phosphate (d18:2) is sphingosine-1-phosphate (d18:2), Butyryl-carnitine is butyrylcarnitine, PE (P-18:0 / 22:5) is ethanolamine acetal phospholipid (P-18:0 / 22:5), PE (18:2 / 18:2) is phosphatidylethanolamine (18:2 / 18:2), PI (18:2_18:2) is phosphatidylinositol (18:2_18:2), SM (d44:4) is sphingomyelin (d44:4), 13-HODE is 13-hydroxy-9Z,11E-octadecadienoic acid, and 9,10-diHOME is 9,10-dihydroxy-12Z-octadecadienoic acid; The α, β1, β2, β3, β4, β5, β6, β7 and β8 are model coefficients; The values of α are -0.15 to -0.01, β1 is 0.5 to 1.5, β2 is -1.0 to -0.01, β3 is -1.0 to -0.01, β4 is -1.5 to -0.5, β5 is 0.8 to 1.8, β6 is -1.0 to -0.1, β7 is 0.1 to 1.5, and β8 is -1.2 to -0.
1.
4. A model for predicting the risk of cerebral hemorrhage, characterized in that, The model is as follows: Cerebral hemorrhage lipid risk score = α + β1*Eicosatrienoyl-carnitine + β2*PI (18:2_18:2) + β3*TAG (58:3)-FA (22:1) + β4*DAG (20:4 / 18:2) + β5*PC (18:2 / 18:2) + β6*LPC (24:0)-sn2 + β7*13-HODE + β8*9,10-diHOME; Wherein, Eicosatrienoyl-carnitine is eicosatrienoylcarnitine, PI (18:2_18:2) is phosphatidylinositol (18:2_18:2), TAG (58:3)-FA (22:1) is triacylglycerol (58:3)-fatty acyl (22:1), DAG (20:4 / 18:2) is diacylglycerol (20:4 / 18:2), PC (18:2 / 18:2) is phosphatidylcholine (18:2 / 18:2), LPC (24:0)-sn2 is lysophosphatidylcholine (24:0)-sn2, 13-HODE is 13-hydroxy-9Z,11E-octadecadienoic acid, and 9,10-diHOME is 9,10-dihydroxy-12Z-octadecadienoic acid; The α, β1, β2, β3, β4, β5, β6, β7 and β8 are model coefficients; The values of α are -5 to -1, β1 is 0.1 to 1.0, β2 is 0.8 to 1.8, β3 is 0.1 to 1.0, β4 is -1.0 to -0.1, β5 is -1.0 to -0.1, β6 is -1.0 to -0.1, β7 is 0.1 to 1.0, and β8 is -1.0 to -0.
1.
5. A method for constructing a model to predict ischemic stroke and / or cerebral hemorrhage, characterized in that, include: Lipidome data of plasma samples are obtained, and the lipidome data is preprocessed. The preprocessed lipidome data were screened using forward selection and optimal subset analysis to obtain the optimal lipid subset associated with ischemic stroke and / or cerebral hemorrhage. Based on the optimal lipid subset, a prediction model is constructed using logistic regression.
6. The construction method according to claim 5, characterized in that, The acquisition of lipidomics data includes: detecting lipids in plasma samples using chromatography, mass spectrometry, or a combination thereof, and the detection process requires correction for experimental bias and batch effects.
7. The construction method according to claim 5, characterized in that, The data preprocessing includes at least one of the following: lipid feature data processing, missing value processing, batch effect correction, variant feature removal, or standardization.
8. The construction method according to claim 5, characterized in that, The screening criteria for the optimal lipid subset include: Lipid pairs with a Pearson correlation coefficient greater than 0.8 were excluded; Using the forward selection method, lipids with high contribution increments were selected by ranking them according to the area under the ROC curve. At this time, the model AUC exceeded 0.
9. The optimal subset analysis method is used to compare the Bayesian information criterion values of different lipid subsets and select the lipid subset with the smallest criterion value.
9. The diagnostic factor composition for cerebrovascular diseases according to claim 1 or 2, the model for predicting the risk of ischemic stroke according to claim 3, the model for predicting the risk of cerebral hemorrhage according to claim 4, or the model constructed by the method according to any one of claims 5 to 8, for use in the preparation of products for predicting the risk of cerebrovascular diseases, or for use in the preparation of products for diagnosing cerebrovascular diseases.
10. A computer program product, characterized in that, It includes computer-readable instructions that, when executed on an electronic device, cause the electronic device to perform the predictive function of the model as described in claim 3 or 4, output the ischemic stroke lipid risk score and / or cerebral hemorrhage lipid risk score, or perform the model construction method as described in any one of claims 5 to 8.
11. An electronic device, characterized in that, It includes at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program so that the electronic device can realize the predictive function of the model as described in claim 3 or 4, output the ischemic stroke lipid risk score and / or cerebral hemorrhage lipid risk score, or realize the model construction method as described in any one of claims 5 to 8.
12. A computer storage medium, characterized in that, The storage medium carries one or more computer programs, which, when executed by an electronic device, enable the electronic device to perform the predictive function of the model as described in claim 3 or 4, output the ischemic stroke lipid risk score and / or cerebral hemorrhage lipid risk score, or implement the model construction method as described in any one of claims 5 to 8.