A method for identifying cardiogenic stroke based on clinical and genetic characteristics and application thereof
Patent Information
- Application Number
- CN202611060859.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-16
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-07-16
AI Technical Summary
目前开发的CES分类器存在如下局限性:(1)模型均基于欧美人群的横断面数据开发,预后评估缺失;(2)研究样本量相对较小(141至1598例),人群代表性不足;(3)变量选择依赖于影像学数据或超声心动图,而这些数据通常难以获取,或者在病因不明的卒中患者中无法准确评估;(4)变量选择包括自我报告的疾病史和用药史,这些指标往往会引入回忆偏倚,并且跨队列应用判别性能降低
(1)本发明先通过LASSO正则化的logistic回归模型从数据集中的40多个临床指标和遗传特征中筛选出15个与CES高相关且稳定的预测变量集,再利用该预测变量集对logistic回归模型进行训练,从而确定各变量的权重,得到识别精度较高的分类器。采用不同种子数重复多次LASSO优化变量选择程序,降低了变量间的相关性并克服了单次LASSO结果的随机性,最终模型去掉了冗余的变量,仅包含少量预测变量,便于临床应用,对隐匿性CES患者的早期识别与治疗具有重要意义。
Smart Images

Figure CN122552108B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biomedical technology, specifically relating to a method and application for identifying cardioembolic stroke based on clinical and genetic characteristics. Background Technology
[0002] The etiological classification of ischemic stroke is crucial for clinical anticoagulation or antithrombotic therapy. Currently, the cause of ischemia remains undetermined in nearly 25% of ischemic stroke patients. The annual recurrence rate of this type of stroke of unknown cause is approximately 4.5-5%, and poor clinical care leads to a poor prognosis. Recent studies have found evidence of cardioembolic involvement in some strokes of unknown cause and reported that some patients can benefit from anticoagulants. Cardioembolic stroke (CES) accounts for about one-quarter of ischemic strokes and is one of the most severe subtypes of ischemic stroke, characterized by high acute mortality, poor functional prognosis, high risk of recurrence, and a tendency to transform into hemorrhagic stroke. Accurate diagnosis of CES is essential for guiding anticoagulation therapy and improving patient outcomes. Therefore, strategies are needed to help identify subgroups of stroke patients of unknown cause who are most likely to have cardioembolic involvement and could benefit from anticoagulation therapy.
[0003] Previously, researchers developed several CES classifiers based on medical history and imaging data to distinguish CES from non-CES. This also includes published clinical risk scores for cardiovascular disease and atrial fibrillation to differentiate cardioembolic events, such as CHA2DS2-VASc (…). Chest 2010;137(2):263-272), CHARGE-AF ( J Am Heart Assoc . 2013;2(2):e000102) and EHR-AF ( JACC Clin Electrophysiol . 2019;5(11):1331-1341) model. The currently developed CES classifiers have the following limitations: (1) The models are all developed based on cross-sectional data of European and American populations, and the prognostic assessment is lacking; (2) The sample size of the studies is relatively small (141 to 1598 cases), and the population representativeness is insufficient; (3) The selection of variables depends on imaging data or echocardiography, which are usually difficult to obtain, or cannot be accurately assessed in stroke patients with unknown etiology; (4) The selection of variables includes self-reported medical history and medication history, which often introduce recall bias and reduce discriminative performance when applied across cohorts.
[0004] In summary, current research limitations in variable selection and predictive outcome assessment restrict the clinical application of these models. Therefore, it is of great significance to employ easier-to-use and more accurately measured predictive factors to identify CES patients among those with unexplained stroke and to confirm whether they can benefit from early treatment. Summary of the Invention
[0005] The purpose of this invention is to provide a method and application for identifying cardioembolic stroke based on clinical and genetic characteristics, so as to improve the accuracy of identifying occult CES in patients with unexplained stroke.
[0006] To achieve the above objectives, according to a first aspect of the present invention, a method for constructing a classifier for cardioembolic stroke is provided, comprising: S1. Collect multimodal data on age, sex, clinical data and genetic characteristics of ischemic stroke patients as predictive variables, and construct a dataset with the corresponding etiology; S2. Based on the sample set of known causes of ischemia in the dataset, a LASSO-regularized logistic regression model is trained. The optimal regularization parameter λ is selected through cross-validation strategy, and finally a sparsed set of predictive variables is obtained to achieve multimodal variable screening. S3. Construct a training set by combining the predicted variable set with the known causes of ischemia in patients with cardioembolic stroke. Use a logistic regression model to train the variable set weights in the training set, perform multimodal data fusion, and define the optimal classification threshold for the predicted probability based on the maximum value of the Youden index of the model in the training set. If the predicted probability is greater than the optimal classification threshold, it is cardioembolic stroke; otherwise, it is judged as non-cardioembolic stroke, thus obtaining the cardioembolic stroke classifier.
[0007] Furthermore, in step S2, in order to reduce the correlation between variables and overcome the randomness of a single LASSO result, the LASSO optimization variable selection procedure is repeated multiple times with different seed numbers. The optimal λ value, which is within one standard error range of the maximum observation area under the receiver operating characteristic curve, is selected as the optimal regularization parameter λ corresponding to each cross-validation strategy.
[0008] Furthermore, the clinical data include: smoking status, alcohol consumption, height, weight, body mass index, blood glucose, blood pressure, blood lipids, hematological characteristics, coagulation tests, kidney and liver function, and inflammatory factors; The genetic characteristics include: genotype of a single variant site associated with CES, mitochondrial DNA copy number, and genetic risk scores for cardioembolic stroke and atrial fibrillation; The construction of the dataset involves assigning an output label of 1 to the cause of cardiac stroke and assigning an output label of 0 to other known causes that are not cardiac stroke.
[0009] Furthermore, in step S2, the sparsed predictive variable set includes the age, sex, and clinical data of ischemic stroke patients, including smoking status, systolic blood pressure, creatinine, high-sensitivity C-reactive protein, aspartate aminotransferase, platelet count, absolute value of monocyte population, prothrombin time, D-dimer, triglycerides, and homocysteine. Or it could include the age, sex, and genetic risk score of atrial fibrillation in patients with ischemic stroke.
[0010] Furthermore, the sparsed predictive variable set includes age, sex, and smoking status in clinical data, systolic blood pressure, creatinine, high-sensitivity C-reactive protein, aspartate aminotransferase, platelet count, absolute monocyte population, prothrombin time, D-dimer, triglycerides, homocysteine, and atrial fibrillation genetic risk score.
[0011] Furthermore, the genetic risk score in the genetic trait is calculated through the following steps: S1.1 SNP filtering was performed on the pooled genome-wide association study (GWAS) data of East Asian and European samples of cardiac stroke and atrial fibrillation, and multi-trait GWAS joint analysis was performed. The genetic correlation between traits was used to weight the CES and atrial fibrillation trait-specific GWAS to increase the effective sample size, and multi-trait integrated pooled data of cardiac stroke and atrial fibrillation weighted by the covariance model were obtained respectively. S1.2 Based on the integrated summary data of multiple traits of cardiac stroke and atrial fibrillation, the Bayesian multigene prediction method is used to couple the genetic structure shared by different populations to estimate the specific variation weights of East Asian and European populations. The variation weights of the two populations are then integrated through inverse variance weighted meta-analysis to construct cross-population genetic risk assessment models for cardiac stroke and atrial fibrillation respectively. S1.3 Based on the genotype data of the target population, the genetic risk of the trait is quantified into digital molecular markers by weighting and summing the more than 1 million common variants in HapMap3 using the variant weights of the cross-population genetic risk assessment model, and the genetic risk scores of cardiogenic stroke and atrial fibrillation are calculated.
[0012] Furthermore, the formula for calculating the predicted probability is as follows:
[0013] Where, P(Y=1|X1,…,+X) m The value represents the probability of developing a cardiac stroke, where Y=1 indicates a cardiac stroke, and m is the number of predictive variables selected for calculating the incidence of cardiac stroke. Represents the first of all samples m The values of the predictor variables, Indicates the first m The weights of the predictor variables are estimated by logistic regression.
[0014] According to a second aspect of the present invention, a method for identifying cardioembolic stroke is provided, comprising: acquiring data corresponding to a predictor variable set of ischemic stroke patients, inputting a cardioembolic stroke classifier obtained by any of the above construction methods, outputting a predicted probability of cardioembolic stroke, and classifying a predicted probability greater than the optimal classification threshold as cardioembolic stroke, otherwise classifying it as non-cardioembolic stroke.
[0015] According to a third aspect of the present invention, a system for identifying cardioembolic stroke is provided, employing the above-described method for identifying cardioembolic stroke, comprising: The predictor set acquisition module is used to acquire data corresponding to the predictor set of patients with ischemic stroke. The identification module is used to input the data corresponding to the predicted variable set into the cardioembolic stroke classifier obtained by any of the above construction methods, output the predicted probability of cardioembolic stroke, and judge it as cardioembolic stroke if the predicted probability is greater than the optimal classification threshold; otherwise, it is judged as non-cardioembolic stroke.
[0016] According to a fourth aspect of the present invention, a computer-readable storage medium is provided, on which a computer program is stored, wherein when the computer program is executed by a processor, it implements the method for constructing a cardiac stroke classifier as described in any of the preceding claims, or implements the method for identifying cardiac stroke as described in the preceding claims.
[0017] In summary, compared with the prior art, the above-described technical solutions conceived by this invention mainly possess the following technical advantages: (1) This invention first uses a LASSO-regularized logistic regression model to select 15 highly correlated and stable predictive variables from more than 40 clinical indicators and genetic characteristics in the dataset. Then, the logistic regression model is trained using this predictive variable set to determine the weight of each variable and obtain a classifier with high recognition accuracy. By repeating the LASSO process multiple times with different seed numbers, the variable selection procedure is optimized, reducing the correlation between variables and overcoming the randomness of a single LASSO result. The final model removes redundant variables and contains only a small number of predictive variables, which is convenient for clinical application and is of great significance for the early identification and treatment of patients with occult CES.
[0018] (2) This invention incorporates age, gender, smoking, drinking, physical examination indicators and laboratory test indicators, and excludes self-reported medical history, family history and other indicators to avoid recall bias; it also excludes imaging indicators that are difficult to identify in patients with unexplained stroke, making it easier to apply clinically to patients with unexplained stroke in order to identify occult CES, improve clinical care and prognosis.
[0019] (3) Atrial fibrillation is an important risk factor for CES, but it is often occult or paroxysmal in patients with unexplained stroke, making it difficult to identify CES in patients with unexplained stroke. This invention uses a genetic risk scoring model of atrial fibrillation to accumulate and weight more than 1 million variants, and uses digital molecular markers in the form of genetic risk scores to replace imaging data to quantify the history of atrial fibrillation disease, further improving the accuracy of diagnosis based on clinical factors.
[0020] (4) The sample size of disease cases is often lower than that of body measurement indicators, resulting in lower statistical power of genome-wide association analysis. Therefore, by integrating population-specific GWAS summary data of CES and atrial fibrillation using the correlation between traits, the limitation of single-trait GWAS data sample size is overcome, potentially improving the sample size and statistical power of the target trait. The integration of cross-population genetic information can utilize the linkage disequilibrium diversity and genetic correlation of populations to integrate the large European dataset with the relatively small East Asian dataset, overcoming the problem of insufficient genetic data research in East Asian populations and further improving the predictive effect of genetic risk scores for underrepresented East Asian populations. At the same time, inverse variance weighted meta-analysis integrates the specific posterior SNP effects of multiple populations into a cross-population SNP effect. The inverse variance weighting method does not require independent validation sets to train the cross-population integration weights, and can achieve cross-population integration in small samples lacking independent validation, solving the problem of the current East Asian population cohort being too small.
[0021] (5) This invention applies the cardioembolic stroke identification model to identify potential CES cases in stroke of unknown cause. It quantifies the risk of CES in stroke of unknown cause by combining clinical indicators with genetic characteristics, suggesting that clinicians need to conduct further, more detailed examinations on high-risk patients to improve the clinical identification rate of CES. Simultaneously, it allows for targeted examination of imaging data required for CES diagnosis in high-risk patients, rationally allocating clinical resources. Identified occult CES can also benefit from early anticoagulation therapy, reducing stroke recurrence and death. The cardioembolic stroke identification method based on clinical and genetic characteristics provided by this invention can provide a reference for the identification of other stroke subtypes and has significant application prospects in the precision diagnosis and treatment of chronic diseases with complex etiologies such as stroke. Attached Figure Description
[0022] Figure 1This is a flowchart illustrating the construction method of the method for identifying cardioembolic stroke according to the present invention.
[0023] Figure 2 This is a schematic diagram of the construction process of the method for identifying cardioembolic stroke and the CES assessment system for patients with unexplained stroke, provided in the embodiments of the present invention.
[0024] Figure 3 A flowchart illustrating the research process of an embodiment of the present invention is shown.
[0025] Figure 4 The comparison of genetic traits between CES and non-CES is shown. (A) is a Manhattan plot of CES GWAS (646 CES cases, 4714 non-CES cases). (B), (C), and (D) show the distribution of mitochondrial DNA copy number (mtDNA-CN), CES, and atrial fibrillation genetic risk score (PRS) in CES and non-CES samples, respectively.
[0026] Figure 5 The paper shows how to build CES classifiers using various machine learning models and compare their performance on a test dataset.
[0027] Figure 6 Optimal cutoff values for ROC curves of different classifier models.
[0028] Figure 7 The performance comparison of the STROMICS CES classifier on the test dataset is shown. (A) is the ROC curve of the prediction model, (B) is the precision-recall curve of the prediction model, (C) is the calibration curve, and (D) is the decision curve analysis.
[0029] Figure 8 This shows the SHAP interpretation of variable importance in the optimal STROMICS CES classifier model. (A) Distribution of SHAP features for each variable, (B) Ranking of the importance of variables included in the model, and (C) and (D) Examples of feature contributions in a CES case and a non-CES case, respectively. Abbreviations: PRS, Polygenic Risk Score; AF, Atrial Fibrillation; INR, International Normalized Ratio; PLT, Platelet Count; TG, Triglycerides; SBP, Systolic Blood Pressure; AST, Aspartate Aminotransferase; PT, Prothrombin Time; hsCRP, High-Sensitivity C-Reactive Protein; MONO, Absolute Monocyte Count; Hcy, Homocysteine.
[0030] Figure 9Prognostic assessment of stroke patients with unknown etiology using STROMICS CES classifier model 3 is shown. (A) One-year cumulative mortality curves for patients with occult CES and non-CES predicted by model 3. Hazard ratios (HRs) were estimated by the Cox model and adjusted for age, stroke history, and NIHSS score at discharge. x The axis represents the follow-up time. y The axis represents cumulative all-cause mortality. Shaded areas represent 95% confidence intervals. (B) Model 3 predicted occult CES patients ( n = 1067) and (C) non-CES patients ( n In the study of 3698 patients with unexplained stroke who received anticoagulation therapy, the proportion of patients with ΔmRS < 0 at discharge was included in the study. ΔmRS represents the difference between the mRS score at 3 months and / or 12 months after discharge and the mRS score at admission. Data are expressed as the proportion of individuals with each ΔmRS value among those who received and did not receive anticoagulation therapy. Indicates statistical test P <0.05.
[0031] Figure 10 This paper presents the prognostic assessment of stroke patients with unknown etiology after stratification using the published model EHR-AF. (A) One-year cumulative mortality curves for occult CES and non-CES patients predicted by the EHR-AF model. Hazard ratios (HRs) were estimated by the Cox model and adjusted for age, stroke history, and NIHSS score at discharge. x The axis represents the follow-up time. y The axis represents cumulative all-cause mortality. Shaded areas represent 95% confidence intervals. (B) and (C) show the occult CES predicted by the EHR-AF model, respectively. n =1750) and non-CES ( n = 3015) The proportion of stroke patients with unexplained causal infection who received anticoagulation therapy compared to those who did not receive anticoagulation therapy with ΔmRS < 0. Indicates statistical test P <0.05. Detailed Implementation
[0032] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.
[0033] like Figure 1This invention provides a method for constructing a stroke classifier (STROMICS CES classifier), comprising: S1. Collect multimodal data such as age, gender, clinical data and genetic characteristics of ischemic stroke patients as predictive variables, and construct a dataset with the corresponding etiology; S2. Based on the sample set of known ischemic causes in the dataset, a LASSO-regularized logistic regression model is trained. The optimal regularization parameter λ is selected through a 10-fold cross-validation strategy, and finally a sparsed set of predictive variables is obtained to achieve multimodal variable screening. S3. Construct a training set by combining the predicted variable set with information on whether patients with known ischemic causes have cardioembolic stroke. Use a logistic regression model to train the variable set weights within the training set, perform multimodal data fusion, and define an optimal classification threshold for the predicted probability based on the maximum Youden index value of the model on the training set. If the predicted probability is greater than this optimal classification threshold, the stroke is classified as cardioembolic; otherwise, it is classified as non-cardioembolic stroke, thus obtaining the stroke classification model. The formula for calculating the predicted probability is shown below:
[0034] in, P (Y=1|X1,…,+X m The value represents the probability of developing a cardiac stroke, where Y=1 indicates a cardiac stroke, and m is the number of predictive variables selected for calculating the incidence of cardiac stroke. Represents the first of all samples m The values of the predictor variables, Indicates the first m The weights of the predictor variables are estimated by logistic regression.
[0035] The clinical data includes: smoking status, alcohol consumption, height, weight, body mass index, blood glucose, blood pressure, blood lipids, hematological characteristics, coagulation tests, kidney and liver function, coagulation markers, and inflammatory factors. The skewed distribution indicators in the clinical data undergo normalization transformation and missing value imputation. Specifically, this includes: ① deleting continuous variables with a missing percentage greater than 25%; ② performing a logarithmic transformation on right-skewed variables; ③ designating outliers exceeding 8 standard deviations from the mean as missing values; ④ imputing missing values using the predicted mean matching method, averaging the five imputation iterations, with each iteration running 10 times.
[0036] The genetic characteristics include: genotype of a single variant site associated with CES, mitochondrial DNA copy number, and genetic risk scores for cardioembolic stroke and atrial fibrillation. These genetic characteristics are obtained based on whole-genome sequencing data from the target population. Specifically, whole-genome sequencing data is obtained by extracting leukocytes from samples and performing whole-genome sequencing (WGS) using the BGISEQ-500 platform. To calculate the genetic characteristics, this invention excludes samples with low call rates (genotype quality <20 in >5% of samples) or those deviating from Havenberg equilibrium. P <1×10 6 Genetic variations were selected from 22 common variations (minor allele frequency >1%) on autosomes for subsequent implementation.
[0037] Individual variants associated with CES were identified through genome-wide association analysis (GWAS) in patients with stroke of known cause. Specifically, logistic regression in PLINK 2.00 was used to perform GWAS on CES cases and non-CES cases (large artery atherosclerotic stroke and small vessel occlusion stroke). Sex, age, and the top ten principal components were included as covariates in the model. The significance threshold for GWAS was set at P < 5 × 10⁻⁶. 8 By requiring that the major variants within each locus be >500 kb apart and not in linkage disequilibrium with other major variants ( r 2 A locus is defined as a single locus with a value <0.1. A second logistic regression analysis is performed, conditioned on the significant SNPs identified in the first genome-wide association analysis, to identify any potential secondary signals other than those SNPs.
[0038] The mitochondrial DNA copy number was quantitatively analyzed using the fastMitoCalc program on the whole-genome sequencing data of the samples. This program defines the mitochondrial DNA copy number as the ratio of mitochondrial DNA coverage to autosomal coverage, multiplied by 2. Linear regression analysis was then used to subtract the influence of platelet and neutrophil ratios on the mitochondrial DNA copy number, and the residuals were used for subsequent analysis.
[0039] The genetic risk score is calculated through the following steps: S1.1. SNP filtering was performed on the GWAS pooled data of East Asian and European samples of cardiac stroke and atrial fibrillation, and multi-trait GWAS joint analysis was performed. The genetic correlation between traits was used to weight the CES and atrial fibrillation trait-specific GWAS to increase the effective sample size, and multi-trait integrated pooled data of cardiac stroke and atrial fibrillation weighted by the covariance model were obtained respectively. S1.2 Based on the integrated summary data of multiple traits of cardiac stroke and atrial fibrillation, the Bayesian multigene prediction method is used to couple the genetic structure shared by different populations to estimate the specific variation weights of East Asian and European populations. The variation weights of the two populations are then integrated through inverse variance weighted meta-analysis to construct cross-population genetic risk assessment models for cardiac stroke and atrial fibrillation respectively. S1.3. Based on the genotype data of the target population, the genetic risk of traits is quantified into digital molecular markers by weighting and summing the more than one million common variants in HapMap3 using the variant weights of the cross-population genetic risk assessment model. The genetic risk scores for cardioembolic stroke and atrial fibrillation are calculated using the following formulas: .
[0040] Among them, PRS i Representative of individuals i Genetic risk score, Representative variation j Estimation of genetic effect size (i.e., weight) of multi-trait cross-population integration. Representing the individual i Medium variation j The effect allele count (0, 1, or 2), m It represents the total number of genetic variations.
[0041] Preferably, in step S1.1, the multi-trait GWAS joint analysis employs the multi-trait analysis method for genome-wide association studies (MTAG). Nat Genet (2018;50(2):229-237). First, GWAS pooled data of CES and atrial fibrillation were obtained from public databases. Then, population-specific genome-wide association pooled data of CES and atrial fibrillation were integrated using the correlation between traits to generate trait-specific SNP effect estimates weighted by the covariance model, thereby improving the potential sample size and statistical power of the target trait. Multi-trait integrated pooled data of CES and atrial fibrillation were output separately through multi-trait integration.
[0042] SNP (single nucleotide polymorphism) filtering includes: GWAS pooled data based on CES and atrial fibrillation in East Asian and European populations, excluding SNPs with a minor allele frequency of less than 1%, SNPs with a sample size below 75% of the 90th percentile of the SNP sample size distribution, and SNPs with abnormal effect size estimation.
[0043] Preferably, in step S1.2, the aggregated data based on multi-trait integration analysis is used to perform cross-population integration of East Asian and European populations, and the Bayesian multigene prediction method is used to estimate the variation weights of the integrated East Asian and European populations. Nat Genet . 2022;54(5):573-580), constructing a cross-population PRS for CES and atrial fibrillation. This method couples shared genetic information from different populations through shared continuous contractions and preserves linkage disequilibrium diversity from different samples through variant-specific local contractions, thereby improving the estimation accuracy of variant effect sizes. The population-matched linkage disequilibrium reference panel was derived from the UK Biobank genotype data, which included common genetic variants of HapMap3 common to both the 1000G Genome and the UK Biobank genotype datasets.
[0044] This method automatically learns all model parameters from genome-wide association study pooled data for effect size estimation in specific populations. It then integrates population-specific posterior SNP effects into a single population using inverse variance-weighted meta-analysis with a Gibbs sampler. The inverse variance-weighted approach eliminates the need for independent validation sets to train cross-population integration weights, enabling cross-population integration even with small samples lacking independent validation, thus addressing the current issue of excessively small East Asian population cohort sizes.
[0045] By integrating CES and GWAS data of atrial fibrillation from the same population into a novel combination and further integrating data from different populations across populations, this study overcomes the shortcomings of insufficient sample size in the aggregated data of single-trait and single-population disease-specific genome-wide association studies. It makes full use of the genetic information of large-scale correlated traits and the genetic information of large European populations, thereby enhancing the predictive effect of PRS in the underrepresented East Asian population.
[0046] In step S2, based on the obtained 40 clinical indicators and 4 genetic characteristics, a LASSO-regularized logistic regression model is used to predict CES in the training set. 100 LASSO analyses with 10-fold cross-validation are run with different seed numbers. Each time, the optimal result is selected from the variable set with the optimal λ value within one standard error range of the largest observed area under the receiver operating characteristic curve (AUC). This results in a sparse matrix of variables and their coefficients of 100×m (feature dimension). The optimal subset of variables consists of those with non-zero coefficients >80% in the 100 LASSO analyses, thus obtaining a highly correlated and stable set of predictive variables for CES, which is then used for subsequent modeling.
[0047] Preferably, in step S3, the STROMICS CES classifier mainly includes three models: Model 1: Age + Gender + Selected Genetic Trait; Model 2: Age + Gender + Selected Clinical Characteristics; Model 3: Age + Gender + Selected Genetic Trait + Selected Clinical Trait.
[0048] This invention also provides a model building system for assessing the probability of CES in cases of stroke of unknown cause, comprising: Data collection module: used for clinical feature collection and quality control, blood sample collection, sequencing and quality control; Genetic trait extraction and calculation module: used to extract individual SNPs based on whole genome sequencing data, calculate mitochondrial DNA copy number and PRS score, where PRS score is processed through multi-trait genetic data integration and analysis module and cross-population PRS construction module; Clinical and genetic feature selection module: Based on the candidate clinical features and genetic variables, and using LASSO to select features related to CES, the module selects variables through a coefficient sparse matrix formed by 100 LASSO operations, removes redundant features, and facilitates rapid acquisition of prediction results in clinical applications. The module for weight estimation and classification threshold determination constructs a logistic regression model for stroke samples with known causes based on the above optimal feature set, and defines the optimal classification threshold for the predicted probability according to the maximum value of the Youden exponent of the model.
[0049] The present invention also provides a system for assessing the probability of CES in cases of stroke of unknown cause, including The predictor set acquisition module is used to acquire data corresponding to the predictor set of patients with ischemic stroke. The identification module is used to input the data corresponding to the predicted variable set into the stroke classification model obtained by the construction method described above, and output the predicted probability of cardioembolic stroke. If the predicted probability is greater than the optimal classification threshold, it is judged as cardioembolic stroke; otherwise, it is judged as non-cardioembolic stroke. When the probability is greater than the set threshold, it indicates that the embolism in the stroke sample of unknown cause may be caused by cardioembolic stroke, and further detailed examination is needed to check for cardioembolic emboli in order to initiate early anticoagulation therapy.
[0050] The predictor variables include the following data: age (years); gender (male); current smoking status; systolic blood pressure (mmHg); creatinine (μmol / L); high-sensitivity C-reactive protein (mg / L); aspartate aminotransferase (U / L); and platelet count (×10⁻⁶). 9 / L; absolute value of monocyte population, unit ×10 9 / L; International Normalized Ratio (INR); Prothrombin Time (s); D-dimer (μg / mL); Triglycerides (mmol / L); Homocysteine (μmol / L); PRS AFThe probability of a sample is calculated based on the CES probability system described above, thereby assessing the probability of a sample with occult CES in a stroke of unknown cause.
[0051] This invention integrates detectable clinical factors with genetic information with fixed attributes from birth through multimodal data fusion to construct a highly accurate classifier. This not only fills the research gap in the Chinese population but can also be used to effectively identify occult CES in patients with unexplained stroke, thereby guiding clinical treatment decisions and improving patient prognosis. It overcomes the shortcomings of existing CES classifier construction methods, particularly the lack of studies based on Chinese population cohorts, the omission of birth-fixed genetic factors, reliance on self-reported medical history prone to recall bias, and high missing rate imaging features in remote areas.
[0052] The present invention will be further illustrated by specific embodiments below.
[0053] Example 1 like Figure 2 This embodiment utilizes the longitudinal dataset from the Chinese National Stroke Registry Study-III (CNSR-Ⅲ) to explain the STROMIC CES classifier construction method and system. The design process of this embodiment can be found in [link to documentation]. Figure 3 As shown, this classifier integrates 40 traditional clinical features and 4 genomic features, including genetic variations associated with CES, mitochondrial DNA copy number, and PRS of related phenotypes. This invention uses LASSO regularization to screen for predictive variables highly correlated with CES and employs logistic regression to estimate the weights, resulting in a model for identifying CES that contains a few predictive variables and possesses high discriminative power. This model is ultimately applied to the classification of stroke of unknown cause, and is expected to facilitate the identification of occult CES in stroke of unknown cause, with the ultimate goal of improving the treatment and prognosis of these patients.
[0054] The specific steps are as follows: Step S1: Acquisition of clinical features and whole-genome sequencing data This embodiment utilizes the longitudinal dataset from the China National Stroke Registry Study-III (CNSR-III) for model development and validation. The CNSR-III cohort is a nationwide prospective cohort study designed to investigate the etiology of stroke in the Chinese population. Between August 2015 and March 2018, the study recruited 15,166 patients diagnosed with acute ischemic cerebrovascular events from 201 hospitals. Trained staff collected demographic information, medical history, family history, physical examination results, and laboratory test results from patients' admission records. Biological samples were collected immediately after patient enrollment. The working group monitored the prognosis of all samples during a 12-month follow-up period.
[0055] The genetic data of the CNSR-III cohort were obtained by performing whole-genome sequencing (WGS) on 10,914 samples using the BGISEQ-500 platform, i.e., the STROMICS Genome Research Project. Cell Discov ;9(1):75;2023).
[0056] Step S2: Quality control of clinical features and whole-genome sequencing data This embodiment excluded 673 samples with substandard genotype data: 170 individuals were removed due to library construction failure and microbial contamination; 452 individuals were removed due to substandard genotype quality such as low coverage (<80%); 13 individuals were removed due to X chromosome heterozygosity abnormalities; and 38 individuals within the second degree of kinship were removed (PI_HAT>0.125) based on kinship analysis. A total of 10,241 quality-controlled samples were included in this invention. To conduct genome-wide association studies and calculate genetic risk, this invention excluded samples with low recall rates (genotype quality <20 in more than 5% of samples) or those deviating from Havenberg equilibrium (P<1×10⁻⁶). 6 Genetic variations were selected from 22 common variations (minor allele frequency >1%) on autosomes for subsequent implementation.
[0057] Ischemic stroke was classified according to the TOAST (Org 10172 trial in the treatment of acute stroke) criteria. Of the 10,241 individuals, 5,360 (52.3%) had a known cause of stroke: 646 (6.3%) were classified as CES, 2,546 (24.9%) as large artery atherosclerosis (LAA), and 2,168 (21.2%) as small artery occlusion (SAO). Of the remaining 4,881 stroke patients, 116 (1.1%) had other causes, while 4,765 (46.5%) had stroke of unknown cause.
[0058] This embodiment incorporates demographic information, physical examination results, and laboratory variables at admission to the CNSR-III cohort, including age, sex, smoking status (yes / no), alcohol consumption status (yes / no), height, weight, body mass index, blood glucose, blood pressure, blood lipids, hematological characteristics, coagulation tests, renal and liver function, and inflammatory factors. Preprocessing of clinical features included: ① deleting continuous variables with a missing percentage greater than 25%; ② performing a logarithmic transformation on right-skewed variables; ③ designating outliers exceeding 8 standard deviations from the mean as missing values; ④ imputing missing values using predicted mean matching, averaging the imputation results across 5 iterations, with each iteration running 10 times. After quality control, this embodiment ultimately retained 40 clinical features, including age, sex, smoking, alcohol consumption, body measurements, and laboratory indicators, for CES classifier construction (Table 1, where data represents the average of each indicator).
[0059] Table 1. Baseline characteristics of CES and non-CES patients in the CNSR-III cohort.
[0060] Note: Continuous variables are expressed as mean (standard deviation, SD) or median (interquartile range, IQR), and categorical variables are expressed as frequency (percentage). Continuous variables are tested using a two-sample t-test or the Mann-Whitney U test, and binary variables are tested using a chi-square (χ²) test.
[0061] Step S3: Multi-trait integration of CES and GWAS pooled data for atrial fibrillation. The PRS construction described in this invention is based on the multi-trait integration of CES and GWAS pooled data for atrial fibrillation. In this embodiment, the GWAS pooled statistics for CES are derived from the GWAS catalog (Genome-wide Association Study Outcomes Database), including large-scale CES genome-wide association study pooled data from two populations: 1) 238,168 East Asian samples, containing 926 CES cases (GCST90104546); 2) 1,245,612 European samples, containing 10,804 CES cases (GCST90104541). The large-scale GWAS pooled data for atrial fibrillation in this embodiment comes from: 1) the East Asian samples, consisting of 150,272 samples from the Japan Biobank, containing 9,826 atrial fibrillation cases; 2) the European samples, consisting of 1,840,341 samples, containing 228,926 atrial fibrillation cases (GCST90624412).
[0062] In this embodiment, the SNPs in the above-mentioned GWAS summary data are first filtered to exclude SNPs with a minor allele frequency of less than 1%, SNPs with a sample size lower than 75% of the 90th percentile of the SNP sample size distribution, and SNPs with abnormal effect size estimation.
[0063] This embodiment employs the multi-trait analysis for genome-wide association studies (MTAG) (Nat Genet. 2018;50(2):229-237), which uses the correlation between traits to weight marginal effects and integrates population-specific GWAS data of CES and atrial fibrillation into multi-trait summary data to improve the potential sample size and the accuracy of effect estimation.
[0064] Step S4: Construction of multi-trait cross-population PRS for CES and atrial fibrillation Based on the results of the multi-trait integration analysis in step S3 above, this embodiment uses the Bayesian multigene prediction method (Nat Genet. 2022;54(5):573-580) to estimate the specific variation weights of East Asian and European populations. Then, through cross-population meta-integration of the weights of the two populations, a cross-population PRS (PRS) for CES and atrial fibrillation is constructed. CES PRS AF This embodiment, based on the entire CNSR-III cohort sample (714 cases), uses a logistic regression model to evaluate the standardized PRS. AF The association with a history of atrial fibrillation was assessed. Similarly, the PRS was evaluated based on a sample of 5360 stroke patients with known etiologies from the CNSR-III cohort. CES and PRS AF Predict CES performance and compare it with PRS for two specific populations (by MTAG respectively). CES.EAS and MTAG CES.EUR The generated data were compared, as shown in Table 2. Simultaneously, the cross-population PRS (Population Regression Syndrome) was calculated by integrating the GWAS aggregated data from the five CES populations using a Bayesian multigene prediction method. CES.multi () is used for comparison and evaluation.
[0065] Step S5, Obtain and evaluate the genetic characteristics of the sample: Extract individual variant sites associated with CES, calculate mitochondrial DNA copy number and PRS. Single variant site associated with CES: This example performed a genome-wide association analysis on 646 CES cases and 4714 non-CES cases in the CNSR-III cohort. The results showed a significant signal at the 4q25 locus ( Figure 4 A peak association was detected at rs72900155 (A). P = 1.16×10-12 OR [C allele] = 0.61, 95% confidence interval: 0.54–0.70), neighboring genes are PITX2 The allele frequency in CES was 0.381, lower than that in the large artery atherosclerosis type (0.502) and the small vessel occlusion type (0.490). λ GC = 0.908, indicating that the CES GWAS data is of high quality and there is no obvious genome expansion.
[0066] Mitochondrial DNA copy number: In this embodiment, mitochondrial DNA copy number quantification was performed on WGS data from 10,241 patients in the CNSR-III cohort. Results showed that mtDNA-CN was significantly lower in the CES group than in the non-CES group ( Figure 4 (B)
[0067] PRS: This embodiment constructs a multi-trait cross-population PRS. CES and PRS AF As shown in Table 2, among the 10,241 patients in the CNSR-III cohort, PRS... AF In predicting atrial fibrillation, PRS was superior to three published polygenic risk scores for atrial fibrillation, with an odds ratio (OR) of 1.96 (1.81–2.12) and an AUC of 0.676 (0.655–0.697). In 5360 stroke patients with a definite etiology, PRS... CES It outperformed population-specific multi-trait PRS and cross-population PRS in predicting stroke, with an OR of 1.54 (1.41-1.67) and an AUC of 0.619 (0.595-0.643), but was slightly lower than PRS. AF (OR=1.67 [1.53-1.41]; AUC=0.639 [0.615-0.662]).
[0068] The distribution of PRS in known-cause stroke samples from the CNSR-III cohort is as follows: Figure 4 CD results show PRS CES and PRS AF The CES group had a significantly higher rate than the non-CES group.
[0069] Table 2 shows the prediction performance of PRS in the CNSR-III cohort.
[0070] Step S6, CES Feature Selection and Classifier Construction In this embodiment, three CES models, namely the STROMICS CES classifier, were constructed and tested in 5360 patients with known-cause stroke from the STROMICS genomic study data in the CNSR-III cohort. During the feature selection and model construction process, the training and test sets were first divided in a 2:1 ratio. Multiple machine learning models were trained on the training set and evaluated on the test set. Figure 5 As shown, after preliminary evaluation of various machine learning methods, LASSO (Least Absolute Shrinkage Operator) regularized logistic regression was ultimately selected because it demonstrated excellent performance in preliminary experiments and reduced model redundancy.
[0071] Finally, the LASSO method was used to screen genetic and clinical characteristics separately in the training set. The `cv.glmnet` function from the `glmnet` package in R was used to fit a LASSO-regularized logistic regression model to predict CES, and 10-fold cross-validation (repeated 100 times) was performed to optimize the hyperparameter λ. The standard for the optimal λ value for each result was: the area under the receiver operating characteristic curve (AUC) within one standard error of the maximum observation value. Variables with an 80% repetition rate in the 100-repeated model were ultimately used in the model construction.
[0072] In this embodiment, during the construction of the three models of the STROMICS CES classifier, a linear combination of the selected predictive factors is performed, and logistic regression is used to estimate the coefficients in the training set. The variables included in the three models of the STROMICS CES classifier and their estimated weights in the training set are shown in Table 3. Model 1 includes age, gender, and PRS. AF Model 2 includes 14 variables such as age, gender, and smoking, while Model 3 includes all the variables from Model 1 and Model 2, totaling 15 variables.
[0073] Table 3. Variables and coefficients included in the three models of the STROMICS CES classifier.
[0074] Step S7: Calculate the probability of an individual having CES in the test set samples and verify the model's predictive ability. In the model evaluation described in this embodiment, each model was evaluated using discriminative power (receiver operating characteristic [ROC] curve), area under the precision-recall curve (AUPR), calibration curve (calibration slope and intercept), and clinical utility (decision curve analysis) calculated on the test dataset. The roc.test function was used, based on 10... 4The statistical significance of differences in model AUC was assessed using a one-sided Z-test with replacement-guided sampling. The calibration slope and intercept were calculated by comparing the observed and predicted probabilities of the CES. The test dataset was divided into 10 groups based on the predicted probabilities, and the observed probabilities for each group were derived using the proportions of the CES. Decision curve analysis was performed using the R package rmda to evaluate the clinical utility of each model. The invention was compared with three published AF risk models—CHA2DS2-VASc, CHARGE-AF, and EHR-AF scores—on the test dataset and in 5360 stroke cases with known etiologies. The invention was evaluated based on the maximum Youden index of each model in the training dataset (e.g., ...). Figure 6 The classification threshold for predictions in the new dataset was defined. Several metrics were used to evaluate the performance of each model, including accuracy, sensitivity (recall), specificity, Cohen's Kappa score, and Brier score.
[0075] like Figure 7 As shown, Model 3 outperforms Models 1 and 2. Model 1 has an AUC of 0.728 (0.690-0.765), Model 2 has an AUC of 0.774 (0.739-0.808), while Model 3, combining the variables of Models 1 and 2, achieves an AUC of 0.800 (0.767-0.833). Model 3 also demonstrates superior AUPR, calibration curve, and decision curve analysis results. As shown in Table 4, Model 3 outperforms Model 3 in Brier score, accuracy, sensitivity (recall), specificity, and Cohen's Kappa index.
[0076] Table 4 shows the performance of different classifiers on the test set and the dataset of stroke with all known causes.
[0077] This embodiment uses the R function kernelshap to calculate the Shapley additive interpretation (SHAP value) of the variables included in the best CES classifier (Model 3) to understand the importance of features and the cumulative impact of different data modalities on prediction. Figure 8 As shown in AB, PRS AF The primary predictor is PRS, followed by the international normalized ratio (INR), age, sex, and platelet count (PLT). AFScores, INR, creatinine, absolute monocyte population, high-sensitivity C-reactive protein, D-dimer, prothrombin time, and age were associated with higher SAP values, indicating that their predictive utility could increase the likelihood of CES. Conversely, higher platelet counts, triglycerides, systolic blood pressure, and homocysteine were associated with lower SAP values, thus decreasing the predictive likelihood of CES. Women were more likely to predict CES than men. Furthermore, this example provides one CES sample and one non-CES sample to illustrate the effect of different variable levels on the outcome. Figure 8 Medium CD).
[0078] Example 2 Identifying occult CES in patients with unexplained stroke using an optimal classifier This implementation used the optimal CES classifier model 3 to divide 4765 patients with unexplained stroke in the CNSR-III cohort into a latent CES group and a non-CES group. In patients with unexplained stroke, this invention compared baseline characteristics, follow-up outcomes, and anticoagulation therapy efficacy between the predicted latent CES group and the non-CES group. Baseline characteristics included demographics, National Institutes of Health Stroke Scale (NIHSS) score, medical history, and medication status at admission. Follow-up outcomes included recurrence rate, mortality, and changes in the modified Rankin Scale (mRS) after discharge (ΔmRS). ΔmRS represents the difference between the mRS at 3 months and 12 months after discharge and the mRS at admission. Cumulative mortality was calculated using the survfit.coxph function, and the predicted hazard ratio (HR) for the CES group relative to the non-CES group was estimated using a Cox model, adjusted for age, stroke history, and NIHSS at discharge. To assess the predictive efficacy of anticoagulation therapy in patients with stroke of unknown cause, stroke patients were divided into anticoagulation and non-anticoagulation groups based on their admission records, and the changes in the modified Rankin Scale (mRS) at 3 months and 12 months of follow-up after discharge (ΔmRS) were compared.
[0079] As shown in Table 5, the clinical characteristics and prognosis of patients with unexplained stroke predicted as CES are similar to those of patients with ischemic stroke of known etiology. Specifically, patients with unexplained stroke predicted as CES have lower rates of smoking, alcohol consumption, and family history of stroke compared to those without CES; they also have higher ages at admission, higher NIHSS scores, and higher prevalence of atrial fibrillation; the proportion of patients receiving anticoagulation therapy during hospitalization, the mortality rate at follow-up, and the mean remission rate (mRS) are significantly higher in the CES group compared to those without CES. Cases predicted as occult CES have worse functional recovery at follow-up (ΔmRS<0) compared to patients without CES. At 12 months post-discharge, the proportion of patients with ΔmRS<0 in the predicted CES group was 51.92%, significantly lower than that in the non-CES group (57.76%). The mortality rate in the predicted CES group was 6.75%, significantly higher than that in the non-CES group (2.06%). Figure 9 As shown in Figure A, after adjusting for age, stroke history, and NIHSS at discharge, the mortality rate of cases predicted as occult CES was significantly higher than that of non-CES cases, with a hazard ratio of 1.59 (1.08–2.32). Figure 9 As shown in Figures B and C, among patients predicted to have occult CES, 156 (14.62%) received anticoagulation therapy at admission. At 12 months post-discharge, 60% of patients on anticoagulation therapy showed disability recovery (ΔmRS<0), significantly higher than those on non-anticoagulation therapy. Among patients predicted to have non-CES, 263 (7.11%) received anticoagulation therapy at admission. Changes in mRS showed that the proportion of patients on anticoagulation therapy with disability recovery (ΔmRS<0) was comparable to that on non-anticoagulation therapy. In contrast, as... Figure 10 As shown, the EHR-AF model failed to identify the difference in prognostic benefit of anticoagulation therapy between its CES group and the non-CES group in patients with unexplained stroke.
[0080] This invention is the first to construct a CES classifier in the Chinese population. The above results fully demonstrate the superior performance of the CES classifier of this invention, and its reliability in identifying occult CES from samples of stroke of unknown cause.
[0081] Table 5. Comparison of features in the CNSR-III cohort divided into CES and non-CES groups.
[0082] Note: Continuous variables are expressed as mean (standard deviation, SD) or median (interquartile range, IQR), and categorical variables are expressed as frequency (percentage). Continuous variables are tested using a two-sample t-test or the Mann-Whitney U test, and binary variables are tested using a chi-square (χ²) test. ΔmRS represents the difference between the mRS score at 3 months and 12 months post-discharge and the mRS score at admission (mRS...). discharge - mRS admission ).
[0083] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for constructing a classifier for cardioembolic stroke, characterized in that, include: S1. Collect age, sex, clinical data and genetic characteristics of patients with ischemic stroke as predictive variables, and construct a dataset with the corresponding etiology; the clinical data include: smoking status, drinking status, height, weight, body mass index, blood glucose, blood pressure, blood lipids, hematological characteristics, coagulation test, kidney function and liver function, and inflammatory factors; The genetic characteristics include: genotype of a single variant site associated with cardioembolic stroke, mitochondrial DNA copy number, and genetic risk scores for cardioembolic stroke and atrial fibrillation; the genetic risk scores are calculated through the following steps: S1.
1. SNP filtering was performed on the genome-wide association study (GWAS) data of East Asian and European samples of cardiac stroke and atrial fibrillation. A joint analysis of multi-trait GWAS studies was conducted. The genetic correlation between traits was used to weight the trait-specific GWAS studies of cardiac stroke and atrial fibrillation to increase the effective sample size. Multi-trait integrated summary data of cardiac stroke and atrial fibrillation were obtained by weighting the data by the covariance model. S1.2 Based on the integrated summary data of multiple traits of cardiac stroke and atrial fibrillation, the Bayesian multigene prediction method is used to couple the genetic structure shared by different populations to estimate the specific variation weights of East Asian and European populations. The variation weights of the two populations are then integrated through inverse variance weighted meta-analysis to construct cross-population genetic risk prediction models for cardiac stroke and atrial fibrillation respectively. S1.3 Based on the genotype data of the target population, the genetic risk of the trait is quantified into digital molecular markers by weighted summation of the variation weights of the cross-population genetic risk prediction model, and the genetic risk scores of cardiogenic stroke and atrial fibrillation are calculated. The construction of the dataset includes setting the output label corresponding to the cause of cardiac stroke as 1, and setting the output label corresponding to other known causes that are not cardiac stroke as 0. S2. Based on the sample set of known causes of ischemia in the dataset, a LASSO-regularized logistic regression model is trained, and the optimal regularization parameter λ is selected through cross-validation strategy to finally obtain a sparsed set of predictive variables. S3. Construct a training set by combining the set of predictive variables with the known causes of ischemic stroke (whether the patient's cause is cardioembolic stroke). Use a logistic regression model to train the weights of the predictive variable set. Define the optimal classification threshold for the predicted probability based on the maximum value of the Youden index on the training set. If the predicted probability is greater than the optimal classification threshold, it is cardioembolic stroke; otherwise, it is judged as non-cardioembolic stroke, thus obtaining the cardioembolic stroke classifier.
2. The method for constructing a cardioembolic stroke classifier according to claim 1, characterized in that, In step S2, the optimal λ value that is within one standard error range of the maximum observation value under the receiver operating characteristic curve is selected as the optimal regularization parameter λ for each cross-validation strategy.
3. The method for constructing a cardioembolic stroke classifier according to claim 1, characterized in that, In step S2, the sparsed predictive variable set includes the age, sex, and clinical data of ischemic stroke patients, including smoking status, systolic blood pressure, creatinine, high-sensitivity C-reactive protein, aspartate aminotransferase, platelet count, absolute value of monocyte population, prothrombin time, D-dimer, triglycerides, and homocysteine. Or it could include the age, sex, and genetic risk score of atrial fibrillation in patients with ischemic stroke.
4. The method for constructing a cardioembolic stroke classifier according to claim 3, characterized in that, In step S2, the sparsed predictive variable set includes age, sex, and clinical data such as smoking status, systolic blood pressure, creatinine, high-sensitivity C-reactive protein, aspartate aminotransferase, platelet count, absolute value of monocyte population, prothrombin time, D-dimer, triglycerides, and homocysteine, as well as a genetic risk score for atrial fibrillation from a genetic perspective.
5. The method for constructing a cardioembolic stroke classifier according to any one of claims 1-4, characterized in that, The formula for calculating the predicted probability is as follows: Where, P(Y=1|X1,…,+X) m The value represents the probability of developing a cardiac stroke, where Y=1 indicates a cardiac stroke, and m is the number of predictive variables selected for calculating the incidence of cardiac stroke. Represents the first of all samples m The values of the predictor variables, Indicates the first m The weights of the predictor variables are estimated by logistic regression.
6. A method for differentiating cardiogenic stroke, characterized in that, include: Obtain the data corresponding to the predictor variable set of patients with ischemic stroke, input the cardiogenic stroke classifier obtained by the construction method described in any one of claims 1-5, output the predicted probability of cardiogenic stroke, if the predicted probability is greater than the optimal classification threshold, it is judged as cardiogenic stroke, otherwise it is judged as non-cardiogenic stroke.
7. A system for differentiating cardiogenic stroke, characterized in that, The method for differentiating cardiogenic stroke according to claim 6 includes: The predictor set acquisition module is used to acquire data corresponding to the predictor set of patients with ischemic stroke. The identification module is used to input the data corresponding to the predicted variable set into the cardioembolic stroke classifier obtained by the construction method according to any one of claims 1-5, output the predicted probability of cardioembolic stroke, and judge it as cardioembolic stroke if the predicted probability is greater than the optimal classification threshold, otherwise judge it as non-cardioembolic stroke.
8. A computer-readable storage medium, characterized in that, It stores a computer program, which, when executed by a processor, implements the method for constructing a cardiac stroke classifier as described in any one of claims 1-5, or the method for identifying cardiac stroke as described in claim 6.
Citation Information
Patent Citations
Genetic markers for risk management of atrial fibrillation and stroke
CN102449165A
Device for predicting attack risk of cerebral apoplexy and application
CN118638915A