Cerebral arterial thrombosis diagnosis model based on neutrophil extracellular trap genes
By integrating NETs-related genes and immune infiltration characteristics, a high-sensitivity ischemic stroke diagnosis model is constructed, which solves the problems of low diagnostic lag and accuracy in the prior art, and achieves more efficient diagnosis and individualized treatment.
Patent Information
- Application Number
- CN202510463889.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-22
AI Technical Summary
The prior art has problems of diagnostic lag, low diagnostic accuracy and failure to fully reflect complex immune mechanisms in the diagnosis of ischemic stroke, and there is a lack of in-depth research on neutrophil extracellular trapping nets (NETs) in ischemic stroke.
By integrating NETs-related genes and immune infiltration characteristics, a high-sensitivity ischemic stroke diagnostic model is constructed, a multiomics data fusion and machine learning algorithm are used to screen core genes, a Nomo graph scoring system is constructed, and a single-cell sequencing and immune infiltration analysis is combined to identify key regulatory factors and targeted drugs.
It improves the diagnostic sensitivity and accuracy of ischemic stroke, provides new targets for immunotherapy, broadens the understanding of pathological mechanisms, and improves clinical applicability and diagnostic performance.
Smart Images

Figure CN120356653A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of clinical medicine, and specifically relates to an ischemic stroke diagnosis model based on neutrophil extracellular trap genes. Background Art
[0002] Stroke, commonly known as a stroke, is divided into two types: ischemic stroke and hemorrhagic stroke. Ischemic stroke (IS) refers to the general term for necrosis of brain tissue caused by stenosis or occlusion of the blood supply arteries of the brain (carotid artery and vertebral artery) and insufficient blood supply to the brain. It has the characteristics of high incidence, disability rate, recurrence rate, and mortality rate, and is the leading cause of death among Chinese residents.
[0003] Currently, in the diagnosis and prognosis evaluation of ischemic stroke, cranial CT and MRI scans have a diagnostic lag and are difficult to capture early pathological changes in the acute phase; the biomarkers in biomarker detection are single (such as MMP9), and cannot comprehensively reflect the complex immune mechanism of IS, resulting in low diagnostic accuracy; gene research is mostly based on simple screening of differentially expressed genes (DEGs), without integrating multi-omics data, and without using machine learning to optimize the model, there are great limitations.
[0004] Moreover, traditional pathological mechanism research focuses on macroscopic processes such as atherosclerosis and thrombosis, lacking in-depth analysis of the immune microenvironment; neutrophil extracellular traps (NETs) are a special structure formed after neutrophil necrosis or apoptosis. Studies in cardiovascular diseases have shown that NETs are related to various diseases, and are more closely related to the pathogenesis of diseases such as coronary artery disease and myocardial infarction.
[0005] However, the role of NETs in IS has not been fully explored at present, lacking the correlation analysis with immune cell infiltration. In view of this, the present invention is specifically proposed. Summary of the Invention
[0006] The purpose of the present invention is to provide a highly sensitive IS diagnosis model by integrating NETs-related genes and immune infiltration characteristics to solve the above problems.
[0007] To achieve the above purpose, the present invention provides the following technical solutions: An ischemic stroke diagnosis model based on neutrophil extracellular trap genes, characterized by including the following steps: Step S1: Obtain the peripheral blood gene expression dataset of ischemic stroke patients from a public gene expression database, and perform data preprocessing to eliminate batch effects; Step S2: Screen neutrophil extracellular trap (NETs)-related genes, intersect them with differentially expressed genes (DEGs) to obtain NETs-IS candidate genes, and verify the causal association between NETs-related genes and ischemic stroke and its subtypes through Mendelian randomization analysis; Step S3: Use machine learning algorithms to screen out several core genes from the NETs-IS candidate genes; Step S4: Integrate the core gene expression data using logistic regression to construct a nomogram scoring system as a diagnostic model for ischemic stroke; Step S6: Use the immune infiltration analysis tool CIBERSORT to analyze the immune cell composition of ischemic stroke patients to further understand the mechanism of action of NETs-related genes in ischemic stroke; perform single-cell sequencing verification, and confirm the enriched expression of core genes in specific immune cells through multi-omics data integration; Step S6: Use the immune infiltration analysis tool CIBERSORT to analyze the immune cell composition of ischemic stroke patients to further understand the mechanism of action of NETs-related genes in ischemic stroke; perform single-cell sequencing verification, and confirm the enriched expression of core genes in specific immune cells through multi-omics data integration; Step S7: Construct mRNA-TF and mRNA-miRNA regulatory networks to identify key regulatory factors, and screen for targeted drugs through the DGIdb database to provide guidance for the individualized treatment of ischemic stroke.
[0008] Furthermore, the NETs-related genes are screened through comprehensive analysis of GeneCards, GWAS Catalog, and literature, totaling 349; the differentially expressed genes (DEGs) are obtained by comparing the gene expression data of ischemic stroke patients and healthy control individuals.
[0009] Furthermore, the core genes include C5AR1, MMP9, LRG1, TREM1, NFIL3, PGLYRP1, and PADI4.
[0010] Furthermore, the machine learning algorithm in Step S3 is the Boruta algorithm, which is a wrapper feature selection method based on random forests.
[0011] Furthermore, the Mendelian randomization analysis in Step S2 includes the following steps: 1. Data acquisition and instrumental variable screening: Obtain the genetic summary data of IS and its subtypes from the publicly available genome-wide association study (GWAS) database; Screen single nucleotide polymorphisms (SNPs) significantly associated with gene expression from the cis - expression quantitative trait locus (cis - eQTL) and cis - protein expression quantitative trait locus (cis - pQTL) databases as instrumental variables. The logFC threshold for differential gene screening is logFC > 1, p < 0.05; 2. Two - sample MR analysis: Use inverse - variance weighting (IVW), MR Egger regression, and weighted median to evaluate the causal effects of candidate genes; Quantify the association strength between gene expression levels and disease risk through the odds ratio (OR) and 95% confidence interval (CI); 3. Sensitivity analysis and validation: Use Cochran's Q - test to evaluate the heterogeneity of instrumental variables; Exclude pleiotropy bias through the MR Egger intercept test.
[0012] The beneficial effects of the present invention are as follows: High sensitivity, good diagnostic effect, convert complex gene expression data into a visual scoring tool, and improve clinical applicability; integrate gene expression, immune infiltration, and single - cell sequencing data through multi - omics data fusion to comprehensively analyze the pathological mechanism of IS, and screen high - specificity gene combinations through the Boruta algorithm and CytoHubba, breaking through the limitations of traditional single - gene research. At the same time, for the first time, it reveals that NETs affect the progression of IS by regulating neutrophil infiltration, providing new targets for immunotherapy, and showing a significant improvement compared with the prior art.
[0013] The above description is only an overview of the technical solution of the present invention. In order to understand the technical means of the present invention more clearly and be able to implement it according to the content of the specification, the following describes in detail with preferred embodiments of the present invention in conjunction with the accompanying drawings. Brief description of the drawings
[0014] Figure 1 It is a schematic structural diagram of an ischemic stroke diagnosis model based on neutrophil extracellular trap genes shown in an embodiment of the present invention; Figure 2 It is the relative expression level of the gene mRNA level detected by qRT - PCR technology shown in an embodiment of the present invention. Detailed implementation manners
[0015] The technical solution of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0016] In addition, the technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other. Example 1
[0017] Please refer to Figure 1 , a diagnostic model for ischemic stroke based on neutrophil extracellular trap network genes shown in a preferred embodiment of this application, includes the following steps: Step 1: Data integration and preprocessing Obtain the peripheral blood gene expression datasets (GSE16561, GSE22255, GSE198710) of IS patients from the GEO database, and eliminate batch effects through the Combat algorithm.
[0018] Screen 349 NETs-related genes (from GeneCards, GWAS Catalog, and literature), take the intersection with differentially expressed genes (DEGs) to obtain 10 NETs-IS candidate genes, and verify the causal association between neutrophil extracellular trap (NETs)-related genes and ischemic stroke (IS) and its subtypes through Mendelian Randomization (MR) analysis.
[0019] Step 2: Machine learning to screen key genes Use the Boruta algorithm (based on random forest) to screen out 6 diagnostic genes (C5AR1, MMP9, LRG1, NFIL3, TREM1, PADI4) from 10 candidate genes, and supplement PGLYRP1 by combining CytoHubba analysis to finally determine 7 core genes.
[0020] Step 3: Construct a nomogram diagnostic model Use logistic regression to integrate the 7 gene expression data, construct a nomogram scoring system, and verify the model performance through the ROC curve (AUC = 0.825).
[0021] Verify the model generalization ability with an independent dataset (GSE58294) (AUC = 0.834).
[0022] Step 4: Immune infiltration and single-cell sequencing analysis Use CIBERSORT to analyze the immune cell composition of IS patients and find that CD8+ T cells and NK cells decrease, while monocytes and neutrophils increase.
[0023] The single-cell sequencing data comes from GSE225948, and analyze the enriched expression of 7 genes in specific immune cells (such as neutrophils and basophils).
[0024] Step 5: Regulatory Network and Drug Target Mining Construct mRNA-TF and mRNA-miRNA regulatory networks to identify key regulatory factors (such as hsa-miR-15b-5p).
[0025] Screen for targeted drugs through the DGIdb database (Drug-Gene Interaction database) (such as AVACOPAN targets C5AR1 and PRINOMASTAT targets MMP9).
[0026] Specifically, the Mendelian randomization analysis in step S1 includes the following steps: 1. Data acquisition and instrumental variable screening Genetic exposure data: Obtain the genetic summary data of IS and its subtypes (large artery stroke LAS, cardioembolic stroke CES, small vessel stroke SVS) from publicly available genome-wide association study (GWAS) databases, covering tens of thousands of cases and controls.
[0027] Instrumental variable selection: For candidate NETs-IS genes, screen single nucleotide polymorphisms (SNPs) significantly associated with gene expression (p < 1×10⁻ 5 ) from gene expression quantitative trait locus (cis-eQTL) and protein expression quantitative trait locus (cis-pQTL) databases as instrumental variables, and ensure that they meet the three core assumptions of MR (strong correlation, independence, exclusivity).
[0028] 2. Two-sample MR analysis Adopt five statistical methods, such as inverse variance weighting (IVW), MR Egger regression, and weighted median, to evaluate the causal effects of candidate genes on IS and its subtypes.
[0029] Quantify the association strength between gene expression levels and disease risks through odds ratio (OR) and 95% confidence interval (CI).
[0030] 3. Sensitivity analysis and validation Heterogeneity test: Use Cochran's Q test to evaluate the heterogeneity of instrumental variables (p < 0.05 is considered significant).
[0031] Horizontal pleiotropy test: Exclude pleiotropy bias through the MR Egger intercept test (p < 0.05 is considered to have bias).
[0032] Key results and innovative supports: OLFM4 gene: Significantly reduces the risk of large artery stroke (LAS) at the protein level (MR Egger p = 0.019, OR = 0.929), indicating its protective effect and not being a core gene.
[0033] The PGLYRP1 gene: significantly increases the risk of cardioembolic stroke (CES) (IVW p = 0.034, OR = 1.039), suggesting its potential as a risk marker.
[0034] The remaining candidate genes did not show significant causal associations at the gene expression level. Further screening for high-confidence core genes (such as C5AR1, MMP9, etc.) ensured the specificity of the diagnostic model.
[0035] Technical advantages: Verifying the association between NETs-related genes and IS subtypes at the causal level through MR analysis, breaking through the limitations of traditional correlation studies, providing genetic evidence for screening highly reliable biomarkers, and significantly enhancing the pathological mechanism interpretability and clinical applicability of the model.
[0036] It should be noted that in this example, GWAS and eQTL / pQTL data are both from public databases (such as MEGASTROKE, GTEx), which comply with data use agreements and ethical norms; and the MR analysis follows international general statistical standards (such as the STROBE-MR guidelines) to ensure the reproducibility and scientificity of the results. Example 2
[0037] This example provides an application of detecting specific gene expression profiles in the differential diagnosis of cerebrovascular diseases based on qRT-PCR technology. Please refer to Figure 2 ; In the screening of candidate genes for the pathological mechanism of stroke, four gene markers, MMP9, OLFM4, TREM1, and PGLYRP1, were detected to have statistically significant differential expression (p < 0.05) between the control group and the stroke patient group, and their expression characteristics were specifically associated with the clinical subtypes of stroke (large artery atherosclerosis stroke LAS, cardioembolic stroke CES, small vessel stroke SVS). Specifically, the MMP9 gene showed persistent upregulation relative to the control group in all stroke subtypes, and there was no significant difference in its expression level among the LAS, CES, and SVS subtypes. The OLFM4 gene was specifically detected to have downregulated expression relative to the control group, the CES subtype, and the SVS subtype in the LAS subtype. The PGLYRP1 gene showed a pan-upregulated expression pattern in the stroke population, and a significant increase in expression was detected in the CES subtype relative to the LAS and SVS subtypes. The overexpression characteristic of the TREM1 gene specifically existed in the SVS subtype samples, showing a significant distinction from the control group and other stroke subtypes. It is worth noting that no significant differential expression was detected for the C5AR1, LRG1, NFIL3, and PADI4 genes between the control group and the stroke group and among the various clinical subtypes.
[0038] This discovery indicates that the biomarker combination composed of the above four differentially expressed genes can effectively distinguish the pathological status and subtype classification of stroke, and has important application value in the development of in vitro diagnostic reagents and the field of precision medicine.
[0039] In summary, it can be seen that the AUC of the model shown in the present invention reaches 0.825 in the training set and 0.834 in the validation set, which is significantly better than that of a single biomarker (such as MMP9), and the diagnostic performance is significantly improved. At the same time, novel IS-related genes such as PGLYRP1 and NFIL3 are discovered, broadening the understanding of pathological mechanisms, and potential drugs such as C5AR1 inhibitors and PADI4 antagonists are identified, promoting individualized treatment. In addition, it is also revealed that NETs drive immune cell infiltration through the IL-17 / TNF signaling pathway, providing a theoretical basis for intervention, which is significantly improved compared with the prior art.
[0040] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0041] The above-described embodiments merely represent several implementation manners of the present invention, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the invention patent shall be subject to the appended claims.
Claims
1. An ischemic stroke diagnosis model based on neutrophil extracellular trap network genes, characterized in that, It includes the following steps: Step S1: Obtain the peripheral blood gene expression dataset of ischemic stroke patients from a public gene expression database, and perform data preprocessing to eliminate batch effects; Step S2: Screen genes related to neutrophil extracellular traps (NETs), take the intersection with differentially expressed genes (DEGs), obtain NETs-IS candidate genes, and verify the causal association between genes related to neutrophil extracellular traps (NETs) and ischemic stroke and its subtypes through Mendelian randomization analysis; Step S3: Use machine learning algorithms to screen out several core genes from the NETs-IS candidate genes; Step S4: Integrate the core gene expression data using logistic regression to construct a nomogram scoring system as a diagnostic model for ischemic stroke; Step S5: Verify the model performance through the ROC curve to ensure its high sensitivity and specificity; Step S6: Use the immune infiltration analysis tool CIBERSORT to analyze the immune cell composition of ischemic stroke patients to further understand the mechanism of action of NETs-related genes in ischemic stroke; perform single-cell sequencing verification, and confirm the enriched expression of core genes in specific immune cells through multi-omics data fusion; Step S7: Construct mRNA-TF and mRNA-miRNA regulatory networks, identify key regulatory factors, and screen targeted drugs through the DGIdb database to provide guidance for the individualized treatment of ischemic stroke.
2. The ischemic stroke diagnosis model based on neutrophil extracellular trap genes according to claim 1, characterized in that The NETs-related genes are screened by comprehensively integrating GeneCards, GWAS Catalog and literature, totaling 349; the differentially expressed genes (DEGs) are obtained by comparing and analyzing the gene expression data of ischemic stroke patients and healthy control individuals.
3. The ischemic stroke diagnosis model based on neutrophil extracellular trap genes as described in claim 1, characterized in that, The core genes include C5AR1, MMP9, LRG1, TREM1, NFIL3, PGLYRP1, PADI4.
4. The ischemic stroke diagnosis model based on neutrophil extracellular trap genes according to claim 1, wherein The machine learning algorithm in Step S3 is the Boruta algorithm, which is a wrapper feature selection method based on random forest.
5. The ischemic stroke diagnosis model based on neutrophil extracellular trap genes according to claim 1, characterized in that, The Mendelian randomization analysis in Step S2 includes the following steps:
1. Data acquisition and instrumental variable screening: Obtain the genetic summary data of IS and its subtypes from the public genome-wide association study (GWAS) database; Screen single nucleotide polymorphisms (SNPs) significantly related to gene expression from the gene expression quantitative trait locus (cis-eQTL) and protein expression quantitative trait locus (cis-pQTL) databases as instrumental variables. The logFC threshold for differential gene screening is logFC > 1, p < 0.05; 2. Two-sample MR analysis: Use inverse variance weighting (IVW), MR Egger regression, and weighted median to evaluate the causal effects of candidate genes; Quantify the association strength between gene expression levels and disease risk through the odds ratio (OR) and 95% confidence interval (CI); 3. Sensitivity analysis and verification: Use Cochran's Q test to evaluate the heterogeneity of instrumental variables; Exclude pleiotropy bias through the MR Egger intercept test.