Auxiliary diagnosis system for ischemic stroke based on peripheral blood T cell single cell transcriptome and application of auxiliary diagnosis system
By constructing an auxiliary diagnostic system based on the single-cell transcriptome of peripheral blood T cells, and utilizing nonlinear weighted fusion and distribution sensing, the system addresses the issues of insufficient sensitivity and specificity in the early diagnosis of ischemic stroke, achieving high discriminative power and individualized risk assessment, thereby improving diagnostic accuracy and sensitivity.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-21
AI Technical Summary
Existing technologies lack sensitivity and specificity in the early diagnosis of ischemic stroke, lack precise molecular-level assessment tools, cannot quantify risk signals within immune cells, and traditional biomarkers are insufficient to reflect complex cellular heterogeneity and the contribution of cells at extreme risk.
A peripheral blood T cell single-cell transcriptome-based auxiliary diagnostic system introduces genes corresponding to proteins significantly associated with ischemic stroke, combines nonlinear weighted fusion and distribution sensing, and constructs a single-cell risk scoring model to achieve high-discrimination auxiliary diagnosis.
It improves the diagnostic sensitivity and specificity of ischemic stroke, can identify abnormal immune signals, eliminate interference from non-specific factors between individuals, provide individualized risk assessment, help improve diagnostic accuracy, and support batch sample processing and individualized score output.
Smart Images

Figure CN121905477A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biomedical detection technology, specifically to an auxiliary diagnostic system for ischemic stroke based on peripheral blood T cell single-cell transcriptome and its application. Background Technology
[0002] Ischemic stroke is a major cause of death and disability worldwide, and its early diagnosis and risk prediction are crucial for clinical intervention. In recent years, with the development of single-cell sequencing technology, the molecular characteristics of peripheral blood immune cells are considered to reflect the body's inflammatory state, vascular endothelial function, and stroke-related pathological changes. Among these, T cells, as an important component of peripheral immunity, play a key role in the occurrence and development of stroke.
[0003] Currently, early diagnosis of ischemic stroke in clinical practice mainly relies on imaging examinations (such as MRI and CT) and some serological indicators. However, these methods have the following shortcomings: 1. Insufficient sensitivity and specificity: Conventional inflammatory or thrombotic indicators (such as D-dimer and CRP) lack stroke specificity, making it difficult to accurately identify high-risk groups in the early stages; 2. Lack of precise molecular-level assessment tools: Existing diagnostic protocols do not fully utilize peripheral blood single-cell transcriptomics information and cannot quantify risk signals within immune cells; 3. Single scoring method and lack of individualized adjustment: Traditional biomarkers often use simple weighting or threshold judgments to stratify risk, which is difficult to reflect complex cellular heterogeneity and the contribution of extreme risk cells.
[0004] Therefore, there is an urgent need for a novel diagnostic model based on the single-cell transcriptome of peripheral blood T cells to enable early screening, auxiliary diagnosis, and individualized risk assessment of ischemic stroke. Summary of the Invention
[0005] To address the problems existing in the background technology, the present invention provides an auxiliary diagnostic system for ischemic stroke based on peripheral blood T cell single-cell transcriptome and its application. This auxiliary diagnostic system utilizes peripheral blood T cell single-cell transcriptome sequencing data, introduces the corresponding genes of proteins significantly related to ischemic stroke, and combines nonlinear weighted fusion and distribution sensing to achieve high discrimination ability for ischemic stroke, with high sensitivity and specificity.
[0006] The technical solution of the present invention to solve the above-mentioned technical problems is as follows: In a first aspect, the present invention provides an auxiliary diagnostic system for ischemic stroke based on the single-cell transcriptome of peripheral blood T cells, comprising: Peripheral blood mononuclear cell acquisition module, used to isolate peripheral blood mononuclear cells from peripheral blood samples of subjects; Single-cell RNA sequencing module is used to acquire transcriptome data of a single T cell in peripheral blood mononuclear cells; The data processing module targets the gene encoding a protein molecule that is significantly associated with the occurrence of ischemic stroke. The data processing module is used to obtain the expression level of the target gene in each T cell. The data analysis module includes a single-cell risk scoring model. This module obtains a risk score for each T cell based on the expression level of the target gene in the subject, and then uses distribution-aware recognition to non-linearly weight and fuse all T cell risk scores to obtain an individualized risk score for the subject. The single-cell risk scoring model is constructed as follows: protein molecules significantly associated with ischemic stroke are screened from a database, and their corresponding odds ratios are calculated to determine risk weights, thus constructing the single-cell risk scoring model. The outcome discrimination module categorizes subjects into ischemic stroke negative and positive based on their individualized risk scores.
[0007] By utilizing distribution perception to identify and analyze population distribution characteristics, higher weight coefficients are adaptively assigned to high-risk cell populations, and risk scores of all cells are weighted and fused based on these weight coefficients.
[0008] According to the above scheme, the protein molecules that are significantly associated with the occurrence of ischemic stroke are CD25, SOD2, CXCR3, GPX1, CSF2, IL-17, GPX7, IL-10, PD-1, and HLA-DR.
[0009] According to the above scheme, the risk weights for CD25, SOD2, CXCR3, GPX1, CSF2, IL-17, GPX7, IL-10, PD-1, and HLA-DR are 0.026, 0.031, 0.091, 0.18, 0.029, 0.171, -0.04, -0.101, -0.21, and -0.048, respectively. The single-cell risk scoring model is as follows: Riskscore = (0.026 × CD25 gene expression level) + (0.031 × SOD2 gene expression level) + (0.091 × CXCR3 gene expression level) + (0.18 × GPX1 gene expression level) + (0.029 × CSF2 gene expression level) + (0.171 × IL-17 gene expression level) + (-0.04 × GPX7 gene expression level) + (-0.101 × IL-10 gene expression level) + (-0.21 × PD-1 gene expression level) + (-0.048 × HLA-DR gene expression level).
[0010] According to the above scheme, the data processing module is also used to normalize the transcriptome data of single cells, converting the expression value of each target gene into a standardized signal in the 0-1 range.
[0011] According to the above scheme, the data analysis module nonlinearly maps the individualized risk score to the 0-100 range.
[0012] According to the above scheme, a certain threshold is set in the result discrimination module. If the individualized risk score is greater than or equal to the threshold, it is judged as a positive ischemic stroke, and if the individualized risk score is less than the threshold, it is judged as a negative ischemic stroke.
[0013] According to the above scheme, the threshold in the result discrimination module is set to 50 points.
[0014] According to the above protocol, the number of peripheral blood T cells to be sequenced is 1500-5000.
[0015] Secondly, the present invention provides the application of the above-mentioned auxiliary diagnostic system for ischemic stroke based on peripheral blood T cell single-cell transcriptome in the auxiliary diagnosis of ischemic stroke.
[0016] According to the above scheme, the area under the ROC curve for ischemic stroke based on peripheral blood T cell single-cell transcriptome reaches 0.915.
[0017] The beneficial effects of this invention are: This invention's auxiliary diagnostic system can identify ischemic stroke-related immune abnormalities based on peripheral blood T cell single-cell transcriptome data, effectively overcoming the bottleneck of insufficient sensitivity in early detection of serum biomarkers. Simultaneously, through single-cell-level quantitative analysis, it can eliminate interference from non-specific factors between individuals, achieving real-time and accurate characterization of individual immune status, providing quantitative auxiliary diagnostic evidence for ischemic stroke in clinical practice. The single-cell risk scoring model incorporates weighted summaries of target genes to simultaneously characterize factors that promote and inhibit ischemic stroke, improving diagnostic accuracy. Furthermore, it employs distribution-sensing identification to enhance the contribution of key cells, allowing a smaller number of T cells with significantly abnormal signals to contribute more to the overall score, preventing important cells from being masked by the average value, and further improving diagnostic sensitivity. This auxiliary diagnostic system exhibits high discriminative power for ischemic stroke diagnosis, with a scoring power of 91.5%, a high consistency between predicted probability and actual observation values, high sensitivity and specificity, and provides positive clinical decision-making value at various thresholds.
[0018] The auxiliary diagnostic system of this invention can directly utilize high-throughput single-cell sequencing data, and single-cell data is directly adaptable without the need for additional T-cell subset annotation. It can directly process single-cell transcriptome data, reducing data preprocessing complexity and improving method applicability. It also supports batch sample processing and personalized scoring output. The personalized risk score is mapped to a 0-100 range for easy clinical interpretation and supports clinical stratification and management based on diagnostic results. Furthermore, it can serve as a reference indicator for subsequent research on its correlation with stroke severity and prognosis. Attached Figure Description
[0019] Figure 1 This is a schematic diagram of the Mendelian randomized multi-factor risk forest diagram of the present invention; Figure 2 This is a schematic diagram of cell annotation for single-cell transcriptome sequencing data in this invention; Figure 3 This invention presents a volcano plot of differentially expressed genes in the T cell population based on single-cell transcriptome sequencing data. Figure 4 A violin plot showing the accurate diagnostic assessment and significant differences between the model scoring system used in this invention for healthy volunteers and ischemic stroke patients. Figure 5 The ROC curve of the auxiliary diagnostic system of this invention has an AUC of 0.915. Figure 6 The calibration curve of the auxiliary diagnostic system of the present invention shows the consistency between the system's diagnostic results and the actual occurrence of ischemic stroke, indicating that the diagnostic score has high accuracy and reliability. Figure 7 The decision curve analysis (DCA) of the diagnostic auxiliary model of this invention shows the net clinical decision benefit of the diagnostic score at different thresholds, and the results indicate that the diagnostic score can provide high practical decision value for the diagnosis of ischemic stroke. Detailed Implementation
[0020] The principles and features of the present invention are described below with reference to the accompanying drawings and specific embodiments. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.
[0021] T-cell transcriptome alterations are a core mechanism driving ischemic stroke-related immune disorders. However, current clinically widely relied-upon whole-serum biomarker diagnostic techniques struggle to capture T-cell-specific immune abnormalities, leading to insufficient diagnostic specificity and targeting. Furthermore, some ischemic stroke-related genes are not yet transcribed and translated at high levels in the early stages of the disease, resulting in extremely low concentrations of certain serum protein biomarkers, which are more delayed than transcriptome biomarkers, further complicating early diagnosis. This invention provides an auxiliary diagnostic system for ischemic stroke based on peripheral blood T-cell single-cell transcriptome sequencing data and its application. This system utilizes peripheral blood T-cell single-cell transcriptome sequencing data, introduces genes corresponding to proteins significantly associated with ischemic stroke, and combines nonlinear weighted fusion and distribution sensing to achieve high discriminative ability for ischemic stroke, with high sensitivity and specificity.
[0022] 1. Screening of target genes significantly associated with ischemic stroke and determination of their corresponding weights; construction of a single-cell risk scoring model. We selected 25 protein molecules from the ieugwasr database that have been shown by basic research to be potentially associated with the occurrence of ischemic stroke. Through large-scale Mendelian randomization analysis, we screened out 10 protein molecules (CD25, SOD2, CXCR3, GPX1, CSF2, IL-17, GPX7, IL-10, PD-1, and HLA-DR) that are significantly associated with the occurrence of ischemic stroke, and obtained their corresponding odds ratios (OR). The protein molecules and their corresponding OR values are as follows: CD25 (1.026), SOD2 (1.031), CXCR3 (1.091), GPX1 (1.18), CSF2 (1.029), IL-17 (1.171), GPX7 (0.96), IL-10 (0.899), PD-1 (0.79), and HLA-DR (0.952).
[0023] Mendelian randomized multifactor risk forest diagram Figure 1 As shown, this paper presents 10 protein molecules that are significantly associated with ischemic stroke risk / protection and their OR values, selected by Mendelian randomization from 25 protein molecules that have been shown to be highly associated with ischemic stroke.
[0024] The above 10 proteins are divided into risk factors and protective factors. Define direction:
[0025] Among them, risk factors with an odds ratio (OR) greater than 1 are considered risk factors for the disease; the larger the OR, the stronger the risk contribution. Protective factors with an OR less than 1 are considered protective factors against the disease; the smaller the OR, the stronger the protective effect. Details are as follows: Risk factors: CD25 (1.026), SOD2 (1.031), CXCR3 (1.091), GPX1 (1.18), CSF2 (1.029), IL-17 (1.171) Protective factors: GPX7 (0.96), IL-10 (0.899), PD-1 (0.79), HLA-DR (0.952) Calculate the risk contribution weight for each risk factor and the protection contribution weight for each protection factor separately. Details are as follows: Risk factors: using OR 1 represents the risk contribution weight; the higher the expression value and the larger the OR, the stronger the risk contribution. Protection factor: 1 OR is used as a protective contribution weight; the higher the expression value and the lower the OR, the more obvious the protective effect.
[0026] That is, weight ( )
[0027]
[0028] Single-cell nonlinear scoring Define the softplus function:
[0029] Calculate the nonlinear contribution of a single gene for each cell.
[0030] in Depending on i is a risk factor ( ) or protective factor ( ).
[0031] The encoding genes of the 10 proteins selected above were used as target genes, and the nonlinear contributions of all genes were summed. To form a single-cell risk score .
[0032]
[0033] Specifically, based on the expression level of the target gene and the risk weight, the single-cell risk scoring model is as follows: =(0.026×CD25 gene expression level)+(0.031×SOD2 gene expression level)+(0.091×CXCR3 gene expression level)+(0.18×GPX1 gene expression level)+(0.029×CSF2 gene expression level)+(0.171×IL-17 gene expression level)+(-0.04×GPX7 gene expression level)+(-0.101×IL-10 gene expression level)+(-0.21×PD-1 gene expression level)+(-0.048×HLA-DR gene expression level).
[0034] Next, peripheral blood mononuclear cells were isolated from the subject's peripheral blood samples using the peripheral blood mononuclear cell acquisition module, and transcriptomic data of the peripheral blood mononuclear cells were obtained using the single-cell RNA sequencing module. Finally, the expression levels of genes encoding protein molecules significantly associated with ischemic stroke were obtained using the data processing module. Details are shown in Figures 2-3 below: 2. Sample collection and single-cell sequencing processing Peripheral blood was collected from the subjects, and peripheral blood mononuclear cells (PBMCs) were isolated. T cell transcriptome data were obtained after single-cell RNA sequencing and annotation, without distinguishing T cell subsets. The expression levels of each target gene (CD25, SOD2, CXCR3, GPX1, CSF2, IL-17, GPX7, IL-10, PD-1, and HLA-DR) in each cell were obtained, and the transcriptome data were recorded as an expression matrix. Eic ,in i =1,…,10 are the target genes. c =1,…, Np For individuals p T cells.
[0035] 3. Normalization of single-cell level expression signals The single-cell transcriptome data were standardized by converting the expression value of each target gene into a standardized signal in the 0–1 range, thereby reducing the impact of sequencing depth differences between samples.
[0036] Gene expression signals are mapped to a baseline reference interval in the training set through standardization, denoted as... .
[0037]
[0038] in and These are the minimum and maximum values of the expression of the i-th gene, respectively.
[0039] 4. Single-cell risk score The single-cell risk score of all T cells in the subject was obtained based on the single-cell risk scoring model.
[0040] 5. Individualized risk scoring Through distribution-aware identification, an individualized risk score is calculated for each subject, which is the core individualized raw score using the mean of the top 10%. To enhance the weight of high-risk cells.
[0041]
[0042] Among them, set Individual All T-cell single-cell risk scores The highest 10%.
[0043] 6. Map to 0-100 rating The set of individualized raw scores is calculated from the training set. The original individualized risk score is transformed using a nonlinear mapping. Convert to a final score of 0-100.
[0044]
[0045] in, ,
[0046] 7. Stratification of Subject Diagnosis Based on the individual subject's final score, the subjects were divided into two risk levels to help diagnose whether the subject had ischemic stroke.
[0047]
[0048] Low level: 0≤ A value ≤50 indicates weak T-cell signal associated with ischemic stroke in the subject, which may help determine if the stroke is non-ischemic.
[0049] High grade: 50 < A value ≤100 indicates strong T-cell signals associated with ischemic stroke in the subject, which can help in the diagnosis of ischemic stroke.
[0050] The grading results can serve as a quantitative reference for auxiliary diagnosis, helping to make comprehensive judgments on the tested subjects in clinical practice, and can be combined with imaging or other biomarkers for diagnostic decision-making.
[0051] The auxiliary diagnostic system of this invention was used to test 50 patients diagnosed with ischemic stroke and 50 healthy volunteers.
[0052] The cell annotation diagram of the single-cell transcriptome sequencing data is shown below. Figure 2 The figure shows the annotation of cell types after single-cell sequencing of blood from ischemic stroke patients and healthy volunteers. Single-cell transcriptome sequencing data, T-cell population differentially expressed gene volcano plot as shown in the figure. Figure 3 As shown, the expression levels of the genes corresponding to the 10 ischemic stroke-related target proteins were significantly different in the peripheral blood of ischemic stroke patients and healthy individuals.
[0053] The scoring and diagnostic results obtained using the diagnostic system of this invention were compared with the diagnostic results of diffusion magnetic resonance imaging (DWI), with the DWI results used as the gold standard. The accuracy of the system's diagnostic assessment and the significant differences between healthy volunteers and ischemic stroke patients are illustrated in the violin plot. Figure 4 As shown, the ROC curve is as follows Figure 5 As shown, its AUC is 0.915, and it combines high sensitivity and specificity. The calibration curve is as follows. Figure 6 As shown, the model's predicted probability is consistent with the actual occurrence of ischemic stroke, indicating that the scoring of this auxiliary diagnostic system has high accuracy and reliability. The decision curve analysis (DCA) results are as follows. Figure 7 As shown, the net clinical decision benefit of the diagnostic score at different thresholds is demonstrated, indicating that the score of the auxiliary diagnostic system of the present invention can provide high practical decision value for the diagnosis of ischemic stroke.
[0054] The auxiliary diagnostic system of this invention showed significant differences in scoring between ischemic stroke patients and healthy individuals. ROC curve, calibration curve, and DCA analysis revealed that the scoring system of this invention has high discriminative power for the auxiliary diagnosis of ischemic stroke, achieving a scoring power of 91.5%. Furthermore, the predictive probability highly correlated with actual observations, and it provided positive clinical decision-making value at various thresholds. These results demonstrate that the auxiliary diagnostic system of this invention has high diagnostic discrimination capability and can be prioritized to assist the gold standard in early clinical diagnostic assessment.
[0055] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. An auxiliary diagnostic system for ischemic stroke based on peripheral blood T cell single-cell transcriptome, characterized in that, include: Peripheral blood mononuclear cell acquisition module, used to isolate peripheral blood mononuclear cells from peripheral blood samples of subjects; Single-cell RNA sequencing module is used to acquire transcriptome data of a single T cell in peripheral blood mononuclear cells; The data processing module targets the gene encoding a protein molecule that is significantly associated with the occurrence of ischemic stroke. The data processing module is used to obtain the expression level of the target gene in each T cell. The data analysis module includes a single-cell risk scoring model. This module obtains a risk score for each T cell based on the expression level of the target gene in the subject, and then uses distribution-aware recognition to non-linearly weight and fuse all T cell risk scores to obtain an individualized risk score for the subject. The single-cell risk scoring model is constructed as follows: protein molecules significantly associated with ischemic stroke are screened from a database, and their corresponding odds ratios are obtained to calculate their risk weights. Based on the expression levels of the protein molecules and their risk weights, the single-cell risk scoring model is constructed. The outcome discrimination module categorizes subjects into ischemic stroke negative and positive based on their individualized risk scores.
2. The auxiliary diagnostic system for ischemic stroke based on peripheral blood T cell single-cell transcriptome according to claim 1, characterized in that, The protein molecules that are significantly associated with the occurrence of ischemic stroke are CD25, SOD2, CXCR3, GPX1, CSF2, IL-17, GPX7, IL-10, PD-1, and HLA-DR.
3. The auxiliary diagnostic system for ischemic stroke based on peripheral blood T cell single-cell transcriptome according to claim 2, characterized in that, The risk weights for CD25, SOD2, CXCR3, GPX1, CSF2, IL-17, GPX7, IL-10, PD-1, and HLA-DR were 0.026, 0.031, 0.091, 0.18, 0.029, 0.171, -0.04, -0.101, -0.21, and -0.048, respectively. The single-cell risk scoring model is as follows: Riskscore = (0.026 × CD25 gene expression level) + (0.031 × SOD2 gene expression level) + (0.091 × CXCR3 gene expression level) + (0.18 × GPX1 gene expression level) + (0.029 × CSF2 gene expression level) + (0.171 × IL-17 gene expression level) + (-0.04 × GPX7 gene expression level) + (-0.101 × IL-10 gene expression level) + (-0.21 × PD-1 gene expression level) + (-0.048 × HLA-DR gene expression level).
4. The auxiliary diagnostic system for ischemic stroke based on peripheral blood T cell single-cell transcriptome according to claim 1, characterized in that, The data processing module is also used to normalize the transcriptome data of single cells, converting the expression value of each target gene into a standardized signal in the 0-1 range.
5. The auxiliary diagnostic system for ischemic stroke based on peripheral blood T cell single-cell transcriptome according to claim 1, characterized in that, The data analysis module nonlinearly maps individualized risk scores to the 0-100 range.
6. The auxiliary diagnostic system for ischemic stroke based on peripheral blood T cell single-cell transcriptome according to claim 1, characterized in that, In the result discrimination module, a certain threshold is set. If the individualized risk score is greater than or equal to the threshold, it is judged as a positive ischemic stroke, and if the individualized risk score is less than the threshold, it is judged as a negative ischemic stroke.
7. The auxiliary diagnostic system for ischemic stroke based on peripheral blood T cell single-cell transcriptome according to claim 6, characterized in that, In the result discrimination module, the threshold is set to 50 points.
8. The auxiliary diagnostic system for ischemic stroke based on peripheral blood T cell single-cell transcriptome according to any one of claims 1-7, characterized in that, The number of peripheral blood T cells sequenced was 1500-5000.
9. The application of the peripheral blood T cell single-cell transcriptome-based auxiliary diagnostic system for ischemic stroke as described in any one of claims 1-8 in the auxiliary diagnosis of ischemic stroke.
10. The application according to claim 9, characterized in that, The auxiliary diagnostic system for ischemic stroke based on peripheral blood T cell single-cell transcriptome achieved an area under the ROC curve of 0.915 for ischemic stroke.