Bioinformatics analysis method for identifying gene target points related to prognosis of cerebral hemorrhage

By integrating differential expression analysis, weighted co-expression networks, and protein-protein interaction networks, core gene targets related to the prognosis of cerebral hemorrhage were screened from transcriptome data. This approach addresses the shortcomings of existing screening strategies and achieves highly specific and interpretable gene target identification.

CN121601020BActive Publication Date: 2026-04-24THE FIRST AFFILIATED HOSPITAL OF FUJIAN MEDICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
THE FIRST AFFILIATED HOSPITAL OF FUJIAN MEDICAL UNIV
Filing Date
2026-01-29
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing studies on gene screening related to the prognosis of cerebral hemorrhage mostly employ a single differential expression analysis strategy, which makes it difficult to focus on core targets that truly have prognostic predictive value and mechanistic relevance. Furthermore, the correlation between candidate genes and the core pathological process of cerebral hemorrhage is insufficiently validated, resulting in limited biological interpretability of the screening results.

Method used

By integrating differential expression analysis, weighted co-expression network analysis, and protein interaction network analysis, combined with pathological pathway validation, immune infiltration association analysis, and survival analysis, core gene targets associated with the prognosis of cerebral hemorrhage are systematically screened from transcriptome data, and gene targets with mechanistic interpretability and prognostic predictive value are output.

Benefits of technology

It enables accurate identification of gene targets related to the prognosis of cerebral hemorrhage, improves screening specificity and statistical reliability, and ensures that the output gene targets have expression change characteristics, prognostic predictive value and clinical application value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121601020B_ABST
    Figure CN121601020B_ABST
Patent Text Reader

Abstract

The application discloses a bioinformatics analysis method for identifying a brain hemorrhage prognosis-related gene target point, and the method comprises the following steps: obtaining brain hemorrhage transcriptome expression data, performing quality filtering and batch correction to generate a standardized expression matrix, performing differential expression analysis on the basis of the standardized expression matrix, screening prognosis-associated differential genes in combination with survival regression, constructing a weighted co-expression network to identify a prognosis key gene cluster, screening and verifying a core hub gene through hub degree, taking the intersection of the prognosis-associated differential genes and the core hub gene, filtering and interaction network clustering a core target gene through a pathological pathway, performing immune infiltration correlation analysis and survival analysis on the basis of the core target gene, and outputting the brain hemorrhage prognosis-related gene target point according to an immune correlation marker and a survival correlation index. The application integrates a differential analysis and a network analysis double strategy, confirms mechanism correlation through pathway verification and immune correlation analysis, and the screening result has statistical reliability and biological interpretability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of bioinformatics and medical data mining, and in particular to a bioinformatics analysis method for identifying gene targets related to the prognosis of cerebral hemorrhage. Background Technology

[0002] Intracerebral hemorrhage is a common acute and critical neurological disease with highly variable prognoses. Some patients recover well after treatment, while others suffer severe disability or even death. Accurately identifying key molecular targets affecting the prognosis of intracerebral hemorrhage is crucial for understanding disease mechanisms, developing prognostic prediction models, and screening potential therapeutic targets. In recent years, the widespread adoption of high-throughput sequencing and gene chip technologies has led to a wealth of transcriptomic data related to intracerebral hemorrhage, providing a data foundation for analyzing prognostic factors at the molecular level.

[0003] Existing studies on gene screening related to the prognosis of intracerebral hemorrhage mostly employ a single differential expression analysis strategy, identifying candidate genes by comparing gene expression differences between patients with good and poor prognoses. However, the large number of differentially expressed genes and the high false-positive rate make it difficult to focus on core targets that truly have prognostic predictive value and mechanistic relevance. Some studies have introduced survival analysis to assess the prognostic association strength of genes, but lack a systematic exploration of the functional relationships between genes, failing to reveal the synergistic regulatory patterns of prognostic-related genes. Furthermore, the validation of the association between candidate genes and the core pathological processes of intracerebral hemorrhage is insufficient, limiting the biological interpretability of the screening results. Summary of the Invention

[0004] This invention discloses a bioinformatics analysis method for identifying gene targets related to the prognosis of cerebral hemorrhage. By integrating three strategies—differential expression analysis, weighted co-expression network analysis, and protein interaction network analysis—the method systematically screens core gene targets associated with the prognosis of cerebral hemorrhage from transcriptome data. These targets are then validated in multiple dimensions through pathological pathway verification, immune infiltration association analysis, and survival analysis, resulting in the output of gene targets with mechanistic interpretability and prognostic predictive value.

[0005] The first aspect of this invention proposes a bioinformatics analysis method for identifying gene targets related to the prognosis of cerebral hemorrhage, comprising the following steps:

[0006] Obtain the cerebral hemorrhage transcriptome expression dataset, and perform quality filtering and batch effect correction on the cerebral hemorrhage transcriptome expression dataset to generate a standardized expression matrix;

[0007] Based on the standardized expression matrix, differential expression analysis was performed to extract differentially expressed gene sets by grouping according to prognostic phenotype. The expression-prognostic correlation of the differentially expressed gene sets was then screened to obtain prognostic-associated differentially expressed genes.

[0008] A weighted gene co-expression network is constructed from the standardized expression matrix to generate a co-expression gene cluster set. Based on the co-expression gene cluster set, a prognostic correlation analysis is performed to extract key prognostic gene clusters. The key prognostic gene clusters are then used to perform hub degree-prognostic joint screening to identify core hub genes.

[0009] Intersection screening is performed on the prognostic associated differential genes and the core hub genes to obtain a preliminary screening target set. Pathological pathway association filtering is performed on the preliminary screening target set to obtain mechanism-related targets. Based on the mechanism-related targets, core node clustering of the interaction network is performed to identify core target genes.

[0010] Based on the core target genes, immune infiltration association analysis is performed to obtain immune-related target markers. Survival analysis is performed on the core target genes to obtain survival association indicators. Based on the immune-related target markers and the survival association indicators, prognostic gene targets related to cerebral hemorrhage are output.

[0011] A second aspect of this invention provides a bioinformatics analysis system for identifying gene targets related to the prognosis of cerebral hemorrhage, comprising:

[0012] The data processing unit is used to acquire the cerebral hemorrhage transcriptome expression dataset, and to perform quality filtering and batch effect correction on the cerebral hemorrhage transcriptome expression dataset to generate a standardized expression matrix;

[0013] The differential screening unit is used to perform differential expression analysis based on the standardized expression matrix and grouped by prognostic phenotype to extract differentially expressed gene sets, and to perform expression-prognostic correlation screening on the differentially expressed gene sets to obtain prognostic associated differentially expressed genes.

[0014] The gene identification unit is used to construct a weighted gene co-expression network from the standardized expression matrix to generate a co-expression gene cluster set, perform prognostic correlation analysis based on the co-expression gene cluster set to extract key prognostic gene clusters, and use the key prognostic gene clusters to perform hub degree-prognostic joint screening to identify core hub genes.

[0015] The target mining unit is used to perform intersection screening on the prognostic associated differential genes and the core hub genes to obtain a preliminary screening target set, perform pathological pathway association filtering on the preliminary screening target set to obtain mechanism-related targets, and perform interaction network core node clustering based on the mechanism-related targets to identify core target genes.

[0016] The verification output unit is used to perform immune infiltration association analysis based on the core target genes to obtain immune-related target markers, perform survival analysis on the core target genes to obtain survival association indicators, and output prognostic gene targets for cerebral hemorrhage based on the immune-related target markers and the survival association indicators.

[0017] The beneficial effects of this invention are reflected in the following points: 1. Transcriptome data, after quality filtering and batch effect correction, forms a standardized expression matrix. Based on this, differential expression analysis is performed according to prognostic phenotype groups. Early-related genes and late-related genes are distinguished through univariate survival regression and time-dependent analysis. After risk direction labeling, prognostic-related differentially expressed genes are formed. This combined screening strategy of differential expression and survival analysis ensures that selected genes possess both expression change characteristics and prognostic predictive value, exhibiting higher screening specificity compared to traditional methods that solely rely on fold change thresholds. 2. Weighted co-expression network analysis identifies functionally related gene modules from the perspective of gene co-expression. Key prognostic gene clusters are screened through correlation analysis between characteristic genes and prognostic phenotypes. Based on this, hub genes are screened by combining intra-cluster connectivity, prognostic association strength, and brain tissue expression specificity. Core hub genes are formed through cross-dataset expression consistency verification. This network-driven screening strategy can capture functionally co-expressing genes that are difficult to detect through differential expression analysis and eliminates false positives due to dataset specificity through independent verification. 3. The intersection screening of prognostic differential genes and core hub genes integrates the results of two independent analysis strategies. Pathological pathway association filtering confirms the association between candidate targets and core pathological processes such as neuroinflammation, blood-brain barrier damage, and apoptosis. Protein interaction network clustering further identifies core nodes in functional subnetworks. Immune infiltration association analysis and survival analysis provide dual validation at both the mechanistic and clinical levels. The multiple filtering strategy ensures that the final output of prognostic gene targets for cerebral hemorrhage has statistical reliability, mechanistic interpretability, and clinical application value. Attached Figure Description

[0018] The accompanying drawings illustrate specific examples of the technical solutions described in this invention and, together with the detailed embodiments, form part of the specification, serving to explain the technical solutions, principles, and effects of this invention.

[0019] Figure 1 This is a flowchart illustrating a bioinformatics analysis method for identifying gene targets related to the prognosis of cerebral hemorrhage according to the present invention.

[0020] Figure 2 This is a schematic diagram of the weighted gene co-expression network and module identification of the present invention.

[0021] Figure 3 This is a schematic diagram of the protein interaction network and functional subnetwork of the present invention.

[0022] Figure 4 This is a structural block diagram of a bioinformatics analysis system for identifying gene targets related to the prognosis of cerebral hemorrhage, as described in this invention.

[0023] Wherein: 1-Gene node; 2-Co-expression connection edge; 3-Blue module; 4-Green module; 5-Yellow module; 6-Cyan-green module; 7-Gray module; 8-Module feature gene; 9-Host gene; 10-High connectivity region; 11-Module boundary; 12-Cross-module connection; 13-Seed protein node; 14-First-order interaction partner; 15-Interaction connection edge; 16-Functional subnetwork A; 17-Functional subnetwork B; 18-Core node; 19-Bridging node; 20-Scattered node; 21-Subnetwork boundary; 22-High-density clustering region; 23-Edge weight label. Detailed Implementation

[0024] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0025] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0026] References to "one embodiment" or "some embodiments" in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0027] The technical solutions of the embodiments of this application will be described below.

[0028] like Figure 1 As shown, this embodiment of the invention provides a bioinformatics analysis method for identifying gene targets related to the prognosis of cerebral hemorrhage, including the following steps S110-S150:

[0029] Step S110: Obtain the cerebral hemorrhage transcriptome expression dataset, and perform quality filtering and batch effect correction on the cerebral hemorrhage transcriptome expression dataset to generate a standardized expression matrix.

[0030] Specifically, we obtained transcriptome expression datasets for cerebral hemorrhage. Public gene expression databases are the main source of transcriptome data. The GEO database and ArrayExpress database contain multiple gene expression research datasets related to cerebral hemorrhage. The search keywords for cerebral hemorrhage transcriptome expression datasets were set as disease terms such as cerebral hemorrhage, intracranial hemorrhage, and hemorrhagic stroke, while limiting the sample type to human brain tissue or peripheral blood samples. In the GEO database, the GSE series datasets convert probe signals into gene symbols through platform annotation files. For cerebral hemorrhage transcriptome expression datasets, each gene must correspond to a unique expression value; when multiple probes correspond to the same gene, the maximum value is taken to retain the probe with the strongest signal. The inclusion criteria required that the cerebral hemorrhage transcriptome expression dataset contain at least 20 cerebral hemorrhage patient samples and 10 control samples. Datasets with a sample size of less than 30 samples had a statistical power of less than 80% for differential expression analysis, potentially missing differentially expressed genes of moderate effect size; therefore, datasets with excessively small sample sizes were excluded. Clinical annotation information for the intracerebral hemorrhage transcriptome expression dataset must include patient prognostic status, defined as survival status at 90 days post-onset or a modified Rankin Scale score. Some early GEO datasets only record acute-phase clinical characteristics and lack 90-day follow-up data. Such datasets lacking prognostic information can be used for acute-phase differential expression analysis but cannot be included in prognostic modeling. Microarray platforms primarily use Affymetrix and Illumina, with whole-genome expression microarray data being the preferred choice for the intracerebral hemorrhage transcriptome expression dataset. RNA sequencing data, after TPM or FPKM standardization, is also included in the intracerebral hemorrhage transcriptome expression dataset. The integration of sequencing and microarray data uses ComBat cross-platform calibration. The final intracerebral hemorrhage transcriptome expression dataset includes three parts: an expression matrix, a sample clinical information table, and platform annotation files.

[0031] Quality filtering and batch effect correction were performed on the brain hemorrhage transcriptome expression dataset to generate a standardized expression matrix. The original expression data contained missing values, low-quality probes, and batch-to-batch systematic errors, requiring multiple preprocessing steps before downstream analysis. Genes with missing values ​​exceeding 20% ​​were removed from the dataset, and the remaining genes with a small number of missing values ​​were imputed using the K-nearest neighbor algorithm, with K set to 10. The filtering threshold for low-expression genes was set to ensure that the expression value was above the detection limit in at least 50% of the samples. After low-expression filtering, the number of genes in the brain hemorrhage transcriptome expression dataset was typically reduced from over 20,000 to around 15,000. The generation of the standardized expression matrix first involved quantile standardization of the filtered data to ensure a more uniform distribution of expression values ​​across samples. Since data from different research batches exhibit batch effects, the ComBat algorithm was used to correct batch effects when the brain hemorrhage transcriptome expression dataset contained multiple GEO datasets. This correction process preserved biological variations such as prognostic phenotypes while eliminating technical batch differences. After logarithmic transformation, the expression values ​​in the standardized expression matrix approximate a normal distribution. The transformation formula is log2(x+1), where x is the quantile-standardized expression value, and adding 1 avoids logarithmic problems with zero values. The quality of the standardized expression matrix is ​​verified using principal component analysis. Before correction, samples in the PC1-PC2 scatter plot are clustered according to GSE numbers. When correction is successful, samples should be clustered into good-prognosis and poor-prognosis groups, and the proportion of variance explained by the batch should decrease from more than 30% before correction to less than 5% after correction. The row index of the standardized expression matrix is ​​the gene symbol, the column index is the sample number, and the matrix elements are the standardized gene expression values.

[0032] Step S120: Based on the standardized expression matrix, differential expression analysis is performed to extract differentially expressed gene sets by grouping according to prognostic phenotypes, and expression-prognostic correlation screening is performed on the differentially expressed gene sets to obtain prognostic associated differentially expressed genes.

[0033] Specifically, differential expression analysis was performed to extract differentially expressed gene sets based on prognostic phenotype grouping according to the standardized expression matrix. Prognostic phenotype grouping divided the samples into a favorable prognosis group and a poor prognosis group. Each sample in the standardized expression matrix was assigned to the corresponding group based on its 90-day prognostic status in the clinical annotation information. The favorable prognosis group was defined as patients who survived within 90 days and achieved a functional recovery score of 0-2 on the modified Rankin Scale; the poor prognosis group was defined as patients who died within 90 days or had a functional recovery score of 3-6. The sample size ratio between the two groups should ideally be controlled within the range of 1:1 to 1:3. If the ratio exceeds 1:4, weighted or resampling methods are recommended for correction. After extracting the gene expression values ​​of the two groups from the standardized expression matrix, inter-group comparisons were performed. The differential expression analysis used the limma algorithm to construct a linear model and estimated the statistical significance of gene expression differences using empirical Bayesian methods. The fold change calculation formula is log2FC = mean (poor prognosis group) - mean (good prognosis group), where mean is the mean expression value after log2 transformation, and log2FC is the base-2 logarithmic fold change. A positive value indicates high expression of the gene in the poor prognosis group, and a negative value indicates high expression in the good prognosis group. The screening threshold for the differentially expressed gene set is set as an absolute fold change greater than 1 and a corrected p-value less than 0.05. log2FC > 1 corresponds to a gene expression level difference of more than 2-fold between the two groups. The Benjamini-Hochberg method is used to control the false discovery rate. Genes that pass the threshold screening in the normalized expression matrix are included in the differentially expressed gene set. Typical transcriptome analysis of cerebral hemorrhage can obtain 500-2000 differentially expressed genes, the specific number of which is related to the sample size, disease heterogeneity, and threshold setting. The differentially expressed gene set records four statistical indicators for each gene: gene symbol, fold change, p-value, and corrected p-value, and is stored in descending order of absolute fold change. Genes with positive fold changes in differentially expressed genes are marked as upregulated genes, and genes with negative fold changes are marked as downregulated genes. Upregulated genes correspond to risk gene candidates in subsequent risk direction labeling, while downregulated genes correspond to protective gene candidates.

[0034] In some embodiments, the step of screening the differentially expressed gene set for expression-prognostic correlation to obtain prognostic-associated differentially expressed genes includes: performing univariate survival regression analysis on each gene in the differentially expressed gene set to obtain prognostic association coefficients; performing time-dependent analysis on the prognostic association coefficients to distinguish between early-associated genes and late-associated genes; performing risk orientation labeling on the early-associated genes and the late-associated genes to obtain gene risk attribute labels; and integrating the early-associated genes and the late-associated genes based on the gene risk attribute labels to form prognostic-associated differentially expressed genes.

[0035] Univariate survival regression analysis was performed on each gene in the differentially expressed gene set to obtain prognostic association coefficients. Cox proportional hazards regression is a standard method for survival analysis. The expression value of each gene in the differentially expressed gene set was included as a single covariate in the Cox model for prognostic association assessment. Survival time was defined as the number of days from onset to death or last follow-up. The survival time and survival status of the corresponding samples in the differentially expressed gene set were extracted from clinical annotation information; censored samples were marked as 0, and death events were marked as 1. Univariate Cox regression was fitted to each gene in the differentially expressed gene set, obtaining an independent hazard ratio and regression coefficient estimate for each gene. The calculation process used the maximum likelihood estimation method. The prognostic association coefficient was defined as the beta coefficient of the Cox regression; a positive value indicates that high gene expression is associated with a poor prognosis, and a negative value indicates that high gene expression is associated with a good prognosis. The prognostic association coefficients of each gene in the differentially expressed gene set were assessed for statistical significance using the Wald test; genes with a p-value less than 0.05 were considered significantly associated with prognosis. The absolute value of the prognostic association coefficient reflects the strength of the impact of gene expression changes on prognostic risk; a larger absolute value indicates a stronger prognostic association. The hazard ratio (HR) = exp(beta) represents the relative fold change in risk for each unit increase in gene expression. Proportional hazard hypothesis testing validates the effectiveness of the prognostic association coefficient. A Schoenfeld residual test p-value less than 0.05 suggests that the prognostic effect of the gene may change over time; therefore, the gene's time-dependent effect is re-estimated using a time-varying covariate Cox model. The prognostic association coefficient, along with the hazard ratio, 95% confidence interval, and p-value, constitutes the statistical results of survival analysis for each gene.

[0036] Time-dependent analysis of prognostic association coefficients was conducted to distinguish between early- and late-term associated genes. The prognostic association effect may vary across different follow-up periods; the prognostic association coefficient reflects the average effect over the entire follow-up period and cannot reveal heterogeneity over time. Time-dependent analysis divided the follow-up period into early and late phases, with 30 days as the dividing line. The early phase (0-30 days) corresponds to the acute and subacute phases of intracerebral hemorrhage, during which hematoma absorption and early neurological injury are the main pathological processes. The late phase (after 30 days) corresponds to the recovery period of intracerebral hemorrhage, where neural remodeling and functional compensation become dominant. Prognostic association coefficients were recalculated separately for the early and late phases using piecewise Cox regression models or time-stratified analysis. Regression coefficients and statistical significance were independently estimated for each phase. Genes with significant prognostic association coefficients in the early phase but not in the late phase were classified as early-term associated genes. Changes in the expression of these genes primarily affect survival outcomes during the acute phase of intracerebral hemorrhage and may participate in the pathological process of early brain injury. Genes with significant prognostic correlation coefficients in the late stage but not in the early stage are classified as late-stage associated genes. Changes in the expression of these genes primarily affect the functional prognosis during the recovery period from cerebral hemorrhage and may be involved in the pathological processes of neural repair or secondary injury. Both early-stage and late-stage associated genes must meet the significance requirement of a p-value less than 0.05 for their respective time periods. Genes significant in both time periods are classified according to the time period with the larger absolute value of the prognostic correlation coefficient. Genes with prognostic correlation coefficients of opposite signs in the early and late stages are marked as time-varying effect genes. The risk direction of time-varying effect genes is determined based on the sign of the coefficient in the time period with the larger absolute value. Early-stage associated genes are typically associated with pathological processes such as acute inflammatory response, blood-brain barrier disruption, and cell death, while late-stage associated genes are typically associated with recovery processes such as neural plasticity, vascular remodeling, and glial scar formation.

[0037] Risk direction labeling was performed on early-associated and late-associated genes to obtain gene risk attribute labels. Risk direction determines the biological significance of a gene in prognostic prediction; the direction of prognostic risk change corresponding to the increase in expression of early-associated and late-associated genes needs to be clearly labeled. Genes with positive prognostic association coefficients in early-associated genes are labeled as early-risk genes, indicating that high expression of this gene in the acute phase is associated with poor prognosis and may be involved in harmful processes such as nerve injury, amplified inflammation, or apoptosis. Genes with negative prognostic association coefficients in early-associated genes are labeled as early-protective genes, indicating that high expression of this gene in the acute phase is associated with good prognosis and may be involved in beneficial processes such as neuroprotection, inflammation suppression, or anti-oxidative stress. Late-associated genes are divided into late-risk genes and late-protective genes using the same labeling logic. Late-risk genes may be associated with repair impairment, chronic inflammation, or secondary neurodegeneration, while late-protective genes may be associated with functional recovery, synaptic remodeling, or angiogenesis. Gene risk attribute tags employ a four-category coding system: value 1 represents early risk, 2 represents early protection, 3 represents late risk, and 4 represents late protection. These four categories cover a complete combination of time and risk dimensions. Each gene in both the early-associated and late-associated gene categories is assigned a unique gene risk attribute tag, and this tag information is appended to the corresponding field in the gene annotation table. The statistical distribution of gene risk attribute tags reflects the overall risk composition of genes related to the prognosis of cerebral hemorrhage, and the proportion of the four tag categories reveals the relative abundance of risky and protective genes at different time points.

[0038] Based on gene risk attribute tags, early-stage associated genes and late-stage associated genes are integrated to form prognostic differentially expressed genes. The merging of early-stage and late-stage associated genes constitutes a complete list of prognostically relevant genes, with gene risk attribute tags retaining the time-specificity and risk direction information of each gene. Direct merging of early-stage and late-stage associated genes may result in duplicate genes; for genes significantly associated in both time periods, records of the time period with the larger absolute value of the prognostic association coefficient are retained to avoid duplicate appearances of the same gene in the results. Prognostic differentially expressed genes must simultaneously meet both differential expression and prognostic association conditions; genes that are only differentially expressed but have no significant prognostic association are not included. In the gene risk attribute tags, early-risk genes and late-risk genes are merged to form a risk gene subset, and early-protective genes and late-protective genes are merged to form a protective gene subset. These two subsets constitute the risk classification system for prognostic differentially expressed genes. The number of prognostic differentially expressed genes is typically 30%-50% of the differentially expressed gene set, reflecting that only a portion of the differentially expressed genes have independent prognostic predictive value. The prognostic association differential gene record includes three types of information for each gene: differential expression statistics, survival analysis statistics, and gene risk attribute labels, forming a multi-dimensional gene annotation table. Quality control checks are performed to ensure the ratio of risk genes to protective genes in the prognostic association differential gene record is reasonable; an extremely imbalanced ratio may indicate that sample grouping or analysis parameters need adjustment.

[0039] Step S130: Construct a weighted gene co-expression network from the standardized expression matrix to generate a co-expression gene cluster set. Perform prognostic correlation analysis based on the co-expression gene cluster set to extract key prognostic gene clusters. Use the key prognostic gene clusters to perform hub degree-prognostic joint screening to identify core hub genes.

[0040] A weighted gene co-expression network is constructed from the standardized expression matrix to generate co-expressed gene clusters. For example... Figure 2 As shown, weighted gene co-expression network analysis is a classic method for identifying functionally related gene modules. Genes with similar expression patterns in the normalized expression matrix are clustered into the same module to form co-expression gene clusters. The normalized expression matrix is ​​first filtered to retain genes with the highest expression variation coefficients (top 50%) for network construction. Gene node 1 represents each gene participating in network construction. The correlation between genes is calculated using the Pearson correlation coefficient to form a correlation matrix, and the soft threshold power parameter is selected according to the scale-free network topology criterion. After soft thresholding, the correlation matrix forms an adjacency matrix, which is further converted into a topological overlap matrix to measure the similarity of network connections between genes. Co-expression connection edge 2 connects gene pairs with similar expression patterns, and the edge thickness reflects the co-expression intensity. A hierarchical clustering algorithm clusters genes based on the topological overlap matrix, and a dynamic pruning method identifies gene modules from the clustering tree. In the co-expression gene cluster group, blue module 3, green module 4, yellow module 5, and cyan module 6 represent functionally related gene clusters, and module boundary 11 defines the gene affiliation range of each module. Gray module 7 contains scattered genes that fail to be classified into any functional module; these genes are excluded in subsequent analyses. Module feature gene 8 is marked with a red box and obtained by extracting the first principal component through principal component analysis. Hub gene 9 is located at the center of the high-connectivity region 10, exhibiting the highest intra-cluster connectivity. Cross-module connections 12 are represented by dashed lines, with connection strength lower than intra-module connectivity. Co-expressed gene clusters record the gene member list and average intra-module connectivity for each module; a typical co-expressed gene cluster contains 10-30 non-gray modules.

[0041] In some embodiments, the step of performing prognostic correlation analysis based on the co-expressed gene cluster set to extract key prognostic gene clusters includes: extracting characteristic genes of each gene cluster from the co-expressed gene cluster set; performing prognostic phenotypic correlation analysis on the characteristic genes to obtain gene cluster-prognostic correlation coefficients; performing significant ranking based on the gene cluster-prognostic correlation coefficients to identify highly correlated gene clusters; and performing functional consistency screening on the highly correlated gene clusters to form key prognostic gene clusters.

[0042] Feature genes for each gene cluster are extracted from the co-expressed gene cluster set. The module feature gene is a comprehensive representation of the expression patterns of all genes within the module. For each module in the co-expressed gene cluster set, the first principal component is extracted using principal component analysis as the feature gene for that module. When a module in the co-expressed gene cluster set contains 200 genes, the expression values ​​of these 200 genes in all samples constitute a 200xN submatrix, where N is the number of samples. Principal component analysis reduces the dimensionality of this submatrix; the first principal component captures the direction of maximum gene expression variation within the module, and its score on each sample is the expression value of the feature gene. The length of the feature gene expression vector is equal to the number of samples, and each sample corresponds to one feature gene expression value, which comprehensively reflects the overall expression level of genes within the module in that sample. Feature genes are extracted from all non-gray modules in the co-expressed gene cluster set. If the analysis yields 15 valid modules, 15 feature gene expression vectors are generated. The correlation between feature genes reflects the functional association between modules. Highly correlated modules may participate in similar biological processes; modules with a correlation coefficient exceeding 0.8 can be considered for merging. Feature genes are named using the module color followed by the prefix "ME", such as MEblue for feature genes in the blue module and MEturquoise for feature genes in the turquoise module. The quality of module partitioning of co-expressed gene clusters is assessed by the average correlation between feature genes and gene expression within the module; modules with an average correlation below 0.5 indicate high gene heterogeneity within the module.

[0043] For example, the step of performing prognostic phenotype correlation analysis on the feature genes to obtain gene cluster-prognostic correlation coefficients includes: parsing clinical phenotype associations from the feature genes to extract clinical subtype labels and constructing clinical subtyping vectors; performing stratification on the feature genes according to the clinical subtyping vectors to obtain subtype-level expression data; performing prognostic phenotype correlation tests based on the subtype-level expression data to obtain subtype-specific correlation coefficients; and performing cross-subtype consistency integration on the subtype-specific correlation coefficients to generate gene cluster-prognostic correlation coefficients.

[0044] Clinical subtype labels were extracted from the characteristic genes to construct a clinical subtyping vector. Clinical heterogeneity exists among patients with cerebral hemorrhage, and the prognostic association patterns may differ among different subtypes. The association between characteristic genes and prognosis needs to be analyzed within a subtype hierarchical framework to improve the robustness of the results. Clinical subtypes are classified according to the location, volume, or etiology of the hemorrhage. The clinical information of the samples corresponding to the characteristic genes is extracted from the annotation table and encoded as categorical variables. The hemorrhage location subtype includes four categories: basal ganglia hemorrhage, lobar hemorrhage, thalamic hemorrhage, and cerebellar hemorrhage. The clinical subtype labels are assigned integer codes 1, 2, 3, and 4, respectively. The hemorrhage volume subtype is divided into three categories based on hematoma volume: small hemorrhage, moderate hemorrhage, and large hemorrhage, with 15ml and 30ml as the cutoff thresholds. The corresponding clinical subtype labels are assigned 1, 2, and 3. The clinical subtyping vector is a one-dimensional vector formed by arranging the clinical subtype labels of all samples in sample order. The vector length is equal to the number of samples, and the vector elements are the subtype codes of each sample. The expression vectors of characteristic genes and the clinical subtyping vectors have the same length and sample order, allowing for paired analysis of corresponding elements. The sample size statistics for each subtype in the clinical subtyping vector are used to assess the feasibility of stratified analysis; if the sample size for a subtype is less than 10, the reliability of the analysis results for that subtype decreases. Samples with missing clinical subtype labels are excluded from the stratified analysis, while samples with complete subtype information are retained for subsequent calculations.

[0045] Based on the clinical classification vector, the characteristic genes are stratified to obtain subtype-level expression data. The stratification strategy splits the samples into several subsets according to subtypes. Samples with the same subtype code in the clinical classification vector are grouped into the same subset for independent analysis. The expression vector of the characteristic gene is split according to the subtype code of the clinical classification vector, and the characteristic gene expression values ​​of all samples within each subtype subset are obtained. The subtype-level expression data is stored in a list structure. Each element of the list corresponds to the characteristic gene expression data of a subtype, containing the expression values ​​of all samples in that subtype and the corresponding prognostic status. When the basal ganglia hemorrhage subtype contains 80 samples, the subtype-level expression data contains 80 characteristic gene expression values ​​and 80 prognostic status codes for that subtype. The number of subtypes in the clinical classification vector determines the number of elements in the subtype-level expression data; the four-category hemorrhage site subtype corresponds to four data subsets, and the three-category hemorrhage volume subtype corresponds to three data subsets. The sample size distribution of each subtype subset in the subtype stratified expression data affects the statistical power of subsequent correlation analysis. When the sample size is severely imbalanced, small sample subtypes can be merged. The expression distribution of characteristic genes in different subtypes is visualized using box plots. Significant differences in expression levels between subtypes suggest that this module may be related to subtype-specific biological processes.

[0046] Based on the subtype-stratified expression data, a prognostic phenotypic correlation test was performed to obtain subtype-specific correlation coefficients. Intra-subtype correlation analysis assessed the association between modules and prognosis in a more homogeneous patient population; the correlation coefficient between characteristic genes and prognostic phenotypes was independently calculated for each subtype subset in the subtype-stratified expression data. For the basal ganglia hemorrhage subtype in the subtype-stratified expression data, Pearson correlation coefficients were calculated between the expression values ​​of characteristic genes and the prognostic status codes of patients in that subtype, obtaining subtype-specific correlation coefficients for the basal ganglia subtype. Correlation coefficient calculations were repeated for all subtype subsets in the subtype-stratified expression data, and independent subtype-specific correlation coefficient estimates and corresponding p-values ​​were obtained for each subtype. The sign of the subtype-specific correlation coefficient reflects the direction of association between module expression and prognosis in that subtype; positive values ​​indicate high expression associated with poor prognosis, and negative values ​​indicate high expression associated with good prognosis. Subtypes with smaller sample sizes in the subtype-stratified expression data have wider confidence intervals for their subtype-specific correlation coefficients, and higher uncertainty in point estimates. The consistency of subtype-specific correlation coefficients across different subtypes is a key indicator for assessing the robustness of the prognostic association of modules. Consistent signs of correlation coefficients across subtypes indicate stable association direction across subtypes. Sign reversal in subtype-specific correlation coefficients suggests that the prognostic effect of the module may be modulated by subtype; therefore, the subtype background of patients needs to be considered when applying such modules clinically. The analysis results of subtype-stratified expression data are summarized into a subtype × module correlation coefficient matrix, where rows correspond to subtypes, columns correspond to modules, and elements are subtype-specific correlation coefficients.

[0047] Cross-subtype consistency integration was performed on the subtype-specific correlation coefficients to generate gene cluster-prognostic correlation coefficients. Correlation coefficients from multiple subtypes needed to be integrated into a single module-level prognostic association indicator. The integration of subtype-specific correlation coefficients employed weighted averaging or meta-analysis methods. The weighted average of subtype-specific correlation coefficients used the sample size of each subtype as the weight; subtypes with larger sample sizes contributed more to the integration results. The weighting formula was r_pooled=Σ(n_i×r_i) / Σn_i, where n_i is the sample size of the i-th subtype, and r_i is the corresponding subtype-specific correlation coefficient. Fixed-effects meta-analysis was used to perform Fisher's Z-transformation on the subtype-specific correlation coefficients before weighted pooling. The Z-transformation formula was Z=0.5×ln[(1+r) / (1-r)]. The pooled Z-values ​​were then converted back into correlation coefficients as gene cluster-prognostic correlation coefficients. Heterogeneity of gene cluster-prognostic correlation coefficients was assessed using the I² statistic. An I² greater than 50% indicates significant heterogeneity in subtype-specific correlation coefficients among subtypes, making a random-effects model more appropriate. When all subtype-specific correlation coefficients are positive or all are negative, the cross-subtype consistency of gene cluster-prognostic correlation coefficients is good, and the prognostic association conclusions of this module have high reliability. Confidence intervals for gene cluster-prognostic correlation coefficients were estimated using the Bootstrap resampling method. The 2.5% and 97.5% quantiles obtained from 1000 resampling iterations were used as the boundaries of the 95% confidence interval. For modules with sign reversal in subtype-specific correlation coefficients, their gene cluster-prognostic correlation coefficients were labeled as subtype-dependent associations, indicating that the prognostic effect of this module varies by subtype. Gene cluster-prognostic correlation coefficients, along with heterogeneity indicators and consistency labels, were output for subsequent identification of highly associated gene clusters.

[0048] Highly associated gene clusters are identified by significance ranking based on the gene cluster-prognostic correlation coefficient. The absolute value of the correlation coefficient and statistical significance jointly determine the reliability of the prognostic association of a module. A gene cluster-prognostic correlation coefficient must simultaneously meet the dual criteria of effect size and significance to be considered highly associated. Modules with an absolute value of the gene cluster-prognostic correlation coefficient greater than 0.3 and a corrected p-value less than 0.05 are included in highly associated gene clusters. A correlation coefficient threshold of 0.3 corresponds to a moderate effect size and can be adjusted appropriately according to the research objectives. The number of highly associated gene clusters depends on the characteristics of the dataset and the threshold setting; in typical analyses, approximately 3-8 modules meet the high association criteria, accounting for 20%-40% of all modules. Highly associated gene clusters with positive gene cluster-prognostic correlation coefficients are marked as risk modules, where overall high gene expression within the module is associated with poor prognosis; those with negative correlation coefficients are marked as protective modules, where overall high gene expression within the module is associated with good prognosis. Highly associated gene clusters are ranked according to the absolute value of the gene cluster-prognostic correlation coefficient. Modules ranked higher have stronger prognostic predictive potential, and their pivotal genes are more likely to become core targets. Functional overlap among multiple highly correlated gene clusters was assessed using gene ontology enrichment analysis. Modules with highly overlapping functions may reflect different aspects of the same biological process. The identification results of highly correlated gene clusters were cross-validated with differential expression analysis, and the proportion of differentially expressed genes within a module reflected the degree of consistency between the two analytical methods.

[0049] Functional consistency screening was performed on the highly correlated gene clusters to form key prognostic gene clusters. Functional consistency requires that the genes within a module have intrinsic biological functions, and the biological rationality of highly correlated gene clusters needs to be verified through functional enrichment analysis. Gene ontology enrichment analysis was performed to functionally annotate the gene members of each module in the highly correlated gene clusters, extracting significantly enriched biological processes, molecular functions, and cellular components. Modules in highly correlated gene clusters enriched with biological processes related to cerebral hemorrhage, such as neuroinflammation, apoptosis, blood-brain barrier function, and oxidative stress, have higher mechanistic interpretability. KEGG pathway enrichment analysis was carried out simultaneously, and modules in highly correlated gene clusters enriched with neurodegenerative disease pathways, inflammatory signaling pathways, or cell death pathways were given priority for inclusion in key prognostic gene clusters. The functional consistency score comprehensively considers the number of enriched items, the significance of enrichment, and the functional correlation between items. The score is calculated using a weighted summation method of -log10 (P-value) of enriched items. Modules in highly correlated gene clusters with functional consistency scores in the top 50% were included in key prognostic gene clusters, while modules with scattered functional annotations or lacking a clear biological theme were excluded. Key prognostic gene clusters typically contain 2-5 gene modules with well-defined functions and significant prognostic associations. The pivot genes in these modules are priority candidates for core targets. Each selected module's gene members, prognostic correlation coefficients, enriched functional entries, and functional consistency scores are recorded to form complete module annotation information.

[0050] In some embodiments, the step of using the prognostic key gene cluster to perform hub degree-prognostic joint screening to identify core hub genes includes: extracting intra-cluster connectivity from the prognostic key gene cluster to obtain hub degree ranking; performing high connectivity gene screening based on the hub degree ranking to obtain candidate hub genes; performing joint screening of prognostic association strength and brain tissue expression specificity on the candidate hub genes to obtain high-confidence hub genes; and performing cross-dataset expression consistency verification based on the high-confidence hub genes to form core hub genes.

[0051] The intra-cluster connectivity of the prognostic key gene cluster is extracted to obtain the hub degree ranking. Hub genes are those with the highest intra-module connectivity, occupying a core position in the co-expression network. Hub genes in the prognostic key gene cluster are more likely to be key factors driving module function and prognostic association. Intra-cluster connectivity is defined as the sum of the connection weights between a gene and other genes within the module. The intra-cluster connectivity of each gene in the prognostic key gene cluster is extracted and calculated from the adjacency matrix. When the prognostic key gene cluster contains two modules, blue and green, the intra-cluster connectivity of each gene is calculated independently within each module; genes between modules do not participate in the connectivity calculation between them. The formula for calculating intra-cluster connectivity is kIN=Σa_ij, where a_ij is the adjacency matrix element value between gene i and gene j within the module, summed and iterated through all genes within the module except gene i. The hub degree ranking arranges all genes in the prognostic key gene cluster from high to low intra-cluster connectivity, with the gene with the highest connectivity at the top of the ranking, representing the core hub of the module. Genes within different modules of the prognostic key gene cluster are ranked within each module, with each module generating its own hub degree ranking list. Ranking lists from multiple modules are merged to form a complete hub degree ranking. Intra-cluster connectivity can be further standardized to module membership values, using the formula kME=cor(x_i, ME), where x_i is the expression vector of gene i, ME is the module's characteristic gene, and kME reflects the degree of gene affiliation to the module. The hub degree ranking simultaneously records both the original intra-cluster connectivity and the standardized module membership value; these two metrics are highly correlated but have slightly different emphases.

[0052] Based on the hub ranking, high connectivity genes are screened to obtain candidate hub genes. Connectivity threshold screening selects top-ranking high-connectivity genes from the ranking list, with genes in the top 10% or top 20% of the hub ranking considered as hub candidates. The screening criteria for candidate hub genes can use absolute or relative thresholds. An absolute threshold, such as kME > 0.8, requires the gene to have a correlation greater than 0.8 with the module's characteristic genes. A relative threshold, such as the top 10%, selects the top 10% of genes in each module's connectivity ranking. When the blue module contains 200 genes in the hub ranking, the top 10% strategy selects the 20 genes with the highest connectivity as candidate hub genes; when the green module contains 150 genes, the top 15 genes with the highest connectivity are selected. The number of candidate hub genes is related to the number and size of modules in the prognostic key gene cluster; typical analysis yields 30-100 candidate hub genes. Genes with high connectivity but low module membership values ​​may be multi-module crossover genes; an additional screening condition of kME > 0.7 can be added to exclude these genes from the candidate hub gene pool. In hub ranking, genes with similar rankings have small differences in connectivity, and the inclusion of genes near the threshold has limited impact on the final result. The threshold can be appropriately relaxed to improve recall. For each candidate hub gene, the module to which it belongs, its intra-cluster connectivity, module membership value, and intra-module ranking are recorded, forming a preliminary candidate list of hub genes.

[0053] High-confidence hub genes were obtained by jointly screening candidate hub genes for prognostic association strength and brain tissue expression specificity. Genes with high hubness do not necessarily have strong prognostic associations; therefore, prognostic association screening criteria were combined to identify genes with both network centrality and prognostic predictive value. The prognostic association coefficients of each gene in the candidate hub genes were extracted from the univariate Cox regression results of S120, and genes with a p-value less than 0.05 and a hazard ratio deviating from 1 were retained. The prognostic association strength threshold was set as |beta|>0.3 or HR>1.5 / HR<0.67; genes in the candidate hub genes that meet this threshold have moderate to high prognostic effect sizes. Brain tissue expression specificity was assessed by querying the GTEx database. Genes whose expression levels in brain tissue ranked in the top 30% compared to other tissues were identified as highly expressed genes in brain tissue. The tissue specificity index τ quantifies the breadth of gene expression across different tissues; genes in the candidate hub genes with τ>0.7 exhibit strong tissue-specific expression patterns. High-confidence hub genes must simultaneously meet three criteria: high hub status, significant prognostic association, and brain tissue expression specificity. This triple screening significantly narrows the candidate pool and improves screening specificity. After triple screening, 10-30 genes are typically retained as high-confidence hub genes. These genes are high-priority prognostic target candidates identified in network analysis.

[0054] Core hub genes were formed by cross-dataset expression consistency validation based on the high-confidence hub genes. Analysis results from a single dataset may be affected by sample bias; therefore, the reproducibility of expression patterns for high-confidence hub genes needs to be validated in independent datasets. The validation datasets were selected from the GEO database of brain hemorrhage transcriptome datasets that did not participate in the discovery phase analysis. Expression values ​​of high-confidence hub genes were extracted from the validation datasets and correlated with prognostic phenotypes. Expression consistency was defined as the degree of consistency in the prognostic association direction between the discovery and validation datasets; high-confidence hub genes were considered to have consistent direction if the prognostic association coefficients had the same sign in both datasets. The inclusion criteria for core hub genes required that high-confidence hub genes show consistent prognostic association in at least one independent validation dataset, and that the p-value in the validation dataset be less than 0.1 (a lenient significance standard). Genes that failed validation in the high-confidence hub gene pool may be false positives from the discovery dataset or exhibit dataset-specific expression patterns; these genes were excluded from the final results. Cross-platform validation employed a strategy of discovery using microarray data, validation using sequencing data, or vice versa. Core hub genes needed to show consistency across different technology platforms to eliminate platform-specific bias. The core hub gene records the statistics of each gene during the discovery stage, the statistics during the validation stage, and the final confidence rating. The intersection of the differential genes associated with prognosis is then used to enter the target integration stage.

[0055] Step S140: Perform intersection screening on prognostic differential genes and core hub genes to obtain a preliminary screening target set; perform pathological pathway association filtering on the preliminary screening target set to obtain mechanism-related targets; and perform core node clustering of the interaction network based on mechanism-related targets to identify core target genes.

[0056] Specifically, an intersection screening process is performed on prognostic differentially expressed genes and core hub genes to obtain a preliminary target set. Intersecting prognostic-related genes identified by two independent analysis strategies improves the reliability of the results. Prognostic differentially expressed genes are obtained through integrated screening of differential expression and survival analysis, while core hub genes are obtained through joint filtering of co-expression networks and multiple validation. String matching is performed on the gene symbols of prognostic differentially expressed genes and core hub genes. Genes appearing in both gene lists are included in the preliminary target set. The intersection operation ensures that selected genes possess both differential expression and network hub characteristics. When there are 300 prognostic differentially expressed genes and 12 core hub genes, the intersection typically yields 5-10 genes; the intersection ratio reflects the consistency between the two analysis methods. Each gene in the preliminary target set inherits the risk attribute label of the prognostic differentially expressed genes and the network topology indicators of the core hub genes. These two types of information are merged to form a multi-dimensional gene annotation. An empty intersection indicates a lack of overlap between the results of the two analysis methods. The screening threshold can be appropriately relaxed or a union strategy can be used to expand the candidate range, but the methodological adjustment must be explained during result interpretation. The genes in the initial screening target set are ranked according to the absolute value of the prognostic association coefficient among differentially correlated genes, with the gene with the strongest association ranking first and being the highest priority target candidate. The statistical proportion of risk genes and protective genes among differentially correlated genes in the initial screening target set reflects the risk composition characteristics of the overlapping genes. As an integrated output of differential and network analyses, the initial screening target set needs to be validated through pathological pathway association to confirm its association with the pathogenesis of cerebral hemorrhage.

[0057] In some embodiments, the step of performing pathological pathway association filtering on the initial screening target set to obtain mechanism-related targets includes: performing functional pathway enrichment analysis on the initial screening target set to obtain target-pathway mapping relationships; screening for pathway targets related to neuroinflammation, blood-brain barrier damage, and apoptosis through the target-pathway mapping relationships to obtain pathway-related targets; assessing the pathological mechanism correlation of the pathway-related targets to obtain mechanism correlation scores; and screening for mechanism-related targets with high correlation based on the mechanism correlation scores.

[0058] Functional pathway enrichment analysis was performed on the initial screening target set to obtain target-pathway mapping relationships. Enrichment analysis revealed the biological pathways involved by the genes. The functional attribution information of each gene in the initial screening target set was obtained through pathway database annotation. The KEGG pathway database is the primary source of functional annotation; each gene in the initial screening target set was queried for its KEGG pathway entries. A single gene may participate in multiple pathways, forming a one-to-many mapping from gene to pathway. The Reactome pathway database supplements KEGG; genes lacking KEGG annotation in the initial screening target set can obtain pathway attribution information from Reactome. The annotation results of the two databases are merged to form complete pathway coverage. The target-pathway mapping relationship is stored using a bipartite graph structure. One side of the graph represents genes in the initial screening target set, and the other side represents pathway entries, with edges connecting genes to their respective pathways. When the initial screening target set contains 8 genes involved in a total of 45 pathways, the target-pathway mapping relationship includes 8 gene nodes, 45 pathway nodes, and several connecting edges. Pathway enrichment significance was assessed using a hypergeometric test. Pathways were considered enriched when the number of genes in a particular pathway in the initial screening target set was significantly higher than the random expectation. Pathways with an enrichment p-value less than 0.05 were marked as significantly enriched. The target-pathway mapping relationship simultaneously recorded the enrichment p-value, gene coverage, and pathway functional description for each pathway. Pathways were sorted by enrichment significance to facilitate the identification of core functional themes. When the number of genes in the initial screening target set is small, the statistical power of enrichment analysis is limited; alternative methods such as gene set variation analysis can be used to assess pathway associations.

[0059] This study uses target-pathway mapping to screen for pathways related to neuroinflammation, blood-brain barrier damage, and apoptosis, aiming to identify pathway-associated targets. The core pathological processes of cerebral hemorrhage include neuroinflammatory responses, blood-brain barrier disruption, and neuronal cell death. Pathways related to these three processes in the target-pathway mapping are key areas for screening. Neuroinflammation-related pathways include the NF-κB signaling pathway, TNF signaling pathway, IL-17 signaling pathway, Toll-like receptor signaling pathway, and NOD-like receptor signaling pathway. Gene markers involved in these pathways in the target-pathway mapping are identified as inflammation-associated targets. Blood-brain barrier damage-related pathways include cell adhesion molecule pathways, tight junction pathways, leukocyte transendothelial migration pathways, and matrix metalloproteinase-related pathways. Gene markers involved in these pathways in the target-pathway mapping are identified as barrier-associated targets. Apoptosis-related pathways include apoptosis pathways, necroptosis pathways, ferroptosis pathways, autophagy pathways, and p53 signaling pathways. Gene markers involved in these pathways in the target-pathway mapping are identified as death-associated targets. Pathway-associated targets are defined as genes that participate in at least one of the three pathways mentioned above as initial screening targets. The association tags of the three pathways can be accumulated, and a single gene may simultaneously possess multiple association tags related to inflammation, barriers, and death. Genes that do not participate in any core pathways of cerebral hemorrhage in the target-pathway mapping are excluded from the pathway-associated target list, as these genes may have a weak direct association with the pathological mechanisms of cerebral hemorrhage. The pathway-associated target list records the pathway association type, the specific pathway names involved, and the number of pathways for each gene; genes participating in more pathways have broader functional coverage.

[0060] Mechanism association scores were obtained by assessing the correlation between pathway-associated targets and the pathological mechanisms of intracerebral hemorrhage. The number and importance of pathways jointly determine the degree of mechanism association of genes. For each gene in a pathway-associated target, a comprehensive association score was calculated based on the number and weight of pathways it participates in. A weighted summation model was used to calculate the mechanism association score, with each pathway involved by the gene contributing a certain score. The score value was proportional to the core importance of the pathway in the pathology of intracerebral hemorrhage. Pathway weights were set based on literature evidence. Pathways frequently reported in intracerebral hemorrhage studies, such as the NF-κB signaling pathway and the apoptosis pathway, were assigned a weight of 1.0, while other relevant pathways were assigned weights of 0.5-0.8. When a gene in a pathway-associated target participates in the NF-κB signaling pathway (weight 1.0), the TNF signaling pathway (weight 0.8), and the apoptosis pathway (weight 1.0), the mechanism association score of that gene is 1.0 + 0.8 + 1.0 = 2.8. The breadth of coverage of the three pathological processes serves as a moderating factor for the score. Genes associated with all three processes—inflammation, barrier, and death—receive a 1.5-fold score bonus, while genes associated with only a single process do not receive a bonus. Literature co-occurrence frequency is recorded separately as an auxiliary reference indicator. The number of PubMed publications showing co-occurrence of each gene in the pathway-associated target with keywords such as "cerebral hemorrhage" and "intracranial hemorrhage" reflects the level of research interest. This indicator is not involved in the numerical calculation of the mechanism association score but is used for priority determination when scores are similar. After the mechanism association scores of each gene in the pathway-associated target are calculated, they are sorted from highest to lowest score. The gene with the highest score is most closely associated with the pathological mechanism of cerebral hemorrhage. The distribution characteristics of the mechanism association scores are visualized using a histogram. A clear bimodal distribution of scores indicates that the pathway-associated targets can be clearly divided into two groups: high association and low association.

[0061] Mechanism-related targets are selected based on mechanism association scores. A score threshold is used to classify pathway-related targets into high-association and low-association categories; genes with mechanism association scores exceeding the threshold are included in mechanism-related targets. Thresholds can be set using fixed values ​​or percentiles. A fixed threshold is an example of a mechanism association score > 2.0, while a percentile threshold is an example of selecting genes with scores in the top 50%. The distribution characteristics of the mechanism association scores determine the threshold selection strategy. When the scores are continuously distributed, using percentile thresholds is more robust; when there are obvious discontinuities in the scores, the discontinuities are used as natural thresholds. When the pathway-related targets include 6 genes, the top 50% strategy selects the 3 genes with the highest scores for inclusion in mechanism-related targets. If there is a significant difference in scores between the 3rd and 4th genes, this threshold has a better discriminative effect. The statistical distribution of the three types of pathological association tags for each gene in mechanism-related targets reflects the functional composition of the selected genes; the proportion of inflammation-related, barrier-related, and death-related targets reveals the functional emphasis of mechanism-related targets. The inclusion or exclusion of boundary genes with similar mechanism association scores has a limited impact on the final results; sensitivity analysis can be used to assess the degree of influence of threshold changes on the composition of mechanism-related targets. Mechanism-related targets record the mechanism association score, pathway association type, list of participating pathways, and score ranking for each gene, forming complete mechanism annotation information. The number of genes in the mechanism-related target set typically accounts for 50%-80% of the initial screening target set; genes with weak associations to the mechanism of cerebral hemorrhage are excluded through pathological pathway screening.

[0062] In some embodiments, the step of clustering core nodes of the interaction network based on the mechanism-related targets to identify core target genes includes: constructing an interaction network topology graph by querying a protein interaction database based on the mechanism-related targets; extracting functional subnetworks by performing network density clustering on the interaction network topology graph; performing network propagation and regulatory association integration on each node in the functional subnetwork to obtain node importance scores; and performing sorting and screening based on the node importance scores to form core target genes.

[0063] Based on the aforementioned mechanism, a protein-protein interaction database is queried to construct an interaction network topology graph. For example... Figure 3As shown, protein-protein interaction information reveals the physical or functional relationships between gene products. Genes in mechanism-related targets form a network structure through interactions, which helps identify core functional nodes. The STRING database is the main source of protein-protein interaction information. Proteins corresponding to genes in mechanism-related targets are used as seed protein nodes 13 to retrieve interaction relationships from the database, with an interaction confidence threshold set to 0.4. First-order interaction partners 14 are the direct neighbors of the seed proteins, with an expansion depth of 1 layer. The interaction network topology is represented by a graph data structure. Interaction edges 15 connect associated protein pairs, and edge weights 23 are indicated by line thickness to represent interaction confidence. After MCODE clustering, sub-network regions with significantly higher local connectivity density than the background are identified. Functional sub-networks A16 and B17 represent potential functional modules, and sub-network boundaries 21 are marked with dashed circles. Core node 18 is located at the center of the high-density clustering region 22, possessing the highest intra-cluster connectivity and network propagation influence. Bridging node 19 connects functional subnetwork A16 and functional subnetwork B17, playing a role in cross-module communication. Scattered node 20 is not assigned to any functional subnetwork but is connected to the network. Core node 18 and bridging node 19 in the interaction network topology are preferentially included in the core target genes.

[0064] Functional subnetworks are extracted by performing network density clustering on the interaction network topology. Network clustering divides closely interacting nodes into subgroups, and local regions with higher density in the interaction network topology may represent functionally related protein complexes or signaling modules. The MCODE algorithm is a classic method for protein interaction network clustering. After inputting the interaction network topology into MCODE, it identifies subnetwork regions with significantly higher local connectivity density than the background. MCODE parameter settings include degree threshold, node score threshold, and K-kernel value. The default parameters are degree threshold 2, node score threshold 0.2, and K-kernel value 2. These parameters can be adjusted appropriately according to the network size. The identification result of functional subnetworks is several node subsets, where the interaction density of nodes within each subset is higher than the interaction density between subsets. The size of the subset is usually 3-15 nodes. When the interaction network topology contains 20 nodes, MCODE may identify 2-4 functional subnetworks, each representing a potential functional module. The biological significance of functional subnetworks is verified by functional enrichment analysis of nodes within the subnetwork. Subnetworks enriched to consistent functional themes have high interpretability. The ClusterONE algorithm, as an alternative to MCODE, offers a different clustering perspective. Taking the intersection of the results from both methods improves the robustness of the clustering results. Nodes in the functional subnetwork not covered by any subnetwork are marked as scattered nodes; these nodes may be bridging proteins connecting different functional modules or factors that function independently. When the interaction network topology is too small to form a meaningful clustering structure, the clustering step is skipped, and node importance is directly evaluated across the entire network.

[0065] For example, the step of performing network propagation and regulatory association integration on each node in the functional sub-network to obtain a node importance score includes: performing network propagation analysis on the functional sub-network to obtain a node propagation influence score; screening high-influence nodes through the node propagation influence score; tracing the transcription factor regulatory relationship of the high-influence nodes to obtain the upstream regulatory factor correlation; and integrating the node propagation influence score and the upstream regulatory factor correlation to generate a node importance score.

[0066] Network propagation analysis is performed on functional subnetworks to obtain node propagation influence scores. Network propagation simulates the diffusion process of signals in a network. The propagation influence of each node in a functional subnetwork reflects the range and intensity of its disturbance's impact on the overall network state. The random walk algorithm is a standard method for network propagation analysis. A random walk is performed starting from a node in the functional subnetwork. During the walk, nodes move between adjacent nodes based on edge weights, with the probability of movement proportional to the edge weight. The random walk is restarted, returning to the starting node with probability r after each step. The value of r is set to 0.3 to balance local exploration and global diffusion. The walk iterates until the probability distribution converges. After the restarted random walk starting from node A in the functional subnetwork converges, the steady-state probability of each node reflects the distribution of node A's influence on the entire network. The sum of the steady-state probabilities represents the propagation influence range of node A. The node propagation influence score is defined as the product of the number of nodes with non-zero probabilities after the random walk starting from that node converges and the average probability. A higher score indicates a wider and stronger network influence range for the node. In the functional subnetwork, all nodes take turns as the starting point for random walks, and each node obtains an independent estimate of its node propagation influence score. The PageRank algorithm, as an alternative method, provides a global ranking of node importance; nodes with high PageRank values ​​are frequently pointed to by other important nodes and have high "reputation" in the network. The node propagation influence score is highly correlated with the PageRank value but calculated from a different perspective; the average of the two indicators is used as the comprehensive propagation score to improve the robustness of the assessment. When the functional subnetwork is too small, resulting in insufficient discriminative power in the propagation analysis, first-order neighbors can be included to expand the network size before conducting the propagation analysis.

[0067] High-influence nodes are selected through node propagation influence scores. A propagation influence threshold is used to identify core propagation nodes in the network; nodes with propagation influence scores in the top 50% have strong network influence. The selection threshold for high-influence nodes can be dynamically adjusted based on the score distribution. When the score distribution shows a clear head effect, nodes above the inflection point are selected; when the score distribution is flat, a fixed percentile threshold is used. When the functional subnetwork contains 12 nodes, the top 50% strategy selects the 6 nodes with the highest propagation influence scores as high-influence nodes. Correlation analysis between node propagation influence scores and node degree reveals the network's propagation characteristics. High correlation indicates a degree-driven propagation pattern, while low correlation indicates the existence of degree-independent propagation paths. High-influence nodes are typically located in central or bridging positions in the network visualization diagram, and their topological characteristics corroborate the propagation analysis results. Nodes with high degree centrality but low propagation influence scores may be local hubs, with their influence limited to the neighboring area; nodes with low degree centrality but high propagation influence may be key bridges connecting different functional modules of the network. High-influence nodes record each node's propagation influence score, PageRank value, and influence ranking; these indicators reflect the node's network propagation characteristics. High-influence nodes are used as input for regulatory relationship analysis to further assess their position in the transcriptional regulatory network.

[0068] The correlation of upstream regulatory factors is obtained by tracing the transcription factor regulatory relationships of high-influence nodes. Transcriptional regulatory relationships reveal the upstream control mechanisms of gene expression, and the degree to which high-influence nodes are regulated by transcription factors reflects their hierarchical position in the gene regulatory network. The TRRUST database contains the regulatory relationships between human transcription factors and target genes. Each gene in a high-influence node can be queried for its upstream transcription factor list. A single gene may be regulated by multiple transcription factors. Transcription factor regulatory relationships are divided into two types: activation and inhibition. The number of activating and inhibiting upstream transcription factors is counted for high-influence nodes, and the directionality of the regulatory relationship affects the gene expression regulation pattern. The basic calculation of the correlation of upstream regulatory factors is the logarithmic transformation value of the number of upstream transcription factors, with the formula R_base=log2(N_TF+1), where N_TF is the number of upstream transcription factors, and adding 1 avoids logarithmic calculation of zero values. The importance weight of upstream transcription factors is included in the calculation of the correlation of upstream regulatory factors. Transcription factors known to be associated with cerebral hemorrhage or neurological injury, such as NF-κB, STAT3, and HIF-1α, are given a weight of 2, while other transcription factors are weighted at 1. The final formula for calculating the correlation between upstream regulatory factors is R = R_base × (Σw_i / N_TF), where Σw_i / N_TF is the average weighting coefficient, and R comprehensively reflects the quantity and importance of upstream transcription factors. The ChEA database provides transcription factor-target gene relationships based on ChIP experiments, complementing the text mining results of TRRUST. The intersection of the two databases' regulatory relationships has higher reliability. Genes simultaneously regulated by multiple cerebral hemorrhage-related transcription factors in high-influence nodes may be convergence nodes of pathological pathways, possessing high therapeutic intervention value.

[0069] A node importance score is generated by integrating node propagation influence score and upstream regulatory factor correlation. Network propagation characteristics and regulatory correlation characteristics characterize the biological importance of nodes from different dimensions, and the integration of node propagation influence score and upstream regulatory factor correlation forms a comprehensive evaluation index. The integration formula for the node importance score is I = w_p × P_norm + w_r × R_norm, where I is the node importance score, P_norm is the standardized node propagation influence score, R_norm is the standardized upstream regulatory factor correlation, and w_p and w_r are the weights of the two indicators. The weight setting considers the reliability and importance of the two types of information. The weight w_p for node propagation influence score is set to 0.6 to reflect the fundamental role of network topology, and the weight w_r for upstream regulatory factor correlation is set to 0.4 to reflect the auxiliary value of regulatory relationships. The standardization of node propagation influence score and upstream regulatory factor correlation adopts the min-max method, linearly scaling the original values ​​to the range of 0-1. Specifically, the original value of each node is subtracted from the minimum value among all nodes, and then divided by the difference between the maximum and minimum values. Node importance scores range from 0 to 1, with higher values ​​indicating stronger overall importance. Nodes with scores close to 1 are considered core targets for network analysis. Among high-influence nodes, those with importance scores in the top 30% are considered core node candidates; these nodes possess both network coreness and regulatory relevance. Nodes with high propagation influence scores but low upstream regulatory factor relevance may be network hubs with relatively simple regulatory mechanisms, while nodes with high upstream regulatory factor relevance but low propagation influence may be convergence points in the regulatory network but are relatively peripheral.

[0070] Core target genes are selected through ranking and screening based on node importance scores. Importance ranking arranges nodes in the functional subnetwork from highest to lowest importance score, with higher-ranked nodes exhibiting greater network coreness and regulatory relevance. The screening threshold for core target genes can be a fixed number or a score threshold. A fixed number might be the top 5 nodes with the highest scores, while a score threshold might be nodes with an importance score > 0.6. When a functional subnetwork contains multiple subnetworks, each subnetwork undergoes independent node ranking, selecting 1-2 nodes with the highest scores from each subnetwork to ensure representative genes from each functional module are included. When node importance scores show little difference among multiple nodes, a prognostic association coefficient is introduced as an auxiliary ranking indicator, prioritizing nodes with stronger prognostic associations. The number of core target genes should be controlled within a reasonable range; too many genes lack focus, while too few may miss important targets. Typical analyses select 3-8 genes as core targets. Mechanism-related target genes that are scattered nodes (not included in any functional subnetwork) have their node importance scores calculated solely based on the topological indices of the entire network. Genes ranking in the top 30% of all nodes are also included as core target genes. Core target genes are recorded with their respective functional subnetwork, network topological indices, regulatory association information, and node importance score ranking. After multiple validations including differential analysis, network analysis, pathway screening, and interaction network filtering, core target genes are high-confidence candidates for prognostic gene targets related to cerebral hemorrhage. These are then output for final validation through immune infiltration and survival analysis.

[0071] Step S150: Based on the core target genes, perform immune infiltration association analysis to obtain immune-related target markers, perform survival analysis on the core target genes to obtain survival association indicators, and output prognostic gene targets for cerebral hemorrhage based on immune-related target markers and survival association indicators.

[0072] In some embodiments, the step of obtaining immune-related target markers by performing immune infiltration association analysis based on the core target genes includes: constructing an immune microenvironment association mapping based on the core target genes to obtain an immune cell composition profile; extracting expression feature distributions from the core target genes to obtain a target expression profile; performing correlation analysis on the target expression profile and the immune cell composition profile to obtain a target-immune association matrix; and identifying significantly associated immune cell types for each target based on the target-immune association matrix to generate immune-related target markers.

[0073] Based on the core target genes, an immune microenvironment association mapping was constructed to obtain the immune cell composition profile. Immune cell infiltration participates in neuroinflammation and tissue repair processes after cerebral hemorrhage; association analysis between core target genes and the immune microenvironment can reveal their potential role in immune regulation. Immune cell composition estimation uses a deconvolution algorithm to infer the relative abundance of various immune cell types from transcriptome data, with the standardized expression matrix as the deconvolution input. The CIBERSORT algorithm, based on the LM22 feature matrix, can estimate the relative proportions of 22 immune cell subtypes, including T cell subtypes, B cell subtypes, macrophages M0 / M1 / M2, dendritic cells, and neutrophils. The calculation of the immune cell composition profile involves deconvolution of all samples in the standardized expression matrix, yielding a 22-dimensional immune cell proportion vector for each sample. The dimension of the immune cell composition profile is the number of samples × the number of cell types, and the matrix elements are the estimated cell proportions or enrichment scores. The reliability of the deconvolution results is evaluated using a p-value; a p-value less than 0.05 indicates a good fit of the model to the sample. Differences in immune cell types within the immune cell composition profile between the good-prognosis and poor-prognosis groups were assessed using the Wilcoxon rank-sum test. Cell types with significant differences may be involved in immune processes related to the prognosis of cerebral hemorrhage.

[0074] Expression profiles are obtained by extracting expression feature distributions from the core target genes. The expression levels of target genes are fundamental data for immune association analysis. The expression values ​​of core target genes in each sample are extracted from the standardized expression matrix to form independent expression profile matrices. When the core target genes include six genes, the row vectors corresponding to these six genes are extracted from the standardized expression matrix and combined to form a 6-row × N-column target expression profile matrix. The row index of the target expression profile is the gene symbol of the core target gene, and the column index is consistent with the sample order of the immune cell composition profile to ensure correct sample pairing for subsequent correlation analysis. The expression distribution of each gene in the target expression profile is evaluated using descriptive statistics. Genes with severely skewed expression distributions may be logarithmically transformed before correlation analysis. If a gene in the core target genes is absent in some samples, the missing samples are excluded from the correlation analysis of that gene. Complete matching between the target expression profile and the immune cell composition profile samples is a prerequisite for correlation analysis. After verifying that the column dimensions of the two matrices match the sample order, the correlation analysis stage begins.

[0075] A correlation analysis was performed on the target expression profile and the immune cell composition profile to obtain the target-immune association matrix. The correlation analysis quantified the covariation between target gene expression and immune cell infiltration. A paired correlation coefficient was calculated for each gene in the target expression profile and each immune cell type in the immune cell composition profile. Spearman correlation coefficients are suitable for non-normally distributed data. The paired correlation analysis of the target expression profile and the immune cell composition profile used the Spearman method to improve robustness. When the core target genes included 6 genes and the immune cell composition profile included 22 cell types, the correlation analysis performed 132 paired calculations. The target-immune association matrix stores all correlation results in matrix form, with rows corresponding to core target genes, columns corresponding to immune cell types, and elements representing Spearman correlation coefficients. A statistical significance matrix for the target-immune association matrix was generated simultaneously, and the Benjamini-Hochberg method was used for multiple test correction to control the false discovery rate. The heatmap visualization of the target-immune association matrix uses a red-blue color scheme, with red indicating positive correlation and blue indicating negative correlation. When the core target gene is positively correlated with the pro-inflammatory cell type, it suggests that the gene may be involved in inflammation amplification; when it is positively correlated with the anti-inflammatory cell type, it suggests that the gene may be involved in inflammation resolution or tissue repair.

[0076] Based on the target-immune association matrix, significantly associated immune cell types for each target are identified, generating immune-related target markers. Significance screening extracts statistically significant target-immune pairings from the association matrix; pairings with a corrected p-value less than 0.05 and an absolute correlation coefficient greater than 0.3 are considered significantly associated. Immune-related target markers are stored in a list structure, with one marker entry corresponding to each core target gene. Each entry contains a list of significantly associated immune cell types and their corresponding correlation coefficients. If a gene in the target-immune association matrix has no significantly associated immune cell types, its immune-related target marker is set to "no significant immune association." Immune cell types of the immune-related target markers are sorted by the absolute value of their correlation coefficients, with the most strongly associated cell type listed first. The immune-related target markers also count the number of significantly associated cell types for each core target gene; a higher number indicates a broader immune regulatory effect of the target. Genes that are significantly associated with M1 and M2 macrophages in core target genes but in opposite directions may be involved in macrophage polarization regulation and have the potential to regulate the inflammation-repair balance.

[0077] Survival analysis was performed on core target genes to obtain survival association indicators. Survival analysis quantified the impact of target gene expression on patient prognosis. The prognostic association of each gene in the core target genes was assessed using Kaplan-Meier survival curves and a Cox regression model. The expression values ​​of the core target genes were used to divide the sample into high-expression and low-expression groups based on the median. The Kaplan-Meier method was used to estimate the survival curves for each group, and the differences between the groups were compared using the Log-rank test. The primary component of the survival association indicators was the Log-rank test p-value; a p-value less than 0.05 indicated a significant association between gene expression level and patient survival outcome. Univariate Cox regression was used to analyze the association between continuous expression values ​​of each gene in the core target genes and survival time, obtaining the hazard ratio (HR) and 95% confidence interval. The second component of the survival association indicators was the Cox regression statistic, including the hazard ratio, confidence interval, and Wald test p-value. Multivariate Cox regression incorporated core target gene expression values ​​along with covariates such as age and bleeding volume into the model to assess the independent prognostic predictive value of gene expression. The survival association index summary table includes the Log-rank P-value, univariate HR, multivariate HR, and prognostic association direction for each gene in the core target genes, forming a complete survival analysis results archive.

[0078] Based on the aforementioned immune-related target markers and survival-related indices, prognostic gene targets for cerebral hemorrhage are output. The final output prognostic targets are determined through the integration and screening of the two types of validation results. Immune-related target markers provide supporting evidence at the level of immune mechanisms, while survival-related indices provide validation evidence at the level of clinical prognosis. The inclusion criteria for prognostic gene targets for cerebral hemorrhage require core target genes to simultaneously meet both immune-related and survival-related conditions: the immune-related target marker has at least one significantly associated immune cell type, and the Log-rank P-value of the survival-related indices is less than 0.05. Genes among the core target genes that meet only one condition are retained as candidate targets and marked as pending validation. The priority ranking of prognostic gene targets for cerebral hemorrhage comprehensively considers the number of associated cells of the immune-related target markers and the statistical significance of the survival-related indices. The risk attributes of prognostic gene targets for cerebral hemorrhage are inherited from the gene risk attribute labels of differentially correlated prognostic genes, distinguishing them into two categories: risk targets and protective targets. Complete annotation information for gene targets related to the prognosis of cerebral hemorrhage was integrated from all analysis results from S120 to S150, including differential expression statistics, network topology indicators, pathway enrichment information, immune-associated cell types, and survival analysis statistics. These gene targets, identified through multiple validation screenings including differential expression analysis, co-expression networks, protein-protein interaction networks, pathway associations, immune infiltration, and survival analysis, possess high reliability and biological significance.

[0079] To implement the above-described method embodiments, a bioinformatics analysis method for identifying gene targets related to the prognosis of cerebral hemorrhage is proposed to achieve the corresponding functional and technical effects. See also Figure 4 , Figure 4 This diagram illustrates a structural block diagram of a bioinformatics analysis system 400 for identifying gene targets related to the prognosis of cerebral hemorrhage, as provided in an embodiment of this application. For ease of explanation, only the parts relevant to this embodiment are shown. The bioinformatics analysis system 400 for identifying gene targets related to the prognosis of cerebral hemorrhage, as provided in this embodiment, includes:

[0080] Data processing unit 401 is used to acquire a brain hemorrhage transcriptome expression dataset and perform quality filtering and batch effect correction on the brain hemorrhage transcriptome expression dataset to generate a standardized expression matrix.

[0081] The differential screening unit 402 is used to perform differential expression analysis based on the standardized expression matrix and grouped by prognostic phenotype to extract differentially expressed gene sets, and to perform expression-prognostic correlation screening on the differentially expressed gene sets to obtain prognostic associated differential genes.

[0082] Gene identification unit 403 is used to construct a weighted gene co-expression network to generate a co-expression gene cluster set from the standardized expression matrix, perform prognostic correlation analysis based on the co-expression gene cluster set to extract key prognostic gene clusters, and use the key prognostic gene clusters to perform hub degree-prognostic joint screening to identify core hub genes.

[0083] The target mining unit 404 is used to perform intersection screening on the prognostic associated differential genes and the core hub genes to obtain a preliminary screening target set, perform pathological pathway association filtering on the preliminary screening target set to obtain mechanism-related targets, and perform interaction network core node clustering based on the mechanism-related targets to identify core target genes.

[0084] The verification output unit 405 is used to perform immune infiltration association analysis based on the core target genes to obtain immune-related target markers, perform survival analysis on the core target genes to obtain survival association indicators, and output prognostic gene targets for cerebral hemorrhage based on the immune-related target markers and the survival association indicators.

[0085] The bioinformatics analysis system 400 described above for identifying gene targets related to the prognosis of cerebral hemorrhage can implement one of the bioinformatics analysis methods for identifying gene targets related to the prognosis of cerebral hemorrhage as described in the above method embodiments. The options in the above method embodiments are also applicable to this embodiment and will not be detailed here. The remaining content of this application embodiment can be referred to the content of the above method embodiments, and will not be repeated in this embodiment.

[0086] The above embodiments are not an exhaustive list based on the present invention, and there may be many other embodiments not listed. Any substitutions and improvements made without departing from the concept of the present invention are within the protection scope of the present invention.

Claims

1. A bioinformatics analysis method for identifying gene targets related to the prognosis of cerebral hemorrhage, characterized in that, include: Obtain the cerebral hemorrhage transcriptome expression dataset, and perform quality filtering and batch effect correction on the cerebral hemorrhage transcriptome expression dataset to generate a standardized expression matrix; Based on the standardized expression matrix, differential expression analysis is performed on genes grouped by prognostic phenotype to extract differentially expressed gene sets. Expression-prognostic correlation screening is then performed on these gene sets to obtain prognostic-associated differentially expressed genes. This includes: performing univariate survival regression analysis on each gene in the differentially expressed gene set to obtain prognostic association coefficients; conducting time-dependent analysis on the prognostic association coefficients to distinguish between early-associated and late-associated genes; performing risk orientation labeling on the early-associated and late-associated genes to obtain gene risk attribute labels; and integrating the early-associated and late-associated genes based on the gene risk attribute labels to form prognostic-associated differentially expressed genes. A weighted gene co-expression network is constructed from the standardized expression matrix to generate a co-expression gene cluster set. Based on the co-expression gene cluster set, prognostic correlation analysis is performed to extract key prognostic gene clusters. These key prognostic gene clusters are then used to perform a pivotality-prognostic joint screening to identify core pivot genes. This process includes: extracting intra-cluster connectivity from the key prognostic gene clusters to obtain a pivotality ranking; performing high-connectivity gene screening based on the pivotality ranking to obtain candidate pivot genes; performing a joint screening of prognostic correlation strength and brain tissue expression specificity on the candidate pivot genes to obtain high-confidence pivot genes; and verifying cross-dataset expression consistency based on the high-confidence pivot genes to form core pivot genes. The process involves: performing intersection screening on the prognostic differentially expressed genes and the core hub genes to obtain a preliminary target set; and then performing pathological pathway association filtering on the preliminary target set to obtain mechanism-related targets. This includes: performing functional pathway enrichment analysis on the preliminary target set to obtain target-pathway mapping relationships; screening for pathway targets related to neuroinflammation, blood-brain barrier damage, and apoptosis through the target-pathway mapping relationships to obtain pathway-related targets; assessing the pathological mechanism association degree of the pathway-related targets to obtain a mechanism association score; screening for highly associated targets based on the mechanism association score to form mechanism-related targets; and performing core node clustering of the interaction network based on the mechanism-related targets to identify core target genes. The process of obtaining immune-related target markers through immune infiltration association analysis based on the core target genes includes: constructing an immune microenvironment association mapping based on the core target genes to obtain an immune cell composition profile; extracting expression feature distributions from the core target genes to obtain a target expression profile; performing correlation analysis on the target expression profile and the immune cell composition profile to obtain a target-immune association matrix; identifying significantly associated immune cell types for each target based on the target-immune association matrix to generate immune-related target markers; performing survival analysis on the core target genes to obtain survival association indicators; and outputting prognostic gene targets for cerebral hemorrhage based on the immune-related target markers and the survival association indicators.

2. The method according to claim 1, characterized in that, The step of performing prognostic correlation analysis based on the co-expressed gene clusters to extract key prognostic gene clusters includes: Characteristic genes of each gene cluster were extracted from the co-expressed gene cluster set; Prognostic phenotype correlation analysis was performed on the characteristic genes to obtain the gene cluster-prognostic correlation coefficient; Highly associated gene clusters are identified by significance ranking based on the gene cluster-prognostic correlation coefficient. Functional consistency screening was performed on the highly associated gene clusters to form key prognostic gene clusters.

3. The method according to claim 1, characterized in that, The method of clustering and identifying core target genes in the core node of the interaction network based on the aforementioned mechanism includes: Based on the aforementioned mechanism, a protein interaction database was queried to identify relevant targets, and an interaction network topology was constructed. The network density clustering function is used to extract subnetworks from the interaction network topology graph. For each node in the functional sub-network, network propagation and regulation correlation integration are performed to obtain node importance scores; Based on the node importance score, a sorting and screening process is performed to form core target genes.

4. The method according to claim 2, characterized in that, The step of performing prognostic phenotype correlation analysis on the characteristic genes to obtain gene cluster-prognostic correlation coefficients includes: Clinical phenotype associations are analyzed from the feature genes, clinical subtype labels are extracted, and then a clinical classification vector is constructed. Based on the clinical subtyping vector, the characteristic genes are stratified to obtain subtype-level expression data; Based on the subtype hierarchical expression data, a prognostic phenotypic correlation test was performed to obtain the subtype-specific correlation coefficient; Based on the subtype-specific correlation coefficients, cross-subtype consistency integration was performed to generate gene cluster-prognostic correlation coefficients.

5. The method according to claim 3, characterized in that, The step of performing network propagation and regulation correlation integration on each node in the functional sub-network to obtain a node importance score includes: Network propagation analysis is performed on the functional sub-network to obtain node propagation influence scores; High-influence nodes are selected by spreading influence scores among the nodes. Transcription factor regulatory relationships were traced for the high-influence nodes to obtain the correlation of upstream regulatory factors; A node importance score is generated by integrating the node propagation influence score with the correlation of the upstream regulatory factors.

Citation Information

Patent Citations

  • Method for identifying membranous nephropathy key treatment target by combining multi-omics and Mendel randomization and application thereof

    CN117542406A

  • Application of CHSY3 as a biomarker in the diagnosis and / or treatment of liver cancer

    CN119736397A