A method and system for multi-dimensional mining and knowledge discovery of clinical data of cerebral hemorrhage
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-18
- Publication Date
- 2026-08-11
AI Technical Summary
这些信息分散存储于不同医疗机构的信息平台中,受限于各机构间数据格式和编码体系的差异,难以直接进行跨机构的整合利用
[0007]The beneficial effects of this invention are reflected in the following points: First, by implementing unified cleaning, transformation and normalization processing on multi-center clinical data, the differences in data formats and coding standards between different medical institutions are eliminated. On this basis, cross-dimensional coupling deviation analysis is carried out from dimensions such as level of consciousness, amount of bleeding and location of hematoma. Cooperative and contradictory features between dimensions are distinguished and differentiated fusion weights are assigned accordingly. At the same time, a critical period time window in the course of the disease is set to carry out event association rule mining, and two types of temporal association patterns, namely cascade effect of complications and independent delay effect, are identified. This makes the representation of clinical features and the association analysis of disease events take into account both the complementarity of multi-dimensional information and the hierarchical nature of temporal evolution. Secondly, based on the prognostic outcome features in the association rule knowledge base, a hierarchical standard was set, and the weight coefficient of bleeding site and the site-volume interaction coefficient were introduced as clustering constraints. Prognostic-oriented hierarchical clustering was implemented for the patient population. Based on the obtained patient subgroups, the association parameters of treatment timing and intervention methods were extracted, and the weighted similarity between the plans was calculated. Principal component analysis was used to form the characteristic components of the treatment plan, constructing a mapping relationship from patient clinical characteristics through treatment plan to prognostic outcome. This ensures that the characterization of treatment patterns takes into account both patient heterogeneity and plan differences. Finally, by performing paired interaction effect detection on prognostic influencing factors, the strength of synergistic and antagonistic effects between factors was quantified. Sensitivity correction was performed on high-risk and low-risk disease groups according to the characteristics of patient comorbidities, and individualized factor weight thresholds were generated to screen key prognostic factors for each patient. The factor weights were then verified and corrected by comparison with follow-up data. The final knowledge rules were labeled with evidence-based levels according to accuracy, support, and reproducibility, so that the prognostic analysis can adapt to individual patient differences, and the output decision support results have traceable evidence grading.
Smart Images

Figure CN121905570B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical data analysis technology, and in particular to a method and system for multidimensional mining and knowledge discovery of clinical data on cerebral hemorrhage. Background Technology
[0002] Intracerebral hemorrhage is a common acute and critical condition in neurosurgery, characterized by rapid progression and highly variable prognoses. Clinical practice has resulted in the accumulation of a vast amount of multi-source medical information, including patient vital signs records, imaging assessments, surgical procedure information, and rehabilitation follow-up data. This information is scattered across different medical institutions' information platforms, and due to differences in data formats and coding systems, direct cross-institutional integration and utilization are difficult. Furthermore, the prognosis of intracerebral hemorrhage is influenced by multiple factors, including the patient's admission status, hemorrhage characteristics, treatment strategies, and postoperative management. Relying solely on clinical experience often fails to fully grasp the complex interrelationships between these factors, thus limiting the precision of diagnostic and treatment decisions.
[0003] Existing clinical data analysis methods often focus on the independent assessment of single-dimensional indicators, such as the impact of hemorrhage size or surgical approach on prognosis, lacking the ability to systematically explore the coupling relationships and temporal evolution patterns between different indicators. Furthermore, patients with cerebral hemorrhage exhibit significant individual differences, with patients of varying underlying conditions responding differently to the same treatment regimen. Existing methods are insufficient in identifying individualized risk factors and performing stratified analysis, thus limiting the clinical applicability of the analytical conclusions. Summary of the Invention
[0004] This invention discloses a method and system for multidimensional mining and knowledge discovery of clinical data on cerebral hemorrhage. By standardizing and integrating multi-center clinical data and extracting multidimensional features, the method mines the temporal correlation rules between key events in the course of the disease. Based on this, it implements prognosis-oriented patient grouping and treatment mode analysis, and identifies key prognostic factors by combining factor interaction effect detection and individualized weight adjustment. Finally, a clinical decision support knowledge base with evidence-based level labeling is formed after follow-up data verification.
[0005] The first aspect of this invention proposes a method for multidimensional mining and knowledge discovery of clinical data on cerebral hemorrhage, comprising the following steps: Acquire electronic medical records, surgical records, and follow-up data of patients with multicenter intracerebral hemorrhage, and standardize and integrate the data to establish a clinical data feature set; Multidimensional feature extraction is performed on the clinical data feature set to generate a clinical feature spectrum. Time window association rule mining is performed on the clinical data feature set to determine the temporal association strength distribution. The association pattern screening and evaluation are performed by combining the temporal association strength distribution with the clinical feature spectrum to form an association rule knowledge base. Based on the association rule knowledge base, prognosis-oriented patient subgroup clustering is performed to determine the patient type distribution characteristics. Based on the patient type distribution characteristics, treatment plan feature components are extracted by treatment plan similarity calculation. Treatment pattern atlas is formed by combining the treatment plan feature components with the clinical data feature set. Based on the treatment pattern atlas, a multi-dimensional prognostic correlation analysis is performed to establish a set of prognostic influencing factors. Based on the set of prognostic influencing factors, the individualized factor weight threshold is determined by factor interaction effect detection. Based on the individualized factor weight threshold, the factor importance is ranked to form a combination of key prognostic factors. Based on the follow-up data, the combination of key prognostic factors is revised to form validated knowledge rules. Based on the validated knowledge rules, evidence-based level labeling is implemented to output clinical decision support results, thus completing the clinical data knowledge discovery for cerebral hemorrhage.
[0006] A second aspect of this invention proposes a multi-dimensional mining and knowledge discovery system for clinical data of cerebral hemorrhage, comprising: The data acquisition module is used to acquire electronic medical records, surgical records and follow-up data of patients with multicenter intracerebral hemorrhage, and to standardize and integrate the data to establish a clinical data feature set; The rule mining module is used to extract multidimensional features from the clinical data feature set to generate a clinical feature spectrum, perform time window association rule mining on the clinical data feature set to determine the temporal association strength distribution, and perform association pattern screening and evaluation by combining the temporal association strength distribution with the clinical feature spectrum to form an association rule knowledge base. The pattern discovery module is used to determine the patient type distribution characteristics by performing prognostic-oriented patient subgroup clustering based on the association rule knowledge base, extract treatment plan feature components by calculating treatment plan similarity based on the patient type distribution characteristics, and form a treatment pattern map based on the treatment plan feature components and the clinical data feature set. The factor analysis module is used to perform multi-dimensional prognostic correlation analysis based on the treatment pattern map to establish a set of prognostic influencing factors, determine individualized factor weight thresholds through factor interaction effect detection based on the set of prognostic influencing factors, and perform factor importance ranking based on the individualized factor weight thresholds to form a combination of key prognostic factors. The knowledge verification module is used to modify the combination of prognostic key factors based on the follow-up data to form verified knowledge rules, implement evidence-based level labeling based on the verified knowledge rules, output clinical decision support results, and complete the clinical data knowledge discovery for cerebral hemorrhage.
[0007] The beneficial effects of this invention are reflected in the following points: First, by implementing unified cleaning, transformation and normalization processing on multi-center clinical data, the differences in data formats and coding standards between different medical institutions are eliminated. On this basis, cross-dimensional coupling deviation analysis is carried out from dimensions such as level of consciousness, amount of bleeding and location of hematoma. Cooperative and contradictory features between dimensions are distinguished and differentiated fusion weights are assigned accordingly. At the same time, a critical period time window in the course of the disease is set to carry out event association rule mining, and two types of temporal association patterns, namely cascade effect of complications and independent delay effect, are identified. This makes the representation of clinical features and the association analysis of disease events take into account both the complementarity of multi-dimensional information and the hierarchical nature of temporal evolution. Secondly, based on the prognostic outcome features in the association rule knowledge base, a hierarchical standard was set, and the weight coefficient of bleeding site and the site-volume interaction coefficient were introduced as clustering constraints. Prognostic-oriented hierarchical clustering was implemented for the patient population. Based on the obtained patient subgroups, the association parameters of treatment timing and intervention methods were extracted, and the weighted similarity between the plans was calculated. Principal component analysis was used to form the characteristic components of the treatment plan, constructing a mapping relationship from patient clinical characteristics through treatment plan to prognostic outcome. This ensures that the characterization of treatment patterns takes into account both patient heterogeneity and plan differences. Finally, by performing paired interaction effect detection on prognostic influencing factors, the strength of synergistic and antagonistic effects between factors was quantified. Sensitivity correction was performed on high-risk and low-risk disease groups according to the characteristics of patient comorbidities, and individualized factor weight thresholds were generated to screen key prognostic factors for each patient. The factor weights were then verified and corrected by comparison with follow-up data. The final knowledge rules were labeled with evidence-based levels according to accuracy, support, and reproducibility, so that the prognostic analysis can adapt to individual patient differences, and the output decision support results have traceable evidence grading. Attached Figure Description
[0008] The accompanying drawings illustrate specific examples of the technical solutions described in this invention and, together with the detailed embodiments, form part of the specification, serving to explain the technical solutions, principles, and effects of this invention.
[0009] Figure 1 This is a flowchart illustrating a method for multidimensional mining and knowledge discovery of clinical data related to cerebral hemorrhage, as described in this invention.
[0010] Figure 2 This is a structural block diagram of a multidimensional mining and knowledge discovery system for clinical data of cerebral hemorrhage according to the present invention. Detailed Implementation
[0011] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0012] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0013] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0014] The technical solutions of the embodiments of this application are described below.
[0015] like Figure 1 As shown, this embodiment of the invention provides a method for multidimensional mining and knowledge discovery of clinical data on cerebral hemorrhage, including the following steps S110-S150: Step S110: Obtain electronic medical records, surgical records, and follow-up data of patients with multicenter cerebral hemorrhage, and standardize and integrate the data to establish a clinical data feature set.
[0016] Specifically, electronic medical records, surgical records, and follow-up data of patients with intracerebral hemorrhage from multiple centers were acquired. The multi-center data collection covered patients with intracerebral hemorrhage admitted to the neurosurgery departments of 12 tertiary-level general hospitals between 2018 and 2024. Data sources included hospital information systems, surgical anesthesia systems, and outpatient follow-up systems. Electronic medical records came from the hospital information system, surgical records from the surgical anesthesia system, and follow-up data from the outpatient follow-up system. Electronic medical records contained three types of documents: admission records, progress notes, and discharge summaries. Key data fields in the electronic medical records included patient age, gender, Glasgow Coma Scale score at admission, pupillary response, motor function score, systolic blood pressure, diastolic blood pressure, hematoma volume, hematoma location, involvement of deep structures, degree of brainstem compression, whether rupture into the ventricles, degree of midline shift, and comorbidity information (hypertension, diabetes, coronary heart disease, chronic obstructive pulmonary disease, liver dysfunction, renal dysfunction), etc. Surgical records comprise three types of documents: surgical request form, surgical record form, and anesthesia record form. Key data fields in the surgical records include surgical type (craniotomy for hematoma evacuation, minimally invasive puncture and drainage, conservative treatment), timing of surgery (hours after onset), intraoperative hematoma clearance rate, and surgical duration. Follow-up data is extracted from the outpatient follow-up system and telephone follow-up records. Follow-up time points are set at 1 month, 3 months, 6 months, and 12 months after discharge. Follow-up data includes modified Rankin Scale score, Barthel Index, survival status, rebleeding events, and complication occurrences (pulmonary infection, gastrointestinal bleeding, epileptic seizures, etc.). The modified Rankin Scale score serves as a core indicator for prognostic evaluation, reflecting the patient's functional recovery degree. Electronic medical records, surgical records, and follow-up data for multicenter intracerebral hemorrhage patients are stored in raw data table format, with variations in data format and coding standards among different medical institutions.
[0017] Data standardization and integration were conducted to establish a clinical data feature set. The data standardization process eliminated data format differences and coding inconsistencies between different medical institutions, including three steps: data cleaning, data transformation, and data normalization. Data cleaning identified and processed missing values, outliers, and duplicate records in the raw data. Variables with a missing value ratio below 20% were filled using imputation, while variables with a missing value ratio above 20% were removed. Outlier detection was performed using the interquartile range (IIR). Data transformation unified the data formats and coding standards of different medical institutions to the International Classification of Diseases (ICD). Hematoma location coding uniformly adopted anatomical region coding (basal ganglia, thalamus, lobes, brainstem, cerebellum), and surgical type coding uniformly adopted surgical procedure classification coding. Data normalization converted numerical variables in the clinical data feature set to the same units and numerical ranges. Data integration linked electronic medical records, surgical records, and follow-up data according to the patient master index and time series to form a complete patient medical record trajectory. The patient master index consisted of the medical institution code, patient hospital number, and admission date. The clinical data feature set was constructed by integrating patient medical records into a timeline. The starting point of the timeline was the onset time, and key time nodes included admission time, surgery time, discharge time, and follow-up time. The clinical data feature set consisted of three subsets: baseline features, treatment features, and prognostic features. Baseline features included variables such as age, sex, Glasgow Coma Scale score at admission, pupillary response, motor function score, hematoma volume, hematoma location, involvement of deep structures, degree of brainstem compression, ventricular rupture, midline shift, and comorbidities. Treatment features included variables such as surgical type, timing of surgery, intraoperative hematoma clearance rate, operation duration, and postoperative complications. Prognostic features included variables such as modified Rankin Scale score, Barthel Index, survival status, and rebleeding events.
[0018] Step S120: Perform multi-dimensional feature extraction on the clinical data feature set to generate a clinical feature spectrum; perform time window association rule mining on the clinical data feature set to determine the temporal association strength distribution; and perform association pattern screening and evaluation by combining the temporal association strength distribution with the clinical feature spectrum to form an association rule knowledge base.
[0019] In some embodiments, the step of extracting multidimensional features from the clinical data feature set to generate a clinical feature spectrum includes: extracting features of level of consciousness, bleeding volume, and hematoma location from the clinical data feature set to generate a multidimensional original feature set; performing cross-dimensional coupling bias analysis on the multidimensional original feature set to identify synergistic and contradictory features between dimensions; calculating bias compensation weights based on the synergistic and contradictory features between dimensions to generate hierarchical weight parameters; and performing multidimensional fusion on the hierarchical weight parameters to construct a clinical feature spectrum.
[0020] Multidimensional raw feature sets were generated by extracting features related to level of consciousness, hemorrhage volume, and hematoma location from the clinical data feature set. The feature matrix of the clinical data feature set contains patient samples and feature variables. The multidimensional raw feature sets extracted feature values for three core dimensions from the clinical data feature set. The level of consciousness feature was extracted from the baseline feature subset of the clinical data feature set, including three variables: Glasgow Coma Scale score at admission, pupillary response, and motor function score. The Glasgow Coma Scale score ranges from 3 to 15 points, reflecting the patient's level of consciousness. Pupil response is divided into two categories: presence and absence of pupillary light reflex. The motor function score ranges from 1 to 6 points, reflecting limb motor ability. The hemorrhage volume feature was extracted from the baseline feature subset of the clinical data feature set, including three variables: hematoma volume, degree of midline shift, and whether it has ruptured into the ventricles. The hematoma volume ranges from 5 to 150 ml, reflecting the amount of blood loss. The degree of midline shift ranges from 0 to 20 mm, reflecting the mass effect. Rupture into the ventricles is a binary variable. Hematoma location characteristics were extracted from the baseline feature subset of the clinical data feature set, including three variables: anatomical region coding, involvement of deep structures, and degree of brainstem compression. Anatomical region coding was a five-category variable (basal ganglia, thalamus, cerebral lobes, brainstem, and cerebellum), deep structure involvement was a two-category variable, and the degree of brainstem compression was divided into three levels: no compression, mild compression, and severe compression. The multidimensional original feature sets were organized in the form of feature sub-matrices, corresponding to the three dimensions of level of consciousness, hemorrhage volume, and hematoma location, respectively.
[0021] Cross-dimensional coupling bias analysis was performed on a multidimensional original feature set to identify synergistic and contradictory features between dimensions. The cross-dimensional coupling bias analysis assessed the interactions between features of different dimensions in the multidimensional original feature set. The analytical method identified synergistic and contradictory patterns by constructing an inter-dimensional correlation matrix, and the correlation coefficients were calculated using Pearson correlation coefficient or Spearman rank correlation coefficient. The inter-dimensional correlation matrix calculated the pairwise correlation coefficients of the three dimensions of consciousness level, hemorrhage volume, and hematoma location in the multidimensional original feature set. A negative correlation coefficient between consciousness level and hemorrhage volume indicates a negative correlation; the larger the hemorrhage volume, the lower the consciousness level. This negative correlation represents a synergistic feature between dimensions. The correlation between consciousness level and hematoma location in the multidimensional original feature set exhibited a stratified pattern, with brainstem hemorrhage having a stronger impact on consciousness level. This stratified correlation also represents a synergistic feature between dimensions. The interaction test between hemorrhage volume and hematoma location revealed differences in prognosis for patients with the same hemorrhage volume but different hemorrhage locations. This interaction also represents a synergistic feature between dimensions. Interdimensional contradictory features identified anomalous samples in the multidimensional original feature set where the level of consciousness did not match the amount of bleeding. Severe altered consciousness in some patients with small hemorrhages suggested that the hematoma was located in a critical functional area, while good consciousness in some patients with large hemorrhages suggested strong individual compensatory capacity. Interdimensional synergistic features are feature combinations where the changes in each dimension are consistent and point to the same prognostic outcome, while interdimensional contradictory features are feature combinations where the prognostic directions suggested by each dimension are inconsistent. Cross-dimensional coupling bias analysis identified approximately 80% of the samples as synergistic and approximately 20% as contradictory.
[0022] Hierarchical weight parameters are generated by calculating bias compensation weights based on inter-dimensional collaborative features and inter-dimensional contradictory features. The bias compensation weight calculation adjusts the prognostic predictive contribution of each dimension feature. Collaborative feature samples use standard weights, while contradictory feature samples have their weights adjusted according to the degree of contradiction. Standard weights are assigned to the three dimensions of level of consciousness, hemorrhage volume, and hematoma location, respectively, using w1_standard, w2_standard, and w3_standard, based on the predictive contribution of each dimension and satisfying w1_standard + w2_standard + w3_standard = 1. The weight adjustment for contradictory feature samples uses a bias compensation coefficient, calculated as C_i = 1 + β × (r_i actual - r_i expected) / σ_i, where C_i is the bias compensation coefficient for the i-th dimension, β is the adjustment intensity coefficient set to 0.3, r_i actual is the actual correlation coefficient between the feature and the prognosis, r_i expected is the expected correlation coefficient under the collaborative feature model, and σ_i is the standard deviation of the correlation coefficient. Weights are increased when C_i is greater than 1 and decreased when C_i is less than 1. The formula for adjusting the weights is w_i_adjusted = (w_i_standard × C_i) / Σ(w_j_standard × C_j), where Σ is the normalization factor for summing over the three dimensions. For contradictory feature samples with small hemorrhage but severe impaired consciousness, the weight of consciousness level is increased, and the weight of hemorrhage volume is decreased. For contradictory feature samples with large hemorrhage but good consciousness, the weight of hemorrhage volume is increased, and the weight of consciousness level is decreased. The bias compensation weight calculation considers hematoma location correction: the location weight is increased for brainstem hemorrhage samples, and decreased for lobar hemorrhage samples. The hierarchical weight parameters are divided into two parts: the co-feature weight set {w1_standard, w2_standard, w3_standard} and the contradictory feature weight set {w1_adjusted, w2_adjusted, w3_adjusted}.
[0023] A multi-dimensional fusion of stratified weight parameters is used to construct a clinical feature spectrum. This fusion involves weighting and summing the feature values of three dimensions—level of consciousness, amount of bleeding, and hematoma location—according to the stratified weight parameters to generate a comprehensive feature score for each patient sample. The comprehensive feature score is calculated by multiplying the normalized value of each dimension feature by its corresponding stratified weight parameter and then summing the results. The sum of the weight parameters is normalized to 1 to ensure comparability of comprehensive scores across different patients. For samples with synergistic features, the comprehensive score is calculated using the standard weights in the stratified weight parameters; the three dimensions jointly determine the comprehensive score according to their standard weight proportions. For samples with contradictory features, the comprehensive score is calculated using adjusted weights in the stratified weight parameters. Weight adjustments enhance the contribution of the dominant dimension feature to the comprehensive score while weakening the contribution of the secondary dimension feature, thus more accurately reflecting the actual prognostic risk of the sample. The clinical feature spectrum includes the comprehensive feature score and individual dimension scores for all patient samples, and also labels the synergistic or contradictory feature types of each sample. Feature type labeling facilitates the selection of appropriate association rules for different feature patterns in subsequent processing. The clinical feature spectrum is visualized in the form of score distribution histograms and radar charts. The score distribution histograms show the frequency distribution and central tendency of the comprehensive scores in the clinical feature spectrum, while the radar charts show the relative contribution and balance of each dimension of the clinical feature spectrum. The visualization helps clinicians intuitively understand the distribution patterns of patient features.
[0024] In some embodiments, the step of performing time window association rule mining on the clinical data feature set to determine the temporal association strength distribution includes: extracting key medical event time nodes and event sequence parameters from the clinical data feature set; setting critical period time windows in the disease course based on the key medical event time nodes; obtaining event association temporal patterns by combining the event sequence parameters and the critical period time windows in the disease course; and generating a temporal association strength distribution based on the event association temporal patterns.
[0025] Key medical event time nodes and event sequence parameters were extracted from the clinical data feature set. The patient medical record timeline of the clinical data feature set records the complete sequence of events in each patient's medical history. Key medical event time nodes and event sequence parameters were extracted from the treatment feature subset and prognostic feature subset of the clinical data feature set. Key medical event time nodes include the occurrence time of four categories: admission events, surgical events, complication events, and prognostic events. The key medical event time node for admission events is the admission time (hours after onset). The level of consciousness upon admission is classified into three levels according to the Glasgow Coma Scale: awake (13-15 points), lethargic (9-12 points), and comatose (3-8 points). The volume of hematoma upon admission is classified into three levels according to size: small (<30ml), moderate (30-60ml), and large (>60ml). The key medical event time point for surgical events is the operation time (hours after onset). The surgical type is craniotomy for hematoma evacuation, minimally invasive puncture and drainage, or conservative treatment. Hematoma clearance rate is classified into three levels: complete clearance (>90%), substantial clearance (70-90%), and partial clearance (<70%). The key medical event time point for complication events is the complication occurrence time (days after onset). Complication types include pulmonary infection, gastrointestinal bleeding, seizures, and rebleeding. The key medical event time point for prognostic events is the prognostic assessment time. A good prognosis is defined as a modified Rankin Scale score of 0-2, and a poor prognosis as a score of 3-6. Event sequence parameters record the order and time intervals of events for each patient. The event sequence parameters extract the chronological relationship of key medical event time points from the clinical data feature set, and the time intervals of the event sequence parameters are converted to hours.
[0026] Critical time windows for the course of the disease are established based on key medical event time points. The critical period is defined as the time period in the course of intracerebral hemorrhage where the condition changes drastically and treatment decisions are crucial. The establishment of critical time windows is based on the distribution patterns of key medical event time points and clinicopathophysiological mechanisms. For the hyperacute phase, the critical time window is set at 0 to 6 hours after onset. Key medical events related to hematoma expansion mainly occur within 3 hours of onset (approximately 30% of patients experience hematoma expansion during this period). Surgery during the hyperacute phase can effectively reduce secondary brain injury. For the acute phase, the critical time window is set at 6 to 72 hours after onset. Key medical events related to increased intracranial pressure and brain herniation are concentrated within this critical time window. For the subacute phase, the critical time window is set at 3 to 14 days after onset. Key medical events related to early rehabilitation intervention occur within this critical time window, and early rehabilitation intervention can improve long-term prognosis. The critical time window for the recovery period is set at 14 days to 6 months after onset, with neuroplasticity most active in the first 3 months. The boundaries of the critical time window allow for individual variation adjustment. The critical time window is stored in the form of time intervals: hyperacute phase [0, 6 hours], acute phase [6 hours, 72 hours], subacute phase [3 days, 14 days], and recovery phase [14 days, 180 days].
[0027] Event-related temporal patterns were obtained by combining event sequence parameters with critical period time windows in the disease course. These patterns describe the co-occurrence and chronological relationships of medical events within different critical period time windows. Pattern extraction involves mapping event sequence parameters to critical period time windows and analyzing event co-occurrence frequencies. Analysis of the association pattern between admission and surgical events within the critical period time window of the hyperacute phase revealed that most patients admitted while comatose underwent surgery during the hyperacute phase, while the proportion of hyperacute surgery was lower among patients admitted while conscious. This pattern is described as "admission while comatose → surgery during the hyperacute phase." Analysis of the association pattern between surgical events and complication events within the critical period time window of the acute phase revealed that the incidence of pulmonary infection was higher in patients undergoing craniotomy for hematoma evacuation than in those undergoing minimally invasive puncture and drainage, while the incidence of pulmonary infection was lowest in patients receiving conservative treatment. This pattern is described as "craniotomy → acute-phase pulmonary infection." Association pattern analysis of complication events and prognostic events within the critical time window of the subacute phase revealed that the proportion of patients with moderate to severe pulmonary infection and poor prognosis was significantly higher in the event sequence parameters than in patients without pulmonary infection. This pattern is a "subacute pulmonary infection → poor prognosis" event-related time sequence pattern. The event-related time sequence pattern is stored in the form of an association rule list, which includes five fields: antecedent event, subsequent event, time window, support, and confidence. The support of the "comatose admission → hyperacute phase surgery" event-related time sequence pattern represents the proportion of patients who simultaneously meet the criteria of comatose admission and hyperacute phase surgery, while the confidence represents the proportion of comatose admission patients who undergo surgery during the hyperacute phase.
[0028] For example, generating a time-series correlation strength distribution based on the event correlation time-series pattern includes: calculating the correlation delay duration of the event correlation time-series pattern to obtain a delay time window distribution; identifying cascading effect correlations and independent delay correlations based on the delay time window distribution; sorting the cascading effect correlations and the independent delay correlations according to the delay effect intensity to generate a delay effect sequence; and generating a time-series correlation strength distribution based on the delay effect sequence.
[0029] The distribution of delay time windows was obtained by calculating the associated delay duration of event-related time sequence patterns. The associated delay duration is defined as the time interval between the occurrence of a preceding event and the occurrence of a subsequent event in the event-related time sequence pattern. The delay duration reflects the causal time sequence relationship and the time progression of physiological and pathological mechanisms between events. All delay durations are uniformly converted to hours. For the event-related time sequence pattern of "admission to hospital in a coma → surgery in the hyperacute phase," the delay duration is the interval between the admission time and the surgery time. In patients meeting this event-related time sequence pattern, the delay duration ranges from 0.5 to 6 hours, with a mean delay duration of approximately 2.8 hours and a median delay duration of approximately 2.5 hours. For the event-related time sequence pattern of "craniotomy → acute pulmonary infection," the delay duration is the interval between the surgery time and the diagnosis time of pulmonary infection. In patients undergoing craniotomy, the delay duration for pulmonary infection ranges from 1 to 7 days (equivalent to 24 to 168 hours), with a mean delay duration of approximately 3.5 days and a median delay duration of approximately 3 days. The delay duration of the "subacute pulmonary infection → poor prognosis" event-series correlation pattern is the interval between the diagnosis of pulmonary infection and the time of discharge prognostic assessment, ranging from 5 to 20 days, with a mean delay of approximately 12 days. The delay time window distribution describes the statistical distribution characteristics of the delay duration for each event-series correlation pattern. The delay time window distribution types include concentrated and dispersed distributions. The standard deviation of the concentrated delay time window distribution is less than 30% of the mean, while the standard deviation of the dispersed delay time window distribution is greater than 50% of the mean. The "admission in a coma → hyperacute surgery" event-series correlation pattern has a smaller standard deviation in its delay time window distribution, indicating a concentrated distribution, while the "craniotomy → acute pulmonary infection" event-series correlation pattern has a larger standard deviation in its delay time window distribution, indicating a dispersed distribution.
[0030] The association between complication cascade effects and independent delayed associations were identified based on the distribution of delayed time windows. Complication cascade effect associations are defined as a pattern where one complication event triggers subsequent complication events, manifested as multiple complications occurring consecutively within a short period. For example, the delayed time window distribution of the pattern of pulmonary infection triggering gastrointestinal bleeding is concentrated in the 3-5 day range, indicating a complication cascade effect association. Similarly, the delayed time window distribution of the pattern of pulmonary infection triggering seizures is concentrated in the 1-3 day range, also indicating a complication cascade effect association. Independent delayed associations are defined as a pattern where there is a correlation between the preceding and subsequent events but not through an intermediate complication cascade. Independent delayed associations have longer delay durations and more dispersed delayed time window distributions. The association pattern between surgical type and long-term prognosis has a delay duration of approximately 180 days, with a large standard deviation in the delayed time window distribution, indicating an independent delayed association. The association pattern between admission level of consciousness and long-term prognosis has a delay duration exceeding 90 days, indicating an independent delayed association. The criteria for identifying cascading effects of complications are a delay duration of less than 14 days and a concentrated distribution of delay time windows, while the criteria for identifying independent delay associations are a delay duration of more than 14 days or a dispersed distribution of delay time windows.
[0031] Delayed effect sequences were generated by ranking the cascade effect associations of complications and independent delayed associations according to the strength of the delayed effect. The strength of the delayed effect comprehensively considered three indicators: confidence, support, and delay duration of the association rule. The formula for calculating the effect strength was E = α × confidence + β × support - γ × lg(delay duration + 1), where α, β, and γ are weighting coefficients set to 0.5, 0.3, and 0.2 respectively, lg is the commonly used logarithm (base 10), and the sum of the weighting coefficients is 1.0. The delayed effect strength of the "pulmonary infection → gastrointestinal bleeding" cascade effect association was calculated by substituting the confidence, support, and delay duration into the formula. Similarly, the delayed effect strength of the "surgical type → long-term prognosis" independent delayed association was calculated by substituting the confidence, support, and delay duration into the formula. The delayed effect sequence is sorted from highest to lowest effect strength. The cascade effect association and the independent delayed effect association are also sorted in the delayed effect sequence according to their effect strength. The sorting results show that the cascade effect association receives less negative penalty in the effect strength formula due to its shorter delay duration, while the independent delayed effect association receives a larger negative penalty in the effect strength formula due to its longer delay duration. The association rules with the highest effect strength in the delayed effect sequence are marked as core association rules.
[0032] A time-series association strength distribution was generated based on the delayed-effect sequence. This distribution was used to construct a two-dimensional map with delayed-effect strength as the vertical axis and delay duration as the horizontal axis. Each association rule in the delayed-effect sequence was located in the time-series association strength distribution map according to delay duration and effect strength. Hyperacute-phase association rules in the delayed-effect sequence were located within the delay duration range of 0 to 6 hours. The time-series association strength distribution in this range showed a high average effect strength, indicating this range is a high-intensity association zone in the hyperacute phase. Acute-phase association rules in the delayed-effect sequence were located within the delay duration range of 6 to 72 hours. The time-series association strength distribution in this range showed a moderate average effect strength, indicating this range is a medium-intensity association zone in the acute phase. Subacute-phase association rules in the delayed-effect sequence were located within the delay duration range of 3 to 14 days. The time-series association strength distribution in this range showed a low to medium average effect strength, indicating this range is a low to medium intensity association zone in the subacute phase. The time-series association strength distribution within the delay duration range of 14 to 180 days showed that this range contained several association rules with a low average effect strength, indicating this range is a low-intensity association zone in the recovery phase. The temporal association strength distribution shows a decreasing trend in effect strength with increasing delay time. This trend reflects that earlier events have a more direct and stronger impact on subsequent events, while the association of later events is weakened due to the influence of multiple factors. The temporal association strength distribution identifies outliers in effect strength, with some association rules having effect strengths significantly higher than the average value within the same interval. These super-strong association rules rank high in delayed effect sequences.
[0033] A knowledge base of association rules was formed by combining temporal association strength distribution with clinical feature profiles to perform association pattern screening and evaluation. The temporal association strength distribution contains multiple association rules along with their effect strength and delay duration information. The clinical feature profile includes the comprehensive feature score and dimensional component scores of patient samples. The association pattern screening and evaluation integrates the temporal association strength distribution and clinical feature profile information. Association pattern screening sets effect strength and support thresholds; association rules with effect strength or support below the thresholds are filtered out. After screening, several high-value association rules are retained, covering the four key phases of the disease course: hyperacute, acute, subacute, and recovery. High-value rules are distributed across all time intervals in the temporal association strength distribution. The association pattern evaluation, combined with the clinical feature profile analysis, examines the applicability of association rules in different patient subgroups. Differences exist in the confidence levels of association rules between the severely ill patient subgroup with higher comprehensive feature scores and the mildly ill patient subgroup with lower comprehensive feature scores in the clinical feature profile. Association pattern assessment analyzes the relationship between association rules and prognostic outcomes in the time-series association strength distribution. Association rules with high effect strength are associated with poor prognosis, and the effect strength ranking of the time-series association strength distribution provides a quantitative basis for identifying key prognostic rules. An association rule knowledge base integrates and evaluates high-value association rules and their effect strength and confidence levels in different patient subgroups. The knowledge base includes prognostic outcome characteristics and clinical application scenarios for each rule, and labels the applicable patient types and precautions for each rule.
[0034] Step S130: Based on the association rule knowledge base, prognosis-oriented patient subgroup clustering is performed to determine the patient type distribution characteristics. Based on the patient type distribution characteristics, treatment plan feature components are extracted by calculating the similarity of treatment plans. Treatment pattern atlas is formed by combining the treatment plan feature components with the clinical data feature set.
[0035] In some embodiments, the step of determining patient type distribution characteristics by prognosis-oriented patient subgroup clustering based on the association rule knowledge base includes: determining prognostic stratification criteria based on prognostic outcome characteristics of the association rule knowledge base; setting stratification constraints on bleeding sites to generate a set of constraint parameters for the prognostic stratification criteria; performing hierarchical clustering analysis on the association rule knowledge base based on the set of constraint parameters to construct patient subgroup distributions; and determining patient type distribution characteristics based on the patient subgroup distributions.
[0036] Prognostic stratification criteria were determined based on prognostic outcome features derived from an association rule knowledge base. The association rule knowledge base contains multiple association rules and their relationships with prognostic outcomes. The core indicators for prognostic outcome features were modified Rankin Scale scores extracted from the knowledge base: a score of 0-2 defined as a good prognosis, 3-5 as a poor prognosis, and 6 as death. The prognostic stratification criteria were divided into five levels based on the severity of the prognostic outcome features in the association rule knowledge base: extremely poor prognosis (5-6 points), poor prognosis (4 points), moderate prognosis (3 points), good prognosis (1-2 points), and excellent prognosis (0 points). The prognostic stratification criteria also considered functional independence indicators: a Barthel Index score below 40 defined complete dependence, 40-60 as severe dependence, 60-80 as moderate dependence, 80-95 as mild dependence, and 95-100 as largely independent. The prognostic stratification criteria combine two dimensions: the modified Rankin Scale score and the Barthel Index. (For example, if a patient has a modified Rankin score of 3 but a Barthel Index of only 45, the patient is classified into the poor prognostic stratum rather than the intermediate prognostic stratum due to severe functional dependence.) The very poor prognostic stratum corresponds to a modified Rankin score of 5-6 and a Barthel Index of less than 40, while the excellent prognostic stratum corresponds to a modified Rankin score of 0 and a Barthel Index of 95-100.
[0037] A set of constraint parameters was generated by setting stratification constraints based on the hemorrhage site in the prognostic stratification criteria. The hemorrhage site stratification constraint considers the different impacts of different anatomical regions on prognosis. Significant differences exist in the composition of hemorrhage sites among patients at different prognostic levels within the stratification criteria, with brainstem hemorrhage having the worst prognosis and lobar hemorrhage having a relatively better prognosis. Based on the prognostic stratification criteria, patients with brainstem hemorrhage are preferentially assigned to the extremely poor or poor prognostic level, while patients with lobar hemorrhage are preferentially assigned to the good or intermediate prognostic level. The constraint conditions consider the interaction between hematoma volume and hemorrhage site. Within the same prognostic level, patients with hemorrhage at different sites have different hematoma volume thresholds; cerebellar hemorrhage exceeding 30 ml significantly worsens the prognosis. The constraint parameter set includes a weight coefficient for the hemorrhage site and a site-volume interaction coefficient. The weight coefficients in the constraint parameter set are based on 1.0; increasing the weight coefficient for brainstem hemorrhage indicates a greater prognostic impact, while decreasing the weight coefficient for lobar hemorrhage indicates a lesser prognostic impact. The site-volume interaction coefficient in the constraint parameter set is defined as a correction coefficient for the prognostic impact of hematoma volume. A higher interaction coefficient for cerebellar hemorrhage indicates a stronger negative prognostic impact from increased volume. The constraint parameter set includes correction parameters for ventricular rupture and midline shift. The weighting coefficient is increased when the hematoma ruptures into the ventricle, and the weighting coefficient is increased when the midline shift exceeds 5 mm.
[0038] Hierarchical clustering analysis was performed on the association rule knowledge base based on the constraint parameter set to construct patient subgroup distributions. Hierarchical clustering analysis extracted clinical features from patient samples corresponding to the association rule knowledge base. Weight adjustments based on the constraint parameter set were introduced into the patient similarity calculation. The adjusted similarity calculation formula is d' = d × w_location × (1 + w_volume × interaction coefficient), where d is the original Euclidean distance, w_location is the weight coefficient of the hemorrhage location in the constraint parameter set (baseline value 1.0), w_volume is the normalized value of the hematoma volume, and the interaction coefficient is the location-volume interaction coefficient in the constraint parameter set. The similarity distance of brainstem hemorrhage patients multiplied by a larger weight coefficient in the constraint parameter set makes them more likely to cluster into the subgroup with the worst prognosis. The similarity distance of lobar hemorrhage patients multiplied by a smaller weight coefficient makes them more likely to cluster into the subgroup with the best prognosis. Hierarchical clustering analysis was used to assist clustering based on the difference in effect strength of association rules in the association rule knowledge base across different patient samples. Starting with all patient samples, the two samples or subgroups with the smallest similarity distance were merged each time, and this merging process was repeated until five target subgroups were formed. The patient subgroup distribution records the sample size, average eigenvalue, and prognostic outcome distribution for each subgroup. The vast majority of patients in the extremely severe subgroup had poor prognoses or died, while the proportion of patients with good prognoses was higher in the mild subgroup. The core features of each subgroup were labeled in the patient subgroup distribution. The core features of the extremely severe subgroup were deep coma, massive hemorrhage, brainstem or thalamic location, and ventricular rupture; the core features of the mild subgroup were conscious, small amount of hemorrhage, lobe location, and no ventricular rupture. The patient subgroup distribution was visualized using dendrograms and scatter plots.
[0039] The patient type distribution characteristics were determined based on the patient subgroup distribution. These characteristics integrated sample statistics and clinical feature descriptions of the patient subgroup distributions, with features including subgroup size, mean clinical features, prognostic outcome distribution, and applicability of association rules. The extremely severe subgroup was characterized by a small proportion of patients, extremely low mean Glasgow Coma Scale scores, extremely large mean hematoma volume, and an extremely high proportion of poor prognosis; this subgroup had the worst prognosis among all patient subgroups. The severe subgroup was characterized by a large proportion of patients, low mean Glasgow Coma Scale scores, large mean hematoma volume, and a high proportion of poor prognosis; this subgroup primarily consisted of patients with moderate to severe injuries. The moderate subgroup was characterized by the largest proportion of patients, moderate mean Glasgow Coma Scale scores, moderate mean hematoma volume, and a poor prognosis rate approaching half; this subgroup exhibited the strongest heterogeneity among all patient subgroups. The mild subgroup was characterized by a moderate proportion of patients, high mean Glasgow Coma Scale scores, small mean hematoma volume, and a high proportion of good prognosis. The patient type distribution characteristics of the minimally symptomatic subgroup are: smaller patient size, conscious patient, minor bleeding, and a very high proportion of patients with good prognosis. The patient type distribution characteristics describe the differences in treatment modalities among the subgroups: the extremely severe subgroup has a high surgical rate and a high proportion of craniotomy, while the minimally symptomatic subgroup has a low surgical rate and a high proportion of conservative treatment. The patient type distribution characteristics are output in tables and bar charts.
[0040] In some embodiments, the step of calculating and extracting treatment plan feature components based on the patient type distribution characteristics through treatment plan similarity includes: extracting treatment plan records corresponding to each subgroup from the patient type distribution characteristics to generate a plan candidate set; extracting treatment timing features and intervention method features from the plan candidate set to generate timing intervention association parameters; calculating plan similarity based on the timing intervention association parameters to obtain a weighted similarity matrix; and constructing treatment plan feature components based on the weighted similarity matrix and the patient type distribution characteristics.
[0041] Treatment plan candidate sets were generated by extracting treatment plan records corresponding to each subgroup from the patient type distribution characteristics. The patient type distribution characteristics included treatment pattern statistics for five subgroups, and treatment plan records were extracted from the treatment feature subsets of each subgroup within the patient type distribution characteristics. In the treatment plan records of the extremely severe subgroup, craniotomy accounted for the vast majority, and the patient type distribution characteristics indicated that the average timing of surgery for this subgroup was early after onset, with a high average hematoma clearance rate. In the treatment plan records of the severe subgroup, craniotomy also accounted for the majority, and the patient type distribution characteristics indicated that the average timing of surgery for this subgroup was early to mid-stage after onset. In the treatment plan records of the moderate subgroup, minimally invasive drainage accounted for the highest proportion, while the treatment plan records of the mild and minimally invasive subgroups mainly focused on conservative treatment. The candidate treatment set integrated treatment protocols for five subgroups and identified the dominant treatment modality for each subgroup. The dominant treatment modality for the extremely severe subgroup was ultra-early craniotomy for hematoma evacuation; for the severe subgroup, it was early craniotomy for hematoma evacuation; for the moderate subgroup, it was minimally invasive percutaneous drainage or conservative treatment; and for the mild and minimally symptomatic subgroups, it was conservative treatment. The candidate treatment set recorded the prognostic outcome distribution for each treatment protocol. In the candidate treatment set, the proportion of patients undergoing craniotomy in the extremely severe subgroup was higher than that of patients receiving conservative treatment, but the absolute prognosis was still poor. In the candidate treatment set, the proportion of patients receiving conservative treatment in the minimally symptomatic subgroup was extremely high.
[0042] Timing-intervention correlation parameters were generated by extracting treatment timing and intervention mode characteristics from the candidate protocol set. Treatment timing characteristics were extracted from the surgical timing variable in the candidate protocol set, categorized into four time periods based on time after onset: ultra-early (0-6 hours), early (6-24 hours), delayed (24-72 hours), and late (>72 hours). Intervention mode characteristics were extracted from the surgical type variable in the candidate protocol set, including three methods: craniotomy for hematoma evacuation, minimally invasive percutaneous drainage, and conservative treatment. Craniotomy was further subdivided into conventional craniotomy and decompressive craniectomy. Significant differences were observed in the distribution of treatment timing among the subgroups in the candidate protocol set. The extremely severe subgroup was predominantly ultra-early and early, while the mildly severe subgroup was predominantly late or non-surgical. The timing-intervention correlation parameters quantified the impact of the combination of treatment timing and intervention mode on prognosis. These parameters were calculated by statistically analyzing the good prognosis rate of each timing-intervention combination in the candidate protocol set. Timing-related parameters showed that early craniotomy had the highest prognostic rate, while delayed craniotomy had a lower prognostic rate (surgery within 6-24 hours of onset allows for hematoma removal after stabilization and before secondary damage fully develops, reducing mass effect). Among the timing-related parameters, minimally invasive drainage was less sensitive to surgical timing than craniotomy, and the prognostic rate of ultra-early minimally invasive drainage was not significantly different from that of early minimally invasive drainage. Considering the correction effect of hematoma volume, ultra-early surgery had a more significant prognostic advantage in patients with massive hemorrhage, while the prognosis of conservative treatment was not significantly different from that of surgical treatment in patients with minor hemorrhage.
[0043] A weighted similarity matrix is obtained by calculating the similarity of intervention-timing parameters. The similarity is calculated based on the degree of similarity of these parameters, using cosine similarity with values ranging from 0 to 1. A high similarity in the intervention-timing parameter vectors between the critically ill and severely ill subgroups indicates a high degree of similarity in their treatment timing and intervention patterns, with both subgroups tending towards early craniotomy. A high similarity in the intervention-timing parameter vectors between the moderate and mildly ill subgroups also indicates a high degree of similarity in their treatment patterns, with both subgroups rarely undergoing craniotomy. A low cosine similarity in the intervention-timing parameter vectors between the critically ill and mildly ill subgroups indicates a significant difference in their treatment patterns. The weighted similarity matrix incorporates prognostic outcome weights on top of the cosine similarity. The adjustment principle is to increase the similarity between subgroups with similar prognoses and decrease the similarity between subgroups with large differences in prognoses. The weighted adjustment formula is w_adjustment = w_cosine × exp(-|P1-P2|), where w_cosine is the cosine similarity, P1 and P2 are the proportions of poor prognosis in the two subgroups, and exp is the natural exponential function. The weighted similarity matrix is stored in symmetric form. A diagonal element of 1 indicates that the subgroups are completely similar, while the off-diagonal elements reflect the overall similarity in treatment patterns and prognostic outcomes between the subgroups.
[0044] Treatment plan feature components were constructed based on a weighted similarity matrix and patient type distribution characteristics. These components were extracted from the weighted similarity matrix using principal component analysis (PCA), which reduced the dimensionality of treatment plan similarity across the five subgroups to three principal feature dimensions. The first principal component (PCI) explained the highest proportion of variance; its score was positive for the extremely severe and severe subgroups, and negative for the mild and minimally invasive subgroups, representing the degree of surgical aggression. The second PCI score was positive for the early surgery subgroup and negative for the delayed surgery subgroup, representing the timing of treatment. The third PCI score was positive for the high hematoma clearance rate subgroup, representing the thoroughness of treatment. The treatment plan feature components were then linked to the clinical characteristics of the patient type distribution characteristics, with the PCI scores correlated with the mean of these clinical characteristics. The mean of these clinical characteristics served as an auxiliary dimension for the treatment plan feature components. The treatment regimen for the critically ill subgroup is characterized by highly aggressive, ultra-early intervention, and high thoroughness, a pattern that matches the clinical characteristics of deep coma and massive hemorrhage in the critically ill subgroup. The treatment regimen for the mildly ill subgroup is characterized by conservative, delayed or non-surgical, and low-interventional approaches, a pattern that matches the clinical characteristics of conscious individuals and minor hemorrhage in the mildly ill subgroup.
[0045] A treatment pattern atlas is constructed by combining treatment protocol feature components with clinical data feature sets. The treatment protocol feature components include the projected coordinates of five subgroups across three principal component dimensions. The clinical data feature set includes baseline features, treatment features, and prognostic features of the patient samples. The treatment pattern atlas integrates the treatment protocol feature components and the clinical data feature set to construct a complete mapping relationship from patient characteristics to treatment protocols to prognostic outcomes. The treatment pattern atlas is represented by a three-layer network structure: the input layer consists of key clinical features from the clinical data feature set (Glasgow Coma Scale score, hematoma volume, hematoma location, etc.); the middle layer consists of the three feature dimensions of the treatment protocol feature components (surgical aggressiveness, treatment timing, and treatment thoroughness); and the output layer represents two outcome categories: good prognosis and poor prognosis. The treatment pattern atlas uses edge weights to represent the correlation strength between the treatment protocol feature components and the features in the clinical data feature set. A higher correlation weight between the Glasgow Coma Scale score and the degree of surgical aggressiveness indicates more severe impairment of consciousness and more aggressive surgery; a higher correlation weight between hematoma volume and the degree of treatment thoroughness indicates greater bleeding and a greater need for complete removal. The treatment pattern diagram marks the typical treatment pathways for each subgroup. The typical pathway for the extremely severe subgroup is "deep coma + massive hemorrhage → ultra-early craniotomy → high clearance rate → poor prognosis". The typical pathway for the mild subgroup is "consciousness + small amount of hemorrhage → conservative treatment → no surgical intervention → good prognosis".
[0046] Step S140: Based on the treatment pattern atlas, perform multi-dimensional prognostic correlation analysis to establish a set of prognostic influencing factors. Based on the set of prognostic influencing factors, determine the individualized factor weight threshold through factor interaction effect detection. Based on the individualized factor weight threshold, perform factor importance ranking to form a combination of key prognostic factors.
[0047] Specifically, a multi-dimensional prognostic correlation analysis was conducted based on the treatment pattern atlas to establish a set of prognostic influencing factors. The treatment pattern atlas contains a complete mapping relationship from patient clinical characteristics to treatment plans to prognostic outcomes. The multi-dimensional prognostic correlation analysis identifies key factors affecting prognosis based on the treatment pattern atlas. The multi-dimensional correlation analysis assesses the correlation strength between each factor and prognostic outcome from three perspectives: baseline characteristics, treatment characteristics, and disease events. The correlation strength is quantified using two indicators: correlation coefficient and regression coefficient. The prognostic correlation analysis of the baseline characteristics dimension, based on the input layer features of the treatment pattern atlas, found that the Glasgow Coma Scale score was positively correlated with the rate of good prognosis, hematoma volume was positively correlated with the rate of poor prognosis, and the rate of poor prognosis was significantly higher when the hematoma was located in the brainstem than in other locations. The prognostic correlation analysis of the treatment characteristics dimension, based on the intermediate layer features of the treatment pattern atlas, found that the rate of good prognosis was higher when surgery was performed early after onset than when it was delayed, and the rate of good prognosis was higher when the hematoma clearance rate was higher. Prognostic correlation analysis of the disease course event dimension, based on the edge weight relationships of the treatment pattern atlas, revealed that patients with pulmonary infection had a higher rate of poor prognosis, patients with rebleeding had an extremely high rate of poor prognosis, and patients with early rehabilitation intervention had a better rate of good prognosis. The prognostic influencing factor set integrated the correlation analysis results of the three dimensions, screening for factors with large absolute values of correlation coefficients or significant prognostic differences. The prognostic influencing factor set included key factors such as level of consciousness upon admission, hematoma volume, hematoma location, ventricular rupture, timing of surgery, hematoma clearance rate, pulmonary infection, rebleeding, early rehabilitation intervention, and epileptic seizures. The direction and strength of the correlation between each factor and the prognostic outcome were indicated.
[0048] In some embodiments, determining the individualized factor weight threshold based on the prognostic influencing factor set through factor interaction effect detection includes: performing factor pairing and combination to generate interactive factor pairs based on the prognostic influencing factor set; calculating the interaction effect strength in the interactive factor pairs to generate interaction coefficients; performing patient individual characteristic matching on the interaction coefficients to obtain individualized correction parameters; and determining the individualized factor weight threshold based on the individualized correction parameters.
[0049] Interaction factor pairs were generated based on the prognostic influencing factor set. The prognostic influencing factor set contains several key factors. Factor pairings were performed by selecting two factors from the set, using a full combinatorial approach to ensure no potential interactions were overlooked. The selection of interaction factor pairs was based on two criteria: clinical rationality and statistical significance. Clinical rationality required a potential pathophysiological link between the two factors in the prognostic influencing factor set, while statistical significance required a p-value of less than 0.05 for the interaction term. Consciousness level and hematoma volume formed an interaction factor pair. Both factors in the prognostic influencing factor set pathophysiologically reflect the severity of brain injury; more severe consciousness impairment indicates more severe brain injury, and larger hematoma volume leads to increased intracranial pressure, further aggravating consciousness impairment. Clinical observation showed that both synergistically affect prognosis. Surgical timing and hematoma clearance rate formed an interaction factor pair. Both are therapeutic factors in the prognostic influencing factor set. Early surgery combined with a high clearance rate may produce a synergistic effect, while delayed surgery, even with a high clearance rate, has limited effectiveness due to secondary damage already present. Ventricular rupture and pulmonary infection form an interaction factor pair. Ventricular rupture increases the risk of hydrocephalus, and hydrocephalus requires external ventricular drainage, which increases the risk of infection. After screening the interaction factor pairs, several clinically significant and statistically significant factor pairs were retained. The interaction factor pairs that were screened out included those with no pathophysiological association and those with no significant interaction effect.
[0050] The interaction effect strength is calculated in the interaction factor pairs to generate the interaction coefficient. The interaction effect strength is evaluated by the interaction term coefficient of the logistic regression model. The prediction formula of the regression model is ln(p / (1-p))=β0+β1X1+β2X2+β 12 X1X2, where p is the probability of a poor prognosis, β0 is the intercept term, X1 and X2 are the values of the two factors in the interaction factor pair, β1 and β2 are the main effect coefficients of factor 1 and factor 2, respectively, and β 12 β is the interaction coefficient, also known as the interaction factor. 12 The sign and absolute value of the interaction coefficients reflect the direction and intensity of the interaction, respectively. The regression model of the interaction between consciousness state and hematoma volume shows that the main effect coefficients β1 and β2 for consciousness state and hematoma volume are relatively high, respectively, and the interaction coefficient β... 12 A positive value and significance indicate a positive synergistic effect. The interaction coefficient β between surgical timing and hematoma clearance rate is... 12 A positive value indicates a synergistic effect between early surgery and high clearance rates. The interaction coefficient β of the ventricular rupture-pulmonary infection interaction factor pair. 12A higher interaction coefficient indicates a strong synergistic effect, where both factors contribute to a worsened prognosis. For example, if a patient undergoes external ventricular drainage after ventricular rupture and subsequently develops intracranial infection leading to ventriculitis, the severity of this patient's prognosis exceeds the combined prognostic impact of either ventricular rupture alone or infection alone, demonstrating the amplifying effect of the interaction. The interaction coefficient distinguishes the direction of synergy: a positive interaction coefficient indicates amplified prognostic changes when both factors worsen or improve simultaneously, while a negative interaction coefficient indicates that the two factors cancel each other out when their effects are opposite. A larger absolute value of the interaction coefficient defines a strong interaction, a moderate absolute value a moderate interaction, and a smaller absolute value a weak interaction. Several pairs of interaction factors (strong, moderate, and weak) are identified, with strong interaction pairs given priority in individualized weighting.
[0051] For example, the step of matching the interaction coefficients with individual patient characteristics to obtain individualized correction parameters includes: constructing a disease-stratified interaction sample by classifying and arranging the interaction coefficients according to the characteristics of the patient's comorbidities; identifying the influence pattern of comorbidities on the sensitivity of factors in the disease-stratified interaction sample to determine sensitivity correction conditions; dividing the disease-stratified interaction sample into a high-risk disease group and a low-risk disease group according to the sensitivity correction conditions; and performing sensitivity correction and weight allocation on the high-risk disease group and the low-risk disease group respectively to obtain individualized correction parameters.
[0052] Interaction coefficients were categorized and arranged according to patient comorbidity characteristics to construct stratified disease interaction samples. Patient comorbidity characteristics were extracted from baseline features of the clinical data feature set. Comorbidities included six common diseases: hypertension, diabetes, coronary heart disease, chronic obstructive pulmonary disease (COPD), liver dysfunction, and renal dysfunction. The proportion of patients with hypertension was high, followed by diabetes, COPD, and COPD, decreasing in that order. The stratified interaction samples were divided into three levels based on the type and number of comorbidities: no comorbidities, single comorbidity, and multiple comorbidities. The stratification principle was based on the cumulative effect of comorbidities on prognosis. The interaction coefficients were statistically analyzed for distribution characteristics in each of the three levels of the disease stratified interaction samples. The mean interaction coefficient for consciousness status-hematoma volume was low in the no-comorbidity group, moderate in the single-comorbidity group, and high in the multiple-comorbidity group. The increase in interaction coefficients with increasing number of comorbidities reflects the synergistic effect among enhancing factors related to comorbidities. The disease stratified interactive sample records the distribution of prognostic outcomes for patients in each stratum. The group without comorbidities had a lower rate of poor prognosis, the group with a single comorbidity had a moderate rate of poor prognosis, and the group with multiple comorbidities had a higher rate of poor prognosis.
[0053] In disease-stratified interaction samples, the influence patterns of comorbidities on factor sensitivity were identified to determine sensitivity correction conditions. These patterns were identified by comparing the differences in interaction coefficients across different disease strata within the disease-stratified interaction samples. Analysis of variance was used to compare the mean interaction coefficients of each stratum, and interactions with significant differences were included in sensitivity correction. Hypertensive patients showed significantly higher sensitivity to the hematoma volume-ventricular rupture interaction than non-hypertensive patients in the disease-stratified interaction samples, with a larger difference in interaction coefficients. This difference stems from increased vascular fragility and a higher risk of hematoma expansion due to hypertension, leading to a more rapid increase in intracranial pressure after ventricular rupture. Diabetic patients showed significantly higher sensitivity to the pulmonary infection-rebleed interaction in the disease-stratified interaction samples than non-diabetic patients. This difference arises from the weakened immune function and abnormal coagulation function in diabetic patients, making infection control difficult and increasing the risk of rebleeding. Patients with chronic obstructive pulmonary disease (COPD) showed high sensitivity to the surgical timing-pulmonary infection interaction in the disease-stratified interaction samples. COPD patients have insufficient lung function reserve, resulting in a significantly higher incidence of postoperative pulmonary infection. Sensitivity correction criteria are defined as the criteria for determining comorbidities that trigger the adjustment of interaction coefficients. These criteria include two dimensions: comorbidity type and severity. Specifically, sensitivity correction criteria are applied to hypertension comorbidities triggering hematoma volume-related interactions, diabetes comorbidities triggering infection-related interactions, and chronic obstructive pulmonary disease triggering pulmonary infection-related interactions.
[0054] Based on sensitivity correction criteria, the disease stratification interaction samples were divided into high-risk and low-risk disease groups. The high-risk disease group was defined as patients whose comorbidities had a significant adverse impact on prognosis and whose interaction effects were highly sensitive; the high-risk disease group was defined based on the triggering of the sensitivity correction criteria. The low-risk disease group was defined as patients without comorbidities or whose comorbidities had a minimal impact on prognosis; the low-risk disease group was defined based on the absence of the sensitivity correction criteria. Screening from the disease stratification interaction samples based on sensitivity correction criteria, the high-risk disease group included patients with diabetes, chronic obstructive pulmonary disease, liver dysfunction, or renal dysfunction—comorbidities that directly affect infection control, coagulation function, or multi-organ compensatory capacity—as well as patients with two or more comorbidities. Screening from the disease stratification interaction samples based on sensitivity correction criteria, the low-risk disease group included patients without comorbidities and patients with only hypertension or coronary artery disease as a single comorbidity. The mean interaction coefficient of the high-risk disease group was higher than that of the low-risk disease group, indicating that the high-risk disease group was more sensitive to interaction effects. The rate of poor prognosis was higher in the high-risk disease group than in the low-risk disease group. The difference in prognosis between the two groups was partly attributed to the interaction effect of enhanced comorbidities and partly to the independent adverse effects of the comorbidities themselves. The criteria for classifying the high-risk and low-risk disease groups considered the severity of the comorbidities. Under the sensitivity-adjusted conditions, mild hypertension was classified into the low-risk disease group, and severe hypertension was classified into the high-risk disease group.
[0055] Sensitivity correction and weight allocation were performed separately for high-risk and low-risk disease groups to obtain individualized adjusted parameters. Sensitivity correction adjusted the interaction coefficients for patients in the high-risk disease group, with the correction principle being to increase the coefficients of comorbidity-related interactions to reflect the high sensitivity of this group. In the high-risk disease group, the pulmonary infection-rebleeding interaction coefficient for diabetic patients was adjusted upwards from baseline, based on the pathophysiological mechanisms of immunodeficiency and coagulation abnormalities in diabetic patients. The surgical timing-pulmonary infection interaction coefficient for patients with chronic obstructive pulmonary disease in the high-risk group was also adjusted upwards from baseline. For the low-risk disease group, standard interaction coefficients were used without correction. The interaction coefficients for patients without comorbidities in the low-risk group remained at baseline, and the weight allocation for the low-risk disease group directly used baseline weights. Weight allocation calculated the comprehensive weight of each factor based on the corrected interaction coefficients. The weight calculation formula was W = w_main effect + Σ(w_interaction × interaction coefficient), where W is the comprehensive factor weight, w_main effect is the weight of the independent effect of the factor, and w_interaction is the weight of the interaction term. The individualized correction parameters integrate the sensitivity correction coefficients and weighting results of the high-risk disease group and the low-risk disease group. The average value of the individualized correction parameter vector elements of patients in the high-risk disease group is higher than that of patients in the low-risk disease group. The difference in individualized correction parameters guides individualized prognostic prediction and treatment decisions.
[0056] Individualized factor weight thresholds were determined based on individualized correction parameters. The individualized factor weight threshold was defined as the lowest significant weight of each factor in predicting the prognosis of an individual patient. The threshold was set based on the individualized correction parameters and the strength of the factor's main effect. The main effect strength of consciousness status was high; the individualized correction parameter was increased in the high-risk disease group, and the individualized factor weight threshold for this factor increased accordingly. The main effect strength of hematoma volume was the second highest; the individualized correction parameter was increased more significantly in the high-risk disease group, and the individualized factor weight threshold for this factor also increased accordingly. The main effect strength of surgical timing was high, but the individualized correction parameter varied greatly depending on the type of comorbidity. The increase in the individualized factor weight threshold for surgical timing was greater in patients with chronic obstructive pulmonary disease than in patients with diabetes. An elderly patient with multiple comorbidities had a significantly higher individualized factor weight threshold than a younger patient without comorbidities (e.g., a 75-year-old patient with diabetes, hypertension, and chronic kidney disease had a higher individualized factor weight threshold for hematoma volume, while a 55-year-old patient without comorbidities had a lower individualized factor weight threshold). This difference in thresholds reflects the high sensitivity of high-risk patients to prognostic factors. Individualized factor weight thresholds were generally lower in low-risk disease groups than in high-risk disease groups, while the weights of each factor in individualized adjusted parameters remained at baseline levels for patients without comorbidities. Individualized factor weight thresholds were used to screen key factors that significantly impacted the prognosis of a specific patient; factors with weights exceeding the individualized factor weight threshold were included in the patient's set of key prognostic factors.
[0057] In some embodiments, the step of ranking factors based on individualized factor weight thresholds to form a combination of key prognostic factors includes: dividing the individualized factor weight thresholds into acute option resets and recovery option resets according to disease stages; performing factor importance assessments on the acute option resets and the recovery option resets respectively to obtain a phased factor ranking; extracting high-weight factor sets from the phased factor rankings to determine core influencing factors; and ranking the core influencing factors according to disease priority to form a combination of key prognostic factors.
[0058] Individualized factor weight thresholds were used to generate acute option reset and recovery option reset based on disease course stages. Disease course stages were divided into two phases based on the temporal correlation strength distribution of critical periods: the acute phase and the recovery phase. The acute phase combined with the hyperacute and acute phases covered 0 to 72 hours after onset, while the recovery phase combined with the subacute and recovery phases covered 3 days to 6 months after onset. The merging was based on the similarity of prognostic influencing factors within the same time window and the difference in the importance of prognostic influencing factors across different stages. The acute option reset selected factors with significant prognostic impact during the acute phase from the individualized factor weight thresholds. Key factors in the acute option reset included admission level of consciousness, hematoma volume, hematoma location, ventricular rupture, surgical timing, and hematoma clearance rate. These factors had higher weights in the acute phase according to the individualized factor weight thresholds. In acute option reassessment, the individualized factor weight threshold for level of consciousness was the highest, followed by hematoma volume, and then surgical timing. These three factors constituted the most important prognostic factors in the acute phase. Recovery option reassessment screened the factors with significant prognostic impact during the recovery period from the individualized factor weight thresholds. The key factors in recovery option reassessment were pulmonary infection, rebleeding, epileptic seizures, and early rehabilitation intervention. In recovery option reassessment, pulmonary infection had the highest individualized factor weight threshold, followed by early rehabilitation intervention, and rebleeding was relatively high. Complication prevention and rehabilitation treatment became the main determinants of prognosis during the recovery period.
[0059] Factor importance assessments were performed on both the acute and recovery option reassessments to obtain a phased ranking of factors. The acute-phase factor importance assessment used acute-phase prognostic outcomes (modified Rankin Scale score at discharge) as the target variable. The assessment method employed multivariate logistic regression to calculate standardized regression coefficients for each factor, ensuring comparability of coefficients across different dimensions. In the acute option reassessment, factors were ranked according to their standardized regression coefficients. The ranking results showed that consciousness level ranked first, hematoma volume second, ventricular rupture third, hematoma location fourth, surgical timing fifth, and hematoma clearance rate sixth. Consciousness level ranked highest in the acute option reassessment due to its comprehensive reflection of brain injury severity, while hematoma volume ranked second due to its direct determination of the space-occupying effect. The recovery-phase factor importance assessment used recovery-phase prognostic outcomes (modified Rankin Scale score at 6-month follow-up) as the target variable, and the assessment method also employed multivariate logistic regression. The factors in the recovery option reassessment were ranked according to their standardized regression coefficients. The ranking results showed that pulmonary infection ranked first, early rehabilitation intervention second, rebleeding third, and epileptic seizures fourth. Pulmonary infection ranked highest in the recovery option reassessment due to its prolonging of hospital stay and aggravation of neurological function impairment. The phased factor ranking revealed a phased characteristic: the acute phase was dominated by baseline characteristics and treatment factors, while the recovery phase was dominated by complications and rehabilitation factors. In the phased factor ranking, most of the leading factors in the acute phase were non-interventional or partially interventional, while in the phased factor ranking, most of the leading factors in the recovery phase were preventable or interventional.
[0060] Core influencing factors were determined by extracting high-weight factor sets from each stage of the disease progression. The high-weight factor set was defined as a subset of factors with relatively high weights. The extraction method used a threshold based on weight distribution, ensuring that included factors had a moderate to high prognostic impact. The high-weight factor set for the acute phase was extracted from the acute phase ranking results of the staged factor progression. Factors ranking high in the acute phase, such as level of consciousness, hematoma volume, ventricular rupture, hematoma location, and surgical timing, were included in this set. These factors exhibited high regression coefficients and significant prognostic differences in the acute phase of the staged factor progression. Similarly, the high-weight factor set for the recovery phase was extracted from the recovery phase ranking results. Factors ranking high in the recovery phase, such as pulmonary infection, early rehabilitation intervention, and rebleeding, were included in this set. These factors had a prominent impact on the final prognosis during the recovery phase of the staged factor progression. Core influencing factors were those belonging to the high-weight factor set at at least one stage of the disease progression. The extraction method combined the high-weight factor sets from the acute and recovery phases and removed duplicates. The core influencing factors identified several factors, including level of consciousness, hematoma volume, ventricular rupture, hematoma location, timing of surgery, pulmonary infection, early rehabilitation intervention, and rebleeding. The core influencing factors were divided into acute-phase core factors (level of consciousness, hematoma volume, ventricular rupture, hematoma location, and timing of surgery) and recovery-phase core factors (pulmonary infection, early rehabilitation intervention, and rebleeding). Hematoma volume and timing of surgery, which affect both phases, were labeled as cross-phase factors.
[0061] The key prognostic factors were determined by prioritizing core influencing factors according to disease progression, forming a combination of critical prognostic factors. Disease progression priority was defined as the contribution weight of different disease stages to the final prognostic outcome, with the acute phase having a higher priority than the recovery phase. This priority setting was based on the decisive impact of initial brain injury in the acute phase on prognosis. The phase-by-phase ranking was weighted according to disease progression priority, with the comprehensive weight calculated as W_comprehensive = α × W_acute phase + β × W_recovery phase. Here, W_comprehensive is the comprehensive weight of the core influencing factors, α and β are the weight coefficients of disease progression priority (α + β = 1), and W_acute phase and W_recovery phase are the standardized regression coefficients of the core influencing factors in the acute and recovery phases, respectively. A larger α indicates a greater contribution of the acute phase to prognosis. The prognosis of a comatose patient with brainstem hemorrhage was primarily determined by the acute-phase core influencing factors, with acute-phase factors such as level of consciousness and hematoma location playing a dominant role in the comprehensive weight. Among the core influencing factors, level of consciousness had the highest comprehensive weight, followed by hematoma volume, while pulmonary infection had a lower comprehensive weight but still belonged to the core influencing factors. The prognostic critical factor combination integrates the phased ranking results of core influencing factors. The prognostic critical factor combination is ranked from highest to lowest by comprehensive weight: level of consciousness (1st), hematoma volume (2nd), ventricular rupture (3rd), surgical timing (4th), hematoma location (5th), pulmonary infection (6th), early rehabilitation intervention (7th), and rebleeding (8th). The prognostic critical factor combination marks the intervention window for each factor: the intervention window for level of consciousness and hematoma volume is the acute phase; the intervention window for pulmonary infection, early rehabilitation intervention, and rebleeding is the recovery phase; and the intervention window for surgical timing is from the hyperacute to the acute phase. The marked window periods in the prognostic critical factor combination guide the selection of clinical intervention timing.
[0062] Step S150: Based on follow-up data, the combination of key prognostic factors is revised to form validated knowledge rules. Based on the validated knowledge rules, evidence-based level labeling is implemented to output clinical decision support results, thus completing the clinical data knowledge discovery for cerebral hemorrhage.
[0063] Specifically, a validated knowledge rule is formed by revising the combination of key prognostic factors based on follow-up data. This combination includes core influencing factors such as level of consciousness, hematoma volume, ventricular rupture, surgical timing, hematoma location, pulmonary infection, early rehabilitation intervention, and rebleeding. Follow-up data includes prognostic assessment results at four time points: 1 month, 3 months, 6 months, and 12 months post-discharge. Follow-up data revision involves comparing the predicted results of the key prognostic factor combination with the actual follow-up prognosis to identify prediction biases and adjust factor weights. The prognostic prediction model is constructed based on the core influencing factors of the key prognostic factor combination and their comprehensive weights. The comprehensive weights of each factor in the key prognostic factor combination serve as the initial coefficients of the model. The model input consists of the values of the core influencing factors, and the model output is a probability prediction of a good or bad prognosis. Consistency analysis between the prediction results and follow-up data uses 10-fold cross-validation to assess the model's robustness, comparing the prognostic assessment results at different time points in the follow-up data with the predicted results. Consistency analysis identified systematic patterns of prediction bias. The model showed high accuracy in predicting poor prognosis for critically ill patients, high accuracy in predicting good prognosis for mildly ill patients, but room for improvement in predicting accuracy for moderately ill patients. Follow-up data correction adjusted the factor weights of the prognostic key factor combination to address prediction bias. The weights of early rehabilitation intervention and pulmonary infection in moderately ill patients were increased accordingly. This weight adjustment was based on the sensitivity of moderate patient prognosis to rehabilitation and complication management. The corrected prognostic key factor combination formed a post-validated knowledge rule, expressed in IF-THEN rule form. The antecedent of the rule is the value condition of the core influencing factor, and the consequent is the predicted prognostic outcome. The post-validated knowledge rule contains several rules covering prognostic patterns for different patient subgroups and disease stages. Each rule in the post-validated knowledge rule is labeled with its prediction accuracy and support in the follow-up data.
[0064] Based on validated knowledge rules, evidence-based grading is implemented to output clinical decision support results, completing the knowledge discovery of clinical data related to cerebral hemorrhage. Validated knowledge rules include several IF-THEN rules and their prediction accuracy and support. Evidence-based grading is performed by classifying the validated knowledge rules according to their evidence strength and clinical reliability. The evidence-based grading standard references the Oxford Evidence-Based Medicine Centre's evidence grading system, with four levels: Level 1, Level 2, Level 3, and Level 4. Evidence-based grading considers three dimensions: rule accuracy, rule support, and rule reproducibility. Reproducibility is assessed through independent validation using multi-center data. The dataset is split into training and external validation sets based on medical institutions, and accuracy is calculated separately for each set. A difference between the two sets of accuracy less than a threshold is considered reproducible. Validated knowledge rules with high accuracy, high support, and successful reproducibility verification are labeled as Level 1 evidence-based rules. Validated knowledge rules with moderate accuracy and support and partially validated reproducibility are labeled as Level 2 evidence-based rules. Validated knowledge rules with moderate accuracy but low support and not yet passed independent reproducibility verification are labeled as Level 3 evidence-based rules. Validated knowledge rules with both low accuracy and support or lacking reproducibility verification data are labeled as Level 4 evidence-based rules. Clinical decision support results integrate the evidence-based level-labeled validated knowledge rules to provide clinicians with tiered recommendations: Level 1 rules are strong recommendations, Level 2 rules are recommendations, Level 3 rules are weak recommendations, and Level 4 rules are reference recommendations. Clinical decision support results are output in the form of decision trees and recommendation lists. The decision tree shows the complete reasoning path from patient characteristics to treatment decisions to prognostic prediction, and the recommendation list records the evidence-based level and recommendation strength of each validated knowledge rule. Clinical decision support results provide individualized prognostic prediction reports, which include five parts: basic patient information, core impact factor scores, prognostic risk level, recommended treatment plan, and expected prognostic outcome. Knowledge discovery of clinical data for cerebral hemorrhage completes the entire process of transformation from raw data to knowledge rules to clinical application. The knowledge discovery results include five knowledge products: clinical data feature set, association rule knowledge base, treatment pattern atlas, combination of key prognostic factors, and validated knowledge rules.
[0065] To implement the above-described method embodiments, a method for multidimensional mining and knowledge discovery of clinical data related to cerebral hemorrhage is proposed to achieve the corresponding functional and technical effects. See also... Figure 2 , Figure 2 This diagram illustrates a structural block diagram of a multidimensional mining and knowledge discovery system 200 for clinical data of cerebral hemorrhage provided in an embodiment of this application. For ease of explanation, only the parts relevant to this embodiment are shown. The multidimensional mining and knowledge discovery system 200 for clinical data of cerebral hemorrhage provided in this embodiment of the application includes: Data acquisition module 201 is used to acquire electronic medical records, surgical records and follow-up data of patients with multicenter intracerebral hemorrhage and to standardize and integrate the data to establish a clinical data feature set; Rule mining module 202 is used to extract multi-dimensional features from the clinical data feature set to generate a clinical feature spectrum, perform time window association rule mining on the clinical data feature set to determine the temporal association strength distribution, and perform association pattern screening and evaluation by combining the temporal association strength distribution with the clinical feature spectrum to form an association rule knowledge base; Pattern discovery module 203 is used to determine the patient type distribution characteristics by prognosis-oriented patient subgroup clustering based on the association rule knowledge base, extract treatment plan feature components by treatment plan similarity calculation based on the patient type distribution characteristics, and form a treatment pattern map based on the treatment plan feature components and the clinical data feature set. Factor analysis module 204 is used to perform multi-dimensional prognostic correlation analysis based on the treatment pattern map to establish a set of prognostic influencing factors, determine individualized factor weight thresholds by factor interaction effect detection based on the set of prognostic influencing factors, and perform factor importance ranking based on the individualized factor weight thresholds to form a combination of key prognostic factors. The knowledge verification module 205 is used to modify the combination of prognostic key factors based on the follow-up data to form a verified knowledge rule, and to implement evidence-based level labeling and output clinical decision support results according to the verified knowledge rule, thereby completing the clinical data knowledge discovery for cerebral hemorrhage.
[0066] The aforementioned system 200 for multidimensional mining and knowledge discovery of clinical data related to cerebral hemorrhage can implement the method described in the above embodiments. The options described in the above method embodiments are also applicable to this embodiment and will not be detailed here. The remaining contents of this application's embodiments can be referred to the contents of the above method embodiments, and will not be repeated in this embodiment.
[0067] The above embodiments are not an exhaustive list based on the present invention, and there may be many other embodiments not listed. Any substitutions and improvements made without departing from the concept of the present invention are within the protection scope of the present invention.
Claims
1. A method for multidimensional mining and knowledge discovery of clinical data on cerebral hemorrhage, characterized in that, include: Acquire electronic medical records, surgical records, and follow-up data of patients with multicenter intracerebral hemorrhage, and standardize and integrate the data to establish a clinical data feature set; The process involves extracting multidimensional features from the clinical data feature set to generate a clinical feature spectrum, and then performing time window association rule mining on the clinical data feature set to determine the temporal association strength distribution. This includes: extracting key medical event time nodes and event sequence parameters from the clinical data feature set; setting critical period time windows based on the key medical event time nodes; obtaining event association temporal patterns by combining the event sequence parameters and the critical period time windows; generating a temporal association strength distribution based on the event association temporal patterns; and performing association pattern screening and evaluation using the temporal association strength distribution combined with the clinical feature spectrum to form an association rule knowledge base. Specifically, generating the temporal association strength distribution based on the event association temporal patterns includes: calculating the association delay duration of the event association temporal patterns to obtain a delay time window distribution; identifying cascading effect associations of complications and independent delay associations based on the delay time window distribution; sorting the cascading effect associations of complications and the independent delay associations according to the delay effect strength to generate a delay effect sequence; and generating a temporal association strength distribution based on the delay effect sequence. Based on the association rule knowledge base, prognosis-oriented patient subgroup clustering is performed to determine the patient type distribution characteristics. Based on the patient type distribution characteristics, treatment plan feature components are extracted by treatment plan similarity calculation. Treatment pattern atlas is formed by combining the treatment plan feature components with the clinical data feature set. Based on the treatment pattern atlas, a multi-dimensional prognostic correlation analysis is performed to establish a set of prognostic influencing factors. Based on the set of prognostic influencing factors, the individualized factor weight threshold is determined by factor interaction effect detection. Based on the individualized factor weight threshold, the factor importance is ranked to form a combination of key prognostic factors. Based on the follow-up data, the combination of key prognostic factors is revised to form validated knowledge rules. Based on the validated knowledge rules, evidence-based level labeling is implemented to output clinical decision support results, thus completing the clinical data knowledge discovery for cerebral hemorrhage.
2. The method according to claim 1, characterized in that, The step of extracting multidimensional features from the clinical data feature set to generate a clinical feature spectrum includes: From the clinical data feature set, features of level of consciousness, bleeding volume, and hematoma location are extracted to generate a multidimensional original feature set; Cross-dimensional coupling deviation analysis is performed on the multidimensional original feature set to identify synergistic features and contradictory features between dimensions; Based on the inter-dimensional collaborative features and the inter-dimensional contradictory features, deviation compensation weights are calculated to generate hierarchical weight parameters. The hierarchical weight parameters are fused in multiple dimensions to construct a clinical feature spectrum.
3. The method according to claim 1, characterized in that, The process of determining patient type distribution characteristics through prognosis-guided patient subgroup clustering based on the association rule knowledge base includes: Prognostic stratification criteria are determined based on the prognostic outcome features of the association rule knowledge base. The prognostic stratification criteria are used to set stratification constraints for bleeding sites, generating a set of constraint parameters. Based on the set of constraint parameters, hierarchical clustering analysis is performed on the association rule knowledge base to construct the patient subgroup distribution; The distribution characteristics of patient types are determined based on the distribution of the patient subgroups.
4. The method according to claim 1, characterized in that, The step of extracting feature components of treatment plans based on the distribution characteristics of the patient types and the similarity of the treatment plans includes: From the patient type distribution characteristics, extract the treatment plan records corresponding to each subgroup to generate a candidate plan set; Extract treatment timing features and intervention method features from the candidate scheme set to generate timing intervention correlation parameters; A weighted similarity matrix is obtained by calculating the similarity of the intervention-related parameters based on the timing; Treatment plan feature components are constructed based on the weighted similarity matrix and the patient type distribution characteristics.
5. The method according to claim 1, characterized in that, The step of determining individualized factor weight thresholds based on the prognostic influencing factor set through factor interaction effect detection includes: Based on the set of prognostic influencing factors, interactive factor pairs are generated by performing factor pairing and combination. The interaction effect strength is calculated in the interaction factor pair to generate the interaction coefficient; Individualized correction parameters are obtained by matching the interaction coefficients with individual patient characteristics. The individualized factor weight threshold is determined based on the individualized correction parameters.
6. The method according to claim 1, characterized in that, The process of ranking factors based on their importance according to the individualized factor weight threshold to form a combination of key prognostic factors includes: The individualized factor weight thresholds are divided into acute option reset and recovery option reset according to the disease stage; Factor importance assessments were performed on the acute option reset and the recovery option reset respectively to obtain a phased factor ranking. The core influencing factors are determined by extracting the high-weight factor set for each stage from the phased factor ranking. Based on the core influencing factors, they are sorted in stages according to the priority of the disease course to form a combination of key prognostic factors.
7. The method according to claim 5, characterized in that, The process of obtaining individualized correction parameters by matching the interaction coefficients with individual patient characteristics includes: The interaction coefficients are classified and arranged according to the characteristics of patient comorbidities to construct a disease-stratified interaction sample; Identify the influence patterns of comorbidities on factor sensitivity in the disease-stratified interactive samples to determine sensitivity correction conditions; Based on the sensitivity correction conditions, the disease stratification interaction samples are divided into a high-risk disease group and a low-risk disease group; Sensitivity correction and weight allocation are performed on the high-risk disease group and the low-risk disease group respectively to obtain individualized correction parameters.
8. A system for multidimensional mining and knowledge discovery of clinical data on cerebral hemorrhage, characterized in that, include: The data acquisition module is used to acquire electronic medical records, surgical records and follow-up data of patients with multicenter intracerebral hemorrhage, and to standardize and integrate the data to establish a clinical data feature set; The rule mining module is used to extract multidimensional features from the clinical data feature set to generate a clinical feature spectrum, and to perform time window association rule mining on the clinical data feature set to determine the temporal association strength distribution. This includes: extracting key medical event time nodes and event sequence parameters from the clinical data feature set; setting critical period time windows based on the key medical event time nodes; obtaining event association temporal patterns by combining the event sequence parameters and the critical period time windows; generating a temporal association strength distribution based on the event association temporal patterns; and performing association pattern screening and evaluation by combining the temporal association strength distribution with the clinical feature spectrum to form an association rule knowledge base. Specifically, generating the temporal association strength distribution based on the event association temporal patterns includes: calculating the association delay duration of the event association temporal patterns to obtain a delay time window distribution; identifying cascading effect associations and independent delay associations based on the delay time window distribution; sorting the cascading effect associations and independent delay associations according to the delay effect strength to generate a delay effect sequence; and generating a temporal association strength distribution based on the delay effect sequence. The pattern discovery module is used to determine the patient type distribution characteristics by performing prognostic-oriented patient subgroup clustering based on the association rule knowledge base, extract treatment plan feature components by calculating treatment plan similarity based on the patient type distribution characteristics, and form a treatment pattern map based on the treatment plan feature components and the clinical data feature set. The factor analysis module is used to perform multi-dimensional prognostic correlation analysis based on the treatment pattern map to establish a set of prognostic influencing factors, determine individualized factor weight thresholds through factor interaction effect detection based on the set of prognostic influencing factors, and perform factor importance ranking based on the individualized factor weight thresholds to form a combination of key prognostic factors. The knowledge verification module is used to modify the combination of prognostic key factors based on the follow-up data to form verified knowledge rules, implement evidence-based level labeling based on the verified knowledge rules, output clinical decision support results, and complete the clinical data knowledge discovery for cerebral hemorrhage.
Citation Information
Patent Citations
Hypocephalus correlation analysis method and system based on multi-dimensional data
CN121393933A
Cerebral hemorrhage prognosis analysis system fusing images and clinical data
CN121617639A