A method for identifying mortality risk and safety indicator ranges in ICU patients

By integrating learning and causal inference frameworks, and combining multiple clustering algorithms and prognostic constraints, high-risk subtypes of ICU patients are identified and safe ranges for blood gas indicators are defined. This solves the problems of limited clinical value of subtyping results and unreliable indicator definition in existing technologies, and achieves efficient individualized management.

CN121306572BActive Publication Date: 2026-03-13TIANJIN FIRST CENT HOSPITAL
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-10
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing technologies are insufficient for effectively utilizing early, high-frequency vital sign data from patients to perform dynamic subtype stratification. Traditional clustering analysis methods are ineffective in handling missing values ​​and nonlinear relationships, resulting in poor clinical interpretability of the classification results. Furthermore, the warning thresholds for blood gas indicators do not take into account subgroup differences, making it impossible to accurately identify the mortality risk and safety indicator range of high-risk patients.

Method used

By employing multiple clustering algorithms to fuse consensus clustering and combining it with prognostic difference constraints, high-risk subtypes are identified through a causal inference framework. Furthermore, the safe range of blood gas indicators is analyzed using a causal risk-weighted two-factor interactive algorithm, thus achieving an automated analysis process from data to decision.

Benefits of technology

It improves the stability and clinical prognostic discrimination of clustering results, ensures that the classification results are related to patient mortality, provides accurate monitoring and intervention standards for high-risk patients, and improves the efficiency and accuracy of clinical decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121306572B_ABST
    Figure CN121306572B_ABST
Patent Text Reader

Abstract

This invention relates to the field of medical information technology, specifically disclosing a method for identifying the mortality risk and safety indicator range of ICU patients. The method includes: acquiring vital sign data and blood gas analysis data of patients after admission to the ICU; fusing the clustering results of the vital sign data through consensus clustering with prognostic variability constraints to obtain patient subtypes; selecting high-risk subtypes; performing causal risk-weighted two-factor interactive algorithm analysis on the blood gas analysis data of high-risk subtypes to obtain the safe range of blood gas indicators; and acquiring the blood gas indicators of target patients after admission to the ICU in real time to determine the mortality risk of the target patients. This invention improves the stability of subtyping through ensemble learning and ensures clear clinical prognostic differences in subtyping results through prognostic constraints; it controls confounding bias through a causal inference framework, identifies mortality risk, and accurately defines the risk range of key clinical indicators for high-risk subtype patients.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical information technology, and in particular to a method for identifying the mortality risk and safety indicator range of ICU patients. Background Technology

[0002] A significant challenge in intensive care units has long been the substantial heterogeneity of patients. This means that different patients exhibit vastly different pathophysiological mechanisms, clinical manifestations, disease progression trajectories, and responses to treatment. This heterogeneity limits the effectiveness of current diagnostic and treatment strategies, making it difficult to accurately identify and intervene early in high-risk patients. While some risk stratification methods exist based on static clinical indicators (such as APACHE II and SOFA scores) or single-time-point biomarkers, these methods often fail to capture the dynamic physiological trajectory of patients, struggle to reflect the real-time evolution of the disease, and lack precise prediction capabilities, thus failing to meet the needs of personalized medicine.

[0003] For example, sepsis complicated by thrombocytopenia is a common critical illness in intensive care units, characterized by high patient heterogeneity and mortality. Early and accurate risk stratification is crucial for clinical resource allocation and individualized treatment decisions. Currently, there is a lack of reliable methods for dynamic subtyping using early, high-frequency vital sign data (such as heart rate, blood pressure, respiratory rate, and blood oxygen saturation). Traditional clustering analysis methods (such as K-means) struggle to handle missing values, nonlinear relationships, and complex interactions between variables when processing such time-series data, resulting in poor clinical interpretability and insufficient robustness of the stratification results. Existing unsupervised clustering methods are entirely data-driven, and their stratification results may not be strongly correlated with clinically relevant hard endpoints (such as mortality), potentially producing statistically significant but clinically limited subtypes.

[0004] After identifying high-risk subgroups, the next key challenge is developing precise monitoring and intervention criteria. Taking blood gas analysis as an example, its indicators (such as pH, lactate, and base excess) are core indicators for assessing a patient's oxygenation, ventilation, and tissue perfusion, and are crucial for guiding resuscitation treatment. Partial dependency plots (PDPs) are used to understand the average marginal effect of individual features on model predictions (such as mortality risk). Traditional PDP analysis is based on correlation rather than causation. It assumes that all features are unrelated, which is not valid in clinical practice; for example, lactate levels are closely related to a patient's age and underlying diseases. Ignoring such confounding effects can lead to biased estimates of the risk range of indicators, misleading clinical decisions. Current techniques use PDPs to analyze the relationship between certain blood gas indicators and mortality risk in high-risk subgroups after obtaining patient subtypes, attempting to determine their risk range. However, current clinical warning thresholds for blood gas indicators are mostly general ranges determined based on studies of the overall population, without considering the differences in the intrinsic physiological compensatory mechanisms and risk characteristics of different patient subgroups. Applying universal standards to all patients may result in untimely warnings or excessive interventions for specific high-risk groups.

[0005] Currently, there is no method to further determine the specific and reasonable warning range of blood gas analysis indicators based on the dynamic physiological classification of patients. This technological gap limits the ability to conduct ultra-early and individualized precision management of high-risk patients. Cluster analysis and subsequent risk indicator identification are two separate steps, failing to form a unified analytical framework, resulting in insufficient overall system performance and intelligence. Summary of the Invention

[0006] This invention aims to address the problems of limited clinical value, low clinical prognostic discrimination, and unreliable definition of key indicators in classification results. To this end, this invention provides a method for identifying mortality risk and safety indicator ranges in ICU patients. It improves classification stability through ensemble learning and ensures clear clinical prognostic differences through prognostic constraints. Furthermore, it controls confounding bias through a causal inference framework, identifies mortality risk, and accurately defines the risk range of key clinical indicators for high-risk subtypes. This invention automates the data-to-decision analysis process, improving the efficiency and accuracy of clinical decision-making.

[0007] This invention provides a method for identifying the mortality risk and safety indicator range of ICU patients, and the technical solution adopted is as follows: including:

[0008] S1: Obtain vital signs and blood gas analysis data of patients after admission to the ICU;

[0009] S2: Multiple clustering algorithms are used to cluster vital sign data, and the clustering results are fused by consensus clustering with prognostic variability constraints to obtain patient subtypes;

[0010] S3: Select high-risk subtypes from patient subtypes;

[0011] S4: Perform causal risk weighted two-factor interactive algorithm analysis on the blood gas analysis data of high-risk subtypes to obtain the safe range of blood gas indicators;

[0012] S5: Real-time acquisition of blood gas parameters of the target patient after admission to the ICU, and assessment of the target patient's risk of death based on the safe range of blood gas parameters.

[0013] Furthermore, in step S1, vital sign data and blood gas analysis data of the patient in the first 12 hours after admission to the ICU are obtained.

[0014] Furthermore, step S2 includes:

[0015] Multiple clustering algorithms were used to cluster the vital signs data to obtain multiple clustering results. The multiple clustering results were then subjected to consensus clustering to generate multiple candidate subtypes.

[0016] Calculate the consensus index and prognostic variance of the consensus matrix corresponding to the candidate subtypes; the prognostic variance is calculated based on the p-value of the log-rank test of the survival curves of each subtype of the candidate subtypes.

[0017] Patient subtypes were determined based on concordance indicators and prognostic differences.

[0018] Furthermore, the candidate subtype that maximizes the composite score is selected as the final subtype, thus obtaining the patient subtype;

[0019] Composite score for:

[0020]

[0021] in, For candidate subtypes Corresponding consensus matrix Consistency indicators As weight, For prognostic variability, , For candidate subtypes The p-value.

[0022] Furthermore, clustering algorithms include: group multi-trajectory model, K-means clustering based on dynamic time warping, and K-shape clustering algorithm.

[0023] Furthermore, in step S3, the patient subtypes are validated using a multivariate Cox proportional hazards model to identify high-risk subtypes.

[0024] Furthermore, based on whether the patient died during the observation period, the subtype with the best prognosis was determined among the patient subtypes. Using the subtype with the best prognosis as a reference, the statistical significance of the ICU mortality risk and the mortality risk at 28 days, 90 days, and 365 days for other patient subtypes was verified using a multivariate Cox proportional hazards model. The confounding factors of the multivariate Cox proportional hazards model included age, gender, and SOFA score. High-risk subtypes were determined based on the output results of the multivariate Cox proportional hazards model.

[0025] Furthermore, step S4 includes:

[0026] S41: Using age, sex, SOFA score, and basic laboratory indicators that are related to mortality risk but relatively independent of the target blood gas index as input, a random forest model is used to predict the baseline mortality risk probability, and risk weights are calculated based on the baseline mortality risk prediction probability.

[0027] S42: Calculate the double robust estimator based on blood gas analysis data;

[0028] S43: Calculate the interactive partial dependence value of causal risk weighting based on risk weights and dual robust estimators to determine the safe range of blood gas indicators.

[0029] Furthermore, in step S43, a mortality risk surface is generated based on the interactive partial dependency value, and the region corresponding to the mortality risk at the lowest 5% quantile in the global mortality risk surface is selected as the safe range of blood gas indicators.

[0030] Further blood gas parameters include: pH, base excess, PO2, PCO2, Total CO2, and lactate;

[0031] The safe pH range is 7.32 to 7.64;

[0032] The safe range for PO2 is 25.00~324.32 mmHg;

[0033] The safe range for PCO2 is 21.94~53.74 mmHg;

[0034] The safe range for residual alkali is -7.47 to 23.00 mEq / L;

[0035] The safe range for lactic acid is 0.6~7.49 mmol / L;

[0036] The safe range for total PCO2 is 43.47~56.00 mEq / L.

[0037] The above-described one or more technical solutions in the embodiments of the present invention have at least one of the following technical effects:

[0038] 1. This invention incorporates prognostic variance constraints into the consensus clustering process. By constructing a composite scoring function combining clustering consistency indicators and prognostic variance, it ensures both the mathematical tightness and integration stability of the clustering results while enhancing clinical prognostic discriminability. This design makes the final patient subtyping not only mathematically sound but also accurately distinguishes the mortality risk of different subtypes. Through prognostic constraints, it ensures a strong correlation between subtyping results and patient mortality rates, enabling direct identification of high-risk populations and guiding priority clinical interventions.

[0039] 2. The clustering algorithm of this invention selects a multi-trajectory model (GBMTM), K-means based on dynamic time warping (DTW-Kmeans), and a K-shape clustering algorithm, forming a complementary temporal analysis system that can comprehensively capture vital sign data features from multiple dimensions: GBMTM, as a parametric trajectory modeling method, can accurately identify the overall trend of physiological indicators, such as persistent hypertension and progressive hypotension; DTW-Kmeans is insensitive to the temporal phase shift of physiological signals and can effectively handle the temporal differences in indicators such as heart rate and respiratory rate between individuals; K-shape focuses on the morphological features of physiological waveforms and can identify specific pathological respiratory patterns by ignoring amplitude differences. The combination of these three algorithms avoids the dimensional limitations of a single algorithm, comprehensively covers the physiological data features of trends, temporal differences, and morphology, and significantly improves the comprehensiveness and accuracy of clustering results. The integrated clustering framework of this invention effectively reduces the randomness of a single model, enabling better reproducibility of patient subtypes on different data subsets.

[0040] 3. This invention employs a multivariate Cox proportional hazards model to validate patient subtypes, offering multiple clinical and statistical advantages. By incorporating key covariates such as age, gender, and SOFA score, it eliminates interference from non-target factors, ensuring the validity of causal inference in risk assessment. It validates mortality risk at different time points, including in-ICU, 28 days, 90 days, and 365 days, comprehensively assessing the long-term and short-term prognostic differences among subtypes.

[0041] 4. This invention employs a causal risk-weighted two-factor interactive algorithm, providing precise and targeted technical support for defining the safe range of blood gas analysis indicators for high-risk subtypes. This invention calculates the baseline mortality risk weights of patients using a random forest model, assigning higher weights to data from high-risk patients, making the analysis more focused on high-risk individuals of clinical concern and improving the relevance of the results; the dual robust estimators effectively control confounding bias, ensuring the accuracy of the causal relationship between blood gas indicators and mortality risk. The causal inference framework of this invention effectively isolates the influence of confounding factors, and the risk-weighted mechanism makes the defined results more closely reflect the physiological state of the actual high-risk population, providing safer ranges for indicators with greater clinical reference value.

[0042] 5. This invention utilizes patient data from the first 12 hours after admission to the ICU. The first 12 hours most accurately reflect the initial severity of the disease and the body's compensatory capacity, providing high-quality early physiological data for subsequent classification and risk assessment. Compared to a 6-hour window, 12 hours captures the complete treatment response cycle, avoiding analytical biases caused by incomplete data; compared to a 24-hour window, it provides earlier warning information, meeting clinical requirements for timeliness.

[0043] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0044] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0045] Figure 1 This is a flowchart of the method provided by the present invention.

[0046] Figure 2 This invention provides vital sign trajectories and survival curves for three patient subtypes.

[0047] Figure 3 This invention provides a 3D visualization of the relationship between the lowest and highest blood gas analysis values ​​in the 12 hours prior to ICU admission and ICU mortality for cluster 2.

[0048] Figure 4 This invention provides a 2D heatmap of the lowest and highest pH and PO2 values ​​and ICU mortality values ​​in the first 12 hours after admission to the ICU for cluster 2.

[0049] Figure 5 This invention provides a 2D heatmap of the lowest and highest values ​​of PCO2 and lactate in the first 12 hours before ICU admission and the relationship between ICU mortality and cluster 2.

[0050] Figure 6 This invention provides a 2D heatmap of the minimum and maximum values ​​of alkali residue and Total PCO2 in the 12 hours prior to ICU admission for cluster 2, and the relationship between these values ​​and ICU mortality. Detailed Implementation

[0051] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention. The following embodiments are used to illustrate this invention but should not be used to limit the scope of this invention.

[0052] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0053] The following is combined Figures 1 to 6 The present invention will be further described in detail below, including a method for identifying the mortality risk and safety index range of ICU patients:

[0054] In this embodiment, as Figure 1 As shown, a method for identifying mortality risk and safety indicator ranges in ICU patients is provided, including the following steps:

[0055] S1: Obtain vital signs data and blood gas analysis data of patients after admission to the ICU.

[0056] Sepsis complicated by thrombocytopenia is common in the ICU and significantly increases the difficulty of treatment and the risk of patient death. Therefore, this embodiment extracts patient data that meet the diagnostic criteria for sepsis complicated by thrombocytopenia from the MIMIC-IV clinical database.

[0057] From a clinical perspective, the first 12 hours after a sepsis patient is admitted to the ICU are considered the "golden window" for pathophysiological changes. The physiological trajectory during this period most accurately reflects the initial severity of the disease and the body's compensatory capacity. From a technical standpoint, compared to a shorter time window (such as 6 hours), 12 hours can capture the complete treatment response cycle; and compared to a longer window (such as 24 hours), it can provide early warning information, meeting the clinical requirement for timeliness.

[0058] This embodiment extracts vital sign data and blood gas analysis data for each patient during the first 12 hours after admission to the ICU. Vital sign data are continuous and include, but are not limited to: systolic blood pressure (SBP), diastolic blood pressure (DBP), mean arterial pressure (MBP), heart rate (HR), respiratory rate (RR), and pulse oxygen saturation (SpO2). Blood gas analysis data are the lowest and highest values ​​of the indicators, including: pH (arterial blood acidity / alkalinity), base excess (BE), PO2 (partial pressure of oxygen in arterial blood), PCO2 (partial pressure of carbon dioxide in arterial blood), Total CO2 (total carbon dioxide in blood), and lactate.

[0059] Data preprocessing includes cleaning and interpolation.

[0060] Specifically, the search criteria in this embodiment are: age: 18-89 years; first hospitalization and first ICU admission; length of hospital stay greater than 1 day; length of ICU stay greater than 1 day; survival days greater than 1 day; and the patient was diagnosed with sepsis and thrombocytopenia. 12,126 ICU admissions were extracted from the MIMIC-IV database (V3.1).

[0061] Exclusion criteria were: missing heart rate data, missing respiratory rate data, missing systolic blood pressure data, missing diastolic blood pressure data, missing mean arterial pressure data, missing pulse oximetry data, missing biochemical indicators data, missing blood gas analysis indicators data, and missing data >30%. From 12,126 patients, 9,929 were excluded, and 2,197 patients met the criteria.

[0062] Data on the above 2197 patients who met the diagnostic criteria for sepsis complicated with thrombocytopenia (vital signs and blood gas analysis data in the first 12 hours after admission to the ICU) were extracted from the MIMIC-IV database (V3.1), and missing data were filled using linear interpolation.

[0063] S2: Multiple clustering algorithms are used to cluster vital sign data, and the clustering results are fused by consensus clustering with prognostic variability constraints to obtain patient subtypes.

[0064] This method uses multiple temporal clustering algorithms in parallel to cluster vital sign data.

[0065] Specifically, this embodiment uses three clustering algorithms: Model A: Group Multi-Track Model (GBMTM); Model B: K-means clustering based on Dynamic Time Warping (DTW) (DTW-Kmeans); Model C: K-shape clustering algorithm based on K-shape.

[0066] The three clustering algorithms selected in this embodiment have significant complementary advantages in their technical principles, enabling them to comprehensively capture the characteristics of vital sign data from different dimensions. The group multi-trajectory model, as a parametric trajectory modeling method, excels at capturing the overall changing trends and trajectory shapes of physiological indicators, effectively identifying trend patterns such as persistent hypertension or progressive hypotension. The K-means algorithm based on dynamic time warping is insensitive to the temporal phase shift of physiological signals, effectively handling temporal differences in physiological responses between individuals, and is particularly suitable for capturing different changing patterns of heart rate and respiratory rate. The K-shape clustering algorithm focuses on the morphological features of physiological waveforms, ignoring amplitude differences while identifying specific morphological features such as pathological respiratory patterns. The combination of these three methods forms a comprehensive time-series analysis system. Validation through Bootstrap resampling shows that the clustering stability of the ensemble method is significantly higher than any single model, demonstrating its excellent technical synergy.

[0067] In this embodiment, the vital signs data are clustered using a group multi-trajectory model, K-means clustering based on dynamic time warping, and K-shape clustering algorithm, respectively, to obtain the clustering results of the three models. Then, the three clustering results are subjected to consensus clustering to generate multiple candidate subtypes.

[0068] The specific process is as follows: A co-association matrix is ​​constructed based on the three clustering results. For those with… A dataset of patients was used to construct a... Common correlation matrix Among them, the elements in the co-correlation matrix Indicates the patient and patients The frequency of being classified into the same cluster in all three models.

[0069]

[0070] in, For model A, for patients The cluster allocation results For model B, for patients The cluster allocation results For model C, for patients The cluster allocation results For model A, for patients The cluster allocation results For model B, for patients The cluster allocation results For model C, for patients The cluster allocation results This is an indicator function; its value is 1 when the condition is true, and 0 otherwise.

[0071] Based on the common correlation matrix Multiple candidate fractal schemes are generated using a hierarchical clustering algorithm. The specific process includes: converting the common correlation matrix M into a distance matrix. , Agglomerative hierarchical clustering was applied, and the average linkage method was used to cut the cluster tree at different heights to generate multiple classification schemes with K=2, 3, 4, and 5 clusters, where K is the number of clusters; the cluster allocation results corresponding to each cutting height were recorded as candidate classifications.

[0072] Consensus clustering optimization objective function with constraints: By incorporating prognostic dissimilarity constraints into consensus clustering, the objective is no longer simply to maximize the consistency of the consensus matrix, but rather to maximize a composite scoring function, calculated as follows:

[0073]

[0074] in, For composite scoring, For candidate subtypes Corresponding consensus matrix Consistency indicators As weight, For prognostic variability.

[0075] This embodiment introduces prognostic difference constraints during the consensus process. The log-rank test p-value for the 365-day survival curves of each subtype under each candidate subtyping is calculated. The scheme with the smallest p-value (most significant prognostic difference) is preferentially selected as the final subtyping. The log-rank test is the gold standard method for survival analysis, and its statistical power has been widely validated. The p-value, as a statistical significance indicator, has the characteristic of model independence, ensuring the universality and comparability of prognostic constraints across different clustering algorithms. Minimizing the p-value as the optimization objective directly guides the subtyping scheme that maximizes clinical value, ensuring subtype stability and clear prognostic discriminative ability. In specific implementation, a composite scoring function organically combines clustering quality and prognostic significance, ensuring that prognostic discriminative power is maximized while maintaining clustering quality.

[0076] For any candidate fraction (Generated by consensus clustering), its prognostic variability Calculate using the following steps:

[0077] according to The patients were divided into m subtypes;

[0078] For each subtype, plot its 365-day survival curve;

[0079] The significance level of the differences between the m survival curves is calculated using the log-rank test, and the p-value is obtained, denoted as [p-value]. ,Right now For candidate subtypes The p-value;

[0080] Calculate prognostic variability , .therefore, The smaller, The larger the value, the stronger the prognostic discrimination ability of the candidate subtype.

[0081] In this embodiment, the adjusted RAND index (ARI) is chosen as the consistency metric. The ARI is corrected for "accidental" expected values ​​and is used to measure consistency between different clustering schemes, ensuring ensemble stability. To calculate the consistency between the candidate scheme and the ensemble base, this embodiment calculates the ARI of the candidate scheme with the three base models (GBMTM, DTW-Kmeans, K-shape) and then takes the average.

[0082] It is a pre-defined regularization hyperparameter used to balance the mathematical tightness of fractal typing with the discriminative power of clinical prognosis. In some preliminary experimental or simulation scenarios, A value of 0.52 can achieve a relatively good balance between the two aspects mentioned above. It is not fixed; it can be adjusted based on factors such as specific data characteristics, dataset size, and actual application scenario. For example, when the differences between samples in the dataset are significant, it may be necessary to appropriately increase the size. The value is used to further emphasize the discriminative power of clinical prognosis.

[0083] In patient subtype identification, simply maximizing the consistency of the consensus matrix can only mathematically guarantee a certain compactness of the subtyping, but may ignore its practical clinical significance. A composite scoring function... Consistency indicators Prognostic variability Combined, through Balancing these two aspects ensures that the subtyping is both mathematically compact and clinically discriminative, guaranteeing that the subtyping results are not only mathematically sound but also have practical clinical value. Selecting the subtyping scheme that maximizes the composite score from all candidate schemes allows for the selection of the subtyping scheme with optimal overall performance, considering multiple factors, resulting in stable patient subtypes that are strongly correlated with prognosis. This approach helps achieve the invention's objective of providing a more robust method for patient subtype identification, improving subtyping stability through ensemble learning and ensuring clear clinical prognostic differences through prognostic constraints.

[0084] The composite scoring function is designed based on the principle of multi-objective optimization, comprehensively considering both clustering quality and clinical value. The ARI (Advanced Relationship Index) in the function measures the consistency among different clustering schemes, ensuring the stability of ensemble learning. Transforming the log-rank test p-value into a negative logarithmic form creates a monotonically increasing benefit function, prioritizing the subtypes with the greatest clinical significance. The final weight allocation ensures both the mathematical rationality of the clustering results and emphasizes clinical value orientation, thus balancing mathematical properties with clinical application needs.

[0085] Finally, the candidate subtype with the highest composite score was selected as the final subtype, thus obtaining the patient subtype.

[0086] The patient subtypes calculated in this embodiment include: hypertension with hyperoxia (cluster 1), high heart rate with high respiratory rate and low oxygen (cluster 2), and low heart rate with low respiratory rate and high oxygen (cluster 3). The number of patients in cluster 1 is 527, the number in cluster 2 is 777, and the number in cluster 3 is 893. The vital sign trajectories and survival curves for the three patient subtypes are shown in the figure below. Figure 2 ,in, Figure 2 Figures (a)-(f) show the trajectory diagrams of six vital signs. Figure 2 The graph in (g) is a survival curve.

[0087] S3: Select high-risk subtypes from patient subtypes.

[0088] This embodiment uses a multivariate Cox proportional hazards model to validate patient subtypes and identify high-risk subtypes. Based on whether patients die during the observation period, the subtype with the best prognosis is determined. Using the subtype with the best prognosis as a reference, the statistical significance of ICU mortality risk and 28-day, 90-day, and 365-day mortality risks for other patient subtypes is validated using a multivariate Cox proportional hazards model. Confounding factors in the multivariate Cox proportional hazards model include age, gender, and SOFA score. High-risk subtypes are determined based on the output of the multivariate Cox proportional hazards model.

[0089] (1) The input variables of the multivariate Cox proportional hazards model include:

[0090] Dependent variable (time-event data):

[0091] Time variable: The duration from admission to the ICU to the occurrence of a death or the end of the study, in days;

[0092] Event variable: a binary indicator that records whether a patient dies during the observation period, 1 = death, 0 = censoring.

[0093] Independent variables (core variables and covariates):

[0094] Core independent variable: patient subtype classification; in this embodiment, the low heart rate, low respiratory rate, and high oxygen type with the best prognosis is used as the reference category;

[0095] Create dummy variables for subtypes; for example: Cluster 3 (low heart rate, low respiration, high oxygen): 1 = belongs to cluster 3, 0 = other; Cluster 2 (high heart rate, high respiration, low oxygen): 1 = belongs to cluster 2, 0 = other; Cluster 1: reference category, all dummy variables are 0;

[0096] Covariates (confounding factors selected based on clinical importance): Age (continuous, years); Gender (binary variable: male = 1, female = 0); SOFA score (continuous).

[0097] (2) Model construction and output:

[0098]

[0099] in, Let be the risk function at time t. As the benchmark risk function, As the first dummy variable, As the second dummy variable, For age, For gender, Rate SOFA As the first weight, As the second weight, As the third weight, As the fourth weight, It is the fifth weight.

[0100] Model output includes: Hazard Ratio (HR): the adjusted relative risk value of each variable; 95% confidence interval: the accuracy estimate of the hazard ratio; P-value: to test whether the hazard ratio is statistically significant; and model goodness of fit: evaluated through likelihood ratio test, Score test, etc.

[0101] In this embodiment, based on whether the patient died during the observation period, the subtype with the best prognosis was determined to be the hypertensive hyperoxia type. Figure 2As shown in (g), cluster 1 has the highest survival rate. Then, using cluster 1 as a reference, a multivariate Cox proportional hazards model was used to verify the statistical significance of ICU mortality risk and 28-day, 90-day, and 365-day mortality risks for other patient subtypes (cluster 2 and cluster 3). Kaplan-Meier survival curves were plotted to visually demonstrate the survival differences among patient subtypes. This embodiment adjusted for three confounding factors: age, sex, and SOFA score. Age is an independent prognostic factor for sepsis, reflecting the objective law that physiological reserves decline with age. The sex factor considered the sex differences in immune response and prognosis. The SOFA score, as the quantitative gold standard for assessing organ dysfunction in sepsis, is exponentially correlated with patient mortality. The selection of these factors covered key dimensions such as demographic characteristics, baseline physiological status, and disease severity, while avoiding laboratory indicators that are collinear with blood gas analysis indicators, as well as treatment measures that are mediating variables, ensuring the effectiveness and accuracy of causal inference.

[0102] (3) Output Results: The output results of the multivariate Cox proportional hazards model are shown in Table 1. This table contains the following core information: Subtype risk comparison: Cluster 2 and Cluster 3 relative to Cluster 1 (reference group), including adjusted hazard ratios (HR); 95% confidence interval (95% CI); statistical significance P-value; Time node analysis: verifying the risk of death in the ICU, 28 days, 90 days and 365 days respectively, each time node corresponds to a complete set of Cox regression results. Among them, the factors not adjusted in the coarse model, the factors adjusted in Model 1 are age, gender, and SOFA score; the factors adjusted in Model 2 are age, gender, SOFA score, APSIII score, whether mechanical ventilation is used, whether CPR is used, whether RRT is used and whether CRRT is used.

[0103] The output results show that, with cluster 1 as the reference, the ICU mortality rate and the 28-day / 90-day / 365-day mortality hazard ratio (HR) of patients in cluster 2 are significantly higher than those in the reference group. Therefore, in this embodiment, cluster 2 is identified as a high-risk subtype, that is, the high heart rate, high respiratory rate and low oxygen type is a high-risk subtype.

[0104] Table 1

[0105]

[0106] S4: Perform causal risk weighted two-factor interactive algorithm analysis on the blood gas analysis data of high-risk subtypes to obtain the safe range of blood gas indicators.

[0107] This step targets the identified high-risk subtypes, aiming to accurately define the danger range of blood gas analysis indicators in the first 12 hours after ICU admission. Based on baseline risk prediction and weight allocation using random forest, a causal risk-weighted two-factor interactive algorithm is employed to analyze the causal relationship between the minimum and maximum values ​​of a certain blood gas indicator and the risk of death in the ICU. A 3D curve is obtained after controlling for confounding bias and risk weighting using dual robust estimators. Analysis of this curve reveals that when the indicator exceeds a certain threshold range, the risk of death begins to increase significantly. This threshold range represents the most reasonable danger range for this high-risk subtype, thus determining the safe range for the indicator.

[0108] The specific process is as follows:

[0109] S41: Using a pre-trained random forest model, calculate the baseline mortality risk weight for each individual in the high-risk subtype patient cohort.

[0110] To emphasize high-risk patients in the analysis, this embodiment quantifies the baseline risk of each patient and assigns weights accordingly. This embodiment employs a random forest model, which effectively captures complex nonlinear relationships between variables and is insensitive to outliers. Historical patient data from the same high-risk subtype (cluster 2) is used for model training. Model inputs include age (continuous variable), sex (categorical variable), SOFA score, and basic laboratory indicators related to mortality risk but relatively independent of the target blood gas parameters. Generally, the basic laboratory indicators used in this step include the initial platelet count upon admission and serum creatinine, etc. The target blood gas analysis indicators to be analyzed (such as lactate and pH) are excluded from the basic laboratory indicators to ensure that the predicted baseline risk is independent of the instantaneous changes of these indicators and to avoid introducing bias. Model output: The model's prediction target is a binary 28-day mortality outcome. Therefore, the random forest model actually outputs the predicted mortality risk probability for each patient. Risk weights are calculated based on the predicted mortality risk probability. This weighting mechanism ensures that subsequent analyses can differentiate and focus on patients at different risk levels, assigning higher weights to data from high-risk patients.

[0111] Obtain the baseline mortality risk prediction probability Then, calculate the patient Risk weight The calculation formula is:

[0112]

[0113] in, For patients The baseline mortality risk prediction probability is calculated by a pre-trained random forest model. This is a sensitivity parameter, and its value is adjustable. In this embodiment, =1.5. This formula ensures that patients with a higher risk of death are given a greater weight in subsequent analyses.

[0114] S42: Calculate the double robust estimator based on blood gas analysis data.

[0115] For a given combination of two factors (a, b) for a certain blood gas index, such as the minimum value of lactate... and the highest value For patients Dual robust estimators Defined as:

[0116]

[0117] in, The resulting model is used to predict the outcome given a vector of confounding factors. Expected mortality risk at (e.g., age) and blood gas levels (a, b); This is a propensity score model used to estimate a given... At that time, the patient The combination of blood gas index levels is the joint probability of (a, b); As an indicator function, when the patient The value of the blood gas index is 1 when the actual blood gas index value is equal to (a, b), and 0 otherwise. For patients The lowest value of the actual blood gas index, For patients The highest actual blood gas index value; For patients In the actual 28-day death outcome, 1 person died and 0 people survived.

[0118] S43: Calculate the interactive partial dependence value of causal risk weighting based on risk weights and dual robust estimators to determine the safe range of blood gas indicators.

[0119] Combining the above two parts, for the combination of values ​​(a, b), its final causal risk-weighted interactive partial dependency value The calculation formula is:

[0120]

[0121] in, For the number of patients. Through the entire Calculated on the value grid This generates a causally adjusted, risk-weighted mortality risk surface. Analyzing this surface reveals the lowest 5% quantile for mortality risk within the global range. By identifying the region, the safe range for that blood gas indicator can be accurately found.

[0122] The 3D visualization and 2D heatmap of the lowest and highest blood gas analysis values ​​in the first 12 hours of ICU admission for cluster 2, and the relationship between ICU mortality, are shown below. Figure 3 and Figure 4 As shown, where, Figure 3 In the middle (a), pH is represented. Figure 3 In the middle (b), PO2 is present. Figure 3 In the middle (c), PCO2 is represented. Figure 3 In the middle (d), lactic acid is present. Figure 3 (e) represents the alkali residue. Figure 3 (f) represents Total PCO2; Figure 4 In the middle (a), pH is represented. Figure 4 In the middle (b), PO2 is present. Figure 5 In the middle (a), PCO2 is represented. Figure 5 In (b), lactic acid is present. Figure 6 In the middle (a), the base residue is present. Figure 6 (b) represents Total PCO2.

[0123] according to Figure 3 and Figure 4 The safe range for this indicator is defined as the combination area corresponding to the lowest 5th percentile of the global mortality risk. The safe ranges for each blood gas indicator are as follows:

[0124] The safe pH range is 7.32 to 7.64;

[0125] The safe range for PO2 is 25.00~324.32 mmHg;

[0126] The safe range for PCO2 is 21.94~53.74 mmHg;

[0127] The safe range for residual alkali is -7.47 to 23.00 mEq / L;

[0128] The safe range for lactic acid is 0.6~7.49 mmol / L;

[0129] The safe range for total PCO2 is 43.47~56.00 mEq / L.

[0130] S5: Real-time acquisition of blood gas parameters of the target patient after admission to the ICU, and assessment of the target patient's risk of death based on the safe range of blood gas parameters.

[0131] For target patients admitted to the ICU, their blood gas parameters are monitored in real time. If a certain blood gas parameter exceeds the safe range of its corresponding blood gas parameter at a certain moment, an early warning is triggered, reminding relevant personnel to pay close attention to the target patient.

[0132] This invention successfully identifies high-risk subtypes (cluster 2) and utilizes an innovative two-factor interactive algorithm to find a reasonable and safe dynamic range for lactate levels and other parameters in patients with this subtype over the previous 12 hours. This range is presented in the form of a dual-indicator linkage, representing the ideal physiological state target associated with the lowest mortality risk for patients with this subtype after excluding confounding factors such as age and underlying diseases and focusing on high-risk groups, providing more precise guidance for refined clinical management.

[0133] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for identifying mortality risk and safety indicator ranges in ICU patients, characterized in that, include: S1: Obtain vital signs and blood gas analysis data of patients after admission to the ICU; S2: Multiple clustering algorithms are used to cluster vital sign data, and the clustering results are fused by consensus clustering with prognostic variability constraints to obtain patient subtypes; S3: Select high-risk subtypes from patient subtypes; S4: Perform causal risk weighted two-factor interactive algorithm analysis on the blood gas analysis data of high-risk subtypes to obtain the safe range of blood gas indicators; Step S4 includes: S41: Using age, sex, SOFA score, and basic laboratory indicators that are related to mortality risk but relatively independent of the target blood gas index as input, a random forest model is used to predict the baseline mortality risk probability, and risk weights are calculated based on the baseline mortality risk prediction probability. S42: Calculate the double robust estimator based on blood gas analysis data; S43: Calculate the interactive partial dependence value of causal risk weighting based on risk weights and dual robust estimators to determine the safe range of blood gas indicators; For the combination of two factors (a, b) of blood gas parameters, the interactive partial dependency value of causal risk weighting. The calculation formula is: in, For the number of patients, For patients Risk weights, patient Dual robust estimators Defined as: in, For the result model, For the tendency score model, For indicator functions, For patients The lowest value of the actual blood gas index, For patients The highest actual blood gas index value, For patients The actual outcome was death within 28 days; S5: Real-time acquisition of blood gas parameters of the target patient after admission to the ICU, and assessment of the target patient's risk of death based on the safe range of blood gas parameters.

2. The method for identifying mortality risk and safety indicator ranges in ICU patients as described in claim 1, characterized in that, In step S1, vital sign data and blood gas analysis data of the patient in the first 12 hours after admission to the ICU are obtained.

3. The method for identifying mortality risk and safety indicator ranges in ICU patients as described in claim 1, characterized in that, Step S2 includes: Multiple clustering algorithms were used to cluster the vital signs data to obtain multiple clustering results. The multiple clustering results were then subjected to consensus clustering to generate multiple candidate subtypes. Calculate the consensus index and prognostic variance of the consensus matrix corresponding to the candidate subtypes; the prognostic variance is calculated based on the p-value of the log-rank test of the survival curves of each subtype of the candidate subtypes. Patient subtypes were determined based on concordance indicators and prognostic differences.

4. The method for identifying mortality risk and safety indicator ranges in ICU patients as described in claim 3, characterized in that, The candidate subtype that maximizes the composite score is selected as the final subtype, thus obtaining the patient subtype. Composite score for: in, For candidate subtypes Corresponding consensus matrix Consistency indicators As weight, For prognostic variability, , For candidate subtypes The p-value.

5. A method for identifying mortality risk and safety indicator ranges in ICU patients as described in claim 3 or 4, characterized in that, Clustering algorithms include: group multi-trajectory model, K-means clustering based on dynamic time warping, and K-shape clustering algorithm.

6. The method for identifying mortality risk and safety indicator ranges in ICU patients as described in claim 1, characterized in that, In step S3, the patient subtypes are validated using a multivariate Cox proportional hazards model to identify high-risk subtypes.

7. The method for identifying mortality risk and safety indicator ranges in ICU patients as described in claim 6, characterized in that, Based on whether the patient died during the observation period, the subtype with the best prognosis was determined. Using the subtype with the best prognosis as a reference, the statistical significance of the ICU mortality risk and the mortality risk at 28 days, 90 days, and 365 days for other patient subtypes was verified using a multivariate Cox proportional hazards model. The confounding factors of the multivariate Cox proportional hazards model included age, gender, and SOFA score. High-risk subtypes were determined based on the output of the multivariate Cox proportional hazards model.

8. The method for identifying mortality risk and safety indicator ranges in ICU patients as described in claim 1, characterized in that, In step S43, a mortality risk surface is generated based on the interactive partial dependency value, and the region corresponding to the mortality risk at the lowest 5% quantile in the global mortality risk surface is selected as the safe range of blood gas indicators.

9. The method for identifying mortality risk and safety indicator ranges in ICU patients as described in claim 1, characterized in that, Blood gas parameters include: pH, base excess, PO2, PCO2, Total CO2, and lactate; The safe pH range is 7.32 to 7.64; The safe range for PO2 is 25.00~324.32 mmHg; The safe range for PCO2 is 21.94~53.74 mmHg; The safe range for residual alkali is -7.47 to 23.00 mEq / L; The safe range for lactic acid is 0.6~7.49 mmol / L; The safe range for total PCO2 is 43.47~56.00 mEq / L.

Citation Information

Patent Citations

  • Time trajectory typing method based on peripheral blood lymphocyte counting 72 hours after sepsis patient is transferred into ICU and application of time trajectory typing method based on peripheral blood lymphocyte counting 72 hours after sepsis patient is transferred into ICU

    CN118800460A

  • CVD (Chemical Vapor Deposition) death subgroup identification method and device in combination with causal reasoning and consensus clustering

    CN120260930A