Sepsis death risk assessment method and system based on double-index nonlinear synergy
Patent Information
- Application Number
- CN202611256705.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-19
- Publication Date
- 2026-09-25
AI Technical Summary
1.仅单独分析SHR(应激性高血糖比值)或SII(全身免疫炎症指数)的线性效应,未揭示其与死亡风险的非线性剂量–反应关系;
[0015]本发明实施例中的上述一个或多个技术方案,至少具有如下技术效果之一:
Smart Images

Figure CN122822360A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical information processing technology, and in particular to a method and system for assessing the risk of death from sepsis based on dual-indicator nonlinear synergy. Background Technology
[0002] Sepsis is a leading cause of death in critically ill patients, and early, accurate risk stratification is crucial for optimizing treatment strategies and reducing mortality. Currently used prognostic assessment tools, such as the SOFA score, APACHE II score, and SAPS II score, require the collection of multiple laboratory indicators and clinical parameters (such as bilirubin, creatinine, platelets, and GCS score), making data acquisition cumbersome and difficult to implement in primary care hospitals or emergency departments for rapid assessment. In recent years, SHR and SII, as two simple indicators requiring only routine blood glucose and complete blood counts, have been independently associated with poor prognosis in sepsis patients by multiple studies. However, existing technologies have the following limitations: 1. Analyzing the linear effects of SHR (stress-induced hyperglycemia ratio) or SII (systemic immune inflammation index) alone did not reveal a non-linear dose-response relationship between them and mortality risk; 2. Only simple logistic regression or linear weighting was used, without considering the synergistic amplification interaction between SHR and SII; 3. The model uses statistical indicators such as AUC and accuracy as optimization objectives, outputs abstract probabilities, and attributes no risk to actionable clinical recommendations, resulting in low clinical translation value; 4. The model has poor interpretability, lacks a visual interface, and is difficult to embed into clinical workflows.
[0003] Therefore, there is an urgent need for a highly accurate, probability-calibrated, interpretable, and clinically guiding method for assessing the risk of death from sepsis. Summary of the Invention
[0004] This invention aims to solve the aforementioned problems. To this end, this invention provides a method and system for assessing sepsis mortality risk based on dual-indicator nonlinear synergy. It automatically constructs derived features from SHR and SII that characterize nonlinear effects and synergistic interactions; and builds an ensemble learning model adapted to complex nonlinear relationships to achieve probabilistic posterior calibration, making the predicted probability close to the actual mortality rate; the model output is transformed into risk attribution explanations and structured clinical intervention recommendations, realizing a closed loop of "prediction-explanation-decision". This invention requires only routine laboratory indicators, has strong interpretability, and can directly output clinical recommendations for physician reference, making it suitable for rapid risk stratification in emergency departments and ICUs.
[0005] This invention provides a sepsis mortality risk assessment method based on dual-indicator nonlinear synergy, the technical solution of which is as follows: S1: Obtain the patient's stress-hyperglycemia ratio (SHR), systemic immune-inflammation index (SII), and routine clinical indicators; S2: Construct nonlinear derived features based on SHR and SII, and combine them with conventional clinical indicators to form the original features; S3: Input the original features into the heterogeneous ensemble learning model to calculate the initial probability of death risk; S4: Based on routine clinical indicators, perform hierarchical Bayesian probability calibration on the initial mortality risk probability to obtain the calibrated mortality risk probability; S5: Based on the original features, heterogeneous ensemble learning model and calibrated mortality risk probability, risk levels and clinical recommendations are generated through fuzzification and inference. S6: Output a mortality risk assessment report, which includes the calibrated mortality risk probability, risk level, and clinical recommendations.
[0006] Further routine clinical indicators include: age, SOFA score, lactate, Charlson index, GCS score, length of hospital stay, gender, and AKI score.
[0007] Furthermore, in step S2, based on SHR and SII, derived features are constructed using derived feature construction methods; wherein, the derived feature construction methods include: polynomial basis functions, radial basis functions, quantile piecewise linear transformation, logarithmic-exponential composite transformation, higher-order interaction terms, and nonlinear ratios; Preprocessing routine clinical indicators yields their characteristics; The original features are obtained by concatenating the features of conventional clinical indicators with the derived features.
[0008] Furthermore, when both SHR and SII have more than three measurement data, the derived feature construction method also includes: multi-scale sliding window statistics.
[0009] Furthermore, the heterogeneous ensemble learning model is a cascaded deep forest, whose first layer uses multiple heterogeneous models as base learners to calculate multiple death probabilities based on the original features; Its second layer concatenates the death probability with the original features to form enhanced features; Its third layer employs a meta-learner, using logistic regression and ridge regression regularization to calculate the initial mortality risk probability based on the enhanced features.
[0010] Furthermore, heterogeneous models include random forest, extreme random tree, XGBoost, LightGBM, and SVM.
[0011] Furthermore, in step S4, the AKI subgroup is determined based on the AKI score in routine clinical indicators, and the regression intercept term and regression slope coefficient corresponding to the subgroup are obtained; the initial mortality risk probability is calibrated to obtain the calibrated mortality risk probability. The formula for calculating the probability of death risk is: in, This represents the calibrated probability of death. For the regression intercept term of subgroup g, Let g be the regression slope coefficient for subgroup g. for The corresponding log odds, This represents the initial probability of death.
[0012] Furthermore, step S5 includes: Based on the original features and the heterogeneous ensemble learning model, the total SHAP contribution of each feature to the prediction result is calculated. The calibrated mortality risk probability is fuzzified to obtain the corresponding risk level; the total SHAP contribution of SHR and SII is fuzzified to obtain the corresponding contribution level. Based on a pre-defined fuzzy rule base, clinical recommendations are obtained according to risk level and contribution level.
[0013] Furthermore, in step S5, a waterfall plot is generated based on the total SHAP contribution of each feature to the prediction result; In step S6, the mortality risk assessment report also includes a waterfall chart.
[0014] This invention also provides a sepsis mortality risk assessment system based on dual-index nonlinear synergy, the technical solution of which is as follows: including: The data acquisition module is used to acquire patients' SHR, SII, and routine clinical indicators; The feature generation module constructs nonlinear derived features based on SHR and SII, and combines them with conventional clinical indicators to form the original features. The mortality risk prediction module is used to input the original features into the heterogeneous ensemble learning model and calculate the initial mortality risk probability. The mortality risk calibration module is used to perform hierarchical Bayesian probability calibration on the initial mortality risk probability based on routine clinical indicators, so as to obtain the calibrated mortality risk probability. The mortality risk assessment module is used to generate risk levels and clinical recommendations based on raw features, heterogeneous ensemble learning models, and calibrated mortality risk probabilities through fuzzification and inference. The report generation module is used to output a mortality risk assessment report, which includes a calibrated mortality risk probability, risk level, and clinical recommendations.
[0015] The above-described one or more technical solutions in the embodiments of the present invention have at least one of the following technical effects: 1. This invention constructs high-dimensional derived features by performing nonlinear transformations on SHR and SII, and integrates five heterogeneous models—random forest, extreme random tree, XGBoost, LightGBM, and SVM—through a three-layer cascaded forest structure, covering multiple learning paradigms such as bagging, gradient boosting, and kernel methods. This effectively improves the fitting ability to nonlinear and non-monotonic risk patterns, thereby obtaining more accurate mortality risk prediction results.
[0016] 2. This invention subgroups patients based on the staging of acute kidney injury, and uses a hierarchical Bayesian framework to independently estimate the intercept and slope parameters for each subgroup, correcting the systematic bias of the probability output of the heterogeneous ensemble learning model, and outputting the 95% confidence interval of the calibrated probability of death, providing a reliable measure of uncertainty for clinical decision-making.
[0017] 3. On the one hand, this invention calculates the SHAP value of each base learner using the FastSHAP method and aggregates them with AUC weights to generate a waterfall chart that intuitively displays the marginal contribution of each feature variable to mortality risk. On the other hand, it inputs the calibrated mortality risk probability and the SHAP contribution of SHR and SII into the Mamdani-type fuzzy inference rule engine to automatically generate structured clinical recommendations that match the risk level and key driving factors, providing interpretable and actionable auxiliary decision-making information.
[0018] 4. This invention targets sepsis, a specific disease, using SHR and SII as core dual indicators. Through feature derivation, heterogeneous ensemble learning model prediction, and probabilistic posterior calibration, it brings the predicted probability close to the actual mortality rate, generating risk levels and clinical recommendations. This achieves a closed-loop process from routine laboratory data to decision support recommendations. Furthermore, it requires only routine blood tests and blood glucose levels, eliminating the need for additional sampling or specialized equipment, significantly lowering the clinical implementation threshold for risk assessment.
[0019] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0021] Figure 1 This is a flowchart of the method provided by the present invention.
[0022] Figure 2 This is a comparison chart of the ROC of our method and the benchmark model.
[0023] Figure 3 This is a comparison chart of calibration curves for heterogeneous integrated learning models.
[0024] Figure 4 This is the risk heat map provided by the present invention.
[0025] Figure 5 This is a comparison chart of decision curves for the method provided by this invention.
[0026] Figure 6 This is the SHAP waterfall diagram provided by the present invention. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention. The following embodiments are used to illustrate this invention but should not be used to limit the scope of this invention.
[0028] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0029] The following is combined Figures 1 to 6The present invention will be further described in detail below, including a method and system for assessing the risk of death from sepsis based on dual-index nonlinear synergy: In this embodiment, as Figure 1 As shown, a sepsis mortality risk assessment method based on dual-indicator nonlinear synergy is provided, including the following steps: S1: Obtain the patient's stress hyperglycemia ratio (SHR), systemic immune inflammation index (SII), and routine clinical indicators.
[0030] Obtain SHR and SII within 24 hours of admission for patients with sepsis.
[0031] SHR calculation formula: Stress-induced hyperglycemia ratio (SHR) = acute blood glucose value (mmol / L) / estimated mean blood glucose (eAG), where eAG is converted from HbA1c.
[0032] SII calculation formula: SII = (platelet count × neutrophil count) / lymphocyte count.
[0033] Routine clinical indicators for patients with a history of illness include: age (years), SOFA score, lactate, Charlson index, GCS score, length of hospital stay (days), gender (male / female), AKI score, etc.
[0034] Missing value handling: Chain equation multiple imputation (MICE, 20 iterations) was used. Continuous variables were matched with predicted means, and categorical variables were matched with logistic regression.
[0035] This embodiment acquired the MIMIC-IV database (MIMIC-IV v2.2, sepsis patients, total sample size 1703, 90-day mortality rate 17.03%) and performed outlier and missing value processing. Processing details: Outlier processing: For patients with SHR>3 or SII>90th percentile, Winsorize was used to shrink the values to the 99th percentile. Multiple imputation was performed using the `mice` package in R. Finally, 1703 historical patients with no missing values were obtained, and a dataset was constructed. The dataset was randomly divided into a training set (1192 cases) and a test set (511 cases) in a 7:3 ratio. This dataset was used to train the model used in this method.
[0036] S2: Construct nonlinear derived features based on SHR and SII, and combine them with conventional clinical indicators to form the original features.
[0037] In this embodiment, for each data point in the dataset constructed in step 1, the corresponding original features are constructed in step S2, ultimately obtaining a complete original feature set for model training. Based on the original feature set, this method completes the training of the heterogeneous ensemble learning model, the division of AKI staging subgroups, and the determination of the fuzzification processing parameter table.
[0038] Z-score normalization is applied to SHR and SII (based on a portion of the original feature set corresponding to the training set). Then, based on the normalized SHR and SII, the following seven types of nonlinear derived features are constructed, with specific parameters as follows: (1) Polynomial basis functions (local cubic splines): Three quantiles are determined on the original feature set: the first quantile. Second quantile and the third quantile ; =0.25, =0.5, q3=0.75, using the type 6 quantile algorithm to ensure smooth interpolation. Truncation of the power basis: in, For the first and third cubic spline basis functions, This is a truncation function. As the independent variable, For the second and third cubic spline basis functions, These are the basis functions of the third and third cubic splines.
[0039] Applied to respectively and A total of 6 features were generated; among them, , Each corresponds to three features. For the standardized SHR, This is the standardized SII.
[0040] (2) Radial basis function: K-means clustering is used, with 100 iterations, and seeds 1, 2, and 3 are randomly initialized. Five center points are selected. ,bandwidth Take the average distance from the center to the two nearest centers.
[0041] Gaussian kernel : To each and The calculations generated a total of 10 features.
[0042] (3) Quantile piecewise linearity (piecewise linear effect): The standardized values are divided into four segments according to the quartiles, and an indicator variable is defined: in, This is the indicator function for the first four quantile intervals. This is the indicator function for the second and fourth quantile intervals. This is the indicator function for the third and fourth quantile intervals. This is the indicator function for the fourth quantile interval. The 25th quartile of the independent variable x. The 50th quartile of the independent variable x. The 75th quartile of the independent variable x. This is an indicator function; it takes the value 1 if the condition is true and 0 if the condition is false.
[0043] And construct piecewise linear terms , respectively and The calculations generated a total of 8 features.
[0044] (4) Logarithmic-exponential composite transformation (simulating the double exponential risk basis function): The calculation formula is as follows: in, The SHR features are obtained after composite transformation. These are the SII features after composite transformation. A total of 2 features are generated.
[0045] (5) Higher-order interaction terms: The calculation formula is as follows: in, It is a second-order interaction term. For first and third order interaction terms, For second and third order interaction terms, This is a fourth-order interaction term. A total of four features were generated.
[0046] (6) Nonlinear ratio: The calculation formula is as follows: in, The relative contribution ratio of SHR The relative contribution ratio of SII To find the absolute value, Avoid division by zero resulting in an error. A total of 2 features were generated.
[0047] Thus far, the first six methods have generated a total of 32 features. The seventh method, multi-scale sliding window statistics, will provide 20 features. Multi-scale sliding window statistics is an optional method that only takes effect in scenarios with repeated measurement data (such as continuously monitored blood routine and blood glucose), and is used to capture the time series change patterns of SHR and SII, further enhancing the model's adaptability in dynamic assessment scenarios.
[0048] (7) Multiscale sliding window statistics (this method can be selected for repeated measures): When patients have multiple (at least 3) repeated measurements of SHR and SII (e.g., 0h, 24h, 48h after admission), the following dynamic features (20 dimensions in total) are extracted from the sequence of each indicator: Coefficient of variation (2-dimensional): and The coefficient of variation.
[0049] Linear regression slope (2-dimensional): fitted using the least squares method and The slope changes over time; Exponentially Weighted Moving Average (EWMA, 4-dimensional): The decay factors are set to 0.2 and 0.5, respectively, and calculations are performed accordingly. and The final value of the smoothed sequence; Area under the curve (AUC, 2D): Calculated using the trapezoidal rule. and The area under the time series curve; Forward difference features (4-dimensional): Calculation and The absolute values of the differences between adjacent time points are taken as the maximum and median. Fluctuation clustering features (6 dimensions): and The sequences are divided into three modes: rising, falling, and stationary. The duration, number of transitions, and amplitude of each mode are statistically analyzed.
[0050] Based on the differentiated needs of the pathophysiological mechanism of sepsis, this embodiment uses seven feature construction methods to capture the specific nonlinear relationship between SHR and SII in different clinical scenarios, as shown in Table 1.
[0051] Table 1
[0052] Routine clinical indicators were preprocessed: age, SOFA score, lactate, Charlson index, GCS score, and length of hospital stay were Z-score standardized; gender was converted to 0 / 1 (male = 1, female = 0); AKI scores were converted to two dummy variables (AKI was divided into 3 stages, with the mildest stage 1 as the reference group, generating two dummy variables: moderate renal injury AKI_stage2 and severe renal injury AKI_stage3). Ultimately, the routine clinical indicators contributed 8 dimensions of features.
[0053] Finally, the preprocessed conventional clinical indicator features (8 dimensions), the derived features of SHR and SII (32 or 52 dimensions) are spliced together to obtain the original features.
[0054] S3: Input the original features into the heterogeneous ensemble learning model to calculate the initial mortality risk probability.
[0055] This embodiment selects multiple heterogeneous models as base learners, including random forest, extreme random tree, XGBoost, LightGBM and SVM.
[0056] The principles for selecting heterogeneous models include: (1) Covers multiple learning paradigms: RF (bagging), ExtraTrees (random subspace), XGBoost / LightGBM (gradient boosting), SVM (kernel method), ensuring the generalization ability after integration; (2) Complementary data adaptability: RF and ExtraTrees have a natural resistance to overfitting of high-dimensional nonlinear features (derived features of nonlinearity); (3) Adapting to class imbalance scenarios: XGBoost and LightGBM have built-in weighted loss functions for imbalanced data, which are suitable for class imbalance scenarios with low mortality rates; (4) Model complementarity: SVM combined with RBF kernel can capture non-monotonic decision boundaries, making up for the shortcomings of tree models on smooth boundaries; (5) Significant differences exist between models: The correlation matrix of the predicted values of the above 5 heterogeneous models was calculated by 5-fold cross-validation. The average correlation coefficient was 0.32 (less than the threshold of 0.5), indicating that there are significant differences between models. Integration can effectively reduce the error.
[0057] The reasons for selecting the above 5 heterogeneous models in this embodiment are summarized in Table 2.
[0058] Table 2
[0059] The heterogeneous ensemble learning model adopts a three-layer cascaded forest structure, with the parameters of each layer as follows: First layer: Parallel training of 5 models: (1) Random forest (RF): 500 trees, each tree has unlimited maximum depth, 8 split features, 5 minimum sample size per node, and 1 random seed.
[0060] (2) Extreme random trees: 300 trees, 8 splitting features, and a minimum sample size of 5 per node.
[0061] (3) XGBoost: Tree depth 4, learning rate 0.05, subsampling 0.8, column sampling 0.8, number of iterations 200.
[0062] (4) LightGBM: 64 leaves, 0.05 learning rate, 0.8 feature sampling, 0.8 data sampling, 200 iterations.
[0063] (5) SVM (RBF): Penalty parameter 1.0, kernel parameter 0.1, probability output is True.
[0064] Second layer: Feature re-extraction: The death probabilities output by the five models in the first layer are concatenated with the original feature X to form enhanced features. (45-dimensional or 65-dimensional). The mortality probability is a continuous value, ranging from [0,1], and is limited by the dataset; this is the mortality probability over 90 days.
[0065] The third layer: Meta-learner: Uses logistic regression + ridge regression regularization (L2 penalty coefficient of 0.01), with a binomial cross-entropy loss function, employing the L-BFGS optimizer, and a maximum of 1000 iterations. The meta-learner outputs the initial probability of death. : in, This is the transpose of the feature weight vector of the meta-learner. This is the bias term for the meta-learner.
[0066] The heterogeneous ensemble learning model employs 5-fold cross-validation (maintaining class proportions during partitioning): the training set is divided into 5 equal parts, each part serving as the validation set in turn, and the remaining 4 parts are used to train the base learners. Out-of-bag prediction probabilities are generated for the validation set, and finally, all out-of-bag predictions are used to train the meta-learner. This process avoids overfitting.
[0067] S4: Based on routine clinical indicators, perform hierarchical Bayesian probability calibration on the initial mortality risk probability to obtain the calibrated mortality risk probability.
[0068] The subgrouping method in this embodiment is as follows: grouping is performed based on AKI staging, resulting in 4 groups. The specific subgrouping criteria are as follows: (1) Healthy group: No AKI.
[0069] (2) AKI stage 1 group: serum creatinine increased to 1.5-1.9 times the baseline, or increased by ≥0.3mg / dL, or urine output <0.5mL / kg / h for 6-12h; (3) AKI stage 2 group: serum creatinine increased to 2.0-2.9 times the baseline, or urine output <0.5mL / kg / h lasted for ≥12h; (4) AKI stage 3 group: serum creatinine increases by ≥3 times the baseline, or ≥4.0 mg / dL, or renal replacement therapy is started, or urine output <0.3 mL / kg / h lasts for ≥24 hours.
[0070] Model formula: For each subgroup g, let ,but ,in, for The corresponding log odds, This represents the calibrated probability of death. For the regression intercept term of subgroup g, is the regression slope coefficient of subgroup g.
[0071] Prior distribution: , This reflects the expected slope of the reserve requirement ratio cut to be close to 1, among which, It follows a normal distribution.
[0072] HMC sampling parameters: Using the Stan package in R, 4 chains are set up, each chain undergoes 2000 iterations (the first 1000 iterations are warm-up), the step size is adaptive (target acceptance rate is 0.8), and the maximum tree depth is 10. The posterior median is used as... , Point estimate.
[0073] For new patients, the AKI subgroup g is determined based on the AKI score from step S1, and then the calculation is performed. , , then calculate , Simultaneously, the 2.5% and 97.5% quantiles of the posterior are used as 95% confidence intervals. Finally, the calibrated probability of death is output, along with a 95% confidence interval.
[0074] S5: Based on the original features, heterogeneous ensemble learning model, and calibrated mortality risk probability, risk levels, waterfall plots, and clinical recommendations are generated through fuzzification and inference.
[0075] S5.1: Based on the original features and the heterogeneous ensemble learning model, calculate the total SHAP contribution of each feature to the prediction result.
[0076] S5.11: Based on the original features and the heterogeneous ensemble learning model, calculate the SHAP estimate of each feature of each base learner for the prediction result.
[0077] For the m-th base learner (m=1, 2, 3, 4, 5), the FastSHAP method (based on Monte Carlo sampling) is used to approximate the SHAP value of the prediction result for feature j. .
[0078] The approximate formula for the SHAP value is (the expected form of the Shapley value): in, Let be the SHAP value of feature j of the base learner m on the prediction result. For mathematical expectation, The background sample set is constructed by randomly selecting 200 samples from the training set. for The samples in Let m be the m-th base learner, and let its output be the probability of death. , Features The sample to be explained is replaced with the value of the background sample, while the other features retain their original values. This is a background sample.
[0079] In practice, a sampling approximation is used: in, Let be the SHAP estimate of the prediction result for feature j of base learner m. For the number of Monte Carlo simulations, =50, Features The sample to be explained is replaced with the value of the k-th background sample, while the remaining features retain their original values. From The kth background sample randomly selected from the sample.
[0080] The SHAP value matrix of each base learner was calculated using shapviz+fastshap (50 Monte Carlo simulations). For each base learner... This yields the SHAP matrix, whose dimensions are the same as those of the original features.
[0081] S5.12: Calculate the weighted average SHAP value of each feature on the prediction result based on the SHAP estimate and the AUC weights of the base learner.
[0082] For each feature Calculate the weighted average SHAP value: in, Let be the weighted average SHAP value of feature j for the prediction results, representing the feature j. The marginal contribution to the risk of patient death (measured by log odds), with positive values indicating increased risk and negative values indicating decreased risk. as a base learner The AUC weights.
[0083] Calculation of AUC weights for base learners: On the validation set (out-of-bag predictions after 5-fold cross-validation within the training set), calculate the AUC of each base learner, and then calculate the normalized weights: in, The AUC of the base learner m Let be the AUC of base learner i, where i is the index of the base learner.
[0084] For example, in this embodiment, the AUC of the 5 base learners is The calculated AUC weights are: .
[0085] S5.13: The weighted average SHAP value aggregation yields the total SHAP contribution of each feature to the prediction result.
[0086] Since multiple derived features were constructed for SHR and SII in step S2, the contributions of these derived features need to be aggregated under the original variable names in the final waterfall plot for clinical interpretation.
[0087] SHR's total contribution to SHAP for: in, This is a set of derived features of SHR, with 16 or 26 elements.
[0088] Similarly, for all SII-related derived features Summing yields the total SHAP contribution of SII. .
[0089] Each routine clinical indicator corresponds to only one This can be directly used as the total SHAP contribution of this routine clinical indicator without aggregation.
[0090] Ultimately, 10 SHAP total contributions were obtained, corresponding to SHR, SII, age, SOFA score, lactate, Charlson index, GCS score, length of hospital stay, gender, and AKI score.
[0091] S5.2: Generate a waterfall plot based on the total SHAP contribution of each feature to the prediction results.
[0092] Base value: The log odds of the average probability of death on the training set, i.e. in, As the baseline value, This represents the average mortality probability of all samples in the training set.
[0093] Predicting the log odds: in, It is a logarithmic probability function. This represents the total SHAP contribution of feature j.
[0094] The waterfall plot is represented by a bar chart. Starting from the base, the contribution of each core variable is accumulated sequentially to obtain the predicted log odds, which are then converted into the probability of death risk through a sigmoid transformation.
[0095] S5.3: Input variable fuzzification: Fuzzify the calibrated mortality risk probability to obtain the corresponding risk level; fuzzify the total SHAP contribution of SHR and SII respectively to obtain the corresponding contribution level.
[0096] (1) Fuzzification of the calibrated probability of death: The mapping is applied to three fuzzy sets: low, medium, and high. A trapezoidal membership function is used, with parameters shown in Table 3.
[0097] Table 3 In Table 3, P is the independent variable of the input membership function, that is, the calibrated mortality risk probability to be fuzzified. For the membership function of low-risk fuzzy sets, To obtain the maximum value, To obtain the minimum value.
[0098] Based on Table 3, determine the risk level corresponding to the calibrated mortality risk probability.
[0099] (2) Fuzzification of the total SHAP contribution of SHR and SII: for and Normalization was performed separately. Normalized contribution The calculation formula is: in, The 95th percentile of the total absolute value of SHR contribution from all historical patients in the training set is 1.2 in this embodiment. Similarly, normalization yields... Normalized contribution .
[0100] Three fuzzy sets are defined: weak, medium, and strong, as shown in Table 4.
[0101] Table 4
[0102] According to Table 4, using and The contribution levels of SHR and SII were determined respectively.
[0103] S5.4: Based on a pre-defined fuzzy rule base, clinical recommendations are obtained according to risk level and contribution level.
[0104] The fuzzy rule base (Mamdani type) defines rules, each rule in the form of: IF (minimum value of fuzzy membership degree of the precondition), THEN (conclusion action). Specific rules and conclusions are shown in Table 5.
[0105] Table 5 Note: "—" indicates that this condition is not relevant (any membership degree is acceptable); Clinical recommendations may be adjusted as needed.
[0106] The process of reasoning and defuzzification: (1) Rule trigger strength: for a given input ( ), calculate the minimum membership degree of the preconditions for each rule, and use it as the activation strength of that rule.
[0107] (2) Conflict resolution: The rule with the highest activation intensity is taken as the final output. If multiple rules have the same activation intensity, they are selected according to priority (high risk level > medium risk level > low risk level).
[0108] (3) Output format: Pack the suggested text and confidence scores into a JSON structure.
[0109] S6: Output a mortality risk assessment report, which includes the calibrated mortality risk probability and its 95% confidence interval, risk level, waterfall plot, and clinical recommendations.
[0110] This embodiment is evaluated based on the test set from step S1. The ROC comparison chart is shown below. Figure 2As shown, the AUC of the baseline model (Logistic regression, features: age + gender + SOFA) on the test set is 0.735; the AUC of our method on the test set is 0.856.
[0111] The calibration curves of the heterogeneous ensemble learning model in this embodiment before and after hierarchical Bayes calibration are as follows: Figure 3 After hierarchical Bayesian calibration, the calibration slope increased from 0.577 to 1.108, indicating that the heterogeneous ensemble learning model significantly improved stability and accuracy.
[0112] like Figure 4 As shown, the risk heatmap of this method is plotted with standardized SII on the horizontal axis and standardized SHR on the vertical axis. Other features are fixed at the bit depth of the training set to predict the mortality risk at each point on the grid. Actual sample points from the test set are overlaid (green for survival, red for mortality).
[0113] like Figure 5 As shown in the diagram, the decision curves of this method are compared, and the net benefits of the four curves (this method, the baseline model, full intervention, and no intervention) are calculated. The 95% confidence band is estimated using the bootstrap method (200 times). Within the visible threshold range, the net benefit curve of this method consistently lies above the reference line of the baseline model and the full intervention, indicating significant clinical decision-making value.
[0114] like Figure 6 As shown, the SHAP waterfall plot demonstrates how each characteristic variable individually increases or decreases the current patient's mortality risk, starting from the group's average mortality risk, thus clarifying the risk drivers and achieving interpretability of the model's prediction process.
[0115] This embodiment also provides a sepsis mortality risk assessment system based on dual-index nonlinear synergy, the technical solution of which is as follows: including: The data acquisition module is used to acquire patients' SHR, SII, and routine clinical indicators; The feature generation module constructs nonlinear derived features based on SHR and SII, and combines them with conventional clinical indicators to form the original features. The mortality risk prediction module is used to input the original features into the heterogeneous ensemble learning model and calculate the initial mortality risk probability. The mortality risk calibration module is used to perform hierarchical Bayesian probability calibration on the initial mortality risk probability based on routine clinical indicators, so as to obtain the calibrated mortality risk probability. The mortality risk assessment module is used to generate risk levels and clinical recommendations based on raw features, heterogeneous ensemble learning models, and calibrated mortality risk probabilities through fuzzification and inference. The report generation module is used to output a mortality risk assessment report, which includes a calibrated mortality risk probability, risk level, and clinical recommendations.
[0116] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for assessing the risk of death from sepsis based on dual-indicator nonlinear synergy, characterized in that, include: S1: Obtain the patient's stress hyperglycemia ratio (SHR), systemic immune inflammation index (SII), and routine clinical indicators; S2: Construct nonlinear derived features based on SHR and SII, and combine them with conventional clinical indicators to form the original features; S3: Input the original features into the heterogeneous ensemble learning model to calculate the initial probability of death risk; S4: Based on routine clinical indicators, perform hierarchical Bayesian probability calibration on the initial mortality risk probability to obtain the calibrated mortality risk probability; S5: Based on the original features, heterogeneous ensemble learning model and calibrated mortality risk probability, risk levels and clinical recommendations are generated through fuzzification and inference. S6: Output a mortality risk assessment report, which includes the calibrated mortality risk probability, risk level, and clinical recommendations.
2. The sepsis mortality risk assessment method based on dual-index nonlinear synergy as described in claim 1, characterized in that, Routine clinical indicators include: age, SOFA score, lactate, Charlson index, GCS score, length of hospital stay, gender, and AKI score.
3. A sepsis mortality risk assessment method based on dual-index nonlinear synergy as described in claim 1 or 2, characterized in that, In step S2, based on SHR and SII, derived features are constructed using derived feature construction methods; wherein, the derived feature construction methods include: polynomial basis functions, radial basis functions, quantile piecewise linear transformation, logarithmic-exponential composite transformation, higher-order interaction terms, and nonlinear ratios; Preprocessing routine clinical indicators yields their characteristics; The original features are obtained by concatenating the features of conventional clinical indicators with the derived features.
4. The sepsis mortality risk assessment method based on dual-index nonlinear synergy as described in claim 3, characterized in that, When both SHR and SII have more than three measurement data, the derived feature construction method also includes: multi-scale sliding window statistics.
5. The sepsis mortality risk assessment method based on dual-index nonlinear synergy as described in claim 1, characterized in that, The heterogeneous ensemble learning model is a cascaded deep forest, whose first layer uses multiple heterogeneous models as base learners to calculate multiple death probabilities based on the original features. Its second layer concatenates the death probability with the original features to form enhanced features; Its third layer employs a meta-learner, using logistic regression and ridge regression regularization to calculate the initial mortality risk probability based on the enhanced features.
6. The sepsis mortality risk assessment method based on dual-index nonlinear synergy as described in claim 5, characterized in that, Heterogeneous models include random forest, extreme random tree, XGBoost, LightGBM, and SVM.
7. A sepsis mortality risk assessment method based on dual-indicator nonlinear synergy as described in claim 1 or 2, characterized in that, In step S4, the AKI subgroup is determined based on the AKI score in routine clinical indicators, and the regression intercept and regression slope coefficient corresponding to the subgroup are obtained. The initial mortality risk probability is calibrated to obtain the calibrated mortality risk probability. The formula for calculating the probability of death risk is: in, This represents the calibrated probability of death. For the regression intercept term of subgroup g, Let g be the regression slope coefficient for subgroup g. for The corresponding log odds, This represents the initial probability of death.
8. The sepsis mortality risk assessment method based on dual-index nonlinear synergy as described in claim 1, characterized in that, Step S5 includes: Based on the original features and the heterogeneous ensemble learning model, the total SHAP contribution of each feature to the prediction result is calculated. The calibrated mortality risk probability is fuzzified to obtain the corresponding risk level; the total SHAP contribution of SHR and SII is fuzzified to obtain the corresponding contribution level. Based on a pre-defined fuzzy rule base, clinical recommendations are obtained according to risk level and contribution level.
9. The sepsis mortality risk assessment method based on dual-index nonlinear synergy as described in claim 8, characterized in that, In step S5, a waterfall plot is generated based on the total SHAP contribution of each feature to the prediction result; In step S6, the mortality risk assessment report also includes a waterfall chart.
10. A sepsis mortality risk assessment system based on dual-index nonlinear synergy, characterized in that, A method for assessing sepsis mortality risk based on dual-indicator nonlinear synergy as described in any one of claims 1 to 9, comprising: The data acquisition module is used to obtain the patient's SHR, SII, and routine clinical indicators; The feature generation module is used to construct nonlinear derived features based on SHR and SII, and combine them with conventional clinical indicators to form the original features. The mortality risk prediction module is used to input the original features into the heterogeneous ensemble learning model and calculate the initial mortality risk probability. The mortality risk calibration module is used to perform hierarchical Bayesian probability calibration on the initial mortality risk probability based on routine clinical indicators, so as to obtain the calibrated mortality risk probability. The mortality risk assessment module is used to generate risk levels and clinical recommendations based on raw features, heterogeneous ensemble learning models, and calibrated mortality risk probabilities through fuzzification and inference. The report generation module is used to output a mortality risk assessment report, which includes a calibrated mortality risk probability, risk level, and clinical recommendations.