Sepsis risk assessment model construction method based on clinical detection intention and early warning system
Patent Information
- Application Number
- CN202611025936.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-10
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]有鉴于此,本发明的目的是提供一种基于临床检测意图的脓毒症风险评估模型构建方法及预警系统,以解决现有方法因确诊滞后、稀疏数据下判断困难以及未利用医生检测意图而导致的预警不及时、不准确的问题
[0029]1、发明创造性地构建“动态时间衰减意图矩阵”,将医生的干预行为转化为模型可学习的特征,充分利用了传统方法丢弃的高价值信息,使模型能感知临床关注度的变化,提升了在数据稀疏下的鲁棒性。并且对医师的医疗干预行为进行了连续的数学量化,突出了医疗干预的时效性。
Smart Images

Figure CN122822345A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of medicine and machine learning technology, and in particular to a method for constructing a sepsis risk assessment model. Background Technology
[0002] Sepsis is a disorder of organ function caused by a dysregulation of the body's response to infection, often accompanied by acute inflammation and carrying a high risk of death. The mortality rate for sepsis patients in hospitals is close to 26%, and in intensive care units it exceeds 40%. Furthermore, for every hour of delay in medical intervention, the mortality rate can increase by 7-10%. Therefore, early identification and warning of sepsis can buy valuable time for medical intervention.
[0003] Currently, clinicians primarily rely on continuous monitoring of the Sequential Organ Failure Assessment (SOFA) score and vital signs to diagnose sepsis. However, this method has significant limitations: First, it suffers from delayed diagnosis: diagnosis is usually only made during or after a sepsis attack, wasting precious time for intervention. Second, it is highly dependent on data completeness: it requires continuous monitoring of massive amounts of clinical data, and missing data can severely impact the accuracy of diagnosis. Third, it ignores the intent behind medical intervention: it fails to consider the clinical intent inherent in the physician's act of ordering laboratory tests. When a physician decides to frequently test a certain indicator, it often implies a heightened awareness of potential abnormalities in that indicator; this crucial warning signal is completely ignored in existing methods.
[0004] With the development of artificial intelligence technology, machine learning models have been applied to sepsis prediction. However, existing models typically only process recorded numerical data, simply treating missing data as "noise" to be processed. They fail to effectively encode and utilize the "intent" behind the aforementioned clinical testing behaviors, resulting in insufficient early warning capabilities in data-sparse scenarios and an inability to achieve earlier risk perception based on medical behavior patterns. Summary of the Invention
[0005] In view of this, the purpose of this invention is to provide a method for constructing a sepsis risk assessment model based on clinical testing intent and an early warning system, so as to solve the problems of untimely and inaccurate early warning caused by the delay in diagnosis, difficulty in judgment under sparse data, and failure to utilize the doctor's testing intent in existing methods.
[0006] The present invention provides a method for constructing a sepsis risk assessment model based on clinical testing intent, comprising:
[0007] Step S1: Collect time-series data of various clinical monitoring parameters of the target patient within a preset time interval and integrate them into a raw data matrix;
[0008] Step S2: Construction of Clinical Testing Intent Data: For each clinical data item m, construct a dynamic time decay intent feature at each acquisition time t. Define the decay value of the patient's detection intent for clinical data item m at time t. , This represents the time interval since the most recent clinical data item m was detected; This is the clinical attenuation constant;
[0009] Step S3: Perform preprocessing and standardization on the original data matrix in sequence; the preprocessing includes forward imputation of missing values and logarithmic transformation on clinical data items with skewed data distribution;
[0010] Step S4: Feature Fusion and Model Training: The preprocessed and standardized clinical data are concatenated with the dynamic time decay intent matrix along the feature dimension to form a dual-channel input feature. This dual-channel input feature is then input into a machine learning model for training to obtain a sepsis risk assessment model. The model is trained with the patient's risk of developing sepsis as the prediction target, and a penalty weight coefficient w for misclassification of sepsis-positive samples is introduced into its loss function, where w is greater than 1. The loss function is:
[0011]
[0012] In the formula, L represents the total loss value during model training, and N represents the total number of training samples. This represents the true label of the j-th sample. =1 represents a patient with sepsis. =0 represents a normal patient. This represents the probability of sepsis positivity predicted by the model for the j-th sample. This represents the penalty weighting coefficient.
[0013] Furthermore, the clinical monitoring parameters include vital signs parameters, laboratory test parameters, and demographic parameters.
[0014] Furthermore, the vital signs parameters include at least one of heart rate, blood oxygen saturation, body temperature, mean arterial pressure, and respiratory rate; the laboratory test parameters include at least one of blood urea nitrogen, serum creatinine, glucose, white blood cell count, and platelet count.
[0015] Furthermore, in step S3, the forward imputation of missing values specifically involves: using the most recent valid observation of the same patient to fill subsequent missing values; for missing values for which no previous value is available at the initial time of the patient, using the statistical median of the historical patient group on that data item to fill them.
[0016] Furthermore, in step S3, the clinical data items with skewed data distribution include at least one of mean arterial pressure, blood urea nitrogen, serum creatinine, glucose, white blood cell count, and platelet count.
[0017] Furthermore, in step S3, the standardization process includes performing Z-score standardization on each data item m in the preprocessed clinical data matrix.
[0018] Furthermore, in step S4, the machine learning model is a gradient boosting decision tree model.
[0019] This invention also discloses a method for sepsis risk early warning, comprising:
[0020] Sepsis risk assessment models were constructed using methods such as sepsis risk assessment model construction based on clinical testing intent.
[0021] Obtain current time-series clinical monitoring data for the patient to be evaluated;
[0022] Based on the sepsis risk assessment model, the current data of the patient to be assessed is processed to obtain dual-channel input features, and then the dual-channel data is input into the sepsis risk assessment model to obtain the sepsis risk assessment probability value at the current moment.
[0023] The present invention also discloses a sepsis risk warning system, which includes a computing server for executing the sepsis risk warning method, and at least one warning hardware terminal;
[0024] The early warning hardware terminal includes:
[0025] The data receiving module is used to receive the sepsis risk assessment probability value from the computing server;
[0026] The central control module, connected to the data receiving module, is used to compare the received sepsis risk assessment probability value with a preset risk threshold and generate a corresponding early warning instruction.
[0027] The early warning output module, connected to the central control module, is used to execute the early warning command. The early warning output module includes: a display unit for dynamically displaying the numerical value of the sepsis risk assessment probability value; and an audible and visual alarm unit for triggering different modes of light and sound alarms according to the early warning command.
[0028] The beneficial effects of this invention are:
[0029] 1. A novel "dynamic time-decay intention matrix" was invented and constructed, transforming doctors' intervention behaviors into learnable features for the model. This fully utilizes high-value information discarded by traditional methods, enabling the model to perceive changes in clinical attention and improving robustness under data sparsity. Furthermore, the physicians' medical intervention behaviors were continuously mathematically quantified, highlighting the timeliness of medical interventions.
[0030] 2. By fusing dual-channel features of physiological and intentional data and employing a training strategy that strongly penalizes underreporting of sepsis, the present invention enables the constructed model to achieve high levels of discrimination (AUC), sensitivity, and specificity on public datasets, which are significantly better than traditional single-channel models.
[0031] 3. The sepsis risk early warning system, which combines hardware and software, realizes risk probability assessment and output based on the sepsis risk assessment model. It can push the risk probability calculated by the model to medical staff in real time in the form of sound and light, which greatly shortens the time from the appearance of risk to clinical perception and buys valuable time for early intervention. Attached Figure Description
[0032] Figure 1 This is a flowchart of the method of the present invention;
[0033] Figure 2 This invention provides early warning results for uninfected sepsis patients on the PhysioNet / Computing in Cardiology Challenge 2019 dataset.
[0034] Figure 3 This invention provides early warning results for sepsis-infected individuals on the PhysioNet / Computing in Cardiology Challenge 2019 dataset.
[0035] Figure 4 The ROC curve of the method of this invention compared with other machine learning methods on the PhysioNet / Computing in Cardiology Challenge 2019 dataset is shown.
[0036] Figure 5 This is a confusion matrix diagram of the method of the present invention on the PhysioNet / Computing in Cardiology Challenge 2019 dataset. Detailed Implementation
[0037] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0038] Example 1: The method for constructing a sepsis risk assessment model based on clinical testing intent in this example includes:
[0039] Step S1: Data from 12,267 patients on the PhysioNet / Computing in Cardiology Challenge 2019 dataset were selected for testing. Time-series data of various clinical monitoring parameters of the target patients were collected within a preset time interval and integrated into a raw data matrix. The clinical monitoring parameters include vital signs, laboratory test parameters, and demographic parameters. The specific clinical monitoring parameters are shown in Table 1, including: age, sex, and ICU stay time (demographics); heart rate (HR), oxygen saturation (O2Sat), body temperature (Temp), mean arterial pressure (MAP), and respiratory rate (Resp) (vital signs); blood urea nitrogen (BUN), serum creatinine, glucose, chloride, hematocrit (Hct), hemoglobin (Hgb), white blood cell count (WBC), and platelet count (laboratory tests), totaling 18 data points, collected at a time interval of 1 hour.
[0040] Table 1. Collected Clinical Data
[0041] 1 Demographic data Age: Age 2 Demographic data Gender: gender 3 Demographic data HospAdmTime: The time interval (in hours) between admission and transfer to the ICU. 4 Demographic data ICULOS: ICU stay 5 Demographic data Unit: ICU Ward Type 6 Vital signs data HR: Heart Rate (beats per minute) 7 Vital signs data O2Sat: Pulse Oximetry, pulse oxygen saturation (%) 8 Vital signs data Temp: Temperature, body temperature (degrees Celsius) 9 Vital signs data MAP: Mean Arterial Pressure (mmHg) 10 Vital signs data Resp: Respiration Rate, respiratory rate (breaths per minute). 11 Laboratory test data BUN: Blood Urea Nitrogen (mg / dL) 12 Laboratory test data Creatinine: Serum creatinine (mg / dL) 13 Laboratory test data Glucose: Serum glucose (mg / dL) 14 Laboratory test data Chloride: Chloride ion concentration (mmol / L) 15 Laboratory test data Hct: Hematocrit, the percentage of red blood cells. 16 Laboratory test data Hgb: Hemoglobin (g / dL) 17 Laboratory test data WBC: Leukocyte count. 18 Laboratory test data Platelets: Platelet count
[0042] Step S2: Construction of Clinical Testing Intent Data: For each clinical data item m, construct a dynamic time decay intent feature at each acquisition time t. .
[0043] (1)
[0044] This represents the time interval since the most recent clinical data item m was detected; The clinical attenuation constant is taken as... =0.1. When the doctor first ordered the test... =0, =1, at which point the model's alertness is at its highest; as time progresses, if this indicator is not detected again, The value decays exponentially towards 0, indicating that alertness gradually diminishes.
[0045] Step S3: Perform preprocessing and standardization on the original data matrix in sequence; the preprocessing includes forward imputation of missing values and logarithmic transformation on clinical data items with skewed data distribution.
[0046] The forward imputation of missing values specifically involves: using the most recent valid observation of the same patient to fill subsequent missing values; for missing values for which no previous value is available at the initial time of the patient, using the statistical median of the historical patient group for that data item to fill the missing value, in order to ensure data integrity.
[0047] Logarithmic transformation was performed on clinical data items with skewed distributions, including logarithmic transformation of five key renal function indicators: MAP, BUN, Creatinine, Glucose, WBC, and Platelets (as shown in the formula). As shown in the figure, this is to reduce the impact of extreme values on the model.
[0048] (2)
[0049] In the formula, This represents the value of clinical data m at time t before logarithmic transformation. This represents the value of clinical data m at time t after logarithmic transformation.
[0050] The standardization process includes: performing Z-score standardization on each data point m in the preprocessed clinical data matrix, as shown in the formula... As shown:
[0051] (3)
[0052] In the formula, This represents the standardized value of clinical data m. The value of m in the clinical data before standardization. Let m be the mean of the clinical data across all training data. Let m be the standard deviation of the clinical data across all training data.
[0053] Step S4: Feature Fusion and Model Training: The preprocessed and standardized clinical data are concatenated with the dynamic time-decay intent matrix along the feature dimensions to form a dual-channel input feature. Assuming there are 18 clinical parameters, the feature vector at each time point t becomes 36-dimensional after concatenation: the first 18 dimensions are the standardized physiological parameter values, and the last 18 dimensions are the corresponding binary detection intent data. These 36 dimensions together constitute the "dual-channel input feature," where the first 18 dimensions are the "physiological data channel," and the last 18 dimensions are the "clinical detection intent channel."
[0054] The dual-channel input features are fed into a machine learning model for training to obtain a sepsis risk assessment model. The model is trained with the prediction target of the patient's risk of developing sepsis, and a penalty weight coefficient w, greater than 1, is introduced into its loss function to penalize misclassifications of sepsis-positive samples. The loss function is:
[0055] (4)
[0056] In the formula, L represents the total loss value during model training, and N represents the total number of training samples. This represents the true label of the j-th sample. =1 represents a patient with sepsis. =0 represents a normal patient. This represents the probability of sepsis positivity predicted by the model for the j-th sample. This represents the penalty weighting coefficient.
[0057] In this embodiment, the machine learning model is a gradient boosting decision tree model (XGBoost model). During the training process, the model automatically mines the nonlinear interaction relationships between different clinical variables and their evolutionary characteristics over time by constructing a multidimensional decision tree group.
[0058] Example 3: A sepsis risk early warning method, comprising:
[0059] A sepsis risk assessment model constructed using the method described in Example 1;
[0060] Obtain current time-series clinical monitoring data for the patient to be evaluated;
[0061] Based on the sepsis risk assessment model, the current data of the patient to be assessed is processed to obtain dual-channel input features, and then the dual-channel data is input into the sepsis risk assessment model to obtain the sepsis risk assessment probability value at the current moment. .
[0062] (5)
[0063] In the formula, e is the natural logarithm, and Z t The total score output at time t is the sum of the weights of all decision tree leaf nodes in XGBoost.
[0064] Example 2: The sepsis risk warning system of this example includes a computing server for executing the sepsis risk warning method described in Example 2, and at least one warning hardware terminal;
[0065] The early warning hardware terminal includes:
[0066] The data receiving module is used to receive the sepsis risk assessment probability value from the computing server;
[0067] The central control module, connected to the data receiving module, is used to compare the received sepsis risk assessment probability value with a preset risk threshold and generate a corresponding early warning instruction.
[0068] The early warning output module, connected to the central control module, is used to execute the early warning command. The early warning output module includes: a display unit for dynamically displaying the numerical value of the sepsis risk assessment probability value; and an audible and visual alarm unit for triggering different modes of light and sound alarms according to the early warning command.
[0069] As attached Figure 2 As shown, if the patient If the preset threshold of 50% is not exceeded at any given time, no warning will be triggered. (See attached image) Figure 3 As shown, if the patient If the risk exceeds a preset threshold of 50% at a certain time, the system will trigger a high-risk alarm, indicating to clinicians that the patient is at extremely high risk of sepsis and recommending that blood culture or antibiotic intervention be performed in advance.
[0070] Appendix Figure 2 and Figure 3 Experimental results show that, compared with other machine learning techniques, the sepsis risk assessment model proposed in this invention not only significantly outperforms other models in global discrimination (AUC of 0.975), but also achieves a high sensitivity of 86.4% and a high specificity of 95.7% on the PhysioNet / Computing in Cardiology Challenge 2019 dataset. This demonstrates that the sepsis risk assessment model can keenly capture the features of dual-channel input data and transform the physician's clinical testing intentions into powerful predictive capabilities.
[0071] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for constructing a sepsis risk assessment model based on clinical testing intent, characterized in that: include: Step S1: Collect time-series data of various clinical monitoring parameters of the target patient within a preset time interval and integrate them into a raw data matrix; Step S2: Construction of Clinical Testing Intent Data: For each clinical data item m, construct a dynamic time decay intent feature at each acquisition time t. Define the decay value of the patient's detection intent for clinical data item m at time t. , This represents the time interval since the most recent clinical data item m was detected; This is the clinical attenuation constant; Step S3: Perform preprocessing and standardization on the original data matrix in sequence; the preprocessing includes forward imputation of missing values and logarithmic transformation on clinical data items with skewed data distribution; Step S4: Feature fusion and model training: The preprocessed and standardized clinical data are concatenated with the dynamic time decay intention matrix in the feature dimension to form a dual-channel input feature. The dual-channel input features are fed into a machine learning model for training to obtain a sepsis risk assessment model. The model is trained with the prediction target of the patient's risk of developing sepsis, and a penalty weight coefficient w, greater than 1, is introduced into its loss function to penalize misclassifications of sepsis-positive samples. The loss function is: In the formula, L represents the total loss value during model training, and N represents the total number of training samples. This represents the true label of the j-th sample. =1 represents a patient with sepsis. =0 represents a normal patient. This represents the probability of sepsis positivity predicted by the model for the j-th sample. This represents the penalty weighting coefficient.
2. The method for constructing a sepsis risk assessment model based on clinical testing intent according to claim 1, characterized in that: The clinical monitoring parameters include vital signs, laboratory test parameters, and demographic parameters.
3. The method for constructing a sepsis risk assessment model based on clinical testing intent according to claim 1, characterized in that: The vital signs parameters include at least one of heart rate, blood oxygen saturation, body temperature, mean arterial pressure, and respiratory rate; the laboratory test parameters include at least one of blood urea nitrogen, serum creatinine, glucose, white blood cell count, and platelet count.
4. The method for constructing a sepsis risk assessment model based on clinical testing intent according to claim 1, characterized in that: In step S3, the forward imputation of missing values specifically involves: using the most recent valid observation of the same patient to fill subsequent missing values; for missing values for which no previous value is available at the initial time of the patient, using the statistical median of the historical patient group on that data item to fill them.
5. The method for constructing a sepsis risk assessment model based on clinical testing intent according to claim 1, characterized in that: In step S3, the clinical data items with skewed data distribution include at least one of mean arterial pressure, blood urea nitrogen, serum creatinine, glucose, white blood cell count, and platelet count.
6. The method for constructing a sepsis risk assessment model based on clinical testing intent according to claim 1, characterized in that: In step S3, the standardization process includes performing Z-score standardization on each data item m in the preprocessed clinical data matrix.
7. The method for constructing a sepsis risk assessment model based on clinical testing intent according to claim 1, characterized in that: In step S4, the machine learning model is a gradient boosting decision tree model.
8. A method for early warning of sepsis risk, characterized in that, include: A sepsis risk assessment model constructed using the method described in any one of claims 1-7; Obtain current time-series clinical monitoring data for the patient to be evaluated; Based on the sepsis risk assessment model, the current data of the patient to be assessed is processed to obtain dual-channel input features, and then the dual-channel data is input into the sepsis risk assessment model to obtain the sepsis risk assessment probability value at the current moment.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, it implements the method as described in any one of claims 1-7, or implements the method as described in claim 8.
10. A sepsis risk early warning system, characterized in that: It includes a computing server for performing the method as described in claim 8, and at least one early warning hardware terminal; The early warning hardware terminal includes: The data receiving module is used to receive the sepsis risk assessment probability value from the computing server; The central control module, connected to the data receiving module, is used to compare the received sepsis risk assessment probability value with a preset risk threshold and generate a corresponding early warning instruction. The early warning output module, connected to the central control module, is used to execute the early warning command. The early warning output module includes: a display unit for dynamically displaying the numerical value of the sepsis risk assessment probability value; and an audible and visual alarm unit for triggering different modes of light and sound alarms according to the early warning command.