Ssepsis diagnosis and treatment scheme recommendation method based on reinforcement learning and clinical knowledge
Patent Information
- Application Number
- CN202510788253.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-11-07
AI Technical Summary
Existing reinforcement learning-based sepsis treatment protocols lack comprehensive personalized recommendations for antibiotic combinations and durations, and the models do not adequately represent patient characteristics, making it difficult to meet clinical safety and accuracy requirements.
The problem of antibiotic combination use in sepsis treatment is modeled as a Markov decision process. A DQN model is constructed using this model, and combined with deep reinforcement learning and clinical knowledge, SOFA scores and clinical guidelines are incorporated into the reward function to construct the SAI-DQN model. This model guides the model towards rationally shortening the duration of antibiotic use and ensuring good patient prognosis.
It enhances the medical interpretability and clinical consistency of the model, provides personalized recommendations for antibiotic combination therapy, shortens the duration of antibiotic use, and improves the accuracy and safety of treatment outcomes.
Smart Images

Figure CN120913872A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of artificial intelligence and medical technology, and particularly to a sepsis diagnosis and treatment scheme recommendation method based on reinforcement learning and clinical knowledge. BACKGROUND
[0002] Sepsis is one of the main causes of death worldwide, and treating sepsis patients is extremely challenging. In the treatment scheme for sepsis, clinicians often need to combine patient status, medical test data, patient history data, and integrate various treatment methods, including antibiotic use, fluid therapy, vasopressor drugs, mechanical ventilation, and complication treatment, to combat symptoms such as infection. The way, duration, and dosage of treatment methods will have different effects on patient prognosis. Since most sepsis is caused by bacterial or fungal infection, the most important treatment method is antibiotic therapy, so choosing the best treatment drug, treatment duration, and dosage will achieve better treatment results.
[0003] In recent years, data mining and machine learning analysis have been increasingly widely used in the medical industry. By analyzing relevant medical data using data mining methods, the cost of comparing various intervention measures can be reduced, and information support can be provided for clinical decision-making. Medical clinical decision-making problems are problems that doctors need to make diagnosis and treatment decisions based on patient symptoms, medical history, physical signs, and other information during the diagnosis and treatment process. The characteristics of such problems are high complexity, diverse data sources, high uncertainty, and the need for rapid, personalized, and accurate judgment and processing, and the decision-making results are directly related to the patient's life and health, which is a very complex and high-risk task. With the improvement of computing power and the availability of high-frequency medical data, new computing methods have been used in the medical decision-making process. Artificial intelligence based on machine learning can capture high-complexity pattern data in medicine and effectively help medical decision-makers develop more accurate and personalized treatment plans, improving diagnosis and treatment levels and efficiency.
[0004] Reinforcement learning is a method of machine learning that focuses on how an agent learns from trial and error in the process of interacting with the environment to maximize system performance compared to other methods such as supervised learning and unsupervised learning. Reinforcement learning is a method framework for learning, prediction, and decision-making that can solve sequential decision-making problems, consider long-term return problems, and help optimize strategies, so it is preferred for modeling recommendation systems for treatment plans. Reinforcement learning itself relies on the interaction of an agent with the environment and rewards, and the strategy, such as the best treatment plan, is obtained by the agent interacting with the unknown environment to maximize the cumulative reward. In medical decision-making problems, the agent corresponds to the doctor, the environment corresponds to the patient's physical condition and medical means, and the reward corresponds to the treatment effect. Reinforcement learning data has the characteristics of time series and uncertainty, and machine learning models need to be able to handle these data and pattern characteristics. Reinforcement learning is widely used in the medical field and can help doctors develop personalized and flexible treatment plans, and its autonomous exploration and trial-and-error strategy can help doctors quickly adjust treatment plans and improve decision-making ability and efficiency.
[0005] In recent studies, computational methods using reinforcement learning have been used to evaluate the administration of vasopressor drugs and the recommendation of mechanical ventilation plans for sepsis patients, and the results of related studies also show that reinforcement learning performs well in nonlinear complex dynamic problems, long-term effect sequence decision-making, and personalized recommendation strategies in sequence decision-making medical tasks. However, there are fewer studies on the use of antibiotic treatment combinations and the use of time length and the simulation of intermediate states of the patient treatment process.
[0006] Although some progress has been made in this field, reinforcement learning in medical decision-making problems faces problems such as information data complexity and clinical safety. Information data complexity mainly manifests in the processing of patient information data. Medical data often comes from multiple sources and has various structures and properties, and the amount of data is large, which puts higher requirements on data cleaning, standardization, and processing. Clinical safety requires the full exploitation of information in medical data and the combination of medical professional knowledge, as well as appropriate model construction and data processing methods to ensure comprehensive consideration and accurate measurement of medical problems. In terms of continuous action space, simulation of patient state output, and comprehensive research on antibiotics and other physical treatment methods, existing dynamic diagnosis and treatment plan optimization algorithms based on reinforcement learning mainly focus on the implementation time and dose range selection of treatment plans, and the representation of patient feature information and antibiotic combination plans is not comprehensive enough, and most models cannot observe patient feature changes. In view of this, the transparency of patient feature information and the use of antibiotic treatment combinations in treatment plan recommendation models are still challenging tasks. SUMMARY
[0007] The purpose of the present application is to overcome the shortcomings of the prior art, and propose a sepsis diagnosis and treatment scheme recommendation method based on reinforcement learning and clinical knowledge. First, the use of antibiotic combinations in the sepsis treatment process is modeled as a Markov decision process, and a DQN model is constructed. Second, the deep reinforcement learning based on the value function is combined with sepsis clinical data and medical guideline related content. The reward function refers to the clinical real data, integrates the SOFA score knowledge and the clinical guideline content, and guides the model to make strategy recommendations in the direction of reasonably shortening the antibiotic medication duration and ensuring good prognosis of the patient.
[0008] The present application solves its technical problems by adopting the following technical solutions:
[0009] The sepsis diagnosis and treatment scheme recommendation method based on reinforcement learning and clinical knowledge comprises the following steps:
[0010] Step 1, collect data and preprocess the data;
[0011] Step 2, construct a deep Q network for sepsis anti-infection and train it;
[0012] Step 3, input the preprocessed data into the constructed deep Q network for sepsis anti-infection to obtain a sepsis diagnosis and treatment recommendation scheme.
[0013] Moreover, the step 1 comprises the following steps:
[0014] Step 1.1, collect electronic medical records from the MIMIC-IV dataset;
[0015] Step 1.2, determine whether the electronic medical records in the MIMIC-IV dataset meet the sepsis-3 standard. If yes, proceed to step 1.3, otherwise exclude the electronic medical records;
[0016] Step 1.3, determine whether the information in the electronic medical records is greater than or equal to 18 years old. If yes, proceed to step 1.4, otherwise exclude the electronic medical records;
[0017] Step 1.4, determine whether the information in the electronic medical records is transferred to a hospital. If yes, exclude the electronic medical records, otherwise proceed to step 1.5;
[0018] Step 1.5, select feature items in the electronic medical records according to the missing rate;
[0019] Step 1.6, determine whether the missing rate is less than or equal to 60%. If yes, it constitutes a usable dataset, otherwise exclude the electronic medical records;
[0020] Step 1.7, use a mixed imputation strategy of K-nearest neighbor method and forward filling to impute the missing values of the usable dataset in step 1.6 to obtain a complete dataset;
[0021] Step 1.8, collect the electronic medical records of the eRI data set, and repeat steps 1.2 to 1.7 to obtain the eRI verification set.
[0022] Moreover, the deep Q network for anti-infection of sepsis in step 2 includes an action space generation module, a state space generation module, an SAI-DQN module, and an SAI-DQN training module, wherein the action space generation module and the state space generation module are connected to the SAI-DQN module, and the SAI-DQN module is connected to the SAI-DQN training module.
[0023] Moreover, the SAI-DQN training module includes an agent observation environment module, an experience replay pool, a Q network, a DQN error function, a target Q network, and a reward function, wherein the patient feature vector and the antibiotic sequence in the preprocessed data are input into the agent observation environment module, the agent observation environment module is connected to the experience replay pool and the Q network, the experience replay pool is connected to the reward function and the target Q network, the Q network is connected to the DQN error function and the target Q network, and the target Q network is connected to the DQN error function.
[0024] The agent observation environment module constitutes an environment based on the correspondence between the patient feature vector and the antibiotic sequence. The patient feature vector in the agent observation environment calculates the Q value of each action through the Q value function, so as to select a suitable action (a). The antibiotic combination sequence selected by the agent interacts with the environment, obtains a reward (r) according to the reward function, and obtains a new state (s'), updates the neural network parameters using the DQN algorithm, improves the prediction performance of the Q value function by minimizing the loss of the Q value function, and the experience replay pool includes the memory of the state observation, the selected action, the reward, and the new state; the past state observation, the selected action, the reward, and the new state are used to train the network, to enhance the randomness, generalization, and convergence of the algorithm, and the above steps are repeated, through the interaction with the environment and the update of the network parameters, the DQN algorithm enables the agent to gradually learn to maximize the long-term reward in a given environment.
[0025] Moreover, the reward function includes a reward function R1, a reward function R2, a reward function R3, and a reward function R4, wherein the reward function R1 is motivated to enable the model to understand the prognosis of the patient, and the content is the survival rate of the patient within 90 days, and the reward function R1 is set to survive: +100, and die: -100.
[0026] The reward function R2 is motivated to encourage the model to learn to improve the prognosis of the patient, and the content is the change of the SOFA score of the patient, and the reward function R2 is set to state improvement reward +10, and state deterioration penalty -10.
[0027] The motivation of the reward function R3 is to encourage the model to shorten the antibiotic use time, the content is the medication time, and the setting of the reward function R3 is greater than 9 days, and the penalty is -10;
[0028] The motivation of the reward function R4 is to integrate clinical data, the content is to compare the intermediate state conversion with the data set statistics, and the setting of the reward function R4 is R=Q clinical (s, a, s') * 0.1;
[0029] Wherein, Q clinical (·) represents a clinical knowledge function; s represents a current state, including patient vital signs, laboratory indicators and microbiological data; a represents an evaluation action, that is, a selected antibiotic combination; s' represents a next state, that is, a predicted state after medication.
[0030] Moreover, the specific implementation method of the step 3 is that the preprocessed data is input into the trained deep Q network for sepsis anti-infection, and finally a decision scheme maximizing the expected cumulative reward is obtained.
[0031] The advantages and positive effects of the present application are:
[0032] The present application firstly models the use of antibiotic combination in the treatment of sepsis as a Markov decision process, and uses DQN to construct the model; secondly, the deep reinforcement learning based on the value function is combined with the content related to sepsis clinical data and medical guidelines, the clinical real data are referred to in the reward function, the SOFA score knowledge and the content of clinical guidelines are integrated, and the model is guided to the direction of reasonably shortening the antibiotic use time and ensuring the good prognosis of the patient for strategy recommendation. The model constructed by the present application takes the patient's demographic characteristics, basic vital signs, microbiological culture results and antibiotic use data as the model input, combines the clinical prior knowledge, clinical data and medical guidelines, constructs four reward functions, enhances the medical interpretability and clinical consistency of the SAI-DQN, and uses the SAI-DQN to provide personalized antibiotic combination therapy suggestions for sepsis patients. BRIEF DESCRIPTION OF DRAWINGS
[0033] Figure 1 It is the MIMIC-IV data set processing process of the present application;
[0034] Figure 2 It is a patient feature vector schematic diagram of the present application;
[0035] Figure 3 It is a model framework diagram of the SAI-DQN of the present application;
[0036] Figure 4 It is a state space schematic diagram of patient feature generation of the present application;
[0037] Figure 5The action space diagram for the patient medication sequence generation of the application is shown in the figure;
[0038] Figure 6 The WIS value comparison of Q-Learning, SARSA, SAI-DQN and DQN-1 of the application is shown in the figure;
[0039] Figure 7 The comparison of the medication time of the four models and the clinic of the application is shown in the figure;
[0040] Figure 8 The comparison of the WIS and medication time of DQN and SAI-DQN when different reward functions are set in the application is shown in the figure;
[0041] Figure 9 The comparative analysis histogram of the drug similarity predicted by the clinic and the model under the condition that the initial state of the patient is similar (similarity 80%) in the application is shown in the figure;
[0042] Figure 10 The comparative box plot of the drug similarity predicted by the clinic and the model under the condition that the initial state of the patient is similar (similarity 80%) in the application is shown in the figure;
[0043] Figure 11 The histogram of the similarity of the model drug sequence corresponding to the survival and death of the clinic and the application is shown in the figure;
[0044] Figure 12 The box plot of the similarity of the model drug sequence corresponding to the survival and death of the clinic and the application is shown in the figure. DETAILED DESCRIPTION
[0045] The application will be further described in detail below in combination with the accompanying drawings.
[0046] The sepsis diagnosis and treatment scheme recommendation method based on reinforcement learning and clinical knowledge comprises the following steps:
[0047] Step 1, collect data and preprocess the data.
[0048] Step 1 comprises the following steps:
[0049] Step 1.1, collect electronic medical records from the MIMIC-IV dataset;
[0050] Step 1.2, determine whether the electronic medical records in the MIMIC-IV dataset meet the sepsis-3 standard, if they meet the standard, proceed to step 1.3, otherwise exclude the electronic medical records;
[0051] Step 1.3, determine whether the information in the electronic medical record is greater than or equal to 18 years old, if the condition is met, proceed to step 1.4, otherwise exclude the electronic medical record;
[0052] Step 1.4, judge whether the information in the electronic medical record is transferred to another hospital, if yes, exclude the electronic medical record, otherwise, proceed to step 1.5;
[0053] Step 1.5, select feature items in the electronic medical record according to the missing rate;
[0054] Step 1.6, judge whether the missing rate is less than or equal to 60%, if yes, constitute a usable data set, otherwise, exclude the electronic medical record to screen out the "usable data set";
[0055] Step 1.7, use K-Nearest Neighbor method and forward filling mixed imputation strategy to fill in the missing values of the usable data set to obtain a complete data set which can be directly used for subsequent feature engineering and reinforcement learning training.
[0056] Step 1.8, in order to verify the generalization ability of the recommended model of the scheme, in addition to using the MIMIC-IV database, the electronic ICU Collaborative Research Database (eRI) is also used as an external test set. For the cases in eRI, patients ≥18 years old in ICU are screened according to the Sepsis-3 standard, and patients with incomplete data are excluded; then the missing rate is calculated based on the aforementioned feature list, and samples with a missing rate greater than 60% are removed, and K-Nearest Neighbor method is used to interpolate the remaining missing values; finally, an eRI verification set with the same structure as the MIMIC-IV "usable data set" is generated, which is used to evaluate the recommendation effect and stability of SAI-DQN on heterogeneous hospital data.
[0057] The electronic medical record used in the present application is from the MIMIC-IV dataset, which contains comprehensive clinical information of ICU patients in a large academic medical center, including vital signs, laboratory values, medication information and clinical notes. The present application extracts ICU data of 9982 sepsis patients from the MIMIC-IV dataset. In order to focus on the treatment of sepsis, the present application selects patients diagnosed with sepsis according to the Sepsis-3 standard from the MIMIC-IV dataset. The present application excludes neonatal sepsis patients or patients transferred during treatment. The present application uses time series data of patient demographic data, vital signs (such as blood pressure, etc.) and laboratory results (such as lactic acid, PCO2, etc.), microbiological culture records, medication and treatment process, and diagnosis coding. Figure 1The MIMIC-IV data and the processing process chart. The present application excludes patients with more than 60% missing values of demographic characteristics, laboratory test data or basic vital signs during ICU hospitalization. Finally, the present application retains the demographic characteristics of the patients such as gender, age, weight, and averages the repeated data within 24 hours. The present application selects 37 laboratory items and 15 basic vital signs according to clinical knowledge and missing rate. The present application takes the average of the multiple measurement data of the 37 test items within the same time interval of 24 hours. In addition, the present application also analyzes 15 basic vital signs including heart rate and blood pressure, and performs variance analysis on data such as heart rate which needs to pay attention to the trend of data change.
[0058] Microbial culture results are very important for predicting the final outcome, as they can show the patient's infectious strain and sensitivity to different commonly used antibiotics, helping to select the most effective drugs. During the data screening process, the present application retains the microbial culture results, although the missing rate may exceed 80%. The microbial culture results include information on different bacterial strains and corresponding antibiotic sensitivity, a total of 2047 combinations. The sensitive result is represented by 1, the insensitive result is represented by 0, and the missing data is represented by -1.
[0059] For demographic characteristics, laboratory data and basic vital signs, the present application uses the forward filling method and the K nearest neighbor method to fill in the missing values. If the complete characteristics of the patient are missing, the present application uses the K nearest neighbor method. The data missing filling method used by the present application is mainly: when a certain characteristic of the same patient number is missing, if there is a non-empty value of this patient at other times, the forward filling method is used to fill in, if the value of this item of the patient is all empty, the K nearest neighbor method is used to fill in, where k = 9. Single use of K nearest neighbor or forward filling method is not enough to include the individualization of patients in medical clinical problems. The advantages of the two missing value filling methods are combined. Using the forward filling method can make the filling value more consistent with the individual characteristics of the patient. Using the K nearest neighbor method to fill in the completely missing items of the patient can ensure that the filling value is within the clinically interpretable range and maintains a certain individualization. The daily feature vector of the patient can be represented as Figure 2 .
[0060] Step 2, constructing a deep Q network for sepsis anti-infection and training.
[0061] The deep Q network for sepsis anti-infection in step 2 includes an action space generation module, a state space generation module, a SAI-DQN module and a SAI-DQN training module, wherein the action space generation module and the state space generation module are connected to the SAI-DQN module, and the SAI-DQN module is connected to the SAI-DQN model test module.
[0062] The function of the action space generation module is to construct a set of selectable "antibiotic sequences" according to the drug library preset in the clinical scheme, and to encode the sequence into a discrete action acceptable to the reinforcement learning RL; the function of the state space generation module is to extract a multi-dimensional feature vector of the patient from the MIMIC-IV, including physiological indicators, laboratory test values, medication history and clinical scores, and then normalize and standardize these raw data, and combine the time information to splice into the current time state;
[0063] The function of the SAI-DQN module is to constantly optimize the Q network parameters in the training phase through the double network architecture of the main network and the target network, to incorporate the ∈-greedy exploration strategy of the clinical expert rule and the experience replay and periodic synchronization update, to realize the learning and convergence of the antibiotic administration strategy; the function of the SAI-DQN model test module is to load the optimal parameters trained in the test phase, to execute inference on the state vector of a new patient, to generate the optimal medication action at each time point, and to summarize into a complete personalized recommendation strategy for clinical reference.
[0064] As shown in Figure 3 The SAI-DQN training module includes an agent observation environment module, an experience replay pool, a Q network, a DQN error function, a target Q network and a reward function, wherein the patient feature vector and the antibiotic sequence in the preprocessed data are input into the agent observation environment module, the agent observation environment module is connected to the experience replay pool and the Q network respectively, the experience replay pool is connected to the reward function and the target Q network respectively, the Q network is connected to the DQN error function and the target Q network respectively, and the target Q network is connected to the DQN error function;
[0065] The agent observation environment module constitutes an environment based on the correspondence between the patient feature vector and the antibiotic sequence. The patient feature vector in the agent observation environment calculates the Q value of each action through the Q value function, so as to select a suitable action (a). The antibiotic combination sequence selected by the agent interacts with the environment, obtains a reward (r) according to the reward function and a new state (s'), updates the neural network parameters using the DQN algorithm, improves the prediction performance of the Q value function by minimizing the loss of the Q value function, and the experience replay pool includes the memory of the state observation, the selected action, the reward and the new state; the past state observation, the selected action, the reward and the new state are used to train the network, to enhance the randomness, generalization and convergence of the algorithm, and the above steps are repeated, through the interaction with the environment and the update of the network parameters, the DQN algorithm enables the agent to gradually learn to maximize the long-term reward in a given environment.
[0066] During model testing, SAI-DQN inputs the feature vector of the first day of the patient in the test set and the "final state" position is marked as 0. The trained SAI-DQN generates 1 to n day recommended actions according to the state observation, action selection memory of the experience replay pool. The obtained action obtains different antibiotic combinations according to the clustering category center point. Therefore, the final recommended result is n antibiotic combinations in time sequence.
[0067] Figure 4 A state space diagram is generated for the patient feature vector. The patient feature vector includes 37 laboratory items, 15 basic vital signs, and 2047 microbiological culture results. The microbiological culture results are high-dimensional sparse data. The state space includes the patient's feature vector plus a flag bit representing the final state. The present application represents the final state according to the patient's actual clinical outcome (90-day hospital mortality rate). When the patient's final state is survival, the final state of the patient's antibiotic use is marked as 1. When the patient's final state is death, the final state of the patient's antibiotic use is marked as -1. The rest of the time is marked as 0.
[0068] Figure 5 An action space diagram is generated for the patient's medication sequence. In the patient's daily medication sequence, the sequence value is whether the corresponding antibiotic numbered antibiotic is used on that day. The value is 1 when the antibiotic is used on that day, and 0 when it is not used. The patient's daily antibiotic combination sequence is clustered into 30 classes by k-means clustering, and the obtained 30 classes are used as the action space A of the model input.
[0069] In order to enhance the explainability and clinical consistency of the model decision-making process, the present application combines statistical analysis of the clinical data set with expert knowledge to design reward functions R1-R4. Table 1 lists the content settings of the four reward functions.
[0070] Table 1 SAI-DQN reward function R1-R4 setting
[0071]
[0072] The present application sets four reward functions R1-R4. R1 is the basic reward, taking the outcome of patients in the data set within 90 days as an observation index, rewarding +100 for patients surviving for 90 days and punishing-100 for patients dying. R1 is directed to the prognosis of the patient, so that the model learns towards the survival of the patient. R2 takes the SOFA score of the patient as a standard. The score is graded to increase, indicating that the patient's condition is deteriorating, then punished-10, and vice versa, rewarded +10. R2 is a reward function set according to the state of the patient to make the model learn towards the state change that is beneficial to the prognosis of the patient. R3 is a reward function set according to the average medication duration in the data set. The average medication duration in the data set is about 9 days, so the reward function R3 is set to punish-10 when the medication duration (sequence length) is greater than 9 days. The purpose of setting R3 is to make the model achieve the task goal of using shorter antibiotics. R4 compares the intermediate state transition with the Q table obtained by statistics of the data set, and the Q value obtained by clinical data statistics is multiplied by 0.1 as a local reward, R=Q clinical (s,a,s')*0.1, wherein Q clinical (·) represents a clinical knowledge function; s represents the current state, including patient vital signs, laboratory indicators and microbiological data; a represents the evaluation action, i.e. the selected antibiotic combination; s' represents the next state, i.e. the predicted state after medication. This is to enable the model to set a reward function combined with real clinical data, which incorporates real clinical situations and is closer to clinical facts. The model incorporating four reward settings is called SAI-DQN, i.e. sepsis anti-infection DQN.
[0073] Step 3, input the preprocessed data into the constructed sepsis anti-infection deep Q network to obtain a sepsis diagnosis and treatment recommendation scheme.
[0074] The specific implementation method of step 3 is: input the preprocessed data into the trained sepsis anti-infection deep Q network to finally obtain a decision scheme that maximizes the expected cumulative reward.
[0075] According to the above sepsis diagnosis and treatment scheme recommendation method based on reinforcement learning and clinical knowledge, the effect of the present application is verified by testing.
[0076] The main data missing filling method used by the present application is as follows: when a certain feature of the same patient number is missing, if the patient has non-null values at other times, forward filling method is used for filling; if K nearest neighbors exist in the item, K nearest neighbor method is used for filling. If the patient's item value is all null, K nearest neighbor method is used for filling, where k=9. Whether it is K nearest neighbor method or forward filling method, it is not enough to cover the individualization of patients in medical clinical problems. The advantages of the two missing value filling methods are combined. Using forward filling method can make the filled value more consistent with the change of patient individual characteristics. Using K nearest neighbor method to fill the missing items of patient complete characteristics can ensure that the filled value is within the clinically interpretable range and maintains a certain individualization.
[0077] Table 2 Comparison of feature average values after filling features using different filling methods
[0078]
[0079] Table 3 Comparison of feature standard deviations after filling features using different filling methods
[0080]
[0081] Tables 2 and 3 are the average values and standard deviations of Lactate, White Blood Cells and pO2 when using different filling methods. Among them, the Lactate missing rate is 50.52%, the White Blood Cells missing rate is 6.37%, and the pO2 missing rate is 52.75%. The present application uses mean filling method, forward filling method, K nearest neighbor filling method and mixed method to fill the missing values of Lactate, White Blood Cells and pO2, and compares the filled average and standard deviation. The average and standard deviation obtained by not counting the null values before filling are also included in the table. From Table 2, the present application sees that the average value of the filled data is similar to that before filling. However, as shown in Tables 2 and 3, the standard deviation of the data obtained by using mean filling and forward filling indicates that the dispersion of the filled data is low, i.e., the individualization of patient data is reduced. In clinical practice, the state change of patients often has more discrete nature, therefore, the present application mixes the forward filling and K nearest neighbor filling methods to better ensure the individualization of patient characteristics.
[0082] The deep Q network model constructed by the application is firstly trained and internally verified on the MIMIC-IV data set; then, with the aid of the eRI data set as an independent external test, the same data cleaning and preprocessing process as the training set is adopted to further verify the cross-center applicability and robustness of the anti-infection strategy. The eRI data set includes vital signs, laboratory values, medication records, ICU admission and discharge information and other data types. As a test data set, the application uses the data of 11070 patients from the eRI database, and obtains the permission of the eRI institutional review board. First, patients under the age of 18 and patients transferred in the middle of ICU hospitalization are excluded. Then, the data of patients during the period from the start of using antibiotics to the stop of using in ICU are selected for data screening. The application selects the characteristics of the patients in the MIMIC-IV database to select the characteristics of the patients in the eRI data set. Since the regions and patient conditions of the two data sets are different, the final selected patient characteristics of the eRI data set include 12 antibiotics, 26 laboratory items, 4 basic vital signs and demographic characteristics. In addition, microbiological culture data are also included, obtaining a 311-dimensional data set, which contains the values of different strains and corresponding antibacterial drugs.
[0083] The main reason for selecting the MIMIC-IV data set and the eRI data set is that the two data sets contain patient data in the intensive care unit, including sepsis related. While other data sets are mainly for specific medical fields such as cancer, nutrition, etc. not related to sepsis, therefore, the MIMIC-IV data set and the eRI data set are selected as the training and test data sets in this paper.
[0084] The application selects two basic function reinforcement learning methods as benchmark models. The two methods of Q-Learning and SARSA use discrete state space and action space as input. The state of the patient is clustered using the k-means method to cluster the daily characteristics of the patient, obtaining 500 patient states, plus the survival and death states within 90 days, a total of 502 states. At the same time, the action space is also clustered using the method to cluster the medication series of each patient. The best parameter setting is determined through experiments. The results obtained by the SARSA method are similar to those obtained by the Q-Learning method, but the randomness is reduced, and the exploration risk of the model is limited. The results are more random. Figure 6 The WIS score comparison chart of the SAI-DQN method of the application and the Q-learning, SARSA and DQN-1 methods, wherein DQN-1 is a comparison model using only R1 as the reward function.
[0085] In order to evaluate the performance of the model, Figure 6The WIS values of different models of SAI-DQN, Q-learning, SARSA and DQN-1 are compared as the number of training increases. The data used for testing the clinical, Q-learning, SARSA and DQN-1 models is the MIMIC dataset, and the SAI-DQN model is tested in the MIMIC dataset and the eRI dataset. In this case, DQN-1 only uses the treatment results of patients for 90 days as the reward function. Figure 6 In the figure, the X-axis represents the number of training iterations, and the Y-axis represents the WIS score. The average WIS score of the patient results in the clinical test set obtained by training is 65.0147. As the number of training iterations increases, the performance scores of the four models gradually exceed the clinical trajectory value, and reach a stable state after a certain number of iterations. After sufficient training, the WIS score of the patient medication strategy recommended by the model exceeds the WIS score of the clinical decision, indicating that the model can provide a reference for clinical medication decision-making. The SAI-DQN of the present application can obtain a performance better than the clinical data after only 150,000 training iterations, which indicates that it has strong learning ability. It can be observed that the data corresponding to eRI has a similar structure to the MIMIC data, although the score is lower. However, as the number of training iterations increases, the final result gradually stabilizes and exceeds the clinical WIS score.
[0086] In clinical practice, doctors' medication strategies are generally empirical and conservative. The WIS score result shows the value score of the evaluation action for achieving good results in the test set data statistics. Since in clinical practice, patient conditions are affected by many factors, doctors need to face choices full of risks and challenges. Therefore, the conservatism and empiricism of clinical medication may make the WIS evaluation result worse than the model. However, in practical applications, treatment decisions should be made by considering various factors. The SAI-DQN in this paper can achieve better results in predicting and recommending medication to patients, which indicates that the recommended medication of SAI-DQN has reference value for clinical doctors. At the same time, combined with the experience and judgment of clinical doctors, a safer and more personalized treatment plan can be provided for patients.
[0087] The present application compares the patient trajectory length, i.e. the duration of antibiotic use, of Q-learning, SARSA, SAI-DQN, DQN-1 and eRI datasets. Figure 7 As shown in the figure, the average duration of antibiotic use obtained by clinical and statistical analysis of patient data is 9.4758 days.
[0088] Under all reward function settings, the average duration of antibiotic use by patients in the Q-learning model and the SARSA model was 4.23 and 4.46, respectively. In the SAI-DQN model, the average duration of antibiotic use was 5.51. The SAI-DQN model proposed in this paper shortened the duration of antibiotic use while trying to be consistent with the guidelines. Compared with the baseline DQN-1, the other models incorporated safety exploration and limited the use of the clinical dataset Q value as a behavior reward, thereby guiding the decision-making process of the model closer to typical clinical practice.
[0089] In view of the training data used by the present application, which is derived from the actual antibiotic use records of patients during their stay in the intensive care unit, the internal "prognosis" prediction of the model is not equivalent to the statistical sense of clinical survival rate. The following is intended to explain the difference between the two and how the model learns and evaluates using the "prognosis state".
[0090] The training data used by the present application is directly derived from the actual antibiotic use records of patients during their stay in the ICU, so the outcome observed by the model in the experience replay buffer is defined based on whether the patient survives within 90 days after discharge, rather than directly using the clinical statistical "survival rate" concept. For this reason, the present application refers to the outcome of the patient obtained by the model as "prognosis state": when the patient does not have a death event within 90 days, the prognosis state is recorded as "good"; otherwise, it is recorded as "poor". In the reward function R1, i.e., based on this "prognosis state": a positive reward (such as +100) is given for a good prognosis, and a negative penalty (such as -100) is given for a poor prognosis, to guide the agent to preferentially learn the administration strategy that is conducive to a good prognosis (Table 1 has made general settings for R1-R4).
[0091] Since the data used to train the model is the actual use of antibiotics by patients during their stay in the intensive care unit, the present application cannot refer to the prognosis of patients obtained by the model as "survival rate". The present application uses the prognosis state of the patient to represent the outcome of the patient understood by the model. Table 4 lists the comparison of survival rates between the clinical test set and the three training models, and Table 5 lists the prediction of the model for the clinical survival rate and mortality outcome. In SAI-DQN, the experience replay buffer can provide insight into the prognosis of patients during the training process of the model. However, since the data used to train the model is the actual use of antibiotics by patients during their stay in the intensive care unit, the present application cannot refer to the prognosis of patients understood by the model as "survival rate". Instead, the present application uses the prognosis state of the patient to represent the outcome of the patient understood by the model. Table 5 lists the prediction of the model for the clinical survival rate and mortality outcome. In SAI-DQN, the experience replay buffer can provide insight into the prognosis of patients during the training process of the model.
[0092] Table 4 Comparison of survival rates between the clinical test set and the three models after training multiple models
[0093]
[0094] Table 5 SAI-DQN model prediction results corresponding to clinical survival and death
[0095]
[0096] The present invention uses "good prognosis" or "poor prognosis" to represent the final condition of the patient. According to Table 4, the present invention can see that the model-predicted patient survival rate is between 75-80%, which indicates that the model helps patients to some extent to achieve better prognosis results. According to the data in Table 5, the clinical patient survival rate in the test data set of the present invention is 61.94%, while the model-predicted patient good prognosis rate is 79.80%. According to Table 5, the present invention can see that for patients with clinical results of survival, almost all models predict good future prognosis. At the same time, for patients with clinical prognosis of death, 54.77% of the models predict good future prognosis. According to Table 4 and Table 5, the present invention can conclude that the model's recommendation of antibiotic drug combination and duration can bring more favorable prognosis for patients.
[0097] In order to compare the effects of different reward function settings on model results, the present invention conducted a set of ablation experiments. Table 6 lists the different reward function settings used in the ablation experiments. The SAI-DQN model of the present invention uses the complete reward function. Figure 8 The WIS scores and antibiotic use time statistics of SAI-DQN and DQN-1-7 are shown.
[0098] Table 6 Reward function settings corresponding to the ablation experiment model
[0099]
[0100] The results of the ablation experiments show that the three models of DQN-1, DQN-2 and DQN-4 tend to quickly obtain results, ignoring the changes in patient status, which is not consistent with clinical experience. From Figure 8 it can be seen that DQN-1 has the highest WIS score, but the antibiotic use time is too short, which will lead to the emergence of situations inconsistent with clinical experience. DQN-2 is similar to DQN-1, and the antibiotic use time is also too short. DQN-2 is similar to DQN-1, and its antibiotic use time is too short. On the other hand, although DQN-3 adds a penalty for using time that is too long, compared with DQN-2 and DQN-4, its average recommended medication time is longer. Similarly, the average recommended medication time of DQN-5 and DQN-7 is also longer than DQN-6.
[0101] This invention suggests that incorporating patient SOFA scores into DQN-2 and clinical statistics into DQN-4 may lead the model to learn patients with poor condition and prematurely terminate the trajectory to minimize score deductions for these patients, resulting in a poorer prognosis. In DQN-3, during training, the model learns that patients with prolonged medication use will have higher DQN-3 scores. The model also learns that patients with longer medication use have lower Q-scores than those in DQN-1, DQN-2, and DQN-4. Therefore, the model chooses a more effective action during prediction, which may prolong medication use but has a better prognosis.
[0102] The predicted duration of treatment from DQN-3 to DQN-7 is generally between 4 and 5 days, which is similar to the average duration of treatment based on clinical experience. The average duration of antibiotic use for SAI-DQN is 5.5190 days. Figure 8 As can also be seen in this invention, the WIS score of the SAI-DQN of this invention is also very high. SAI-DQN combines four reward functions, which constrain each other to ensure that the model recommends medication in a direction that leads to a good patient prognosis.
[0103] This invention compares and analyzes the similarity of drug use between clinical data and SAI-DQN simulations. To assess the similarities and differences in antibiotic prescriptions between clinical datasets and SAI-DQN simulations, this invention calculates the similarity of drug sequences.
[0104] To assess the similarity between patient characteristics, this invention normalizes the daily characteristic sequences for each patient and uses the Pearson coefficient to calculate the similarity between any two patient characteristic sequences. First, the characteristic sequences are normalized, and the normalized f is obtained using an outlier normalization method. M ′M、f N The patient characteristic similarity is S. M,N =Pearson(f M ′M,f N ′), where f M ′ represents the characteristics of patient M, f N ′ represents the characteristics of patient N, Pearson() is the similarity calculation method, and S M,N This is the similarity result.
[0105] This invention performs statistical comparisons of drug sequence similarity. Figure 9 and Figure 10 This displays statistical plots and box plots showing the similarity between clinically predicted and model-predicted drug sequences when the initial patient state similarity is greater than 80%. Figure 9 and Figure 10It can be seen that the drug sequence similarity of patients with similar initial states is higher. In clinical practice, the drug sequence similarity of patients with similar initial states is mainly between 60%-90%.
[0106] Figure 9 and Figure 10 “clinical-SAI-DQN” in is the comparison of the similarity between the clinical actual recommendation and the model predicted recommendation of the same patient. From Figure 9 and Figure 10 It can be seen that the similarity of the same patient's clinical actual and model predicted drug combination is between 60%-90%. This result shows that there is a certain degree of similarity between the clinical actual and the model in the antibiotic combination recommendation, and the model can learn the drug use combination strategy from the clinical data. The model SAI-DQN has good clinical consistency, but personalized decision-making may be needed in clinical practice. Due to the limited training data, this may result in the model providing a relatively stable but not personalized enough drug use strategy for patients, and this problem can be solved in the future by building a multi-model decision-making method.
[0107] The present application calculates the drug sequence similarity of clinically surviving or dead patients in the test set. From Figure 11 It can be seen that for clinically model-predicted surviving patients, the similarity of the antibiotic drug sequence recommended by the model is mostly between 60% and 90%. This shows that the drug sequence recommended for surviving patients in clinical practice is consistent with the clinical antibiotic strategy learned by the model.
[0108] For clinically predicted surviving but model predicted dead patients, combined with the results of SAI-DQN predicted patients in Table 4, the situation that the model predicts better than the clinical patients, in Figure 11 (b) the similarity of the drug sequence recommended by the model is mostly between 0%-60%, indicating that the model has some exploration behavior. For clinically dead patients, the model predicts the survival rate of more than half of the patients, and the recommended antibiotic combination is quite different from the antibiotic combination used in clinical practice. For clinically dead patients, the similarity of the drug sequence recommended by the model and the drug sequence used in clinical practice is mostly between 0%-70%. From Figure 11 , Figure 12 It can be seen that the model predicts the drug sequence similarity of clinically surviving patients is higher. However, the exploration behavior of the model may affect the drug sequence recommendation. At the same time, the exploration of the model on the drug use of clinically dead patients can improve their prognosis.
[0109] Table 7 Comparison of clinical actual drug use and model recommended drug use of patient A
[0110]
[0111]
[0112] Table 8 Comparison of actual and model recommended antibiotics for patient B
[0113]
[0114] Tables 7 and 8 show the comparison between the antibiotics actually used by patients A and B in the clinic and the antibiotics recommended by SAI-DQN simulation, respectively. The prognosis of patient A was good in both the clinic and the model simulation. The antibiotic usage showed that the clinician mainly recommended Vancomycin for patient A in the early stage and Ciprofloxacin in the later stage, while recommending two other antibiotics; SAI-DQN also recommended Vancomycin and Ciprofloxacin for patient A, and patient A used Ciprofloxacin in the early stage. Patient B had a poor prognosis in the clinical environment, but a good prognosis in the simulated case. As can be seen from Table b, the clinician used 6 different antibiotics for patient B in 10 days, but patient B had a poor prognosis. The three main antibiotics recommended by SAI-DQN for patient B included Vancomycin, Fluconazole, and Ciprofloxacin.
[0115] In combination with the prognosis results of patient B and Figure 11 (c), the proportion of different (low similarity) decisions recommended by the clinician and SAI-DQN was high, and SAI-DQN could give better recommended medication regimen for patients who died clinically. From the example of patient B, the present application can assume that SAI-DQN can learn from clinical data, optimize the antibiotic treatment regimen of the clinician, and provide personalized treatment according to the patient's condition.
[0116] In combination with the examples of patients A and B, SAI-DQN of the present application can imitate and optimize the medication method of the clinician in combination with the patient's condition. More importantly, SAI-DQN can ensure that the antibiotic usage time is optimized to a certain extent in the case of a good prognosis of the patient, and can provide suggestions for the antibiotic usage regimen for the doctor. However, since the antibiotic usage combination of the present application is an action space generated after clustering, the SAI-DQN model can be improved in subsequent work by starting from the optimization of the antibiotic combination sequence representation.
[0117] It should be emphasized that the embodiments described in the present application are illustrative rather than restrictive, and therefore the present application includes but is not limited to the embodiments described in the specific embodiments, and any other embodiments derived by those skilled in the art from the technical solutions of the present application also fall within the scope of protection of the present application.
Claims
1. A method for recommending a sepsis diagnosis and treatment plan based on reinforcement learning and clinical knowledge, characterized in that: The method comprises the following steps: Step 1, collecting data and pre-processing the data; Step 2, constructing a deep Q network for sepsis anti-infection and training the deep Q network; Step 3, inputting the pre-processed data into the constructed deep Q network for sepsis anti-infection to obtain a sepsis diagnosis and treatment recommendation scheme.
2. The method of claim 1, wherein the method is based on reinforcement learning and clinical knowledge. The step 1 comprises the following steps: Step 1.1, collecting electronic medical records from the MIMIC-IV dataset; Step 1.2, determining whether the electronic medical records in the MIMIC-IV dataset meet the sepsis-3 standard, if yes, proceeding to step 1.3, otherwise excluding the electronic medical records; Step 1.3, determining whether the information in the electronic medical records is greater than or equal to 18 years old, if yes, proceeding to step 1.4, otherwise excluding the electronic medical records; Step 1.4, determining whether the information in the electronic medical records is transferred to a hospital, if yes, excluding the electronic medical records, otherwise proceeding to step 1.5; Step 1.5, selecting feature items in the electronic medical records according to the missing rate; Step 1.6, determining whether the missing rate is less than or equal to 60%, if yes, constituting an available dataset, otherwise excluding the electronic medical records; Step 1.7, using a mixed imputation strategy of K-nearest neighbor method and forward filling to impute the missing values in the available dataset in step 1.6 to obtain a complete dataset; Step 1.8, collecting electronic medical records of the eRI dataset, and repeating steps 1.2 to 1.7 to obtain an eRI validation set.
3. The method of claim 1, wherein the method is based on reinforcement learning and clinical knowledge. The deep Q network for sepsis anti-infection in the step 2 comprises an action space generation module, a state space generation module, an SAI-DQN module and an SAI-DQN training module, wherein the action space generation module and the state space generation module are connected to the SAI-DQN module, and the SAI-DQN module is connected to the SAI-DQN training module.
4. The method of claim 3, wherein the method is based on reinforcement learning and clinical knowledge. The SAI-DQN training module comprises an agent observation environment module, an experience return pool, a Q network, a DQN error function, a target Q network and a reward function, wherein the patient feature vector and the antibiotic sequence in the pre-processed data are input into the agent observation environment module, the agent observation environment module is connected to the experience return pool and the Q network, the experience return pool is connected to the reward function and the target Q network, the Q network is connected to the DQN error function and the target Q network, and the target Q network is connected to the DQN error function; The agent observes the environment module is based on the corresponding relationship between the patient feature vector and the antibiotic sequence to constitute the environment, the agent observes the patient feature vector in the environment to calculate the Q value of each action through the Q value function, so as to select the appropriate action (a). The antibiotic combination sequence selected by the agent is executed to interact with the environment, the reward (r) is obtained according to the reward function and the new state (s') is obtained, the neural network parameters are updated using the DQN algorithm, the prediction performance of the Q value function is improved by minimizing the Q value function loss, the experience replay pool includes the memory of state observation, selected action, reward and new state; The past state observation, selected action, reward and new state are used to train the network, the randomness, generalization and convergence of the algorithm are enhanced, the above steps are repeated, through the interaction with the environment and the network parameter update, the DQN algorithm makes the agent gradually learn to maximize the long-term reward in the given environment.
5. The method of claim 4, wherein the method is based on reinforcement learning and clinical knowledge. The reward function includes reward function R1, reward function R2, reward function R3 and reward function R4, wherein the motivation of reward function R1 is to enable the model to understand the prognosis of the patient, and the content is the survival rate of the patient within 90 days, and the reward function R1 is set to survive: +100, death: -100; The motivation of reward function R2 is to encourage the model to learn to improve the prognosis of the patient, and the content is the change of SOFA score of the patient, and the reward function R2 is set to state improvement reward +10, state deterioration penalty -10; The motivation of reward function R3 is to encourage the model to shorten the antibiotic use time, and the content is the medication time, and the reward function R3 is set to more than 9 days, and the penalty is -10; The motivation of the reward function R4 is to incorporate clinical data, which compares the transition of the intermediate state with the dataset statistics, and the reward function R4 is set as R = Q clinical (s,a,s')*0.1, wherein Q clinical (·) represents a clinical knowledge function; s represents a current state, including patient vital signs, laboratory indicators, and microbiological data; a represents an evaluation action; and s' represents a next state.
6. The method of claim 1, wherein the method is based on reinforcement learning and clinical knowledge. The specific implementation method of step 3 is: input the preprocessed data into the trained deep Q network for sepsis anti-infection, and finally obtain the decision scheme that maximizes the expected cumulative reward.