A Method for Constructing a Reinforcement Learning-Assisted Clinical Decision-Making Model for Sepsis
A reinforcement learning model optimizes sepsis treatment by training an intelligent agent to recommend optimal intravenous fluid and vasopressor doses, addressing data sparsity and improving decision accuracy in sepsis management.
Patent Information
- Application Number
- CN202410945601.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-15
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-07-15
AI Technical Summary
In ICU, the prior art cannot effectively optimize the dosage combination of intravenous infusion and vasopressor, resulting in poor treatment effect, increasing the risk of death in patients, and lacking objective agent decision-making evaluation methods.
Build a clinical decision-making model for sepsis assisted by reinforcement learning, train agents using historical clinical data, optimize dose combinations through feature extraction layer, embedding processing layer and Q value calculation layer, and solve data sparse problems through reward functions and random perturbations to evaluate decision-making effects.
The dosage combination of intravenous infusion and vasopressors is optimized, the treatment effect is improved, the risk of adverse decision-making is reduced, and objective agent decision-making evaluation methods are provided, which improves the accuracy and safety of clinical treatment.
Smart Images

Figure CN118888157B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of artificial intelligence technology, and particularly to a method for constructing a sepsis clinical decision-making model assisted by reinforcement learning. Background Art
[0002] In the ICU, sepsis is a common and serious medical condition. Sepsis is usually caused by infection and is the result of the body's excessive immune response. The most common situation for sepsis patients is insufficient blood pressure, which leads to poor tissue perfusion. Therefore, maintaining blood pressure is a basic idea in the treatment of sepsis. Clinicians usually use two means, intravenous infusion or vasopressor drugs, to maintain the blood pressure of sepsis patients.
[0003] Intravenous infusion aims to increase blood pressure by increasing blood volume, but excessive use will cause tissue fluid accumulation, which in turn affects oxygen delivery; vasopressor drugs elevate blood pressure by constricting blood vessels, but excessive use will cause excessive capillary constriction, affecting tissue perfusion and, in severe cases, leading to organ failure. During the treatment of sepsis, the application of these two means often runs through the whole process. Clinicians need to adjust the dose combination of these two treatment means in real time every once in a while according to the patient's condition to achieve the purpose of gradually improving the patient's condition. Due to the complexity of the action mechanisms and interactions of these two means, the clinical medical community has not reached a consensus on the dose combination of these two means. Inappropriate doses will increase the mortality rate of patients. Therefore, studying the dose combination problem in the treatment of sepsis is of great significance. Summary of the Invention
[0004] The embodiments of the present invention provide a method for constructing a sepsis clinical decision-making model assisted by reinforcement learning, which uses agent learning to assist sepsis clinical decision-making.
[0005] In a first aspect, the embodiments of the present invention provide a method for constructing a sepsis clinical decision-making model assisted by reinforcement learning, including:
[0006] Taking multiple physiological indicators of sepsis patients in historical clinical data as state variables and the dose combination of intravenous infusion and vasopressor drugs actually taken by doctors as action variables, constructing training samples for the reinforcement learning model;
[0007] Building an agent model based on reinforcement learning according to the training samples. The model includes a feature extraction layer, an embedding processing layer, and a Q-value calculation layer; using the training samples, training the model in a reinforcement learning manner, and the trained model is used to recommend the best dose combination according to multiple physiological indicators of the patient at a certain moment.
[0008] Among them, the multiple physiological indicators include fluid balance volume, blood pressure, blood protein, and lactic acid; during the training process, the role of fluid balance volume and blood pressure in dose selection is enhanced through a reward function, while the role of blood protein and lactic acid is inhibited.
[0009] In a second aspect, an embodiment of the present invention provides an electronic device, which includes:
[0010] One or more processors;
[0011] A memory for storing one or more programs,
[0012] When the one or more programs are executed by the one or more processors, the one or more processors implement the method for constructing a reinforcement learning assisted sepsis clinical decision model according to any embodiment.
[0013] In a third aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method for constructing a reinforcement learning assisted sepsis clinical decision model according to any embodiment.
[0014] In summary, the embodiment of the present invention provides a method for constructing a reinforcement learning assisted sepsis clinical decision model. By using reinforcement learning technology, an agent is trained to select a dose combination of intravenous infusion and vasopressor for a patient's state at a certain moment. By learning excellent decisions in historical treatment data, bad decisions are optimized, and an objective evaluation method for auxiliary decisions based on historical treatment data is proposed. In particular, in this embodiment, positive samples are constructed by adding random perturbations in reinforcement learning and jointly trained with negative samples in the same batch to train the reinforcement learning algorithm, solving the problem of data sparsity in the scenario of reinforcement learning assisted sepsis clinical decision-making and improving the stability of the agent. At the same time, in this embodiment, the main factors for constructing the reward function are screened according to the medical logic diagram, solving the problem of inappropriate reward functions in existing work. In addition, this embodiment proposes a method for evaluating the effectiveness of the agent's decision at a single time point and the effectiveness of the agent's sequential decisions, and the effectiveness of the agent's decisions can be judged based on historical clinical data. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0016] Figure 1 It is a flowchart of a method for constructing a reinforcement learning assisted sepsis clinical decision model provided by an embodiment of the present invention;
[0017] Figure 2 It is a schematic structural diagram of a reinforcement learning-assisted sepsis clinical decision-making model provided by an embodiment of the present invention;
[0018] Figure 3 It is a schematic diagram of the logical relationship from the sepsis treatment actions provided according to clinical experience to the treatment outcomes provided by an embodiment of the present invention;
[0019] Figure 4 It is a schematic diagram of a change trend curve for evaluating the decision-making effectiveness of an agent at a single time point provided by an embodiment of the present invention;
[0020] Figure 5 It is a survival curve diagram of different policy groups for evaluating the decision-making effectiveness of an agent's sequential decisions provided by an embodiment of the present invention.
[0021] Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0022] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0023] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be understood as a limitation of the present invention. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.
[0024] In the description of the present invention, it should also be noted that unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0025] With the development of artificial intelligence technology and the improvement of medical informatization, it has become possible to optimize the dose combination in sepsis treatment decision-making by combining artificial intelligence methods and using the data in the intensive care unit. It is a common practice to optimize the sepsis-assisted decision-making process using reinforcement learning methods, but the related methods have not shown ideal optimization effects, and currently no work has proposed a relatively objective method for evaluating the decision-making effects of agents. In addition, when applying reinforcement learning in medical scenarios, the agent cannot interact with the real world and faces the problem of data sparsity. Therefore, this embodiment provides a method for constructing a reinforcement learning-assisted sepsis clinical decision-making model to solve the data sparsity problem and proposes corresponding methods to evaluate the decision-making effects of agents. This model can mine the best decisions for sepsis treatment from historical medical data, assist in identifying inappropriate decisions of clinicians, or provide data-driven auxiliary decision-making support for clinicians (such as inexperienced clinicians). In the following description of this application, the terms deep reinforcement learning algorithm, reinforcement learning algorithm, network, agent, and algorithm have the same meaning.
[0026] Figure 1 It is a flowchart of a method for constructing a reinforcement learning-assisted sepsis clinical decision-making model provided by an embodiment of the present invention. This method is executed by an electronic device, such as Figure 1 shown, and this method specifically includes:
[0027] S110. Using multiple physiological indicators of sepsis patients in historical clinical data as state variables and the dose combination of intravenous infusion and vasopressor actually taken by doctors as action variables, construct training samples of the reinforcement learning model. Among them, the multiple physiological indicators include fluid balance, blood pressure, blood protein, lactic acid, etc.
[0028] As described above, this embodiment will use an agent model based on reinforcement learning to provide clinical auxiliary decision-making. In this step, the physiological state of the patient at a certain moment and the dose combination actually taken by the doctor in this state are first extracted from historical clinical data to jointly form the training samples of the reinforcement learning model. Optionally, the MIMIC-III dataset can be used as historical clinical data.
[0029] The key variables in reinforcement learning include state variables and action variables. Optionally, for state variables, multiple physiological indicators of each sepsis patient at intervals of a set duration can be extracted from historical clinical data, and the multiple physiological indicators at the same moment are used as a state variable. For action variables, the doses of intravenous infusion and vasopressor can be divided into multiple intervals respectively, and the combination of each dose interval jointly constitutes the action variable space; then, the dose combinations actually taken by doctors under each state variable are extracted from historical clinical data, and the interval combination to which each dose combination belongs is used as each action variable. Each state variable and each action variable jointly constitute each training sample.
[0030] In a specific embodiment, data processing can be performed on 37 indicators available in the ICU to construct state variables. Optionally, data normalization is performed on physiological indicators, and 37 normalized indicator data every 4 hours are extracted to jointly form a 37-dimensional vital sign vector, and each vector is a state variable. At the same time, the dose of intravenous infusion and vasopressor can be discretized into intervals. The intravenous infusion (ml / 4h) is divided into 5 intervals according to [0, 0 - 50, 50 - 150, 150 - 500, >500], and the vasopressor (mcg / kg / 4h) is divided into 5 intervals according to [0, 0 - 0.08, 0.08 - 0.2, 0.2 - 0.45, >0.45]. There are a total of 25 combinations of their dose intervals, which are used as the action space for training the reinforcement learning algorithm. The dose interval combinations of intravenous infusion and vasopressor can be represented by a digital combination of 0 - 4 as (0, 0) - (4, 4). Among them, (0, 0) means that the doses of both are 0, and (4, 4) means that both are in the maximum dose interval; the actual dose combinations taken by the doctor under each state variable are extracted from historical clinical data, and the interval combination to which each dose combination belongs is used as each action variable.
[0031] S120. Build an intelligent agent model based on reinforcement learning according to the training samples. The model includes a feature extraction layer, an embedding processing layer, and a Q-value calculation layer.
[0032] In this step, a reinforcement learning model as shown in Figure 1 is built according to the data dimension of the training samples. As shown in Figure 1 , the model includes a feature extraction layer, an embedding processing layer, and a Q-value calculation layer. Among them, the feature extraction layer is used to extract the embedding features of the input data; the embedding processing layer is used to further process the extracted embedding features, continue to learn the depth information related to the model output, and obtain the final embedding features. The Q-value in reinforcement learning refers to the expected return obtained after taking a specific action in a given state, that is, to evaluate the "good or bad" of executing a specific action in the current state; the Q-value calculation layer in the model will evaluate the goodness or badness of each action variable according to the final embedding features, providing a basis for reinforcement learning.
[0033] Among them, the input / output dimensions of the model need to match the training samples. Exemplarily, in the case of 37 state variables and 25 action variables, the input dimension of the first layer of the model should be 37, and the output dimension of the last layer should be 25. Optionally, the entire model can adopt a fully connected structure; among them, the feature extraction layer includes 2 fully connected layers, and the number of nodes in each fully connected layer is 37 -> 128 -> 256 in sequence; the embedding processing layer includes 1 fully connected layer, and the number of nodes is 256 -> 256; the Q-value calculation layer includes 2 fully connected layers, and the number of nodes in each fully connected layer is 256 -> 128 -> 25.
[0034] It should be noted that only the model structure is built in this step, and the parameters in the model will be determined through subsequent training processes. The trained model is used to recommend the optimal sepsis medication plan according to the patient's physiological state at a certain moment.
[0035] S130. Use the training samples to train the model through reinforcement learning. The trained model is used to recommend the best dose combination according to multiple physiological indicators of the patient at a certain moment.
[0036] Based on the above model structure, in this step, the specific parameters of each layer of the model are determined through reinforcement learning. Optionally, during the training process, after the state variables are input into the feature extraction layer, feature embeddings are obtained; after the feature embeddings pass through the embedding processing layer, final embeddings are obtained; the final embeddings will pass through the Q-value calculation layer to obtain the Q-values of each action variable in the action variable space. Then, according to the principle of reinforcement learning, according to the new state variable s' reached after taking the action variable a under the state variable s (for example, the new state variable s' reached 4 hours after a doctor actually takes a certain action variable under the state variable s), the reward function r is calculated; and according to the reward function, the Q-value of the action variable is updated using the following formula:
[0037] Q'(s,a)=(1-α)×Q(s,a)+α×(r+γ×max(Q(s',a * ))) (1)
[0038] Among them, Q(s,a) represents the Q-value of the action variable a under the state variable s, Q'(s,a) represents the updated Q-value of the action variable a under the state variable s, max(Q(s',a * )) represents the largest one among all Q-values under the new state variable s', and α and γ are hyperparameters.
[0039] Finally, according to the difference between the expected value Q'(s,a) and the actual value Q(s,a), the following Q-value loss function is constructed:
[0040]
[0041] Among them, y i Corresponding expected value Q'(s,a), y i Corresponding to the actual value Q(s,a), this loss function combines the advantages of L1 and L2 loss functions: when the error is large, it is similar to L1 loss to reduce the sensitivity of outliers, and when the error is small, it turns to L2 loss, providing a more stable and robust training process. This makes it particularly suitable for processing noisy data or scenarios that require a smoother gradient descent process in prediction tasks. According to this loss function, the network parameters of the feature extraction layer, embedding processing layer, and Q value calculation layer can be updated.
[0042] In one embodiment, in combination Figure 2 , with state s i Take the example to illustrate the process of training a deep reinforcement learning algorithm. Specifically, s i It is a 37-dimensional vector, and after feature extraction, we get s i The feature embedding f(s i ), f(s i ) After the embedding processing layer of the network, the final embedding E(f(s i )). Finally, the network passes E(f(s i )) Calculate the Q value of each action, and the vector Q(f(s i )), for example, there are 25 actions in total, so Q(f(s i )) is a 25-dimensional vector. After the agent training is completed, for a certain state, the agent will select the corresponding action according to the maximum Q value. During the training process, after the agent takes action a from state s, it reaches a new state s' and obtains a reward r; substituting r into formula (1), the Q value can be updated; according to the difference between the updated Q value and the actual Q value, the iterative update of the model parameters can be completed. The remaining details are all existing technologies in the reinforcement learning method and will not be repeated here. This embodiment mainly focuses on the construction of the reward function in reinforcement learning and the response to data sparsity in the sepsis treatment scenario.
[0043] In a specific embodiment, the following operations can be added during reinforcement training to address the problem of data sparsity: After inputting the state variable of any training sample into the feature extraction layer, a feature embedding is obtained; a random perturbation is superimposed on the feature embedding to serve as the feature embedding of the positive sample; other training samples in the same batch as the said any training sample are used as negative samples, and after inputting the state variables of the negative samples into the feature extraction layer, the feature embeddings of the negative samples are obtained. For the convenience of distinction and description, in this embodiment, the feature embedding of the state variable is referred to as the first feature embedding, the feature embedding of the positive sample is referred to as the second feature embedding, and the feature embedding of the negative sample is referred to as the third feature embedding. During model training, by constraining the final embeddings obtained after the first feature embedding and the second feature embedding pass through the embedding processing layer to be close to each other, and the final embeddings obtained after the first feature embedding and the third feature embedding pass through the embedding processing layer to be far from each other, the problem of data sparsity in the sepsis treatment scenario can be addressed.
[0044] Specifically, referring to Figure 2 , for a certain state s i after passing through the feature extraction layer, a feature embedding f(s i ) is obtained. A random perturbation is superimposed on the feature embedding to obtain the feature embedding f'(s i ) of the positive sample; the remaining training data in the same batch are used as negative samples f(s j ); for this state s i and the positive and negative samples, after passing through the embedding processing layer of the network, the final embeddings E(f(s i ))、E(f'(s i )) and E(f(s j )) are obtained respectively. Constraints are added at the final embeddings to prompt the final embedding of the positive sample to be close to this state, and the final embedding of the negative sample to be far from this state, so that the feature extraction layer and the embedding processing layer of the network have the ability to handle the data sparsity situation in the sepsis treatment scenario, making the reinforcement learning algorithm more stable.
[0045] Optionally, the following loss function can be constructed:
[0046]
[0047] where sim(x,y) is a function representing the similarity between x and y, using cosine similarity; τ is a hyperparameter, and N is the number of samples in the batch. After updating the network parameters of the feature extraction layer, the embedding processing layer, and the Q-value calculation layer according to the Q-value loss function, the parameters of the feature extraction layer and the embedding processing layer are adjusted by minimizing L c , that is, the final model parameters at the current training step are obtained.
[0048] In a specific embodiment, to construct a reward function that better conforms to the clinical effect, a logical relationship diagram from action variables to treatment outcomes can be provided to the electronic device, and the electronic device automatically screens the main factors in the reward function according to this diagram. Specifically, Figure 3 is a schematic diagram of the logical relationship from the treatment actions of sepsis to the treatment outcomes provided based on clinical experience. Different legends in the figure are used to distinguish action variables, outcomes, observable variables, and unobservable variables. Among them, each observable variable refers to the physical sign changes that the action variable will cause, the observable variables are the observable variables that can be obtained in the ICU, that is, a part of the above 37 indicators, and the unobservable variables are the observable variables that cannot be directly obtained in the ICU. The arrow direction in the figure represents the influence direction. For example, intravenous infusion affects the fluid balance volume, where the fluid balance volume refers to the total net input fluid volume calculated from the time of admission. Based on the above logical relationship diagram, the observable variables in the figure can be first extracted as the basic factors for constructing the reward function; then, the observable variable with the most arrows pointing to it is selected from the basic factors as the core factor of the reward function, and the proportion of this factor is strengthened. Exemplarily, in combination with Figure 3 , the observable variables include fluid balance volume, blood pressure, blood protein, and lactate, so these 4 items can all be used as the basic factors of the reward function; among them, the number of arrows pointing to the fluid balance volume and blood pressure is the largest. Therefore, the proportion of the fluid balance volume and blood pressure in the dose selection can be increased in the reward function, and the proportion of blood protein and lactate can be suppressed.
[0049] Optionally, the reward function includes:
[0050] R im = a×ΔMAP + b×ΔFB + c×ΔALB + d×ΔLAC(4)
[0051] where R im represents the immediate reward, ΔMAP represents the change in blood pressure brought about by the action variable, ΔFB represents the change in fluid balance volume brought about by the action variable, ΔALB represents the change in blood protein brought about by the action variable, and ΔLAC represents the change in lactate brought about by the action variable. a, b, c, and d are adjustable parameters (i.e., weights) respectively. In this embodiment, the values of a, b, c, and d will be adjusted to increase the effects of blood pressure and fluid balance volume. Optionally, it is generally considered that R imIt is more appropriate that the value ranges in the interval [-1, 1]. Therefore, in this embodiment, the absolute value of ΔMAP every four hours in the clinical data is statistically analyzed, and the median of these absolute values, median{|ΔMAP|}, is taken. Similarly, the absolute values of ΔFB, ΔALB, and ΔLAC every four hours in the clinical data are statistically analyzed respectively, and the medians of these absolute values, median{|ΔFB|}, median{ΔALB}, and median{|ΔLAC|}, are taken respectively. Then, a, b, c, and d are adjusted so that the differences between |a×median{|ΔMAP|}| and |b×median{|ΔFB|}| and 1 are less than the set threshold (i.e., the two are close to 1), and the differences between |c×median{ΔALB}| and |d×median{|ΔLAC|}| and 0.1 are less than the set threshold (i.e., the two are close to 0.1), thereby determining the values of a, b, c, and d. Among them, in order to make each item of data more objective, the clinical data in the median statistics can be new data from the same database as the training samples but different from the training samples.
[0052] Meanwhile, the reward function may further include:
[0053]
[0054] Wherein, R te represents the terminal reward, R is a constant reward value, and Sur represents the survival days of the patient. R te will only be added to the reward function r when the treatment ends (discharge or death). R te is used to guide the balance between immediate treatment and long-term treatment of the medicament. In actual training, the above immediate reward can be adopted, or the combination of the above immediate reward and the terminal reward can be adopted. This embodiment does not make specific restrictions.
[0055] In reinforcement learning, the reward r plays a fundamental guiding role for the agent. Based on the logical relationship diagram in the process of sepsis treatment, this embodiment screens out two important concepts, the fluid balance volume and blood pressure, through the influence relationship arrows, greatly improving the decision-making accuracy of the agent.
[0056] Furthermore, after the agent is trained, this embodiment also provides two implementation methods to evaluate the effect of the strategy recommended by the agent:
[0057] The first implementation method is to evaluate whether the decision (dose combination) given by the agent for the state of the patient at a certain moment is effective. Specifically, it may include the following steps:
[0058] Step 1: For the state variables of a patient at the same moment, extract the dose combinations actually taken by the doctor for this state variable from the historical clinical data, and at the same time use the trained model to recommend the dose combination for this state variable. For the convenience of distinction and description, in this embodiment, the dose combination actually taken by the doctor for this state variable is called the first dose combination, and the dose combination recommended by the trained model for this state energy is called the second dose combination. Perform the above operations for multiple state variables in the historical clinical data respectively, and the first dose combination and the second dose combination of each state variable can be obtained.
[0059] Step 2: Calculate the distances between the first dose combination and the second dose combination of each state variable respectively. The distance is used to characterize the difference between the clinical decision and the agent decision at the same decision moment. Optionally, the decision a t of the dose combination (a ti , a tv ) at time t can be represented as a point on the two-dimensional coordinate plane. Then the Euclidean distance between the coordinates of two different decisions represents the dose difference. Specifically, the distance between the first dose combination a t and the second dose combination b t (b ti , b tv ) can be calculated by the following formula:
[0060]
[0061] where a ti , a tv , b ti and b tv are all integers between 0 and 4.
[0062] Step 3: Group the multiple state variables according to the distances of each state variable, so that the distances of the state variables within the same group are controlled within a set range. That is, the state variables with relatively close distances are grouped into one group. Optionally, since both a t and b t include 25 types, then the number of value types of is also fixed, and the state variables corresponding to the same value of
[0063] Step 4: According to the survival time of each patient in the historical clinical data, calculate the survival rate of patients for each group of state variables. Optionally, for any group of state variables, based on the survival time of the patients corresponding to each state variable within the group, count the number N1 of state variables with a patient survival time exceeding 90 days. Calculate the 90-day survival rate of the patients in this group based on the ratio of N1 to the total number N2 of all state variables within the group. After performing the above operations for all groups, the 90-day patient survival rates of each group can be obtained. Of course, the 90-day survival rate of the patients in this group can also be calculated based on the number of patients with a survival time exceeding 90 days within the group and the total number of patients in the group. Which statistical method to use depends on the number of patients in the group. If there are more patients, the latter method can be used; if the number of patients is small, the former method can be used.
[0064] Step 5: Using the distance and patient survival rate of each group as data points, fit a trend curve of the patient survival rate changing with the distance. Optionally, for any one group, use the average value of the distances of all state variables within the group and the patient survival rate as a data point, and each group corresponds to one data point. In particular, if in Step 1, the state variables corresponding to the same value are grouped together, then the average value here is itself. Based on these data points, a trend curve of the patient survival rate changing with the distance can be fitted. Figure 4 Exemplarily shows a trend line of the 90-day mortality rate changing with the distance, where the abscissa "the policy distance between the doctor and the agent" is the distance 90-day mortality rate = 1 - 90-day survival rate. The original points represent the above data points, and the trend line is the data line fitted based on these data points.
[0065] Step 6: If the greater the distance in the trend curve, the lower the patient survival rate, then it is considered that the best dose combination recommended by the trained model is effective. Referring to Figure 4 , in the trend line, when the decision distance between the doctor and the agent for the same state is greater, the corresponding patient mortality rate is higher (the survival rate is lower). Then it is considered that the decision given by the agent is effective. The second implementation manner is to evaluate whether the decision sequence (dose combination sequence) given by the agent for the entire treatment process of the patient is effective, which can specifically include the following steps:
[0066] Step 1: From the historical clinical data of the same patient from admission to the end of treatment, extract the dose combinations adopted by the doctor at each decision-making moment, and form a decision sequence in chronological order by each dose combination; at the same time, use the trained model to recommend the optimal dose combination of the same patient at each decision-making moment, and form another decision sequence in chronological order by each optimal dose combination. In this embodiment, the process of a patient from admission to discharge or death is regarded as a treatment trajectory. In a treatment trajectory with a length of T, the decisions made by the doctor form a sequence along time, which is called the first decision sequence, and the decisions recommended by the model form another sequence along time, which is called the second decision sequence. After performing the above operations for multiple patients respectively, the first decision sequence and the second decision sequence of each patient can be obtained.
[0067] Step 2: Calculate the distances between the second decision sequences and the first decision sequences of multiple patients respectively. The distances are used to characterize the differences between clinical decisions and agent decisions during the treatment process. Optionally, for each patient, calculate the distance between two decision sequences A and B respectively according to the following formula:
[0068]
[0069] where T represents the number of decision-making moments in the sequence, represents the distance between the dose combinations a t and b t of decision sequences A and B at the same decision-making moment t, and is calculated using formula (6).
[0070] Step 3: Group the multiple patients according to the distance D A-B so that the distances of the patients within the same group are controlled within a set range. That is, group the patients with relatively close distances D A-B into one group.
[0071] Step 4: Fit the survival curves of the patients in each group according to the survival times of the patients within the same group. Among them, the survival curve is a curve with time as the independent variable and patient survival rate as the dependent variable. Optionally, for any group of patients, according to the survival times of the patients in the group, count the number of patients N3 who survive at different treatment days, and divide N3 by the total number of patients N4 in the group, then the patient survival rate at different treatment days can be obtained, and thus the survival curve can be obtained. Perform the above operations for all groups respectively, and the survival curve of each group can be obtained. Figure 5 Exemplarily shows the survival rate curves of each group.
[0072] Step 5: If the greater the distances of each group are and the faster the survival curves of each group decline, it is considered that the optimal dose combinations recommended by the trained model are effective. Combining Figure 5, in the figure, the sequence distance intervals between doctors and agents in each group are shown by different legends in the upper right corner. It can be seen that the greater the distance, the faster the survival rate curve of the corresponding group drops, indicating that the sequence decisions given by the agent are effective.
[0073] The above two evaluation methods can exist alone or simultaneously. Only when the evaluation results show that the strategies and / or strategy sequences given by the model are effective can the agent be used to identify inappropriate decisions of clinicians or provide auxiliary decision-making suggestions for inexperienced clinicians. For example, when the distance between the decisions of a clinician and the agent is greater than the set threshold, a reminder can be given, and the clinician can reconsider whether the dose combination is reasonable based on the reminder. In addition, since the accumulation of clinical data takes a long time, this embodiment uses historical clinical data to verify the effectiveness of the agent's strategy, and the survival rate therein is also statistically obtained based on the real data of historical cases, and the evaluation results are more objective.
[0074] In summary, this embodiment provides a method for constructing a reinforcement learning-assisted sepsis clinical decision-making model. By using reinforcement learning technology, an agent is trained to select the dose combination of intravenous infusion and vasopressor for the state of a patient at a certain moment. By learning excellent decisions in historical treatment data, bad decisions are optimized, and an objective evaluation method for auxiliary decision-making based on historical treatment data is proposed. In particular, in this embodiment, positive samples are constructed by adding random perturbations in reinforcement learning and jointly training the reinforcement learning algorithm with negative samples in the same batch to solve the data sparsity problem in the scenario of reinforcement learning-assisted sepsis clinical decision-making and improve the stability of the agent. At the same time, this embodiment screens and constructs the main factors of the reward function according to the medical logic diagram to solve the problem of inappropriate reward functions in existing work. In addition, this embodiment proposes a method for evaluating the effectiveness of the agent's decision at a single time point and the agent's sequence decision, and the effectiveness of the agent's decision can be judged based on historical clinical data.
[0075] Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. As Figure 6 shown, the device includes a processor 60, a memory 61, an input device 62, and an output device 63; the number of processors 60 in the device can be one or more, Figure 6 taking one processor 60 as an example; the processor 60, the memory 61, the input device 62, and the output device 63 in the device can be connected by a bus or other means, Figure 6 taking the connection by bus as an example.
[0076] The memory 61, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the method for constructing a reinforcement learning-assisted sepsis clinical decision-making model in the embodiments of the present invention. The processor 60 executes various functional applications and data processing of the device by running the software programs, instructions, and modules stored in the memory 61, that is, to implement the above-mentioned method for constructing a reinforcement learning-assisted sepsis clinical decision-making model.
[0077] The memory 61 may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the terminal, etc. In addition, the memory 61 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory 61 may further include a memory remotely set relative to the processor 60, and these remote memories can be connected to the device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0078] The input device 62 can be used to receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the device. The output device 63 may include a display device such as a display screen.
[0079] The embodiments of the present invention also provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method for constructing a reinforcement learning-assisted sepsis clinical decision-making model in any embodiment.
[0080] The computer storage medium of the embodiments of the present invention can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.
[0081] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal may take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.
[0082] The program code contained on a computer-readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wire, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0083] The computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the C language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0084] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.
Claims
1. A method for constructing a sepsis clinical decision-making model assisted by reinforcement learning, characterized in that Including: Taking multiple physiological indicators of sepsis patients in historical clinical data as state variables and the dose combination of intravenous infusion and vasopressor actually taken by doctors as action variables to construct training samples for the reinforcement learning model; Building an agent model based on reinforcement learning according to the training samples, and the model includes a feature extraction layer, an embedding processing layer, and a Q-value calculation layer; Using the training samples to train the model by reinforcement learning, and the trained model is used to recommend the best dose combination according to multiple physiological indicators of the patient at a certain moment; Among them, the multiple physiological indicators include fluid balance, blood pressure, blood protein, and lactic acid; during the training process, the effects of fluid balance and blood pressure in dose selection are enhanced through the reward function, and the effects of blood protein and lactic acid are inhibited; After training the model by the reinforcement learning method, it further includes: For the state variables of the patient at the same moment, extracting the first dose combination actually taken by the doctor from the historical clinical data, and using the trained model to recommend the second dose combination under the state variables; Calculate the distances between the first dose combination and the second dose combination of multiple state variables respectively, where the distances are used to characterize the differences between clinical decisions and agent decisions at a single decision-making moment; specifically, the decision at time t of the dose combination ( , ) is represented as a point on a two-dimensional coordinate plane, and the Euclidean distance between the coordinates of two different decisions represents the dose difference; Grouping the multiple state variables according to the distance, so that the distance within each group of state variables is controlled within a set range; According to the survival time of each patient in the historical clinical data, counting the patient survival rate of each group of state variables; Taking the distance and patient survival rate of each group as data points to fit the change trend curve of the patient survival rate with respect to the distance; If the greater the distance in the change trend curve, the lower the patient survival rate, it is judged that the best dose combination recommended by the trained model is effective.
2. The method according to claim 1, wherein The construction of the training samples by taking multiple physiological indicators of sepsis patients in historical clinical data as state variables and the dose combination of intravenous infusion and vasopressor actually taken by doctors as action variables includes: Extracting multiple physiological indicators of each sepsis patient at regular intervals from the historical clinical data, and taking the multiple physiological indicators at the same moment as state variables; Dividing the doses of intravenous infusion and vasopressor into multiple intervals respectively, and the combination of each dose interval jointly constitutes an action variable space; Extracting the dose combinations actually taken by doctors under each state variable from the historical clinical data, and taking the interval combination to which each dose combination belongs as each action variable; Each state variable and each action variable jointly constitute each training sample.
3. The method according to claim 1, wherein The training of the model by using the training samples through the reinforcement learning method includes: Inputting the state variables in the training samples into the feature extraction layer to obtain feature embeddings; Inputting the feature embeddings into the embedding processing layer to obtain final embeddings; Inputting the final embeddings into the Q-value calculation layer to obtain the Q-values of each action variable in the action variable space; According to the state variable take the action variable and reach the new state variable , calculate the reward function ; According to the reward function, updating the Q-value of the action variable by using the following formula to update the model parameters: ; Among them, represents the Q value of the action variable under the state variable ; represents the updated Q value of the action variable under the state variable ; represents the largest one among all Q values under the new state variable ; is related to which are hyperparameters.
4. The method according to claim 1, wherein The training of the model by using the training samples through the reinforcement learning method includes: After inputting the state variables of any training sample into the feature extraction layer, obtaining the first feature embedding; Superimposing random perturbations on the first feature embedding as the second feature embedding of the positive sample; Use other training samples in the same batch as any of the above-mentioned training samples as negative samples. After inputting the state variables of the negative samples into the feature extraction layer, obtain the third feature embedding. Update the model parameters by constraining the final embeddings obtained after the first feature embedding and the second feature embedding pass through the embedding processing layer to be close to each other, and the final embeddings obtained after the first feature embedding and the third feature embedding pass through the embedding processing layer to be far from each other.
5. The method according to claim 1, characterized in that, The reward function includes: ; Among them, R im represents an immediate reward, Δ MAP represents the blood pressure change brought about by the action variable, Δ FB represents the change in fluid balance brought about by the action variable, Δ ALB represents the change in blood protein brought about by the action variable, Δ LAC represents the change in lactic acid brought about by the action variable; a, b, c and d The value is determined in the following manner: Statistically analyze the | Δ MAP|, | Δ FB|, | Δ ALB|, and | Δ LAC| in the clinical data at each set time interval, and respectively take their medians median{| Δ MAP|}, median{| Δ FB|}, median{ Δ ALB}, and median{| Δ LAC|}; Adjust a, b, c and d so that a × median{| Δ MAP|} and b × median{| Δ FB|} have absolute differences from 1 that are respectively less than the set threshold, and make c × median{ Δ ALB} and d × median{| Δ LAC|} have absolute differences from 0.1 that are respectively less than the set threshold.
6. The method according to claim 5, characterized in that The reward function further includes: Among them, represents the end reward, which is used to be added to the reward function at the end of the treatment. R is a constant reward value. represents the number of days the patient has survived.
7. The method according to claim 1, characterized in that, After training the model by means of reinforcement learning, it further includes: Extract the first dose combination taken by the doctor at each decision-making moment from the historical clinical data of the same patient from admission to the end of treatment. The first decision sequence is composed of the first dose combinations in chronological order. Use the trained model to recommend the best dose combination for the same patient at each decision-making moment. The second decision sequence is composed of the best dose combinations in chronological order. Calculate the distances between the second decision sequences and the first decision sequences of multiple patients respectively. The distances are used to characterize the differences between clinical decisions and agent decisions during the treatment process. Group the multiple patients according to the distances, so that the distances of patients within the same group are controlled within a set range. Fit the survival curves of patients in each group according to the survival times of patients within the same group. Among them, the survival curve takes time as the independent variable and the patient survival rate as the dependent variable. If the greater the distances of each group are, the faster the survival curves of each group decline, it is determined that the best dose combination recommended by the trained model is effective.
8. An electronic device, characterized in that, It includes: One or more processors; A memory for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method for constructing a reinforcement learning-assisted sepsis clinical decision-making model according to any one of claims 1-7.
9. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the program is executed by a processor, it implements the method for constructing a reinforcement learning-assisted sepsis clinical decision-making model according to any one of claims 1-7.
Citation Information
Patent Citations
Method and device for learning sepsis treatment strategy
CN114330566A