Propofol drug delivery method based on deep learning method pharmacokinetics-pharmacodynamics and reinforcement learning and application of propofol drug delivery method
Through the pharmacokinetic-pharmacokinetic model combined with deep learning and reinforcement learning, the problem of how to timely adjust the dose of propofol according to human physiological indicators is solved, and precise control of the depth of anesthesia and personalized treatment is achieved.
Patent Information
- Application Number
- CN202510040957.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-01-10
Smart Images

Figure CN119964657A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent medical treatment, and in particular to a propofol administration method based on pharmacokinetics-pharmacodynamics and reinforcement learning of a deep learning method and an application thereof. Background Art
[0002] Pharmacokinetics (PK) is the science of studying the absorption, distribution, metabolism and excretion of drugs in the body. It is of great significance to the rational use of drugs and the individualization of drug treatment. Pharmacokinetics reveals the mechanism of action of drugs and can characterize the biotransformation process of drugs, such as the absorption, distribution, metabolism and excretion of drugs in the body. By revealing the mechanism of action of drugs, we can better understand the pharmacological effects of drugs, thereby guiding the rational use of drugs. The effect and safety of drug treatment largely depend on the kinetic process of drugs in the body. Pharmacokinetics can measure the blood concentration-time curve of drugs. By analyzing the pharmacokinetic parameters of drugs, the dosage and administration regimen of drugs can be inferred. For example, for drugs with dose dependence, understanding their pharmacokinetic characteristics can reasonably adjust the dosage of drugs to avoid overdose and side effects. In addition, pharmacokinetics can also guide the formulation of drug combination and individualized drug treatment plans to improve the effect of drug treatment. Pharmacokinetics can predict the efficacy and toxicity of drugs by establishing a pharmacokinetic model of drugs. By measuring the concentration of the drug in the body, the maximum effect and half-maximal effect concentration (EC50) of the drug can be estimated, thereby predicting the therapeutic effect of the drug. In addition, pharmacokinetics can indicate the potential toxicity and side effects of the drug, provide an evaluation basis for the clinical application of the drug, and reduce the occurrence of adverse reactions. Pharmacokinetics can reveal the differences between drugs in different individuals and provide a basis for personalized drug therapy. Different individuals have different abilities to metabolize, absorb, and excrete drugs, and these differences may lead to different responses to drugs between individuals. By using pharmacokinetics, we can understand the type of drug-metabolizing enzymes of the drug and judge the individual's ability to metabolize the drug, so as to adjust the drug dosage and administration regimen in an individualized manner and improve the therapeutic effect.
[0003] In clinical medicine, general anesthesia is often required for patients during surgery, and sedatives such as propofol are required. Sedatives such as propofol are dose-dependent, and the dosage and rate of administration will directly affect the effect of anesthesia. Therefore, the dosage and rate of administration need to be controlled within a reasonable range, and the dosage can be dynamically changed according to the depth of anesthesia. When the depth of anesthesia is deeper than the expected depth, the dosage can be reduced in time, and when the depth of anesthesia is shallower than the expected depth, the dosage can be increased in time to ensure that the depth of anesthesia can be stabilized within the expected range. When evaluating the depth of anesthesia of a patient, the most important reference indicator is the Bispectral Index Scale (BIS), which can directly reflect the patient's current level of brain wakefulness. Based on this indicator, it can be determined whether the current patient needs to increase or decrease the dosage of sedatives to ensure that the patient's depth of anesthesia during the operation remains within a reasonable range.
[0004] CN101247809A discloses a method for quantitative administration of propofol prodrug for inducing mild to moderate sedation level, wherein the propofol prodrug dosage required for inducing mild to moderate sedation level of the patient is calculated based on the patient's lean body weight, the dosage commensurate with the patient's weight is determined, and then the coefficient is adjusted based on the age. For example, for a patient aged 60 or older, the dosage required to produce a sedative state or other effect may be about 0.6-0.8 times the dosage required to produce the corresponding effect for a younger patient of the same weight.
[0005] As a subfield of machine learning, reinforcement learning (RL) aims to enhance the behavioral decision-making ability of intelligent agents by using interaction experience and evaluation feedback with the world. Reinforcement learning algorithms can be divided into two categories: value-based reinforcement learning and policy-based reinforcement learning. The proximal policy optimization (PPO) algorithm adopted in this invention is a reinforcement learning algorithm proposed by OpenAI in 2017. It is considered to be the SOTA method in the field of reinforcement learning and one of the most widely applicable algorithms. Policy-based reinforcement learning no longer determines the strategy for selecting actions through a value function, but directly learns the strategy itself, parameterizes the strategy through a set of parameters θ, and optimizes θ through a neural network method. Policy-based reinforcement learning uses parameterized probability distribution π θ (a|s)=P(a|s;θ) replaces the deterministic policy π:s→a in value-based reinforcement learning, sampling different actions from the returned action probability list. Unlike traditional supervised learning methods (which usually rely on one-time, exhaustive, and supervisory signals), reinforcement learning simultaneously handles the sequential decision-making problems of sampling, evaluation, and delayed feedback.
[0006] In summary, how to make timely adjustments to the dosage of the anesthetic drug propofol according to human physiological indicators has become one of the problems that need to be solved urgently in this field. Summary of the invention
[0007] To solve the above technical problems, the present invention provides a propofol administration method based on pharmacokinetics-pharmacodynamics and reinforcement learning of deep learning methods and its application. According to the basic information of the patient such as height, weight, gender, etc., the BIS value can be calculated according to the discrete time series of the input administration rate data.
[0008] To achieve this object, the present invention adopts the following technical solutions:
[0009] In a first aspect, the present invention provides a propofol administration method based on pharmacokinetics-pharmacodynamics and reinforcement learning of a deep learning method, wherein the propofol administration method comprises the following steps:
[0010] (1) collecting basic information of the patient, including height, weight, gender and age, and calculating the patient's lean body mass based on the basic information;
[0011] (2) Based on the collected patient information and the pharmacokinetic properties of propofol, determine the parameters of the pharmacokinetic-pharmacodynamic model, including the volume, clearance rate, and transfer rate of each compartment;
[0012] (3) The pharmacokinetic-pharmacodynamic model includes three compartments, namely the central compartment, the fast peripheral compartment, and the slow peripheral compartment. Using the set parameters, a set of differential equations describing the concentration changes of propofol in the three compartments is constructed;
[0013] The pharmacokinetic-pharmacodynamic model is represented by a differential equation system, as shown in formula (1):
[0014]
[0015] In the formula represent the changes in drug concentration in the central compartment, fast distribution compartment, and slow distribution compartment, respectively; x1(t), x2(t), x3(t) represent the drug concentration in the central compartment, fast distribution compartment, and slow distribution compartment, respectively; u(t) is the infusion rate of the drug, and t represents time;
[0016] Constant K ij (i≠j) represents the drug transfer rate from compartment i to compartment j, where the central compartment is compartment 1, the fast distribution compartment is compartment 2, and the slow distribution compartment is compartment 3. For example, k 12 is the drug transfer rate from the central compartment to the rapid distribution compartment, k 10 is the administration rate of propofol to the central chamber;
[0017] The transfer rate and drug metabolism rate between the various compartments are expressed as follows:
[0018]
[0019] Where V1, V2, and V3 represent the volumes of the central chamber, the rapid distribution chamber, and the slow distribution chamber, respectively; C1, C2, and C3 represent the clearance rates of the central chamber, the rapid distribution chamber, and the slow distribution chamber, respectively;
[0020] The pharmacodynamic model introduces an additional hypothetical effect compartment to represent the site of action of the drug's efficacy. The formula is as follows:
[0021]
[0022] In the formula, is the drug concentration change in the effect compartment, x e (t) is the drug concentration in the effect chamber, x1(t) is the drug concentration in the central chamber, t is the time, and the constant k e is the drug transfer rate from the effect compartment to the central compartment;
[0023] (4) A mapping between intraoperative monitoring information and anesthesia depth is established based on a deep learning method, including using the drug infusion rate and the patient's covariate data as input variables, extracting the time-dependent parameter Lx and the time-continuous parameter Nx through LSTM and Neural ODE_GRU, establishing a nonlinear function based on the Kan network, and obtaining the predicted anesthesia depth by inputting the drug infusion rate, wherein the patient's covariates include the patient's age, gender, height, and weight.
[0024] Constructing an accurate and effective pharmacokinetic model to describe the metabolic process of propofol in the human body is an important basis for the accurate calculation of the anesthetic dosage. There are many methods for constructing a pharmacokinetic model. The present invention uses the three-compartment model proposed by Sheiner in 1979 to construct a pharmacokinetic model of propofol, such as Figure 1 As shown in FIG, the three-compartment pharmacokinetic-pharmacodynamic model consists of two parts: the pharmacokinetic (PK) model and the pharmacodynamic (PD) model, which are used to calculate and predict the optimal dose of a drug.
[0025] Among them, the PK model describes the transfer process and metabolism of drugs between three compartments, in which the fast peripheral compartment represents the direct site of action of the drug, such as blood. The drug concentration in this compartment is also called plasma concentration, while the slow peripheral compartment represents slower-acting tissues, such as muscle and fat.
[0026] In the PK model, u(t) refers to the infusion rate of the drug, usually expressed as u(t) / V1; V i (i=1,2,3) and C i (i=1, 2, 3) is related to the patient's lean body mass (LBM), which in turn depends on each person's height, weight and gender. The pharmacokinetic-pharmacodynamic model provided by the present invention can adjust parameters according to the patient's actual physiological indicators.
[0027] The PD model links plasma concentration to drug effects. Since blood, muscle tissue and organs are not direct manifestations of drug effects, an additional hypothetical effect compartment is introduced to represent the site of action of the drug's efficacy. The relationship between plasma concentration and effect compartment concentration is calculated based on the hypothetical effect compartment. Mathematical equations can be used to model the PD model and convert the effect compartment concentration into BIS for direct observation.
[0028] In the PD model, the constant k e It is the metabolic rate of drugs between the central compartment and the effect compartment, and its value is related to the type of anesthetic drugs.
[0029] The method provided by the present invention inputs the patient's height, weight, gender, and age data, and then inputs the propofol administration rate u(t), and the patient's BIS index can be calculated through the above PKPD model.
[0030] In actual applications, it is found that when an individual changes, a large amount of data is needed to optimize the parameters of the formula. The present invention uses a deep learning method to quickly establish a mapping between intraoperative monitoring information and anesthesia depth using a small amount of data.
[0031] Preferably, the input variable calculation method in step (4) comprises: calculating the concentration change of propofol in the three chambers and the drug concentration, as well as the change and drug concentration in the effect chamber by using the pharmacokinetic-pharmacodynamic model through the patient's covariate and the drug infusion rate, and combining the concentration change and drug concentration in the three chambers within 1 minute, as well as the change and drug concentration in the effect chamber with the patient's covariate to form the input variable X t .
[0032] Preferably, the input variable X t The calculation formula is as follows:
[0033]
[0034] Preferably, the Lx calculation method comprises: tIt is put into the BiLSTM model, and the nonlinear dependency in the time series is captured by BiLSTM. The conversion rules between different anesthesia stages are learned to obtain the time-dependent parameter Lx at this moment.
[0035] The LSTM (Gated Recurrent Unit) model is a recurrent neural network (RNN) architecture, and the calculation formula is as follows:
[0036] i t =σ(W ii x t +b ii +W hi h t-1 +b hi ) Formula (4.2);
[0037] f t =σ(W if x t +b if +W hf h t-1 +b hf ) Formula (4.3);
[0038] g t =tanh(W ig x t +b ig +W hg h t-1 +b hg ) Formula (4.4);
[0039] o t =σ(W io x t +b io +W ho h t-1 +b ho ) Formula (4.5);
[0040] c t =f t *c t-1 +i t *g t Formula (4.6);
[0041] h t =o t *tanh(c t ) Formula (4.7);
[0042] Here (W ii ,W hi ,W fi ,W hf ,W gi ,W hg,W io ,W ho and (b ii ,b hi ,b fi ,b hf ,b gi ,b hg ,b io ,b ho ) are the weight matrix and bias matrix, σ represents the sigmoid activation function, and tanh is the tangent function. t Determines how much of the current input is saved to the unit state; f t Determines how much of the unit state at the previous moment is retained to the current moment; g t Calculate the current state at the moment; o t Determines how much of the cell state is output to the current LSTM output value; c t To update the unit status; h t Send the filtered and updated information in the cell state to the hidden state of the next time step.
[0043] BiLSTM is an extension of the above LSTM model, where two LSTMs are applied to the input data. In the first round, LSTM is applied to the input sequence (i.e., the forward layer), and in the second round, the reverse form of the input sequence is input into the LSTM model (i.e., the backward layer). Applying LSTM twice can improve learning long-term dependencies, thereby improving the accuracy of the model. The calculation formula is as follows:
[0044]
[0045] Here (W f ,W r ) and (B bi ) are the weight matrix and bias matrix.
[0046] Preferably, the Nx calculation method comprises: t Put it into the Neural ODE_GRU model. The ODE_GRU-based model encodes the past trajectory as a variable and obtains the current time continuous parameter Nx. The calculation formula is as follows:
[0047]
[0048] Here GRUCell is the GRU unit, is the unit output, and its hidden variable h t The output is controlled by ODESolve; ODESolve is a numerical differential solver for neural network forward propagation; MLP is a fully connected network that passes through the dependent variable (h t-d ,...,ht ) Calculate Nx at this moment.
[0049] Preferably, the method for calculating the predicted anesthesia depth comprises: inputting Lx, Nx and the patient's covariates into the KAN network, and obtaining the predicted anesthesia depth by inputting the drug infusion rate The calculation formula is as follows:
[0050] Anesthetic_parameters t =[Lx t ,Nx t ,Age,Sex,Height,Weight] formula (4.15);
[0051] KAN(x)=(Φ3°Φ2°Φ1)(x) Equation (4.16);
[0052] Φ(x)=w(b(x)+spline(x)) Formula (4.17);
[0053]
[0054] Here Age, Sex, Height, Weight refer to the patient's age, sex, height, and weight; Φ refers to the different layers of the network; spline(x) represents the spline function. i are the coefficients optimized during training, and B i are the B-spline basis functions defined on the grid.
[0055] We divide the data set into a training set (80%) and a test set (20%) according to the number of surgical cases. The training set is input into the BIS model to obtain training parameters. The test set surgical data is divided internally, and 20% of the data is input into the model to fine-tune the parameters. This allows the model to be closer to the status of the patient being tested during the test process, avoiding the impact of patient differences.
[0056] In the present invention, the LSTM (Gated Recurrent Unit) model is a recurrent neural network (RNN) architecture, which has difficulty learning long-term dependencies. LSTM-based models are an extension of RNNs and can solve the gradient vanishing problem very cleanly. LSTM models essentially extend the memory of RNNs, enabling them to preserve and learn long-term dependencies of inputs. This memory extension is able to remember information for a longer period of time, allowing information to be read, written, and deleted from memory. LSTM memories are called "gated" units, where the word "gate" is inspired by the ability to make decisions to retain or ignore memory information. LSTM models capture important features from the input and preserve this information for a long time. The decision to delete or retain information is based on the weight value assigned to the information during training. Therefore, the LSTM model learns which information is worth retaining or deleting. BiLSTM is an extension of the above LSTM model, in which two LSTMs are applied to the input data. In the first round, LSTM is applied to the input sequence (i.e., the forward layer). In the second round, the reverse form of the input sequence is input into the LSTM model (i.e., the backward layer). Applying LSTM twice can improve learning long-term dependencies and thus improve the accuracy of the model.
[0057] In the present invention, compared with the ODE-based model based on the encoder-decoder structure, the model converted from the typical recurrent model can be predicted online at each time step. In contrast to the standard RNN, the ODE-based RNN learns the dynamics between observations, and ODE is suitable for inferring unknown physical information and can extract continuous features in the data. By extending the hidden state transfer in the RNN to the continuous time dynamics defined by the neural ODE. In our model, the transition of the hidden state and the transition between the latent states are calculated using GRU and Neural ODE.
[0058] In a second aspect, the present invention provides a propofol delivery device based on pharmacokinetics-pharmacodynamics and reinforcement learning of a deep learning method, wherein the propofol delivery device is used to perform the propofol delivery method described in the first aspect.
[0059] Preferably, the device comprises:
[0060] A vital sign data acquisition module is used to collect basic information of the patient, including height, weight, gender and age, and calculate the patient's lean body mass based on the basic information;
[0061] The drug administration control analysis module is used to determine the parameters of the pharmacokinetic-pharmacodynamic model, construct the pharmacokinetic-pharmacodynamic model, link the propofol drug effect concentration with the clinically observed anesthesia depth indicator BIS, and establish a mapping between intraoperative monitoring information and anesthesia depth based on deep learning methods.
[0062] Preferably, the drug administration control analysis module is specifically used to perform the following steps:
[0063] (2) Based on the collected patient information and the pharmacokinetic properties of propofol, determine the parameters of the pharmacokinetic-pharmacodynamic model, including the volume, clearance rate, and transfer rate of each compartment;
[0064] (3) The pharmacokinetic-pharmacodynamic model includes three compartments, namely the central compartment, the fast peripheral compartment, and the slow peripheral compartment. Using the set parameters, a set of differential equations describing the concentration changes of propofol in the three compartments is constructed;
[0065] The pharmacokinetic-pharmacodynamic model is represented by a differential equation system, as shown in formula (1):
[0066]
[0067] In the formula represents the drug concentration changes in the central chamber, fast distribution chamber and slow distribution chamber, respectively; x1(t), x2(t), x3(t) represent the drug concentrations in the central chamber, fast distribution chamber and slow distribution chamber, respectively; u(t) is the drug infusion rate; the constant K ij (i≠j) represents the drug transfer rate from compartment i to compartment j, and t represents time;
[0068] The transfer rate and drug metabolism rate between the various compartments are expressed as follows:
[0069]
[0070] Where V1, V2, and V3 represent the volumes of the central chamber, the rapid distribution chamber, and the slow distribution chamber, respectively; C1, C2, and C3 represent the clearance rates of the central chamber, the rapid distribution chamber, and the slow distribution chamber, respectively;
[0071] The pharmacodynamic model introduces an additional hypothetical effect compartment to represent the site of action of the drug's efficacy. The formula is as follows:
[0072]
[0073] In the formula, is the drug concentration change in the effect compartment, x e (t) is the drug concentration in the effect chamber, x1(t) is the drug concentration in the central chamber, t represents time, and the constant k e is the drug transfer rate from the effect compartment to the central compartment;
[0074] (4) A mapping between intraoperative monitoring information and anesthesia depth is established based on a deep learning method, including using the drug infusion rate and the patient's covariate data as input variables, extracting the time-dependent parameter Lx and the time-continuous parameter Nx through LSTM and Neural ODE_GRU, establishing a nonlinear function based on the Kan network, and obtaining the predicted anesthesia depth by inputting the drug infusion rate, wherein the patient's covariates include the patient's age, gender, height, and weight.
[0075] In a third aspect, the present invention provides a computer-readable storage medium having a computer program or instructions stored thereon, wherein the computer program or instructions are used to enable a processor to implement the propofol administration method as described in the first aspect when executed.
[0076] In a fourth aspect, the present invention provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program or instructions, and when the computer program or instructions are executed by the processor, the steps in the propofol administration method described in the first aspect are implemented.
[0077] Compared with the prior art, the present invention has the following beneficial effects:
[0078] The present invention provides a propofol administration method based on pharmacokinetic-pharmacodynamics and reinforcement learning of a deep learning method, using a PKPD model to simulate the metabolic process of propofol in the human body, extracting time dependency and dynamic continuous information through LSTM and Neural OED, and then establishing a nonlinear relationship between the propofol drug effect concentration and the clinically observed anesthesia depth index BIS through a Kan network. When performing accurate dosage calculation, a reinforcement learning algorithm is used, and the patient's current BIS, the patient's historical BIS value, and the administration rate in the previous period are used as the state of the environment of the reinforcement learning algorithm, and the administration rate at the next moment is used as the action of the reinforcement learning algorithm. The intelligent agent constructed by the reinforcement learning algorithm can gradually optimize its own action selection strategy in the process of interaction with the environment, and adjust the strategy in time according to the change of the environment, which is very suitable for handling non-standardized complex problems. Rapidly establish the mapping between intraoperative monitoring information and anesthesia depth, and adjust the strategy in time according to the change of the environment, which is suitable for handling non-standardized complex problems, and adjust the dosage and administration scheme of the drug in a personalized manner to improve the treatment effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] Figure 1 This is the structural diagram of the PKPD model.
[0080] Figure 2 It is a diagram of the propofol administration process of the present invention.
[0081] Figure 3This is a diagram for dividing the data set of the present invention.
[0082] Figure 4 This is a diagram of the BIS prediction results and PPF drug administration results of a certain surgery during the test process of the present invention.
[0083] Figure 5 This is a diagram of the control results of a certain surgery during the testing process of the present invention. DETAILED DESCRIPTION
[0084] To further illustrate the technical means and effects of the present invention, the present invention is further described below in conjunction with the embodiments and drawings. It should be understood that the specific implementation methods described herein are only used to explain the present invention, rather than to limit the present invention.
[0085] Example 1
[0086] This embodiment provides a method for administering propofol based on pharmacokinetic-pharmacodynamic and reinforcement learning, such as Figure 1 and Figure 2 As shown, the following steps are included:
[0087] (1) Construct a PKPD model and use this model as the environment of the reinforcement learning algorithm to facilitate the training of the intelligent agent. The PK model formula is as follows:
[0088]
[0089] In the formula represents the drug concentration changes in the central chamber, fast distribution chamber and slow distribution chamber, respectively; x1(t), x2(t), x3(t) represent the drug concentrations in the central chamber, fast distribution chamber and slow distribution chamber, respectively; u(t) is the drug infusion rate; the constant K ij (i≠j) represents the drug transfer rate from compartment i to compartment j, and t represents time;
[0090] The transfer rate and drug metabolism rate between the various compartments are expressed as follows:
[0091]
[0092] Where V1, V2, and V3 represent the volumes of the central chamber, the rapid distribution chamber, and the slow distribution chamber, respectively; C1, C2, and C3 represent the clearance rates of the central chamber, the rapid distribution chamber, and the slow distribution chamber, respectively;
[0093] The PD model formula is as follows:
[0094]
[0095] In the formula, is the drug concentration change in the effect compartment, xe (t) is the drug concentration in the effect chamber, x1(t) is the drug concentration in the central chamber, t is the time, and the constant k e0 is the drug transfer rate from the effect compartment to the central compartment.
[0096] (2) Collect the patient's basic information, including height, weight, gender and age, and calculate the patient's lean body mass using a specific formula. This parameter is crucial for the subsequent pharmacokinetic model parameter setting.
[0097] (3) Based on the collected patient information and the pharmacokinetic properties of propofol, the parameters of the three-compartment model were determined, including the volume, clearance rate, and transfer rate of each compartment. Using the set parameters, a set of differential equations describing the concentration changes of propofol in the three compartments (central compartment, rapid distribution compartment, and slow distribution compartment) was constructed to ensure accurate simulation of drug concentration.
[0098] (4) A mapping between intraoperative monitoring information and anesthesia depth is established based on a deep learning method, including using the drug infusion rate and the patient's covariate data as input variables, extracting the time-dependent parameter Lx and the time-continuous parameter Nx through LSTM and Neural ODE_GRU, and then inputting Lx, Nx and the patient's covariates into the Kan network to establish a nonlinear function. The patient's covariates include the patient's age, gender, height, and weight.
[0099] Wherein, step (4) specifically comprises the following steps:
[0100] The pharmacokinetic-pharmacodynamic model is used to calculate the concentration changes of propofol in the three compartments and the drug concentration, as well as the changes and drug concentration in the effect chamber, through the patient's covariates and the drug infusion rate. The concentration changes and drug concentrations in the three compartments within 1 minute, as well as the changes and drug concentrations in the effect chamber, are combined with the patient's covariates to form the input variable X t ;
[0101]
[0102] X t Put it into the BiLSTM model, capture the nonlinear dependency in the time series through BiLSTM, and learn the conversion rules between different anesthesia stages to obtain the time-dependent parameter Lx at this moment;
[0103] The LSTM (Gated Recurrent Unit) model is a recurrent neural network (RNN) architecture, and the calculation formula is as follows:
[0104] i t =σ(W ii x t +b ii +Whi h t-1 +b hi ) Formula (4.2);
[0105] f t =σ(W if x t +b if +W hf h t-1 +b hf ) Formula (4.3);
[0106] g t =tanh(W ig x t +b ig +W hg h t-1 +b hg ) Formula (4.4);
[0107] o t =σ(W io x t +b io +W ho h t-1 +b ho ) Formula (4.5);
[0108] c t =f t *c t-1 +i t *g t Formula (4.6);
[0109] h t =o t *tanh(c t ) Formula (4.7);
[0110] BiLSTM is an extension of the above LSTM model, where two LSTMs are applied to the input data. In the first round, LSTM is applied to the input sequence (i.e., the forward layer), and in the second round, the reverse form of the input sequence is input into the LSTM model (i.e., the backward layer). Applying LSTM twice can improve learning long-term dependencies, thereby improving the accuracy of the model. The calculation formula is as follows:
[0111]
[0112] At the same time, X t Put it into the Neural ODE_GRU model. The ODE_GRU-based model encodes the past trajectory as a variable and obtains the current time continuous parameter Nx. The calculation formula is as follows:
[0113]
[0114] N x (t) = MLP(h t-d ,...,h t ) formula (4.14);
[0115] Lx, Nx and the patient's covariates are input into the KAN network, and the predicted anesthesia depth is obtained by inputting the drug infusion rate. The calculation formula is as follows:
[0116] Anesthetic_parameters t =[Lx t ,Nx t ,Age,Sex,Height,Weight] formula (4.15);
[0117] KAN(x)=(Φ3°Φ2°Φ1)(x) Equation (4.16);
[0118] Φ(x)=w(b(x)+spline(x)) Formula (4.17);
[0119]
[0120] Test Example 1
[0121] This test example applies the propofol administration method provided in Example 1. First, the data is selected, and the age range is 20-60 years old. The intraoperative information records the real-time BIS and propofol administration rate. Then, the BIS and propofol administration are aligned in the time dimension, and the latest recorded time point is selected as the data starting point. Then, the PKPD model is used to calculate the drug concentrations in different organs of the patient according to the patient's age, height, weight and gender. Then, the sliding window technology is used with a 60s window length and a 1s step length to pre-process the drug concentrations in different parts of the propofol administration rate and the per-second changes of the two, and the BIS at the end of the window is used as the label. The data set is randomly divided according to the amount of surgery (the order is not disrupted within each surgery), and divided according to the training set and the test set being equal to 8:2. The data set division is as follows Figure 3 As shown. The model is trained using the training set. After the training is completed, the test set is used for testing. The model is fine-tuned using the first 20% of the data for each surgery, and the learning rate is adjusted to 10% of the original. After the training is completed, the remaining test set is tested. The test result MAE is 8.9, which is currently a more advanced algorithm. The visualization of a surgery in the test set is shown below. Figure 4As shown, we can see that the entire trend of BIS is almost fitted. The reinforcement learning part uses the Proximal Policy Optimization (PPO) proximal policy optimization algorithm, and the environment that interacts with the agent uses the BIS model and the three-compartment pharmacokinetic model. The observation space of the environment is BIS(t-2), BIS(t-1), BIS(t), PPF_v(t-1), the action space of the environment is the propofol rate (PPF_v(t)), and the reward function is set to The reward here is the difference between the current BIS value (BIS=50) and the current BIS value of reinforcement learning, which is normally distributed. Then the training and testing process is carried out. Figure 5 As the test results show, it can be seen that the target value can be maintained well.
[0122] In summary, the present invention provides a propofol administration method based on pharmacokinetics-pharmacodynamics and reinforcement learning of deep learning methods, uses a PKPD model to simulate the metabolic process of propofol in the human body, extracts time dependence and kinetic continuous information through LSTM and Neural OED, and then establishes a nonlinear relationship between the propofol drug effect concentration and the clinically observed anesthesia depth index BIS through a Kan network, quickly establishes a mapping between intraoperative monitoring information and anesthesia depth, and adjusts the strategy in time according to environmental changes. It is suitable for dealing with non-standardized complex problems, and can adjust the drug dosage and administration regimen in a personalized manner to improve the treatment effect.
[0123] The applicant declares that the above is only a specific implementation mode of the present invention, but the protection scope of the present invention is not limited thereto. Those skilled in the art should understand that any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention shall fall within the protection scope and disclosure scope of the present invention.
Claims
1. A method for administering propofol based on pharmacokinetic-pharmacodynamics and reinforcement learning based on deep learning methods, characterized in that: The propofol administration method comprises the following steps: (1) collecting basic information of the patient, including height, weight, gender and age, and calculating the patient's lean body mass based on the basic information; (2) Based on the collected patient information and the pharmacokinetic properties of propofol, determine the parameters of the pharmacokinetic-pharmacodynamic model, including the volume, clearance rate, and transfer rate of each compartment; (3) The pharmacokinetic-pharmacodynamic model includes three compartments, namely the central compartment, the fast peripheral compartment, and the slow peripheral compartment. Using the set parameters, a set of differential equations describing the concentration changes of propofol in the three compartments is constructed; The pharmacokinetic-pharmacodynamic model is represented by a differential equation system, as shown in formula (1): In the formula represents the drug concentration changes in the central chamber, fast distribution chamber and slow distribution chamber, respectively; x1(t), x2(t), x3(t) represent the drug concentrations in the central chamber, fast distribution chamber and slow distribution chamber, respectively; u(t) is the drug infusion rate; the constant K ij (i≠j) represents the drug transfer rate from compartment i to compartment j, and t represents time; The transfer rate and drug metabolism rate between the various compartments are expressed as follows: Where V1, V2, and V3 represent the volumes of the central chamber, the rapid distribution chamber, and the slow distribution chamber, respectively; C1, C2, and C3 represent the clearance rates of the central chamber, the rapid distribution chamber, and the slow distribution chamber, respectively; The pharmacodynamic model introduces an additional hypothetical effect compartment to represent the site of action of the drug's efficacy. The formula is as follows: In the formula, is the drug concentration change in the effect compartment, x e (t) is the drug concentration in the effect chamber, x1(t) is the drug concentration in the central chamber, t is the time, and the constant k e is the drug transfer rate from the effect compartment to the central compartment; (4) A mapping between intraoperative monitoring information and anesthesia depth is established based on a deep learning method, including using the drug infusion rate and the patient's covariate data as input variables, extracting the time-dependent parameter Lx and the time-continuous parameter Nx through LSTM and Neural ODE_GRU, establishing a nonlinear function based on the Kan network, and obtaining the predicted anesthesia depth by inputting the drug infusion rate, wherein the patient's covariates include the patient's age, gender, height, and weight.
2. The method for administering propofol according to claim 1, characterized in that: The input variable calculation method in step (4) includes: The pharmacokinetic-pharmacodynamic model is used to calculate the concentration changes of propofol in the three compartments and the drug concentration, as well as the changes and drug concentration in the effect chamber, through the patient's covariates and the drug infusion rate. The concentration changes and drug concentrations in the three compartments within 1 minute, as well as the changes and drug concentrations in the effect chamber, are combined with the patient's covariates to form the input variable X t .
3. The method for administering propofol according to claim 2, characterized in that: The input variable X t The calculation formula is as follows:
4. The method for administering propofol according to any one of claims 1 to 3, characterized in that: The Lx calculation method comprises: t It is put into the BiLSTM model, and the nonlinear dependency in the time series is captured by BiLSTM. The conversion rules between different anesthesia stages are learned to obtain the time-dependent parameter Lx at this moment.
5. The method for administering propofol according to any one of claims 1 to 4, characterized in that: The Nx calculation method comprises: t Put it into the Neural ODE_GRU model. The ODE_GRU-based model encodes the past trajectory into variables and obtains the current time continuous parameter Nx.
6. The method for administering propofol according to any one of claims 1 to 5, characterized in that: The method for calculating the predicted anesthesia depth includes: inputting Lx, Nx and the patient's covariate into a KAN network, and obtaining the predicted anesthesia depth by inputting the infusion rate of the drug.
7. A propofol drug delivery device based on deep learning methods of pharmacokinetics-pharmacodynamics and reinforcement learning, characterized in that: The propofol administration device is used to perform the propofol administration method according to any one of claims 1 to 6.
8. The device according to claim 7, characterized in that The device comprises: A vital sign data acquisition module is used to collect basic information of the patient, including height, weight, gender and age, and calculate the patient's lean body mass based on the basic information; The drug administration control analysis module is used to determine the parameters of the pharmacokinetic-pharmacodynamic model, build a pharmacokinetic-pharmacodynamic model, link the propofol drug effect concentration with the clinically observed anesthesia depth index BIS, and establish a mapping between intraoperative monitoring information and anesthesia depth based on deep learning methods; Preferably, the drug administration control analysis module is specifically used to perform the following steps: (2) Based on the collected patient information and the pharmacokinetic properties of propofol, determine the parameters of the pharmacokinetic-pharmacodynamic model, including the volume, clearance rate, and transfer rate of each compartment; (3) The pharmacokinetic-pharmacodynamic model includes three compartments, namely the central compartment, the fast peripheral compartment, and the slow peripheral compartment. Using the set parameters, a set of differential equations describing the concentration changes of propofol in the three compartments is constructed; The pharmacokinetic-pharmacodynamic model is represented by a differential equation system, as shown in formula (1): In the formula represents the drug concentration changes in the central chamber, fast distribution chamber and slow distribution chamber, respectively; x1(t), x2(t), x3(t) represent the drug concentrations in the central chamber, fast distribution chamber and slow distribution chamber, respectively; u(t) is the drug infusion rate; the constant K ij (i≠j) represents the drug transfer rate from compartment i to compartment j, and t represents time; The transfer rate and drug metabolism rate between the various compartments are expressed as follows: Where V1, V2, and V3 represent the volumes of the central chamber, the rapid distribution chamber, and the slow distribution chamber, respectively; C1, C2, and C3 represent the clearance rates of the central chamber, the rapid distribution chamber, and the slow distribution chamber, respectively; The pharmacodynamic model introduces an additional hypothetical effect compartment to represent the site of action of the drug's efficacy. The formula is as follows: In the formula, is the drug concentration change in the effect compartment, x e (t) is the drug concentration in the effect chamber, x1(t) is the drug concentration in the central chamber, t represents time, and the constant k e is the drug transfer rate from the effect compartment to the central compartment; (4) A mapping between intraoperative monitoring information and anesthesia depth is established based on a deep learning method, including using the drug infusion rate and the patient's covariate data as input variables, extracting the time-dependent parameter Lx and the time-continuous parameter Nx through LSTM and Neural ODE_GRU, establishing a nonlinear function based on the Kan network, and obtaining the predicted anesthesia depth by inputting the drug infusion rate, wherein the patient's covariates include the patient's age, gender, height, and weight.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program or instruction, and the computer program or instruction is used to enable a processor to implement the propofol administration method according to any one of claims 1 to 6 when executed.
10. A computer device, characterized in that: The computer device comprises a memory and a processor, wherein the memory stores a computer program or instruction, and when the computer program or instruction is executed by the processor, the steps in the propofol administration method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Methods of dosing propofol prodrugs for inducing mild to moderate levels of sedation
CN101247809A
Perioperative anesthesia depth monitoring system based on deep learning
CN116982937A
Method for establishing propofol induced drug dosage model of painless gastrointestinal endoscope
CN118072978A
The control of dosing amount for microemulsion and long chain triglyceride propofol on inducion and maintenance of canine anesthesia
KR1020100104695A
Prediction of pharmacokinetic curves
US20240321406A1
Cited By
Controlled drug injection monitoring system and armlet
CN120753606A
Ai modeling closed-loop anesthesia management method and device, storage medium and terminal equipment
CN121747824A