Insulin infusion decision system and method combining reinforcement learning with metabolic simulation
By combining reinforcement learning and metabolic simulation, an individual metabolic model is constructed, and the insulin infusion strategy is optimized. This solves the problem of inaccurate insulin infusion dosage in existing technologies, realizes real-time response to patients' metabolic characteristics and ensures safety, and improves the real-time performance and interpretability of blood glucose control.
Patent Information
- Application Number
- CN202511053570.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2045-07-30
AI Technical Summary
Current technologies cannot accurately determine insulin infusion dosage in diabetes treatment, leading to the risk of blood glucose fluctuations and difficulty in responding to changes in patient status in real time, affecting self-management ability and safety.
By combining reinforcement learning and metabolic simulation, an individual metabolic model is constructed. Through data acquisition and processing, dynamic updates, path simulation, strategy control, monitoring constraints, decision analysis, and injection control modules, the insulin infusion strategy is optimized, and the dosage is adjusted in real time to respond to the patient's condition.
It achieves accurate reflection of patients' metabolic characteristics, improves the targeting and safety of insulin decisions, enhances the real-time response capability of stable blood glucose control, and reduces the risk of model aging and error accumulation.
Smart Images

Figure CN120565097B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of clinical decision assistance, in particular to an insulin infusion decision system and method combining reinforcement learning and metabolic simulation. BACKGROUND
[0002] Diabetes, especially type 1 diabetes, is a typical metabolic chronic disease, whose core feature is the deficiency of insulin secretion function in the body, resulting in the loss of blood glucose regulation ability. In order to maintain stable blood glucose, patients need to rely on exogenous insulin injection, but how to accurately determine the daily and hourly infusion dose is still a core problem in current treatment. In current clinical practice, the formulation of insulin dose still mainly depends on the experience of doctors, the manual records of patients and the preset calculation formula. Although this kind of method is simple in structure and easy to implement, it has significant limitations in dealing with real-time changes of patient physiological state and multi-factor coupled feedback mechanism, often cannot respond to blood glucose fluctuations in time and accurately, easily leads to risks such as hypoglycemia or hyperglycemia, in addition, the traditional method is difficult to adapt to the changes of patient state in daily environment, requires high self-management ability of patients and heavy burden.
[0003] Through retrieval, Chinese patent No. CN116705230A discloses an MDI decision system and method with insulin sensitivity adaptive estimation, which, although the blood glucose control algorithm with insulin sensitivity estimation is used to complete the pre-meal and basal insulin dose recommendation, the local database is established to realize data storage and data visualization, the necessary information of patients is collected to assist decision-making, and the system integration design is realized for the needs of patients, cannot accurately reflect the metabolic characteristics of patients, and cannot realize real-time perception and response of patient state, has the risk hidden danger caused by model aging or error accumulation, and reduces the explainability and safety guarantee of clinical decision-making; therefore, the present application proposes an insulin infusion decision system and method combining reinforcement learning and metabolic simulation. SUMMARY
[0004] The purpose of the present application is to solve the defects in the prior art, and to propose an insulin infusion decision system and method combining reinforcement learning and metabolic simulation.
[0005] In order to achieve the above-mentioned purpose, the present application adopts the following technical solutions:
[0006] The insulin infusion decision system combining reinforcement learning and metabolic simulation comprises a collection and processing module, a metabolic simulation module, a dynamic updating module, a path simulation module, a strategy control module, a monitoring constraint module, a decision analysis module, a correction processing module, a training optimization module and an injection control module.
[0007] The collection processing module is configured to receive and pre-process continuous glucose monitoring data, insulin injection records, diet intake information, body weight, and exercise intensity of the patient;
[0008] The metabolic simulation module is configured to construct and initialize an individual metabolic model according to the pre-processed real-time physiological data;
[0009] The dynamic updating module is configured to correct parameters of the individual metabolic model according to the real-time physiological data;
[0010] The path simulation module is configured to simulate and predict glucose-insulin metabolic processes under different insulin input strategies according to the corrected individual metabolic model, and generate physiological variable change trends of the patient;
[0011] The strategy control module is configured to adjust a basal insulin infusion rate at the current time according to the physiological variable change trends of the patient, and optimize a long-term insulin infusion strategy;
[0012] The monitoring constraint module is configured to set key safety thresholds according to clinical knowledge and simulation data, and to block actions that cause hypoglycemia in real time, while suppressing drastic changes in insulin dosage;
[0013] The decision analysis module is configured to identify the influence of the long-term insulin infusion strategy on the change in the patient's blood glucose level, and to generate a metabolic intervention report with clinical interpretability;
[0014] The correction processing module is configured to identify the meta-factors of potential intervention failure according to the simulation prediction results, and to correct the dosage of the strategy output;
[0015] The training optimization module is configured to use simulation-strategy adversarial training to improve the authenticity of simulation prediction and the robustness of the long-term insulin infusion strategy in different scenarios;
[0016] The injection control module is configured to output a final insulin injection scheme according to the output of the short-term control module and the strategy learning module that have passed safety verification.
[0017] As a further scheme of the present application, the specific steps of constructing and initializing the individual metabolic model by the metabolic simulation module are as follows:
[0018] S1.1: Aligning the continuous glucose monitoring data, insulin injection records, diet intake information, body weight, and exercise intensity according to the time stamps of the collected data to construct a unified input vector set , wherein, represents the time point, represents the blood glucose value at the time point, represents insulin injection dose at the time point, represent carbohydrate intake at the time point, represent physical activity intensity indicator at the time point;
[0019] S1.2: According to the real-time input vector set, the basal glucose utilization rate of the patient in the postprandial glucose decline segment is calculated, the natural metabolic capacity after removing the influence of insulin is analyzed, the glucose consumption rate induced by exercise is calculated by analyzing the blood glucose change of the patient during exercise, and the individual exercise sensitivity of the corresponding patient is obtained according to the glucose consumption rate of the patient;
[0020] S1.3: The blood glucose decline trend after the patient injects insulin in the resting state is analyzed, the basal insulin glucose-lowering efficiency of the patient is calculated, the interference of carbohydrates and exercise is excluded, the basal insulin sensitivity of the patient is obtained, and then the blood glucose rise rate of the patient is monitored before the morning fasting, without injecting insulin or ingesting carbohydrates, and the hepatic glucose output rate of the patient is calculated;
[0021] S1.4: Combining the basal glucose utilization rate, the exercise-induced glucose consumption rate, the basal insulin glucose-lowering efficiency, and the hepatic glucose output rate measured in multiple time periods, the model parameters are fused by using the least square method to generate model parameters, so as to establish the individual metabolic model corresponding to each patient.
[0022] As a further scheme of the present application, the specific calculation formula of the basal glucose utilization rate in S1.2 is as follows:
[0023] ;
[0024] In the formula, represent the basal glucose utilization rate; represent the postprandial glucose decline value; represent the carbohydrate conversion coefficient, the value range is 10-18; represent the total amount of carbohydrates ingested in the current time interval; represent the duration of blood glucose decline;
[0025] The specific calculation formula of the exercise-induced glucose consumption rate in S1.2 is as follows:
[0026] ;
[0027] In the formula, represent the glucose consumption rate under unit activity intensity; represent the blood glucose decline value during exercise; represent the exercise metabolism correction factor; represent the average METs value during activity; represent the total duration of the exercise;
[0028] The specific calculation formula of the basal insulin hypoglycemic efficiency in S1.3 is as follows:
[0029] ;
[0030] In the formula, represent the blood glucose reduction rate under the action of unit insulin; represent the average value of blood glucose in the time interval before injection; represent the average value of blood glucose in the time interval after injection; represent the current insulin injection dose; represent the duration of insulin activity;
[0031] The specific calculation formula of the patient's hepatic glucose output rate in S1.3 is as follows:
[0032] ;
[0033] In the formula, represent the hepatic glucose output rate; represent the blood glucose value at the end of the fasting period; represent the blood glucose value at the beginning of the fasting period; represent the observation period;
[0034] The specific calculation formula of the parameter fusion in S1.4 is as follows:
[0035] ;
[0036] In the formula, represent the individual metabolic model parameters generated by fusion; represent the parameter vector to be estimated. Including , , and ; represent the total number of parameter vectors to be estimated; represent the simulation output value estimated by the current parameters; represent the true observation value.
[0037] As a further scheme of the present application, the specific steps of the dynamic updating module correcting the individual metabolic model parameters are as follows:
[0038] S2.1: The dynamic updating module receives the new real-time physiological data collected by the collection processing module at the current time, generates a blood glucose prediction value for the patient under the current state based on the current individual metabolic model parameters, and then calculates the loss value between the model prediction value and the real-time physiological data through a residual loss function;
[0039] S2.2: Obtain the sensitivity of individual metabolic model output to parameters through automatic differentiation or numerical estimation, and calculate the gradient of the current loss value to each parameter in the current individual metabolic model according to the sensitivity information, and update the individual metabolic model parameters in real time using the incremental gradient descent method;
[0040] S2.3: Real-time record the parameter information of each period of the individual metabolic model, and the number of parameter updates, when the number of parameter updates reaches the preset number threshold, take the current individual metabolic model parameters as the initial value, and calculate the prior probability distribution of the initial value through the multi-dimensional Gaussian distribution , wherein represents the prior probability distribution of the initial value , represents the initial mean vector , and represents the initial covariance matrix
[0041] S2.4: Randomly collect multiple groups of candidate parameters from the current prior probability distribution , wherein represents the candidate parameter index, and the loss value of the individual metabolic model at the current moment is calculated according to each group of parameters, and the distribution of the current loss function to each individual metabolic model parameter is fitted by using the Gaussian process, so as to establish the corresponding surrogate model
[0042] S2.5: Based on the surrogate model and the observed sample loss value, update the posterior predictive distribution of each candidate parameter, calculate the expected improvement value of each candidate parameter according to the posterior predictive distribution of the current surrogate model, and select the candidate parameter with the highest expected improvement value as the individual metabolic model parameter used at the next moment in the candidate parameters
[0043] S2.6: Add the updated individual metabolic parameters and the corresponding loss value to the historical model data set to update the surrogate model input set, and re-count the number of parameter updates, and wait for the next round of individual metabolic model parameter update.
[0044] As a further scheme of the present application, the specific calculation formula of the residual error function in S2.1 is as follows:
[0045]
[0046] In the formula, represents the prediction error of the individual metabolic model at the moment; represents the actual physiological data observed at the moment; represents the predicted value of the individual metabolic model for the physiological data under the current parameters;
[0047] The specific calculation formula of the incremental gradient descent method described in S2.2 is as follows:
[0048] ;
[0049] Where, Representative Individual metabolic model parameters at the moment; Representative Individual metabolic model parameters at the moment; Representative The learning rate at each moment; Represents the gradient of the residual function with respect to the parameters; Represents the residual loss function value;
[0050] The specific form of the proxy model described in S2.4 is as follows:
[0051] ;
[0052] Where, Represents the current loss function About parameters Gaussian distribution; Representative parameters mean function; Represents the kernel function, and its specific calculation formula is , represents the length scale hyperparameter of the kernel function, which is used to control the degree of smoothing; represents candidate model parameters;
[0053] The specific calculation formula for the posterior predictive distribution described in S2.5 is as follows:
[0054] ;
[0055] Where, represent Time parameters The predicted mean of the location; represent Time parameters The prediction variance of the location; Represents the existing parameter-loss pair. After each round of proxy model parameter update, update ,in represents the total number of parameter-loss pairs;
[0056] The specific calculation formula for the expected improvement value described in S2.5 is as follows:
[0057] ;
[0058] in,
[0059] ;
[0060] wherein, represents the expected improvement value under the parameter ; represents the current minimum loss value; represents the cumulative distribution function of the standard normal distribution; represents the probability density function of the standard normal distribution; represents the expected improvement value under the parameter ; represents the standard normal distribution.
[0061] As a further scheme of the present application, the specific steps of the strategy control module adjusting the basal insulin infusion rate at the current time to optimize the long-term insulin infusion strategy are as follows:
[0062] S3.1: The short-term control module establishes and initializes the actor network and the critic network, the short-term control module receives the real-time physiological data of the patient, and the established patient physiological state is calculated through the actor network, and the basal insulin infusion rate at the time is calculated according to the patient physiological state using the formula , and then the even reward of the path simulation module simulating the physiological variable change trend in the preset time period after insulin injection is calculated, wherein, represents the actor network function, represents the actor network parameter, and the specific form of the patient physiological state is as follows:
[0063] ;
[0064] wherein, represents the patient physiological state at the time , represents the blood glucose concentration at the time , represents the insulin concentration in the blood at the time , represents the muscle tissue sugar uptake rate at the time , represents the liver glucose output rate at the time , represents the glucagon level at the time ;
[0065] The specific calculation formula of the even reward is as follows:
[0066] ;
[0067] wherein, represents instant reward at the moment; represents simulated blood glucose value at the moment; represents target blood glucose value; represents baseline rate variation range at the moment; and represents weighting factor, respectively used for balancing glycemic control accuracy and action smoothness;
[0068] S3.2: According to the basal insulin infusion rate at each moment and the corresponding physiological state of the patient, a plurality of state-action pairs are established, and each group of state-action pairs is input into the critic network, and the critic network is used to calculate the expected cumulative reward of the corresponding patient in the current physiological state , the insulin infusion rate is taken , the cumulative expected reward obtained within a preset time period;
[0069] S3.3: Based on the current actor network and critic network, delayed update copies of the actor network and critic network are respectively created, and the target Q value is calculated according to the two groups of delayed update copies, and the MSE function is used as the critic network loss function, and the TD error corresponding to the critic network is calculated based on the cumulative expected reward and the target Q value, and the TD error is back propagated in the critic network to obtain the gradient of the TD error on each parameter of the critic network, and the gradient descent method is used to optimize the parameters of the delayed update copy of the critic network;
[0070] S3.4: Based on the chain rule, the gradient corresponding to each parameter of the critic network is input into the actor network to obtain the gradient information of the critic network gradient on the parameters of the actor network, and the parameters of the actor network are updated through the obtained gradient information, and the Adam optimizer is used to update the parameters of the delayed update copy of the actor network;
[0071] S3.5: The alternative update of the parameters of the delayed update copies of the actor network and the critic network is repeatedly performed until the change value of the target Q value after multiple iterations converges to a preset range, and then the parameters of the actor network and the critic network are updated using the final delayed update copy parameter values;
[0072] S3.6: input the current patient physiological state into the actor network, generate each group of insulin infusion rates in different time periods, and select the basal insulin infusion rate with the highest target Q value in the critic network as the optimal action for the corresponding time period, and adjust the basal insulin infusion rate at the current time based on the selected optimal actions in different time periods.
[0073] The insulin infusion decision-making method combining reinforcement learning and metabolic simulation has the following specific steps:
[0074] I. Collect real-time physiological data of the patient, calculate the basal metabolic parameters of the corresponding patient, and construct and real-time correct the individual metabolic model of the patient;
[0075] II. Based on the individual metabolic model, simulate the glucose-insulin response process under different insulin administration strategies, and predict the change trend of the patient's physiological variables;
[0076] III. According to the predicted change trend of the patient's physiological variables, establish the current optimal insulin infusion strategy, and set the safety threshold during operation according to the clinical knowledge and simulation data;
[0077] IV. According to the safety threshold, block the abnormal action in the current insulin infusion strategy, and readjust the current insulin infusion strategy;
[0078] V. Analyze the influence of the current insulin infusion strategy on the change of the patient's blood glucose level, and generate a metabolic intervention report for the corresponding patient;
[0079] VI. Continuously monitor the patient's state, and identify the meta-factor of potential intervention failure by comparing the simulation path with the actual reaction, and correct the insulin infusion strategy dose.
[0080] As a further scheme of the present application, the specific steps for identifying the meta-factor of potential intervention failure in step VI are as follows:
[0081] S4.1: Calculate the individual metabolic model at the current time, the distance deviation of the simulation results of the glucose-insulin response process under different insulin infusion strategies and the expected metabolic path, set an abnormal trigger value δ, if the distance deviation > δ, it is judged that the current insulin infusion strategy deviates, and the cause-effect analysis is triggered;
[0082] S4.2: According to the medical knowledge and metabolic system modeling experience, select the physiological variables involved in blood glucose regulation to construct a set of physiological variables, and according to the clinical and physiological priori knowledge, time series granger causality test and PC algorithm, identify the causal direction relationship between each physiological variable;
[0083] S4.3: taking each physiological variable as a node, taking the causal direction relationship as an edge, connecting each node through the corresponding edge based on the connection relationship between each physiological variable to construct a variable causal graph, and modeling the causal mechanism of each node in the variable causal graph using a linear regression function to generate the generating function of the corresponding node;
[0084] S4.4: verifying the performance indicators of the structural identifiability of the variable causal graph, the independence between variables, and the counterfactual reasoning ability, and if there is any performance indicator that does not meet the preset threshold, adjusting the causal graph or the generating function through backtracking, and then constructing the corresponding structural causal model according to the final causal graph structure and each generating function;
[0085] S4.5: when the current insulin infusion strategy is judged to deviate, based on the structural causal model, backtracking the variables in the path that have causal connectivity with the strategy variables, identifying each meta-factor node, and calculating the causal strength of each meta-factor node to the abnormal state using the mediation effect decomposition method;
[0086] S4.6: taking the meta-factor node with the largest causal strength as the potential failure key point, and based on the variable information corresponding to the potential failure key point, solving the inverse control problem to regulate the corresponding strategy variable to generate a minimum disturbance dose correction scheme, and adjusting the current insulin infusion strategy to make the simulated metabolic path tend to the expected trajectory again.
[0087] As a further scheme of the present application, the specific calculation formula of the distance deviation in S4.1 is as follows:
[0088] ;
[0089] In the formula, represents the overall path deviation degree at time t; represents the number of variables in the metabolic pathway; represents the importance weight of the i-th variable; represents the value of the i-th metabolic variable in the individual simulation model; represents the value of the i-th metabolic variable in the expected path;
[0090] The specific calculation formula of the mediation effect decomposition method in S4.5 is as follows:
[0091] ;
[0092] In the formula, represents the indirect causal effect of the variable as a mediator; represents the output metabolic indicator; and respectively represent the current and contrast dose strategy; represent the intervention operator.
[0093] Compared with the prior art, the beneficial effects of the present application are:
[0094] The present application constructs a unified input vector according to the patient's continuous blood glucose monitoring, diet, weight, exercise and other data, estimates the individual basal metabolism characteristics, and generates an individual metabolism model by least squares fusion. Then, the model parameters are updated in real time using incremental gradient descent, and the parameters are sampled from the prior distribution by combining Bayesian optimization, fitting the loss function proxy model, selecting the parameter with the maximum expected improvement value to continuously optimize the model precision. Based on the updated individual metabolism model, the basal insulin infusion rate is adjusted through the deep deterministic policy gradient algorithm, and the Actor-Critic network is alternately trained to maximize the future reward. The output period optimal insulin strategy is output. When the simulation path of the individual metabolism model deviates from the expectation, a structural causal model is established to identify the meta-variable that causes the deviation, and the intervention control strategy variable is adjusted. The system can accurately reflect the patient's own metabolic characteristics, make the insulin decision more targeted, realize the automatic adjustment of the model, maintain real-time perception and response to the patient's state, avoid the risk of model aging or error accumulation, enhance long-term control ability, improve control efficiency while ensuring blood glucose stability, enhance system fault tolerance, improve the explainability and safety of clinical decision-making, and enhance the trust of doctors. BRIEF DESCRIPTION OF DRAWINGS
[0095] The accompanying drawings are included to provide a further understanding of the present application, and constitute a part of the specification, which together with the embodiments of the present application serve to explain the present application, and do not constitute a limitation of the present application.
[0096] Figure 1 The system block diagram of the insulin infusion decision system combining reinforcement learning and metabolic simulation proposed by the present application;
[0097] Figure 2 The flowchart of the insulin infusion decision method combining reinforcement learning and metabolic simulation proposed by the present application. DETAILED DESCRIPTION
[0098] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments.
[0099] Embodiment 1, refer to Figure 1, the insulin infusion decision system combining reinforcement learning and metabolic simulation, comprising a collection and processing module, a metabolic simulation module, a dynamic updating module, a path simulation module, a strategy control module, a monitoring constraint module, a decision analysis module, a correction processing module, a training optimization module, and an injection control module;
[0100] The collection and processing module is configured to receive and pre-process continuous blood glucose monitoring data, insulin injection records, dietary intake information, body weight, and exercise intensity of the patient.
[0101] Specifically, the continuous blood glucose monitoring data, insulin injection records, dietary intake information, body weight, and exercise intensity are aligned according to the timestamps of collection to construct a unified input vector set , wherein, represents the first time point, represents the time point, represents the time point, represents the time point, represents the time point, According to the real-time input vector set, the glucose utilization rate of the patient in the postprandial blood glucose decline segment is calculated to remove the natural metabolic capacity after the influence of insulin, the glucose consumption rate induced by exercise during the patient's exercise is calculated, and the individual exercise sensitivity of the patient is obtained according to the glucose consumption rate of the patient.
[0102] It should be further pointed out that the specific calculation formula of the glucose utilization rate is as follows:
[0103] ;
[0104] In the formula, represents the glucose utilization rate; represents the postprandial blood glucose decline value; represents the carbohydrate exchange factor, with a value range of 10-18; represents the total amount of carbohydrates ingested in the current time interval; represents the duration of blood glucose reduction;
[0105] The specific calculation formula of exercise-induced glucose consumption rate is as follows:
[0106] ;
[0107] In the formula, represents the glucose consumption rate under unit activity intensity; represents the blood glucose reduction value during exercise; represents the exercise metabolism correction factor; represents the average METs value during activity; represents the total exercise duration;
[0108] The specific calculation formula of basal insulin glucose-lowering efficiency is as follows:
[0109] ;
[0110] In the formula, represents the blood glucose reduction rate under unit insulin action; represents the average blood glucose value in the time interval before injection; represents the average blood glucose value in the time interval after injection; represents the current insulin injection dose; represents the insulin activity duration;
[0111] The specific calculation formula of patient hepatic glucose output rate is as follows:
[0112] ;
[0113] In the formula, represents the hepatic glucose output rate; represents the blood glucose value at the end of the fasting period; represents the blood glucose value at the start of the fasting period; represents the observation time period;
[0114] The specific calculation formula of the parameter fusion described in S1.4 is as follows:
[0115] ;
[0116] In the formula, represents the individual metabolic model parameters generated by fusion; represents the parameter vector to be estimated. Including , , and ; Represents the total number of parameter vectors to be estimated; Represents the simulation output value estimated by the current parameters; represents the true observed value.
[0117] The dynamic update module modifies the individual metabolic model parameters based on real-time physiological data.
[0118] Specifically, the dynamic update module receives the new real-time physiological data collected by the acquisition and processing module at the current moment, and uses the individual metabolic model simulation to generate a predicted value of the patient's blood glucose in the current state based on the current individual metabolic model parameters. Then, the loss value between the model predicted value and the real-time physiological data is calculated through the residual loss function. The sensitivity of the individual metabolic model output to the parameters is obtained through automatic differentiation or numerical estimation, and based on the sensitivity information, the gradient of the current loss value for each parameter in the current individual metabolic model is calculated, and the individual metabolic model parameters are updated in real time using the incremental gradient descent method. The parameter information of each time period of the individual metabolic model and the number of parameter updates are recorded in real time. When the number of parameter updates reaches the preset threshold, the current individual metabolic model parameters are used as the initial value, and the prior probability distribution of the initial value is calculated through the multi-dimensional Gaussian distribution. ,in Represents the initial value The prior probability distribution of represents the initial mean vector, Represents the initial covariance matrix, randomly sampling multiple sets of candidate parameters from the current prior probability distribution ,in Represents the candidate parameter index, and calculates the loss value of the individual metabolic model at the current moment based on each set of parameters. Use Gaussian process to fit the distribution of the current loss function for each individual metabolic model parameter to establish the corresponding proxy model. Based on the proxy model and the observed sample loss value, update the posterior prediction distribution of each candidate parameter. According to the posterior prediction distribution of the current proxy model, calculate the expected improvement value of each candidate parameter. Among the candidate parameters, select the candidate parameter with the highest expected improvement value as the individual metabolic model parameter to be used at the next moment. Add the updated individual metabolic parameters and the corresponding loss values to the historical model data set to update the proxy model input set, and re-count the number of parameter updates, waiting for the next round of individual metabolic model parameter updates.
[0119] In this embodiment, the specific calculation formula of the residual function is as follows:
[0120] ;
[0121] Where, Representative Individual metabolic model prediction error at the moment; represent the actual physiological data observed at the time instant; represent the predicted value of the physiological data by the individual metabolic model at the time instant under the current parameters;
[0122] The specific calculation formula of the incremental gradient descent method is as follows:
[0123] ;
[0124] wherein, represent the individual metabolic model parameters at the time instant; represent the individual metabolic model parameters at the time instant; represent the learning rate at the time instant; represent the gradient of the residual function with respect to the parameters; represent the residual loss function value;
[0125] The specific form of the surrogate model is as follows:
[0126] ;
[0127] wherein, represent the current loss function about the parameters of the Gaussian distribution; represent the parameter mean function; represent the kernel function, and the specific calculation formula is , represent the length scale hyperparameter of the kernel function, which is used to control the smoothness; represent the candidate model parameters;
[0128] The specific calculation formula of the posterior predictive distribution is as follows:
[0129] ;
[0130] wherein, represent the predicted mean at the time instant parameter position; represent the predicted variance at the time instant parameter position; represent the existing parameter-loss pair, and after each round of surrogate model parameter update, the is updated, wherein represent the total number of parameter-loss pairs;
[0131] The specific calculation formula for the expected improvement value is as follows:
[0132] ;
[0133] in,
[0134] ;
[0135] Where, Representative parameters Expected improvement under Represents the current minimum loss value; represents the cumulative distribution function of the standard normal distribution; represents the probability density function of the standard normal distribution; Representative parameters The standard deviation of the prediction under Represents the standard normal distribution.
[0136] The pathway simulation module simulates and predicts the glucose-insulin metabolic process under different insulin infusion strategies based on the revised individual metabolic model, and generates the corresponding patient's physiological variable change trend; the strategy control module is used to adjust the current basal insulin infusion rate based on the patient's physiological variable change trend and optimize the long-term insulin infusion strategy.
[0137] Specifically, the short-term control module establishes and initializes the actor network and the critic network. The short-term control module receives the patient's real-time physiological data and establishes the patient's physiological state. , and then through the actor network, using the formula , according to the patient's physiological status calculate Basal insulin infusion rate at all times ,Then the calculation path simulation module simulates the immediate reward of the physiological variable change trend within the preset time period after insulin injection, Represents the actor network function, Represents the actor network parameters, establishes multiple sets of state-action pairs based on the basal insulin infusion rate at each moment and the corresponding patient's physiological state, and inputs each set of state-action pairs into the critic network, and calculates the corresponding patient's current physiological state through the critic network. Take this insulin infusion rate , the cumulative expected reward obtained within the preset time period, based on the current actor network and critic network, create delayed update copies of the actor network and critic network respectively, and calculate the target Q value based on the two sets of delayed update copies, and use the MSE function as the critic network loss function, and calculate the TD error corresponding to the critic network based on the cumulative expected reward and the target Q value, and backpropagate the TD error in the critic network to obtain the gradient of the TD error for each parameter of the critic network, use the gradient descent method to optimize the delayed update copy parameters of the critic network, and based on the chain rule, input the gradient corresponding to each parameter of the critic network into the actor network to obtain the gradient of the critic network for the actor network. The gradient information of the r network parameters is obtained, and the parameters of the actor network are updated by the obtained gradient information. The Adam optimizer is used to update the delayed update replica parameters of the actor network. The alternating update of the delayed update replica parameters of the actor network and the critic network is repeated until the target Q-value change value after multiple iterations converges to the preset range. Then, the final delayed update replica parameter value is used to update the actor network and critic network parameters respectively. The current patient's physiological state is input into the actor network to generate various groups of insulin infusion rates in different time periods. The basal insulin infusion rate with the highest target Q value in the critic network is selected as the optimal action for the corresponding time period. The basal insulin infusion rate at the current moment is adjusted based on the selected optimal actions for different time periods.
[0138] It should be further explained that the specific manifestations of the patient's physiological state are as follows:
[0139] ;
[0140] in, represent The patient's physiological state at all times, represent Blood sugar concentration at any moment, represent The blood insulin concentration at each moment, represent Muscle tissue sugar uptake rate at that time, represent The glycogen output rate at each moment, represent Glucagon levels at any given moment;
[0141] Even if the specific calculation formula of the reward is as follows:
[0142] ;
[0143] Where, representing instant reward at the moment; representing simulated blood glucose value at the moment; representing target blood glucose value; representing basal rate variation range at the moment; and representing weighting factors, respectively, for balancing glycemic control accuracy and action smoothness.
[0144] The monitoring constraint module sets key safety thresholds in operation according to clinical knowledge and simulation data, and blocks actions that cause hypoglycemia in real time, while suppressing drastic changes in insulin dose; the decision analysis module is used to identify the influence of long-term insulin infusion strategy on the change of patient's blood glucose level, and generate a metabolic intervention report with clinical interpretability; the correction processing module identifies the meta-factor of potential intervention failure according to the simulation prediction result, and corrects the dose of the strategy output; the training optimization module uses simulation-strategy adversarial training to improve the authenticity of simulation prediction and the robustness of long-term insulin infusion strategy in different scenarios; the injection control module outputs the final insulin injection scheme according to the output of the short-term control module and the strategy learning module that have passed safety verification.
[0145] Embodiment 2, with reference to Figure 2 , the insulin infusion decision-making method combining reinforcement learning and metabolic simulation, the specific steps of the analysis method are as follows:
[0146] Collect real-time physiological data of the patient, calculate the basal metabolic parameters of the corresponding patient, and construct and real-time correct the individual metabolic model of the patient.
[0147] Based on the individual metabolic model, simulate the glucose-insulin response process under different insulin administration strategies, and predict the trend of changes in patient physiological variables.
[0148] According to the predicted trend of changes in patient physiological variables, establish the current optimal insulin infusion strategy, and set the safety threshold in operation according to clinical knowledge and simulation data.
[0149] According to the safety threshold, block abnormal actions in the current insulin infusion strategy, and readjust the current insulin infusion strategy.
[0150] Analyze the influence of the current insulin infusion strategy on the change of patient's blood glucose level, and generate a metabolic intervention report for the corresponding patient.
[0151] Continuously monitor the patient's state, and identify the meta-factor of potential intervention failure by comparing the simulation path with the actual reaction, and correct the dose of the insulin infusion strategy.
[0152] Specifically, the distance deviation of the simulation result of the glucose-insulin reaction process under different insulin infusion strategies from the expected metabolic path is calculated for the individual metabolic model at the current time, an abnormal trigger value δ is set, and if the distance deviation > δ, it is judged that the current insulin infusion strategy deviates, the causal analysis is triggered, the physiological variables participating in blood glucose regulation are selected according to the medical knowledge and metabolic system modeling experience to construct a physiological variable set, and then the causal direction relationship between each physiological variable is identified according to the clinical and physiological prior knowledge, the time series Granger causality test and the PC algorithm, each physiological variable is taken as a node, the causal direction relationship is taken as an edge, each node is connected through the corresponding edge based on the connection relationship between the physiological variables, a variable causal graph is constructed, and a linear regression function is used to model the causal mechanism of each node in the variable causal graph to generate a generating function corresponding to the node. The performance indicators of the structure identifiable, the independence between variables and the counterfactual reasoning ability of the variable causal graph are verified, and if any performance indicator does not meet the preset threshold, the causal graph or the generating function is adjusted through backtracking. Then, according to the final causal graph structure and each generating function, a corresponding structural causal model is constituted. When the current insulin infusion strategy is judged to deviate, the variables in the path that have causal connectivity with the strategy variables are identified based on the structural causal model, each meta-factor node is identified, and the causal strength of each meta-factor node to the abnormal state is calculated using the mediation effect decomposition method. The meta-factor node with the largest causal strength is taken as the potential failure key point, and based on the variable information corresponding to the potential failure key point, an inverse control problem is solved to control the corresponding strategy variable to generate a minimum disturbance dose correction scheme, and the current insulin infusion strategy is adjusted to make the simulation metabolic path tend to the expected trajectory again.
[0153] In this embodiment, the specific calculation formula of the distance deviation is as follows:
[0154] ;
[0155] In the formula, represents the distance deviation of the simulation result of the glucose-insulin reaction process under different insulin infusion strategies from the expected metabolic path at the current time; represents the overall path deviation degree at the current time; represents the number of variables in the metabolic path; represents the importance weight of the i-th variable; represents the value of the i-th metabolic variable in the individual simulation model; represents the value of the i-th metabolic variable in the expected path; represents the value of the i-th metabolic variable in the expected path; The specific calculation formula of the mediation effect decomposition method is as follows:
[0156]
[0157] ;
[0158] In the formula, representative variables indirect causal effects when acting as intermediaries; representative output metabolic indicators; and representing current and comparative dose strategies, respectively; representative intervention operators.
Claims
1. An insulin infusion decision system that combines reinforcement learning with metabolic simulation, characterized in that, The system comprises a collection processing module, a metabolism simulation module, a dynamic updating module, a path simulation module, a short-term control module, a strategy learning module, a monitoring constraint module, a decision analysis module, a correction processing module, a training optimization module, and an injection control module. The metabolism simulation module constructs and initializes an individual metabolism model according to the processed real-time physiological data. The dynamic updating module corrects the parameters of the individual metabolism model according to the real-time physiological data. The path simulation module simulates and predicts the glucose-insulin metabolism process under different insulin input strategies according to the corrected individual metabolism model, and generates the physiological variable change trend of the corresponding patient. The strategy control module is used to adjust the basal insulin infusion rate at the current time and optimize the long-term insulin infusion strategy according to the physiological variable change trend of the patient. The correction processing module identifies the meta-factor of potential intervention failure according to the simulation prediction result, and corrects the dose of the strategy output. The specific steps of constructing and initializing the individual metabolism model by the metabolism simulation module are as follows: S1.1: Aligning the continuous glucose monitoring data, insulin injection records, dietary intake information, body weight, and exercise intensity each real-time physiological data according to the time stamp collected to construct a unified input vector set wherein, represents the first time point, represents the second time point, represents the blood glucose value at the first time point, represents the blood glucose value at the second time point, represents the insulin injection dose at the first time point, represents the insulin injection dose at the second time point, represents the carbohydrate intake at the first time point, represents the carbohydrate intake at the second time point, represents the body activity intensity index at the first time point, represents the body activity intensity index at the second time point. S1.2: According to the real-time input vector set, the basal utilization rate of glucose of the patient in the postprandial blood glucose decline segment is calculated to remove the natural metabolism capacity after removing the influence of insulin, the exercise-induced glucose consumption rate is calculated by analyzing the blood glucose change of the patient during exercise, and the individual exercise sensitivity of the corresponding patient is obtained according to the glucose consumption rate of the patient; S1.3: The blood glucose decline trend of the patient after injecting insulin in a resting state is analyzed, the basal insulin glucose-lowering efficiency of the patient is calculated, the interference of carbohydrates and exercise is excluded, the basal insulin sensitivity of the patient is obtained, and then the blood glucose rise rate of the patient is monitored before the morning fasting, without injecting insulin or ingesting carbohydrates, and the hepatic glucose output rate of the patient is calculated; S1.4: The model parameters are fused by using the least squares method to establish the individual metabolism model of each patient by combining the measured basal utilization rate of glucose, exercise-induced glucose consumption rate, basal insulin glucose-lowering efficiency, and hepatic glucose output rate in multiple periods; The specific steps of correcting the parameters of the individual metabolism model by the dynamic updating module are as follows: S2.1: The dynamic updating module receives the new real-time physiological data collected by the collection processing module at the current time, generates the blood glucose prediction value of the patient under the current state by using the individual metabolism model based on the current individual metabolism model parameters, and then calculates the loss value between the model prediction value and the real-time physiological data by using the residual loss function; S2.2: The sensitivity of the individual metabolism model output to the parameters is obtained by automatic differentiation or numerical estimation, the gradient of the current loss value to each parameter in the current individual metabolism model is calculated according to the sensitivity information, and the individual metabolism model parameters are updated in real time by using the incremental gradient descent method; S2.3: record the parameter information of each period of the individual metabolic model in real time, and the number of parameter updates, when the number of parameter updates reaches a preset number threshold, take the current individual metabolic model parameter as an initial value, and calculate the prior probability distribution of the initial value through a multi-dimensional Gaussian distribution wherein represents the initial value prior probability distribution, represents the initial mean vector, represents the initial covariance matrix; S2.4: Randomly sample multiple groups of candidate parameters from the current prior probability distribution wherein represent the index of candidate parameters, and the loss value of the individual metabolic model at the current time is calculated according to each group of parameters, and the distribution of the current loss function for each individual metabolic model parameter is fitted by using a Gaussian process to establish the corresponding proxy model; S2.5: Based on the proxy model and the observed sample loss value, the posterior predictive distribution of each candidate parameter is updated, the expected improvement value of each candidate parameter is calculated according to the posterior predictive distribution of the current proxy model, and the candidate parameter with the highest expected improvement value is selected as the individual metabolism model parameter used at the next time. S2.6: Add the updated individual metabolic parameters and the corresponding loss value to the historical model data set to update the agent model input set, and re-count the number of parameter updates, and wait for the next round of individual metabolic model parameter update; The specific steps of the strategy control module for adjusting the basal insulin infusion rate at the current time and optimizing the long-term insulin infusion strategy are as follows: S3.1: The short-term control module establishes and initializes the actor network and the critic network. The short-term control module receives the patient's real-time physiological data and establishes the patient's physiological state. , and then through the actor network, using the formula , according to the patient's physiological status calculate Basal insulin infusion rate at all times ,Then the calculation path simulation module simulates the immediate reward of the physiological variable change trend within the preset time period after insulin injection, Represents the actor network function, Represents the actor network parameters. The patient's physiological state is specifically expressed as follows: ; wherein, representing the patient's physiological state at the time instant, representing the blood glucose concentration at the time instant, representing the blood insulin concentration at the time instant, representing the muscle tissue glucose uptake rate at the time instant, representing the hepatic glucose output rate at the time instant, representing the glucagon level at the time instant; The specific calculation formula of the reward is as follows: ; In the formula, represent instant reward at the time; represent simulated blood glucose value at the time; represent target blood glucose value; represent basic rate variation range at the time; and represent weighting factors, respectively, for balancing glycemic control accuracy and motion smoothness; S3.2: According to the basal insulin infusion rate at each time point and the corresponding patient physiological state, a plurality of state-action pairs are established, and each group of state-action pairs is input into the critic network, and the critic network is used to calculate the expected reward of the corresponding patient in the current physiological state The insulin infusion rate is taken as follows The cumulative expected reward obtained within a preset time period; S3.3: Based on the current actor network and critic network, create delayed update copies of the actor network and critic network respectively, and calculate the target Q value based on the two sets of delayed update copies. At the same time, the MSE function is used as the critic network loss function, and the TD error corresponding to the critic network is calculated based on the cumulative expected reward and the target Q value. The TD error is backpropagated in the critic network to obtain the gradient of the TD error on each parameter of the critic network. The delayed update copy parameters of the critic network are optimized using the gradient descent method; S3.4: Based on the chain rule, input the gradient of each parameter of the critic network into the actor network to obtain the gradient information of the critic network gradient on the actor network parameters, and update the parameters of the actor network through the obtained gradient information, and update the parameters of the delayed update copy of the actor network using the Adam optimizer; S3.5: The parameters of the delayed update copies of the actor network and the critic network are alternately updated repeatedly until the change value of the target Q value after multiple iterations converges to a preset range. Then, the final delayed update copy parameter values are used to update the parameters of the actor network and the critic network; S3.6: Input the current physiological state of the patient into the actor network to generate each group of insulin infusion rates in different time periods, and select the basal insulin infusion rate with the highest target Q value in the critic network as the optimal action for the corresponding time period. Adjust the basal insulin infusion rate at the current time based on the selected optimal actions in different time periods.
2. The insulin infusion decision system that combines reinforcement learning with metabolic simulation of claim 1, wherein, The acquisition and processing module is configured to receive and preprocess the continuous glucose monitoring data, insulin injection records, dietary intake information, body weight, and exercise intensity of the patient. The monitoring and constraint module sets safety thresholds based on clinical knowledge and simulation data, and blocks actions that may cause hypoglycemia in real time, while suppressing drastic changes in insulin dosage. The decision analysis module is configured to identify the influence of the long-term insulin infusion strategy on the patient's blood glucose level changes, and generate a metabolic intervention report with clinical interpretability. The training and optimization module uses simulation-strategy countermeasures training to improve the authenticity of simulation prediction and the robustness of long-term insulin infusion strategy in different scenarios. The injection control module outputs the final insulin injection scheme based on the output of the safety-verified short-term control module and the strategy learning module.
3. The insulin infusion decision system that combines reinforcement learning with metabolic simulation of claim 1, wherein, The specific calculation formula of the residual error function in S2.1 is as follows: ; wherein representing the individual metabolic model at the prediction error of the individual metabolic model at the representing the actual physiological data observed at the representing the individual metabolic model at the representing the current predicted value of the physiological data by the individual metabolic model at the The specific calculation formula of the incremental gradient descent method in S2.2 is as follows: ; wherein represent the individual metabolic model parameters at the time instant; represent the individual metabolic model parameters at the time instant; represent the learning rate at the time instant; represent the gradient of the residual function with respect to the parameters; represent the residual loss function value; The specific form of the proxy model described in S2.4 is as follows: ; wherein, represents the current loss function regarding the parameters of the Gaussian distribution; represents the parameters of the mean function; represents the kernel function, whose specific formula is , represents the length scale hyperparameter of the kernel function, which is used to control the smoothness; represents the candidate model parameters; The specific calculation formula of the posterior predictive distribution described in S2.5 is as follows: ; wherein represent time parameter predicted mean of the position; represent time parameter predicted variance of the position; represent the existing parameter-loss pairs, and update wherein represent the total number of parameter-loss pairs; The specific calculation formula of the expected improvement value described in S2.5 is as follows: ; Wherein, ; wherein a representative parameter a desired improvement value under the condition a representative current minimum loss value a cumulative distribution function of a standard normal distribution a probability density function of a standard normal distribution a representative parameter a predicted standard deviation under the condition a standard normal distribution.
4. The insulin infusion decision method combining reinforcement learning and metabolic simulation according to any one of claims 1 to 3, characterized in that, The specific steps of the analysis method are as follows: I. Collect real-time physiological data of the patient, calculate the basal metabolic parameters of the corresponding patient, and construct and real-time correct the individual metabolic model of the patient; II. Based on the individual metabolic model, simulate the glucose-insulin response process under different insulin administration strategies, and predict the trend of the patient's physiological variables; III. According to the predicted trend of the patient's physiological variables, establish the current optimal insulin infusion strategy, and set the safety threshold during operation according to the clinical knowledge and simulation data; IV. According to the safety threshold, block the abnormal action in the current insulin infusion strategy, and adjust the current insulin infusion strategy; V. Analyze the influence of the current insulin infusion strategy on the change of the patient's blood glucose level, and generate a metabolic intervention report for the corresponding patient; VI. Continuously monitor the patient's state, and identify the meta-factor of potential intervention failure by comparing the simulation path with the actual response, and correct the insulin infusion strategy dose.
5. The method of claim 4, wherein the insulin infusion decision is determined by combining the reinforcement learning and the metabolic simulation. The specific steps of identifying the meta-factor of potential intervention failure described in step VI are as follows: S4.1: Calculate the distance deviation of the individual metabolic model at the current time between the simulation results of the glucose-insulin response process under different insulin infusion strategies and the expected metabolic path, set an abnormal trigger value δ, if the distance deviation > δ, it is judged that the current insulin infusion strategy deviates, and the causal analysis is triggered; S4.2: According to the medical knowledge and metabolic system modeling experience, select the physiological variables involved in blood glucose regulation to construct a set of physiological variables, and according to the clinical and physiological prior knowledge, time series Granger causality test and PC algorithm, identify the causal direction relationship between each physiological variable; S4.3: Take each physiological variable as a node, take the causal direction relationship as an edge, and based on the connection relationship between each physiological variable, connect each node through the corresponding edge to construct a variable causal graph, and use a linear regression function to model the causal mechanism of each node in the variable causal graph to generate a generating function for the corresponding node; S4.4: Verify the performance indicators of the structural identifiability, independence between variables and counterfactual reasoning ability of the variable causal graph, if any of the performance indicators do not meet the preset threshold, adjust the causal graph or generating function through backtracking, and then according to the final causal graph structure and each generating function, construct a corresponding structural causal model; S4.5: When the current insulin infusion strategy is judged to deviate, based on the structural causal model, backtrack the variables in the path that have causal connectivity with the strategy variables, identify each meta-factor node, and calculate the causal strength of each meta-factor node to the abnormal state using the mediation effect decomposition method; S4.6: Take the meta-factor node with the largest causal strength as the potential failure key point, and based on the variable information corresponding to the potential failure key point, solve the inverse control problem to regulate the corresponding strategy variable, generate a minimum disturbance dose correction scheme, and adjust the current insulin infusion strategy to make the simulation metabolic path tend to the expected trajectory again.
6. The method of claim 5, wherein the method is implemented by a computer system. The specific calculation formula of the distance deviation in S4.1 is as follows: ; wherein represents the overall path deviation at the time instant; represents the number of variables in the metabolic pathway; represents the importance weight of the th variable; represents the value of the th metabolic variable in the individual simulation model; represents the value of the th metabolic variable in the expected pathway; The specific calculation formula of the mediation effect decomposition method in S4.5 is as follows: ; wherein representing variables indirect causal effects as intermediaries; representing output metabolic indicators; and representing current and contrast dose strategies, respectively; representing intervention operators.
Citation Information
Patent Citations
MDI decision system and method with insulin sensitivity adaptive estimation
CN116705230A
Task offloading method and apparatus, scheduling optimization method and apparatus, electronic device, and storage medium
WO2022242468A1