Insulin infusion decision-making system and method combining reinforcement learning and metabolism simulation
By combining the insulin infusion decision system with reinforcement learning and metabolic simulation, an individual metabolic model is constructed and the insulin infusion strategy is optimized using the Actor-Critic network, the problem of difficult to accurately determine the insulin infusion dose in the existing technology is solved, real-time response and safety guarantee for the patient's physiological status is achieved, and the efficiency and interpretability of blood sugar control are improved.
Patent Information
- Application Number
- CN202511053570.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-07-30
AI Technical Summary
In the treatment of diabetes, especially type 1 diabetes, it is difficult to accurately determine the insulin infusion dose and cannot respond to changes in patients' physiological status in real time, resulting in the risk of hypoglycemia or hyperglycemia. It is difficult for traditional methods to continuously adapt to changes in patients' status, reducing the interpretability and safety guarantee of clinical decisions.
Combining the insulin infusion decision system for reinforcement learning and metabolic simulation, individual metabolic model is constructed through the acquisition and processing module, metabolic simulation module, dynamic update module, path simulation module, strategy control module, monitoring and constraint module, decision analysis module, correction processing module and training optimization module, individual metabolic model is constructed, in real time optimized insulin infusion strategy, and the basic insulin infusion rate is adjusted using the Actor-Critic network, and potential intervention failure meta-causes are identified through causal analysis to achieve automated adjustment and real-time response of the model.
It realizes an accurate reflection of the patient's metabolic characteristics, enhances the targeted and real-time perception of insulin decisions, avoids model aging or accumulation of errors, improves the efficiency and safety of blood sugar control, and enhances the interpretability and safety of clinical decisions.
Smart Images

Figure CN120565097A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of clinical decision support, and in particular to an insulin infusion decision-making system and method combining reinforcement learning and metabolic simulation. Background Art
[0002] Diabetes, especially type 1 diabetes, is a typical metabolic chronic disease. Its core feature is the defective insulin secretion function in the body, which leads to the loss of blood sugar regulation ability. In order to maintain blood sugar stability, patients need to rely on exogenous insulin injections, but how to accurately determine the daily and momentary infusion doses remains a core problem in current treatment. In current clinical practice, the formulation of insulin dosages still mainly relies on the doctor's experience, the patient's manual records, and preset calculation formulas. Although this type of method is simple and easy in structure, it has significant limitations in responding to real-time changes in the patient's physiological state and multi-factor coupling feedback mechanisms. It often fails to respond to blood sugar fluctuations in a timely and accurate manner, which can easily lead to risk events such as hypoglycemia or hyperglycemia. In addition, traditional methods are difficult to continuously adapt to changes in the patient's state in daily environments, and they place high demands on the patient's self-management ability and are a heavy burden.
[0003] After searching, Chinese patent number CN116705230A discloses an MDI decision-making system and method with adaptive insulin sensitivity estimation. Although this invention uses a blood glucose control algorithm with insulin sensitivity estimation to complete pre-meal and basal insulin dosage recommendations, establishes a local database for data storage and data visualization, collects necessary patient information to assist in decision-making, and implements a system integration design oriented to patient needs, it cannot accurately reflect the patient's own metabolic characteristics and cannot perceive and respond to the patient's status in real time. There are risks caused by model aging or error accumulation, which reduces the interpretability and safety of clinical decision-making. To this end, we propose an insulin infusion decision-making system and method that combines reinforcement learning and metabolic simulation. Summary of the Invention
[0004] The purpose of the present invention is to address the defects in the prior art and to propose an insulin infusion decision-making system and method that combines reinforcement learning with metabolic simulation.
[0005] In order to achieve the above object, the present invention adopts the following technical solutions:
[0006] An insulin infusion decision-making system that combines reinforcement learning and metabolic simulation includes an acquisition and processing module, a metabolic simulation module, a dynamic update module, a path simulation module, a strategy control module, a monitoring and constraint module, a decision analysis module, a correction processing module, a training optimization module, and an injection control module.
[0007] The acquisition and processing module is used to receive and pre-process the patient's continuous blood glucose monitoring data, insulin injection records, dietary intake information, weight and exercise intensity and other real-time physiological data;
[0008] The metabolic simulation module constructs and initializes an individual metabolic model based on the processed real-time physiological data;
[0009] The dynamic update module amends individual metabolic model parameters based on real-time physiological data;
[0010] The pathway simulation module simulates and predicts the glucose-insulin metabolic process under different insulin input strategies based on the modified individual metabolic model, and generates the corresponding patient's physiological variable change trend;
[0011] The strategy control module is used to adjust the current basal insulin infusion rate according to the changing trend of the patient's physiological variables and optimize the long-term insulin infusion strategy;
[0012] The monitoring and constraint module sets key safety thresholds in operation based on clinical knowledge and simulation data, and blocks actions that trigger hypoglycemia in real time, while suppressing drastic changes in insulin dosage;
[0013] The decision analysis module is used to identify the impact of long-term insulin infusion strategies on changes in patient blood glucose levels and generate a clinically interpretable metabolic intervention report;
[0014] The correction processing module identifies the factors causing potential intervention failure based on the simulation prediction results and performs dosage correction on the strategy output;
[0015] The training optimization module uses simulation-strategy adversarial training to improve the authenticity of simulation predictions and the robustness of long-term insulin infusion strategies in different scenarios;
[0016] The injection control module outputs the final insulin injection plan based on the output of the short-term control module and the strategy learning module that have passed safety verification.
[0017] As a further solution of the present invention, the specific steps of constructing and initializing the individual metabolic model in the metabolic simulation module are as follows:
[0018] S1.1: Align continuous blood glucose monitoring data, insulin injection records, dietary intake information, weight, and exercise intensity according to the collected timestamps to construct a unified input vector set ,in, Representative A point in time, represent Blood glucose value at the time point, represent Insulin injection dose at time point, represent Carbohydrate intake at a given time point, represent Physical activity intensity index at a given time point;
[0019] S1.2: Based on the real-time input vector set, calculate the patient's basal glucose utilization rate during the post-meal blood glucose drop phase, which is the natural metabolic capacity after removing the influence of insulin. Then, analyze the patient's blood glucose changes during exercise to calculate the exercise-induced glucose consumption rate. Based on the patient's glucose consumption rate, obtain the patient's individual exercise sensitivity.
[0020] S1.3: Analyze the patient's resting blood glucose reduction trend after insulin injection, calculate the patient's basal insulin glucose-lowering efficiency, eliminate the interference of carbohydrates and exercise, and obtain the patient's basal insulin sensitivity. Then, monitor the patient's blood glucose rise rate in the morning on an empty stomach, before insulin injection or carbohydrate intake, and calculate the patient's glycogen output rate.
[0021] S1.4: Combining the basal glucose utilization rate, exercise-induced glucose consumption rate, basal insulin glucose-lowering efficiency, and glycogen output rate measured over multiple time periods, the least squares method is used to fuse the parameters to generate model parameters to establish an individual metabolic model for each patient.
[0022] As a further embodiment of the present invention, the specific calculation formula for the basal glucose utilization rate described in S1.2 is as follows:
[0023] ;
[0024] Where, represents the basal glucose utilization rate; Represents the drop in blood sugar after a meal; represents the carbon-water conversion factor, ranging from 10 to 18; Represents the total amount of carbohydrates consumed in the current time interval; Represents the duration of blood sugar drop;
[0025] The specific calculation formula for the exercise-induced glucose consumption rate described in S1.2 is as follows:
[0026] ;
[0027] Where, Represents the glucose consumption rate per unit activity intensity; Represents the drop in blood sugar during exercise; stands for exercise metabolism correction factor; represents the average METs value during the activity; Represents the total duration of exercise;
[0028] The specific calculation formula for the glucose-lowering efficiency of basal insulin described in S1.3 is as follows:
[0029] ;
[0030] Where, It represents the rate of decrease of blood sugar under the action of one unit of insulin; represents the average blood glucose value in the time interval before injection; Represents the average blood sugar level during the time interval after injection; Represents the current insulin injection dose; Represents the duration of insulin activity;
[0031] The specific calculation formula for the patient's glycogen output rate in S1.3 is as follows:
[0032] ;
[0033] Where, represents the rate of glycogen output; Blood glucose level representing the end of the fasting period; Blood glucose level representing the start of the fasting period; represents the observation period;
[0034] The specific calculation formula for parameter fusion described in S1.4 is as follows:
[0035] ;
[0036] Where, represents the individual metabolic model parameters generated by the fusion; Represents the parameter vector to be estimated. 、 、 as well as ; Represents the total number of parameter vectors to be estimated; Represents the simulation output value estimated by the current parameters; represents the true observed value.
[0037] As a further solution of the present invention, the specific steps of the dynamic update module to correct the individual metabolic model parameters are as follows:
[0038] S2.1: The dynamic update module receives the new real-time physiological data collected by the acquisition and processing module at the current moment. Based on the current individual metabolic model parameters, it uses the individual metabolic model simulation to generate a predicted blood glucose value for the patient in the current state. It then calculates the loss between the model prediction value and the real-time physiological data using the residual loss function.
[0039] S2.2: Obtain the sensitivity of the individual metabolic model output to the parameters through automatic differentiation or numerical estimation, and based on this sensitivity information, calculate the gradient of the current loss value with respect to each parameter in the current individual metabolic model, and use the incremental gradient descent method to update the individual metabolic model parameters in real time;
[0040] S2.3: Real-time recording of the parameter information of each period of the individual metabolic model and the number of parameter updates. When the number of parameter updates reaches the preset threshold, the current individual metabolic model parameters are used as the initial values, and the prior probability distribution of the initial values is calculated through multidimensional Gaussian distribution. ,in Represents the initial value The prior probability distribution of represents the initial mean vector, represents the initial covariance matrix;
[0041] S2.4: Randomly sample multiple sets of candidate parameters from the current prior probability distribution ,in Represents the candidate parameter index, and calculates the loss value of the individual metabolic model at the current moment based on each set of parameters. Use Gaussian process to fit the distribution of the current loss function for each individual metabolic model parameter to establish the corresponding proxy model;
[0042] S2.5: Based on the surrogate model and the observed sample loss value, update the posterior predictive distribution of each candidate parameter. Calculate the expected improvement value of each candidate parameter based on the posterior predictive distribution of the current surrogate model. Select the candidate parameter with the highest expected improvement value as the individual metabolic model parameter to be used at the next moment.
[0043] S2.6: Add the updated individual metabolic parameters and corresponding loss values to the historical model data set to update the proxy model input set, and re-count the number of parameter updates, waiting for the next round of individual metabolic model parameter updates.
[0044] As a further solution of the present invention, the specific calculation formula of the residual function in S2.1 is as follows:
[0045] ;
[0046] Where, Representative Individual metabolic model prediction error at the moment; Representative Actual physiological data observed at all times; Represents the current The predicted value of the individual metabolic model for the physiological data under the parameters;
[0047] The specific calculation formula of the incremental gradient descent method described in S2.2 is as follows:
[0048] ;
[0049] Where, Representative Individual metabolic model parameters at the moment; Representative Individual metabolic model parameters at the moment; Representative The learning rate at each moment; Represents the gradient of the residual function with respect to the parameters; Represents the residual loss function value;
[0050] The specific form of the proxy model described in S2.4 is as follows:
[0051] ;
[0052] Where, Represents the current loss function About parameters Gaussian distribution; Representative parameters mean function; Represents the kernel function, and its specific calculation formula is , represents the length scale hyperparameter of the kernel function, which is used to control the degree of smoothing; represents candidate model parameters;
[0053] The specific calculation formula for the posterior predictive distribution described in S2.5 is as follows:
[0054] ;
[0055] Where, represent Time parameters The predicted mean of the location; represent Time parameters The prediction variance of the location; Represents the existing parameter-loss pair. After each round of proxy model parameter update, update ,in represents the total number of parameter-loss pairs;
[0056] The specific calculation formula for the expected improvement value described in S2.5 is as follows:
[0057] ;
[0058] in,
[0059] ;
[0060] Where, Representative parameters Expected improvement under Represents the current minimum loss value; represents the cumulative distribution function of the standard normal distribution; represents the probability density function of the standard normal distribution; Representative parameters The standard deviation of the prediction under Represents the standard normal distribution.
[0061] As a further solution of the present invention, the specific steps of the strategy control module adjusting the current basal insulin infusion rate and optimizing the long-term insulin infusion strategy are as follows:
[0062] S3.1: The short-term control module establishes and initializes the actor network and the critic network. The short-term control module receives the patient's real-time physiological data and establishes the patient's physiological state. , and then through the actor network, using the formula , according to the patient's physiological status calculate Basal insulin infusion rate at all times ,Then the calculation path simulation module simulates the immediate reward of the physiological variable change trend within the preset time period after insulin injection, Represents the actor network function, Represents the actor network parameters. The patient's physiological state is specifically expressed as follows:
[0063] ;
[0064] in, represent The patient's physiological state at all times, represent Blood sugar concentration at any moment, represent The blood insulin concentration at each moment, represent Muscle tissue sugar uptake rate at that time, represent The glycogen output rate at each moment, represent Glucagon levels at any given moment;
[0065] Even if the specific calculation formula of the reward is as follows:
[0066] ;
[0067] Where, represent Instant rewards of the moment; represent Blood glucose value after simulation at each moment; represents the target blood glucose value; represent The fluctuation range of the base rate at the moment; and Represents weighting factors, which are used to balance sugar control accuracy and action smoothness;
[0068] S3.2: Based on the basal insulin infusion rate at each moment and the corresponding patient's physiological state, multiple sets of state-action pairs are established, and each set of state-action pairs is input into the critic network. The critic network calculates the corresponding patient's current physiological state. Take this insulin infusion rate , the cumulative expected rewards obtained within the preset time period;
[0069] S3.3: Based on the current actor network and critic network, create delayed update replicas of the actor network and critic network respectively, and calculate the target Q value based on the two sets of delayed update replicas. At the same time, use the MSE function as the critic network loss function, and calculate the corresponding TD error of the critic network based on the cumulative expected reward and the target Q value. The TD error is back-propagated in the critic network to obtain the gradient of the TD error with respect to each parameter of the critic network, and the parameters of the delayed update replica of the critic network are optimized using the gradient descent method;
[0070] S3.4: Based on the chain rule, the gradients of the critic network parameters are fed into the actor network to obtain the gradients of the critic network with respect to the actor network parameters. The actor network parameters are updated using the obtained gradients, and the Adam optimizer is used to update the delayed update replica parameters of the actor network.
[0071] S3.5: Repeatedly update the delayed update replica parameters of the actor network and the critic network until the target Q value change after multiple iterations converges to a preset range. Then, use the final delayed update replica parameter value to update the actor network and critic network parameters respectively.
[0072] S3.6: Input the current patient physiological state into the actor network to generate different groups of insulin infusion rates in different time periods. The basal insulin infusion rate with the highest target Q value in the critic network is selected as the optimal action for the corresponding time period. The basal insulin infusion rate at the current moment is adjusted based on the selected optimal actions for different time periods.
[0073] The insulin infusion decision-making method combines reinforcement learning and metabolic simulation. The specific steps of this analysis method are as follows:
[0074] Ⅰ. Collect the patient's real-time physiological data and calculate the patient's basal metabolic parameters, while building and modifying the patient's individual metabolic model in real time;
[0075] II. Based on the individual metabolic model, simulate the glucose-insulin reaction process under different insulin administration strategies and predict the changing trend of patients' physiological variables;
[0076] III. Establish the optimal insulin infusion strategy based on the predicted trends of the patient's physiological variables and set operational safety thresholds based on clinical knowledge and simulation data;
[0077] IV. Based on the safety threshold, block abnormal actions in the current insulin infusion strategy and readjust the current insulin infusion strategy;
[0078] V. Analyze the impact of the current insulin infusion strategy on the patient's blood glucose level changes and generate a metabolic intervention report for the corresponding patient;
[0079] VI. Continuously monitor the patient's status and compare the simulated pathway with the actual response to identify potential causes of intervention failure and modify the insulin infusion strategy and dosage.
[0080] As a further embodiment of the present invention, the specific steps of identifying the potential causes of intervention failure in step VI are as follows:
[0081] S4.1: Calculate the individual metabolic model at the current moment and the distance deviation between the simulated results of the glucose-insulin reaction process under different insulin infusion strategies and the expected metabolic pathway. Set an abnormal trigger value δ. If the distance deviation is greater than δ, the current insulin infusion strategy is judged to be deviated, triggering causal analysis.
[0082] S4.2: Based on medical knowledge and metabolic system modeling experience, select physiological variables involved in blood glucose regulation to construct a set of physiological variables. Then, based on clinical and physiological prior knowledge, time series Granger causality tests, and the PC algorithm, identify the causal direction relationship between each physiological variable.
[0083] S4.3: Consider each physiological variable as a node and the causal direction relationship as an edge. Based on the connection relationship between the physiological variables, connect each node through the corresponding edge to construct a variable causal graph. Use the linear regression function to model the causal mechanism of each node in the variable causal graph to generate a generating function for the corresponding node.
[0084] S4.4: Verify the structural identifiability, inter-variable independence, and counterfactual reasoning performance indicators of the variable causal diagram. If any of the performance indicators do not meet the preset threshold, adjust the causal diagram or generating function by backtracking. Then, construct the corresponding structural causal model based on the final causal diagram structure and each generating function.
[0085] S4.5: When the current insulin infusion strategy is judged to be deviated, based on the structural causal model, trace back the variables in the path that have a causal connection with the strategy variable, identify each meta-factor node, and use the mediation effect decomposition method to calculate the causal strength of each meta-factor node on the abnormal state;
[0086] S4.6: The meta-factor node with the greatest causal strength is taken as the potential failure key point. Based on the variable information corresponding to the potential failure key point, the corresponding strategy variables are regulated by solving the inverse control problem to generate a dosage correction plan with minimum disturbance, and the current insulin infusion strategy is adjusted to make the simulated metabolic path return to the expected trajectory.
[0087] As a further solution of the present invention, the specific calculation formula of the distance deviation described in S4.1 is as follows:
[0088] ;
[0089] Where, represent The overall path deviation at the moment; represents the number of variables in the metabolic pathway; Representative The importance weight of each variable; Represents the individual simulation model metabolic variable values; Represents the expected path metabolic variable values;
[0090] The specific calculation formula of the mediation effect decomposition method described in S4.5 is as follows:
[0091] ;
[0092] Where, Representative variables indirect causal effects when acting as a mediator; represents output metabolic index; as well as represent current and comparative dosing strategies, respectively; Represents the intervention operator.
[0093] Compared with the prior art, the present invention has the following beneficial effects:
[0094] This method constructs a unified input vector based on a patient's continuous blood glucose monitoring, diet, weight, and exercise data to estimate individual basal metabolic characteristics. This is then integrated into an individual metabolic model using the least squares method. Incremental gradient descent is then used to update model parameters in real time. Bayesian optimization is combined to sample parameters from a prior distribution and fit a surrogate loss function. Parameters with the highest expected improvement are selected to continuously optimize model accuracy. Based on the updated individual metabolic model, a deep deterministic policy gradient algorithm is used to adjust the basal insulin infusion rate. Actor-Critic networks are alternately trained to maximize future rewards, outputting the optimal insulin strategy for the time period. When the simulation path of the individual metabolic model deviates from expectations, a structural causal model is established to identify the meta-variables causing the deviation. Policy variables are then manipulated through inverse intervention. This approach accurately reflects the patient's metabolic characteristics, making insulin decisions more targeted. This approach enables automated model adjustment, maintains real-time awareness and response to patient status, avoids the risks of model aging or error accumulation, and enhances long-term control capabilities. This improves control efficiency while ensuring blood glucose stability, enhances system fault tolerance, improves the interpretability and security of clinical decisions, and strengthens physician trust. BRIEF DESCRIPTION OF THE DRAWINGS
[0095] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.
[0096] Figure 1 This is a system block diagram of the insulin infusion decision-making system that combines reinforcement learning and metabolic simulation proposed in the present invention;
[0097] Figure 2 This is a flowchart of the insulin infusion decision-making method combining reinforcement learning and metabolic simulation proposed in the present invention. DETAILED DESCRIPTION
[0098] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0099] Example 1, with reference to Figure 1, an insulin infusion decision-making system combining reinforcement learning and metabolic simulation, including acquisition and processing module, metabolic simulation module, dynamic update module, path simulation module, strategy control module, monitoring constraint module, decision analysis module, correction processing module, training optimization module and injection control module;
[0100] The acquisition and processing module is used to receive and pre-process the patient's continuous blood glucose monitoring data, insulin injection records, dietary intake information, weight, and exercise intensity, as well as other real-time physiological data; the metabolic simulation module constructs and initializes the individual metabolic model based on the processed real-time physiological data.
[0101] Specifically, continuous blood glucose monitoring data, insulin injection records, dietary intake information, weight, and exercise intensity are aligned according to the collected timestamps to construct a unified input vector set. ,in, Representative A point in time, represent Blood glucose value at the time point, represent Insulin injection dose at time point, represent Carbohydrate intake at a given time point, represent The physical activity intensity index at the time point is used to calculate the patient's basal glucose utilization rate in the post-meal blood sugar drop period based on the real-time input vector set, in order to remove the natural metabolic capacity after the influence of insulin. The blood sugar changes during exercise are analyzed, and the exercise-induced glucose consumption rate is calculated. Based on the patient's glucose consumption rate, the individual exercise sensitivity of the corresponding patient is obtained. The blood sugar drop trend after insulin injection in the patient's resting state is analyzed, and the patient's basal insulin lowering efficiency is calculated. The interference of carbohydrates and exercise is eliminated to obtain the patient's basal insulin sensitivity. Afterwards, in the early morning on an empty stomach, before insulin injection or carbohydrate intake, the patient's blood sugar rise rate is monitored, and the patient's liver glucose output rate is calculated. The basal glucose utilization rate, exercise-induced glucose consumption rate, basal insulin lowering efficiency and liver glucose output rate measured in multiple time periods are combined, and the least squares method is used to perform parameter fusion to generate model parameters to establish an individual metabolic model corresponding to each patient.
[0102] It should be further explained that the specific calculation formula for the basic utilization rate of glucose is as follows:
[0103] ;
[0104] Where, represents the basal glucose utilization rate; Represents the drop in blood sugar after a meal; represents the carbon-water conversion factor, ranging from 10 to 18; Represents the total amount of carbohydrates consumed in the current time interval; Represents the duration of blood sugar drop;
[0105] The specific calculation formula for exercise-induced glucose consumption rate is as follows:
[0106] ;
[0107] Where, Represents the glucose consumption rate per unit activity intensity; Represents the drop in blood sugar during exercise; stands for exercise metabolism correction factor; represents the average METs value during the activity; Represents the total duration of exercise;
[0108] The specific calculation formula for the glucose-lowering efficiency of basal insulin is as follows:
[0109] ;
[0110] Where, It represents the rate of decrease of blood sugar under the action of one unit of insulin; represents the average blood glucose value in the time interval before injection; Represents the average blood sugar level during the time interval after injection; Represents the current insulin injection dose; Represents the duration of insulin activity;
[0111] The specific calculation formula for the patient's glycogen output rate is as follows:
[0112] ;
[0113] Where, represents the rate of glycogen output; Blood glucose level representing the end of the fasting period; Blood glucose level representing the start of the fasting period; represents the observation period;
[0114] The specific calculation formula for parameter fusion described in S1.4 is as follows:
[0115] ;
[0116] Where, represents the individual metabolic model parameters generated by the fusion; Represents the parameter vector to be estimated. 、 、 as well as ; Represents the total number of parameter vectors to be estimated; Represents the simulation output value estimated by the current parameters; represents the true observed value.
[0117] The dynamic update module modifies the individual metabolic model parameters based on real-time physiological data.
[0118] Specifically, the dynamic update module receives the new real-time physiological data collected by the acquisition and processing module at the current moment, and uses the individual metabolic model simulation to generate a predicted value of the patient's blood glucose in the current state based on the current individual metabolic model parameters. Then, the loss value between the model predicted value and the real-time physiological data is calculated through the residual loss function. The sensitivity of the individual metabolic model output to the parameters is obtained through automatic differentiation or numerical estimation, and based on the sensitivity information, the gradient of the current loss value for each parameter in the current individual metabolic model is calculated, and the individual metabolic model parameters are updated in real time using the incremental gradient descent method. The parameter information of each time period of the individual metabolic model and the number of parameter updates are recorded in real time. When the number of parameter updates reaches the preset threshold, the current individual metabolic model parameters are used as the initial value, and the prior probability distribution of the initial value is calculated through the multi-dimensional Gaussian distribution. ,in Represents the initial value The prior probability distribution of represents the initial mean vector, Represents the initial covariance matrix, randomly sampling multiple sets of candidate parameters from the current prior probability distribution ,in Represents the candidate parameter index, and calculates the loss value of the individual metabolic model at the current moment based on each set of parameters. Use Gaussian process to fit the distribution of the current loss function for each individual metabolic model parameter to establish the corresponding proxy model. Based on the proxy model and the observed sample loss value, update the posterior prediction distribution of each candidate parameter. According to the posterior prediction distribution of the current proxy model, calculate the expected improvement value of each candidate parameter. Among the candidate parameters, select the candidate parameter with the highest expected improvement value as the individual metabolic model parameter to be used at the next moment. Add the updated individual metabolic parameters and the corresponding loss values to the historical model data set to update the proxy model input set, and re-count the number of parameter updates, waiting for the next round of individual metabolic model parameter updates.
[0119] In this embodiment, the specific calculation formula of the residual function is as follows:
[0120] ;
[0121] Where, Representative Individual metabolic model prediction error at the moment; Representative Actual physiological data observed at all times; Represents the current The predicted value of the individual metabolic model for the physiological data under the parameters;
[0122] The specific calculation formula of the incremental gradient descent method is as follows:
[0123] ;
[0124] Where, Representative Individual metabolic model parameters at the moment; Representative Individual metabolic model parameters at the moment; Representative The learning rate at each moment; Represents the gradient of the residual function with respect to the parameters; Represents the residual loss function value;
[0125] The specific form of the proxy model is as follows:
[0126] ;
[0127] Where, Represents the current loss function About parameters Gaussian distribution; Representative parameters mean function; Represents the kernel function, and its specific calculation formula is , represents the length scale hyperparameter of the kernel function, which is used to control the degree of smoothing; represents candidate model parameters;
[0128] The specific calculation formula for the posterior predictive distribution is as follows:
[0129] ;
[0130] Where, represent Time parameters The predicted mean of the location; represent Time parameters The prediction variance of the location; Represents the existing parameter-loss pair. After each round of proxy model parameter update, update ,in represents the total number of parameter-loss pairs;
[0131] The specific calculation formula for the expected improvement value is as follows:
[0132] ;
[0133] in,
[0134] ;
[0135] Where, Representative parameters Expected improvement under Represents the current minimum loss value; represents the cumulative distribution function of the standard normal distribution; represents the probability density function of the standard normal distribution; Representative parameters The standard deviation of the prediction under Represents the standard normal distribution.
[0136] The pathway simulation module simulates and predicts the glucose-insulin metabolic process under different insulin infusion strategies based on the revised individual metabolic model, and generates the corresponding patient's physiological variable change trend; the strategy control module is used to adjust the current basal insulin infusion rate based on the patient's physiological variable change trend and optimize the long-term insulin infusion strategy.
[0137] Specifically, the short-term control module establishes and initializes the actor network and the critic network. The short-term control module receives the patient's real-time physiological data and establishes the patient's physiological state. , and then through the actor network, using the formula , according to the patient's physiological status calculate Basal insulin infusion rate at all times ,Then the calculation path simulation module simulates the immediate reward of the physiological variable change trend within the preset time period after insulin injection, Represents the actor network function, Represents the actor network parameters, establishes multiple sets of state-action pairs based on the basal insulin infusion rate at each moment and the corresponding patient's physiological state, and inputs each set of state-action pairs into the critic network, and calculates the corresponding patient's current physiological state through the critic network. Take this insulin infusion rate , the cumulative expected reward obtained within the preset time period, based on the current actor network and critic network, create delayed update copies of the actor network and critic network respectively, and calculate the target Q value based on the two sets of delayed update copies, and use the MSE function as the critic network loss function, and calculate the TD error corresponding to the critic network based on the cumulative expected reward and the target Q value, and backpropagate the TD error in the critic network to obtain the gradient of the TD error for each parameter of the critic network, use the gradient descent method to optimize the delayed update copy parameters of the critic network, and based on the chain rule, input the gradient corresponding to each parameter of the critic network into the actor network to obtain the gradient of the critic network for the actor network. The gradient information of the r network parameters is obtained, and the parameters of the actor network are updated by the obtained gradient information. The Adam optimizer is used to update the delayed update replica parameters of the actor network. The alternating update of the delayed update replica parameters of the actor network and the critic network is repeated until the target Q-value change value after multiple iterations converges to the preset range. Then, the final delayed update replica parameter value is used to update the actor network and critic network parameters respectively. The current patient's physiological state is input into the actor network to generate various groups of insulin infusion rates in different time periods. The basal insulin infusion rate with the highest target Q value in the critic network is selected as the optimal action for the corresponding time period. The basal insulin infusion rate at the current moment is adjusted based on the selected optimal actions for different time periods.
[0138] It should be further explained that the specific manifestations of the patient's physiological state are as follows:
[0139] ;
[0140] in, represent The patient's physiological state at all times, represent Blood sugar concentration at any moment, represent The blood insulin concentration at each moment, represent Muscle tissue sugar uptake rate at that time, represent The glycogen output rate at each moment, represent Glucagon levels at any given moment;
[0141] Even if the specific calculation formula of the reward is as follows:
[0142] ;
[0143] Where, represent Instant rewards of the moment; represent Blood glucose value after simulation at each moment; represents the target blood glucose value; represent The fluctuation range of the base rate at the moment; and Represents weighting factors, which are used to balance sugar control accuracy and action smoothness.
[0144] The monitoring and constraint module sets key safety thresholds in operation based on clinical knowledge and simulation data, and blocks actions that cause hypoglycemia in real time, while suppressing drastic changes in insulin dosage; the decision analysis module is used to identify the impact of long-term insulin infusion strategies on changes in patients' blood glucose levels and generate clinically interpretable metabolic intervention reports; the correction processing module identifies potential causes of intervention failure based on simulation prediction results and makes dosage corrections to strategy outputs; the training optimization module uses simulation-strategy adversarial training to improve the authenticity of simulation predictions and the robustness of long-term insulin infusion strategies in different scenarios; the injection control module outputs the final insulin injection plan based on the outputs of the short-term control module and strategy learning module that have passed safety verification.
[0145] Example 2, reference Figure 2 , an insulin infusion decision-making method combining reinforcement learning and metabolic simulation. The specific steps of this analysis method are as follows:
[0146] Collect patients' real-time physiological data, calculate the corresponding patients' basal metabolic parameters, and build and modify the patients' individual metabolic models in real time.
[0147] Based on the individual metabolic model, the glucose-insulin reaction process under different insulin administration strategies is simulated, and the changing trend of the patient's physiological variables is predicted.
[0148] Based on the predicted trend of changes in the patient's physiological variables, the current optimal insulin infusion strategy is established, and the safety threshold during operation is set based on clinical knowledge and simulation data.
[0149] Based on the safety threshold, abnormal actions in the current insulin infusion strategy are blocked and the current insulin infusion strategy is readjusted.
[0150] Analyze the impact of the current insulin infusion strategy on the patient's blood glucose level changes and generate a metabolic intervention report for the corresponding patient.
[0151] Continuously monitor patient status and compare simulated pathways with actual responses to identify potential intervention failures and adjust insulin infusion strategies and dosages.
[0152] Specifically, the individual metabolic model at the current moment is calculated, and the distance deviation between the simulation results of the glucose-insulin reaction process under different insulin infusion strategies and the expected metabolic pathway is set. The abnormal trigger value δ is set. If the distance deviation is greater than δ, the current insulin infusion strategy is judged to be deviated, triggering causal analysis. Based on medical knowledge and metabolic system modeling experience, physiological variables involved in blood glucose regulation are selected to construct a set of physiological variables. Then, based on clinical and physiological prior knowledge, time series Granger causality test and PC algorithm, the causal direction relationship between each physiological variable is identified. Each physiological variable is used as a node, and the causal direction relationship is used as an edge. Based on the connection relationship between each physiological variable, each node is connected through the corresponding edge to construct a variable causal graph. The causal mechanism of each node in the variable causal graph is modeled using a linear regression function to generate a generating function for the corresponding node. Verify the structural identifiability of the variable causal graph, independence between variables, and counterfactual reasoning ability performance indicators. If there is any performance indicator that does not meet the preset threshold, the causal graph or generating function is adjusted by backtracking. Then, based on the final causal graph structure and each generating function, the corresponding structural causal model is constructed. When the current insulin infusion strategy is judged to be deviated, based on the structural causal model, the variables in the backtracking path that have a causal connectivity relationship with the strategy variables are identified, and the mediating effect decomposition method is used to calculate the causal strength of each meta-factor node to the abnormal state. The meta-factor node with the largest causal strength is regarded as the potential failure key point. Based on the variable information corresponding to the potential failure key point, the corresponding strategy variables are regulated by solving the inverse control problem to generate a dosage correction plan with minimum disturbance, and the current insulin infusion strategy is adjusted to make the simulated metabolic path return to the expected trajectory.
[0153] In this embodiment, the specific calculation formula of the distance deviation is as follows:
[0154] ;
[0155] Where, represent The overall path deviation at the moment; represents the number of variables in the metabolic pathway; Representative The importance weight of each variable; Represents the individual simulation model metabolic variable values; Represents the expected path metabolic variable values;
[0156] The specific calculation formula of the mediation effect decomposition method is as follows:
[0157] ;
[0158] Where, Representative variables indirect causal effects when acting as a mediator; represents output metabolic index; as well as represent current and comparative dosing strategies, respectively; Represents the intervention operator.
Claims
1. An insulin infusion decision-making system combining reinforcement learning and metabolic simulation, characterized by: It includes acquisition and processing module, metabolic simulation module, dynamic update module, pathway simulation module, short-term control module, strategy learning module, monitoring constraint module, decision analysis module, correction processing module, training optimization module and injection control module; The metabolic simulation module constructs and initializes an individual metabolic model based on the processed real-time physiological data; The dynamic update module amends individual metabolic model parameters based on real-time physiological data; The strategy control module is used to adjust the current basal insulin infusion rate according to the changing trend of the patient's physiological variables and optimize the long-term insulin infusion strategy; The correction processing module identifies the factors causing potential intervention failures based on the simulation prediction results and performs dosage correction on the strategy output.
2. The insulin infusion decision-making system combining reinforcement learning and metabolic simulation according to claim 1, characterized in that: The acquisition and processing module is used to receive and pre-process the patient's continuous blood glucose monitoring data, insulin injection records, dietary intake information, weight and exercise intensity and other real-time physiological data; The pathway simulation module simulates and predicts the glucose-insulin metabolic process under different insulin input strategies based on the modified individual metabolic model, and generates the corresponding patient's physiological variable change trend; The monitoring and constraint module sets operational safety thresholds based on clinical knowledge and simulation data, and blocks actions that trigger hypoglycemia in real time, while suppressing drastic changes in insulin dosage; The decision analysis module is used to identify the impact of long-term insulin infusion strategies on changes in patient blood glucose levels and generate a clinically interpretable metabolic intervention report; The training optimization module uses simulation-strategy adversarial training to improve the authenticity of simulation predictions and the robustness of long-term insulin infusion strategies in different scenarios; The injection control module outputs the final insulin injection plan based on the output of the short-term control module and the strategy learning module that have passed safety verification.
3. The insulin infusion decision-making system combining reinforcement learning and metabolic simulation according to claim 1, characterized in that: The specific steps of constructing and initializing the individual metabolic model in the metabolic simulation module are as follows: S1.1: Align the real-time physiological data of continuous blood glucose monitoring data, insulin injection records, dietary intake information, weight, and exercise intensity according to the collected timestamps to construct a unified input vector set ,in, Representative A point in time, represent Blood glucose value at the time point, represent Insulin injection dose at time point, represent Carbohydrate intake at a given time point, represent Physical activity intensity index at a given time point; S1.2: Based on the real-time input vector set, calculate the patient's basal glucose utilization rate during the post-meal blood glucose drop phase, which is the natural metabolic capacity after removing the influence of insulin. Then, analyze the patient's blood glucose changes during exercise to calculate the exercise-induced glucose consumption rate. Based on the patient's glucose consumption rate, obtain the patient's individual exercise sensitivity. S1.3: Analyze the patient's resting blood glucose reduction trend after insulin injection, calculate the patient's basal insulin glucose-lowering efficiency, eliminate the interference of carbohydrates and exercise, and obtain the patient's basal insulin sensitivity. Then, monitor the patient's blood glucose rise rate in the morning on an empty stomach, before insulin injection or carbohydrate intake, and calculate the patient's glycogen output rate. S1.4: Combining the basal glucose utilization rate, exercise-induced glucose consumption rate, basal insulin glucose-lowering efficiency, and glycogen output rate measured over multiple time periods, the least squares method is used to fuse the parameters to generate model parameters to establish an individual metabolic model for each patient.
4. The insulin infusion decision system combining reinforcement learning and metabolic simulation according to claim 3, characterized in that: The specific steps of the dynamic update module to correct the individual metabolic model parameters are as follows: S2.1: The dynamic update module receives the new real-time physiological data collected by the acquisition and processing module at the current moment. Based on the current individual metabolic model parameters, it uses the individual metabolic model simulation to generate a predicted blood glucose value for the patient in the current state. It then calculates the loss between the model prediction value and the real-time physiological data using the residual loss function. S2.2: Obtain the sensitivity of the individual metabolic model output to the parameters through automatic differentiation or numerical estimation, and based on this sensitivity information, calculate the gradient of the current loss value with respect to each parameter in the current individual metabolic model, and use the incremental gradient descent method to update the individual metabolic model parameters in real time; S2.3: Real-time recording of the parameter information of each period of the individual metabolic model and the number of parameter updates. When the number of parameter updates reaches the preset threshold, the current individual metabolic model parameters are used as the initial values, and the prior probability distribution of the initial values is calculated through multidimensional Gaussian distribution. ,in Represents the initial value The prior probability distribution of represents the initial mean vector, represents the initial covariance matrix; S2.4: Randomly sample multiple sets of candidate parameters from the current prior probability distribution ,in Represents the candidate parameter index, and calculates the loss value of the individual metabolic model at the current moment based on each set of parameters. Use Gaussian process to fit the distribution of the current loss function for each individual metabolic model parameter to establish the corresponding proxy model; S2.5: Based on the surrogate model and the observed sample loss value, update the posterior predictive distribution of each candidate parameter. Calculate the expected improvement value of each candidate parameter based on the posterior predictive distribution of the current surrogate model. Select the candidate parameter with the highest expected improvement value as the individual metabolic model parameter to be used at the next moment. S2.6: Add the updated individual metabolic parameters and corresponding loss values to the historical model data set to update the proxy model input set, and re-count the number of parameter updates, waiting for the next round of individual metabolic model parameter updates.
5. The insulin infusion decision-making system combining reinforcement learning and metabolic simulation according to claim 4, characterized in that: The specific calculation formula of the residual function described in S2.1 is as follows: ; Where, Representative Individual metabolic model prediction error at the moment; Representative Actual physiological data observed at all times; Represents the current The predicted value of the individual metabolic model for the physiological data under the parameters; The specific calculation formula of the incremental gradient descent method described in S2.2 is as follows: ; Where, Representative Individual metabolic model parameters at the moment; Representative Individual metabolic model parameters at the moment; Representative The learning rate at each moment; Represents the gradient of the residual function with respect to the parameters; Represents the residual loss function value; The specific form of the proxy model described in S2.4 is as follows: ; Where, Represents the current loss function About parameters Gaussian distribution; Representative parameters mean function; Represents the kernel function, and its specific calculation formula is , represents the length scale hyperparameter of the kernel function, which is used to control the degree of smoothing; represents candidate model parameters; The specific calculation formula for the posterior predictive distribution described in S2.5 is as follows: ; Where, represent Time parameters The predicted mean of the location; represent Time parameters The prediction variance of the location; Represents the existing parameter-loss pair. After each round of proxy model parameter update, update ,in represents the total number of parameter-loss pairs; The specific calculation formula for the expected improvement value described in S2.5 is as follows: ; in, ; Where, Representative parameters Expected improvement under Represents the current minimum loss value; represents the cumulative distribution function of the standard normal distribution; represents the probability density function of the standard normal distribution; Representative parameters The standard deviation of the prediction under Represents the standard normal distribution.
6. The insulin infusion decision-making system combining reinforcement learning and metabolic simulation according to claim 4, characterized in that: The specific steps of the strategy control module adjusting the current basal insulin infusion rate and optimizing the long-term insulin infusion strategy are as follows: S3.1: The short-term control module establishes and initializes the actor network and the critic network. The short-term control module receives the patient's real-time physiological data and establishes the patient's physiological state. , and then through the actor network, using the formula , according to the patient's physiological status calculate Basal insulin infusion rate at all times ,Then the calculation path simulation module simulates the immediate reward of the physiological variable change trend within the preset time period after insulin injection, Represents the actor network function, Represents the actor network parameters. The patient's physiological state is specifically expressed as follows: ; in, represent The patient's physiological state at all times, represent Blood sugar concentration at any moment, represent The blood insulin concentration at each moment, represent Muscle tissue sugar uptake rate at that time, represent The glycogen output rate at each moment, represent Glucagon levels at any given moment; Even if the specific calculation formula of the reward is as follows: ; Where, represent Instant rewards of the moment; represent Blood glucose value after simulation at each moment; represents the target blood glucose value; represent The fluctuation range of the base rate at the moment; and Represents weighting factors, which are used to balance sugar control accuracy and action smoothness; S3.2: Based on the basal insulin infusion rate at each moment and the corresponding patient's physiological state, multiple sets of state-action pairs are established, and each set of state-action pairs is input into the critic network. The critic network calculates the corresponding patient's current physiological state. Take this insulin infusion rate , the cumulative expected rewards obtained within the preset time period; S3.3: Based on the current actor network and critic network, create delayed update replicas of the actor network and critic network respectively, and calculate the target Q value based on the two sets of delayed update replicas. At the same time, use the MSE function as the critic network loss function, and calculate the corresponding TD error of the critic network based on the cumulative expected reward and the target Q value. The TD error is back-propagated in the critic network to obtain the gradient of the TD error with respect to each parameter of the critic network, and the parameters of the delayed update replica of the critic network are optimized using the gradient descent method; S3.4: Based on the chain rule, the gradients of the critic network parameters are fed into the actor network to obtain the gradients of the critic network with respect to the actor network parameters. The actor network parameters are updated using the obtained gradients, and the Adam optimizer is used to update the delayed update replica parameters of the actor network. S3.5: Repeatedly update the delayed update replica parameters of the actor network and the critic network until the target Q value change after multiple iterations converges to a preset range. Then, use the final delayed update replica parameter value to update the actor network and critic network parameters respectively. S3.6: Input the current patient physiological state into the actor network to generate different groups of insulin infusion rates in different time periods. The basal insulin infusion rate with the highest target Q value in the critic network is selected as the optimal action for the corresponding time period. The basal insulin infusion rate at the current moment is adjusted based on the selected optimal actions for different time periods.
7. An insulin infusion decision-making method combining reinforcement learning and metabolic simulation, for implementing the function of an insulin infusion decision-making system combining reinforcement learning and metabolic simulation as described in any one of claims 1-6, characterized in that: The specific steps of this analysis method are as follows: Ⅰ. Collect the patient's real-time physiological data and calculate the patient's basal metabolic parameters, while building and modifying the patient's individual metabolic model in real time; II. Based on the individual metabolic model, simulate the glucose-insulin reaction process under different insulin administration strategies and predict the changing trend of patients' physiological variables; III. Establish the optimal insulin infusion strategy based on the predicted trends of the patient's physiological variables and set operational safety thresholds based on clinical knowledge and simulation data; IV. Based on the safety threshold, block abnormal actions in the current insulin infusion strategy and readjust the current insulin infusion strategy; V. Analyze the impact of the current insulin infusion strategy on the patient's blood glucose level changes and generate a metabolic intervention report for the corresponding patient; VI. Continuously monitor the patient's status and compare the simulated pathway with the actual response to identify potential causes of intervention failure and modify the insulin infusion strategy and dosage.
8. The insulin infusion decision-making method combining reinforcement learning and metabolic simulation according to claim 7, characterized in that: The specific steps for identifying potential intervention failure factors in step VI are as follows: S4.1: Calculate the individual metabolic model at the current moment and the distance deviation between the simulated results of the glucose-insulin reaction process under different insulin infusion strategies and the expected metabolic pathway. Set an abnormal trigger value δ. If the distance deviation is greater than δ, the current insulin infusion strategy is judged to be deviated, triggering causal analysis. S4.2: Based on medical knowledge and metabolic system modeling experience, select physiological variables involved in blood glucose regulation to construct a set of physiological variables. Then, based on clinical and physiological prior knowledge, time series Granger causality tests, and the PC algorithm, identify the causal direction relationship between each physiological variable. S4.3: Consider each physiological variable as a node and the causal direction relationship as an edge. Based on the connection relationship between the physiological variables, connect each node through the corresponding edge to construct a variable causal graph. Use the linear regression function to model the causal mechanism of each node in the variable causal graph to generate a generating function for the corresponding node. S4.4: Verify the structural identifiability, inter-variable independence, and counterfactual reasoning performance indicators of the variable causal diagram. If any of the performance indicators do not meet the preset threshold, adjust the causal diagram or generating function by backtracking. Then, construct the corresponding structural causal model based on the final causal diagram structure and each generating function. S4.5: When the current insulin infusion strategy is judged to be deviated, based on the structural causal model, trace back the variables in the path that have a causal connection with the strategy variable, identify each meta-factor node, and use the mediation effect decomposition method to calculate the causal strength of each meta-factor node on the abnormal state; S4.6: The meta-factor node with the greatest causal strength is taken as the potential failure key point. Based on the variable information corresponding to the potential failure key point, the corresponding strategy variables are regulated by solving the inverse control problem to generate a dosage correction plan with minimum disturbance, and the current insulin infusion strategy is adjusted to make the simulated metabolic path return to the expected trajectory.
9. The insulin infusion decision-making method combining reinforcement learning and metabolic simulation according to claim 8, characterized in that: The specific calculation formula for the distance deviation mentioned in S4.1 is as follows: ; Where, represent The overall path deviation at the moment; represents the number of variables in the metabolic pathway; Representative The importance weight of each variable; Represents the individual simulation model metabolic variable values; Represents the expected path metabolic variable values; The specific calculation formula of the mediation effect decomposition method described in S4.5 is as follows: ; Where, Representative variables indirect causal effects when acting as a mediator; represents output metabolic index; as well as represent current and comparative dosing strategies, respectively; Represents the intervention operator.
Citation Information
Patent Citations
Safeguarding measures for a closed-loop insulin infusion system
CA2884999A1
Strategy network training method, insulin infusion scheme generation method and electronic equipment
CN114300090A
Hyperparameter learning, intelligent recommendation, keyword and multimedia recommendation method and device
CN114329167A
MDI decision system and method with insulin sensitivity adaptive estimation
CN116705230A
Insulin pump fault diagnosis system and method based on continuous blood glucose monitoring
CN117612692A
Cited By
Clinical aid decision analysis method and system based on electronic medical record
CN120748697A
Clinical Decision Support Analysis Methods and Systems Based on Electronic Medical Records
CN120748697B
Anti-coagulation regulation and control intelligent management system and method for extracorporeal circulation
CN121071633A
Intelligent Management System and Method for Anticoagulation Regulation in Extracorporeal Circulation
CN121071633B
Metabolism simulation method and system for hyperuricemia organ chip model
CN121709010A