Using Q learning for determining insulin doses and related systems, methods, and devices

By adjusting CR and CF in real time through the Q-learning algorithm and the nearest neighbor Q-learning algorithm, the problem of inaccurate insulin dosage calculation is solved, better glucose control is achieved, hypoglycemia events are reduced, and the quality of life of diabetic patients is improved.

CN120677532APending Publication Date: 2025-09-19BIGFOOT BIOMEDICAL INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480011669.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-02-10
Filing Date
2024-02-09
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing technologies make it difficult to effectively track and adjust the carbohydrate ratio (CR) and correction factor (CF) of diabetic patients, resulting in inaccurate insulin dosage calculation and affecting glucose control.

Method used

The Q-learning algorithm and the nearest neighbor Q-learning algorithm, combined with the reinforcement learning method, track the changes in CR and CF in real time, and achieve more accurate insulin calculation by adjusting the insulin dose of meal bolus and correction bolus.

Benefits of technology

It significantly improves the accuracy and stability of glucose control, reduces hypoglycemia events, and improves the quality of life of diabetic patients.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120677532A_ABST
    Figure CN120677532A_ABST
Patent Text Reader

Abstract

Determination of bolus doses of insulin and / or bolus doses and related systems, methods, and devices are disclosed. A method of determining a dietary bolus, a corrective bolus, and / or an underlying insulin dose may include tracking a change in a carbohydrate ratio (CR) using a Q-learning algorithm, tracking a change in a correction factor (CF) using a nearest neighbor Q-learning algorithm, and determining a dietary bolus dose in response to the tracked CR and the tracked CF.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 484,387, filed February 10, 2023, which is incorporated herein by reference in its entirety. Technical Field

[0003] The present disclosure generally relates to determining a bolus dose of rapid-acting insulin and / or a basal dose of long-acting insulin, and more particularly to using reinforcement learning to determine a bolus dose of rapid-acting insulin and / or a basal dose of long-acting insulin for use as part of insulin therapy for treating diabetes. Background Art

[0004] Diabetes mellitus is a chronic metabolic disorder caused by the inability of the pancreas to produce sufficient amounts of the hormone insulin, rendering the person's metabolism unable to adequately absorb sugars and starches. The inability to absorb these carbohydrates sometimes leads to hyperglycemia, the presence of excess glucose in the blood plasma. Hyperglycemia is associated with a variety of serious symptoms and life-threatening long-term complications, such as dehydration, ketoacidosis, diabetic coma, cardiovascular disease, chronic kidney failure, retinal damage, and nerve damage with the risk of limb amputation.

[0005] Typically, permanent therapy is necessary for maintaining appropriate glucose levels within normal limits. Maintaining appropriate glucose levels is typically achieved by regularly supplying insulin to patients with diabetes (PWD). Maintaining appropriate glucose levels may create a significant cognitive burden for PWD (or caregiver), and affect many aspects of the life of PWD. For example, among other things, the cognitive burden of PWD may be owing to tracking meals and continuously checking and slightly correcting glucose levels, etc. The adjustment of glucose levels by PWD may include using insulin, tracking insulin dosage and glucose, deciding how much insulin to use, the frequency of using insulin, where to inject insulin, and how to arrange the time of insulin dosage with respect to meals and / or glucose fluctuations.

[0006] Treatment plans can be difficult to implement due to (among other things) differences in how different individuals respond to treatment, as well as fluctuations in individuals' own responses to treatment. Summary of the Invention

[0007] The present disclosure provides one or more computer-readable storage media and a method for determining a bolus insulin dose for a meal, as defined in the accompanying claims. Disclosed are systems, methods, and apparatus for determining a bolus insulin dose and / or a bolus insulin dose. A method for determining a bolus insulin dose for a meal, a correction bolus insulin dose, and / or a basal insulin dose may include tracking changes in a carbohydrate ratio (CR) using a Q-learning algorithm, tracking changes in a correction factor (CF) using a nearest neighbor Q-learning algorithm, and determining a bolus insulin dose for a meal in response to the tracked CR and the tracked CF. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] While the present disclosure concludes with claims that particularly point out and distinctly claim specific embodiments, the various features and advantages of embodiments within the scope of this disclosure may be more readily ascertained from the following description when read in conjunction with the accompanying drawings, wherein:

[0009] Figure 1 is a flow chart illustrating a method of determining a prandial bolus, correction bolus insulin dose, and / or basal dose according to some embodiments.

[0010] Figure 2 is a flow chart illustrating a method for learning a carbohydrate ratio (CR) according to some embodiments.

[0011] Figure 3 is a flow chart illustrating a method for learning a correction factor (CF) according to some embodiments.

[0012] Figure 4 is a flow chart illustrating a method for estimating a bolus dose and / or a basal dose according to some embodiments.

[0013] Figure 5 is a block diagram of the architecture of an advanced bolus calculator for PWDs receiving multiple daily injection (MDI) therapy, according to some embodiments.

[0014] Figure 6 is a block diagram of a medical treatment system according to some embodiments.

[0015] Figure 7 The percentage of time spent in the target range, below 4.0 mmol / L and above 10.0 mmol / L under the nominal scenario is shown (mean ± SD).

[0016] Figure 8 The percentage time spent in the target range, below 4.0 mmol / L and above 10.0 mmol / L under the variance scenario is shown (mean ± SD).

[0017] Figure 9is a block diagram of circuits that, in some embodiments, may be used to implement the various functions, operations, actions, processes and / or methods disclosed herein. DETAILED DESCRIPTION

[0018] In the following detailed description, reference is made to the accompanying drawings, which form a part hereof, and in which are shown by way of illustration specific examples of embodiments in which the present disclosure may be practiced. These embodiments are described in sufficient detail to enable one of ordinary skill in the art to practice the present disclosure. However, other embodiments implemented herein may be utilized, and changes in structure, materials, and processes may be made without departing from the scope of the present disclosure.

[0019] The illustrations presented herein are not intended to be actual views of any particular method, system, device, or structure, but are merely idealized representations for describing embodiments of the present disclosure. In some cases, similar structures or components in the various figures may retain the same or similar numbering for the convenience of the reader; however, similar numbering does not necessarily mean that the structures or components are identical in size, composition, configuration, or any other attributes.

[0020] The following description may include examples to help enable those skilled in the art to practice the disclosed embodiments. The use of the terms "exemplary," "by way of example," and "for example" means that the relevant description is illustrative, and although the scope of the present disclosure is intended to encompass examples and legal equivalents, the use of such terms is not intended to limit the embodiments or the scope of the present disclosure to the specified components, steps, features, functions, etc.

[0021] It will be readily understood that the components of the embodiments as generally described herein and illustrated in the accompanying drawings may be arranged and designed in a wide variety of different configurations. Accordingly, the following description of various embodiments is not intended to limit the scope of the present disclosure, but rather is merely representative of various embodiments. Although various aspects of the embodiments may be presented in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise indicated.

[0022] In addition, the specific implementations shown and described are merely examples and, unless otherwise specified herein, should not be construed as the only way to implement the present disclosure. Components, circuits, and functions may be shown in block diagram form so as not to obscure the present disclosure with unnecessary details. On the contrary, the specific implementations shown and described are merely exemplary and, unless otherwise specified herein, should not be construed as the only way to implement the present disclosure. Additionally, the block definitions and the logical divisions between the various blocks are typical of a particular implementation. It will be readily apparent to one of ordinary skill in the art that the present disclosure can be practiced through many other partitioning solutions. In most cases, details about timing considerations, etc. have been omitted where such details are unnecessary for obtaining a complete understanding of the present disclosure and are within the capabilities of those of ordinary skill in the relevant art.

[0023] Those skilled in the art will appreciate that information and signals can be represented using any of a variety of different technologies and techniques. For clarity of representation and description, some figures may illustrate signals as a single signal. Those skilled in the art will appreciate that a signal can represent a bus of signals, where the bus can have a variety of bit widths, and that the present disclosure can be implemented on any number of data signals, including a single data signal.

[0024] The various illustrative logic blocks, modules, and circuits described in conjunction with the embodiments disclosed herein may be implemented or executed using a general-purpose processor, a special-purpose processor, a digital signal processor (DSP), an integrated circuit (IC), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic device, discrete gate circuits or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general-purpose processor (which may also be referred to herein as a main processor or simply a host) may be a microprocessor, but in an alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. A general-purpose computer including a processor is considered a special-purpose computer, while a general-purpose computer is configured to execute computing instructions (e.g., software code) associated with the embodiments of the present disclosure.

[0025] Embodiments can be described in terms of a process depicted as a flow chart, a flow diagram, a structure diagram, or a block diagram. Although a flow chart can describe an operational action as a sequential process, many of these actions can be performed in another order, in parallel, or substantially concurrently. Additionally, the order of the actions can be rearranged. A process can correspond to a method, a thread, a function, a process, a subroutine, a subprogram, other structures, or a combination thereof. In addition, the methods disclosed herein can be implemented in hardware, software, or both. If implemented in software, the function can be stored or transmitted on a computer-readable medium as one or more instructions or codes. Computer-readable media include both computer storage media (for example, but not limited to, non-transient computer-readable media) and communication media (including any medium that facilitates the transmission of a computer program from one place to another).

[0026] Utilize any reference to element such as " first ", " second " etc. designation herein and do not limit the quantity or order of those elements, unless such restriction is clearly stated.On the contrary, these designations can be used as a convenient method of distinguishing two or more elements or element instances in this article.Therefore, reference to the first and second elements does not mean that only two elements can be adopted there, or that the first element must precede the second element in some way.Additionally, unless otherwise stated, a group of elements can include one or more elements.

[0027] As used herein, the term "substantially" in reference to a given parameter, attribute, or condition means and includes to the extent that one of ordinary skill in the art would understand that the given parameter, attribute, or condition is met with a small degree of variance (such as, for example, within acceptable manufacturing tolerances). For example, depending on the particular parameter, attribute, or condition that is substantially met, the parameter, attribute, or condition may be met by at least 90%, by at least 95%, or even by at least 99%.

[0028] Some patients with diabetes (PWDs) who receive multiple insulin injections use a carbohydrate ratio (CR) and correction factor (CF) to determine mealtime insulin and correction boluses. Due to physiological changes in individual responses to insulin, the physiological characteristics of PWDs, as represented by CR and CF, change over time. Therefore, tracking changes in PWDs' CR and CF is relevant to calculating insulin boluses.

[0029] Various embodiments disclosed herein implement a novel learning method that uses Q-learning (e.g., but not limited to, a model-free reinforcement learning method) to track CR (e.g., but not limited to, the optimal CR) and uses nearest neighbor Q-learning to track CF (e.g., but not limited to, the optimal CF). The learning approach was compared with the run-to-run and Herrero et al. algorithms (presented in P. Herrero et al., “Method for automatic adjustment of an insulin bolus calculator: In silico robustness evaluation under intra-day variability,” Computer Methods and Programs in Biomedicine, vol. 119, No. 1, pp. 1-8, 2018, and P. Herrero et al., “Advanced Insulin Bolus Advisor Based on Run-To-Run Control and Case-Based Reasoning,” IEEE J Biomed Health Inform, vol. 19, No. 3, pp. 1087-96, 2015) over an 8-week period using a validated simulator under realistic scenarios created with suboptimal CR and CF values, carbohydrate counting errors, and random meal sizes at random times. From week 1 to week 8, the learning algorithm increased the percentage of time spent in the target glucose range (3.9 to 10.0 mmol / L) from 51% to 64%, compared to 61% and 58% using the run-by-run and Herrero et al. algorithms, respectively. The learning method reduced the percentage of time spent below 4.0 mmol / L from 9% to 1.9%, compared to 3.4% and 2.3% using the run-by-run and Herrero et al. algorithms, respectively. Thus, the disclosed learning method can improve glucose control in PWD.

[0030] Type 1 diabetes is characterized by autoimmune destruction of pancreatic β islet cells that secrete insulin (a hormone that regulates blood glucose levels via inhibition of hepatic glucose production and promotion of glucose utilization by cells). Insulin replacement therapy to achieve strict glucose control reduces macrovascular and microvascular complications, but hypoglycemia remains a major obstacle to achieving strict glucose targets, and most patients with type 1 diabetes still have suboptimal glucose control. Multiple daily injections (MDIs) and continuous subcutaneous insulin infusion via an insulin pump remain the standard of care for PWD, with MDI becoming the most commonly used therapy due to its lower cost and ease of acquisition.

[0031] As used herein, term " MDI therapy " refers to use multiple (such as, but not limited to, three to four times) insulin (comprising long-acting and quick-acting forms of insulin) injection every day.While long-acting insulin controls background glucose metabolism (background glucose metabolism) throughout the day and night, quick-acting insulin is used for postprandial glucose control and for the correction of hyperglycemia.Although some people who accept MDI therapy use fixed dose and ratio to be used for meal and correction push, the more accurate method of calculating quick-acting insulin dosage is by carbohydrate ratio (CR) and correction factor (CF, also referred to as " insulin sensitivity factor ").CR determines the carbohydrate grams covered by 1 unit insulin push.CF determines the glucose level amount that 1 unit insulin pushes and declines.However, known these CR and CF fluctuate between day and day due to individual physiological changes to the response of insulin.This represents one of the obstacles realizing optimal glucose control, makes PWD and its medical care provider use periodic adjustment to CR and CF.

[0032] Several methods have been proposed to automatically optimize the CR of MDI therapy. Herrero et al. combined a run-by-run framework with a case-based reasoning method, which solves the current situation by recalling similar past situations (in " Advanced Insulin Bolus Advisor Based on Run-To-Run Control and Case-Based Reasoning " by P.Herrero et al., IEEE J Biomed Health Inform, Vol. 19, No. 3, p. 1087-96, 2015). This algorithm was clinically tested in a 6-week feasibility study of 10 adults (M.Reddy et al. " Clinical safety and feasibility of the advanced bolus calculator for type 1 diabetes based on case-based reasoning: a 6-week nonrandomized single-arm pilot study ", Diabetes Technol & Ther, Vol. 18, No. 8, p. 487-493, 2016). Patek et al. developed an algorithm that estimates glucose flux from glucose data and then retrospectively simulates glucose trajectories under different insulin treatments to select the best one (SD Patek et al., “Retrospective optimization of daily insulin therapy parameters: control subject to a regenerative disturbance process,” in Proceedings of the International Federation of Automatic Control Conference, 2016, Vol. 49, No. 7, pp. 773-778).This algorithm was tested in a short-term (48-hour) randomized clinical study of 24 adults (MD Breton et al. "Continuous Glucose Monitoring and Insulin Informed Advisory System with Automated Titration and Dosing of Insulin Reduces Glucose Variabilityin Type 1 Diabetes Mellitus", Diabetes Technol Ther, Vol. 20, No. 8, pp. 531-540, 2018). Tyler et al. developed an artificial intelligence-based decision support system and tested it on retrospective data collected from 25 adults over 28 days (NS Tyler et al. "Anartificial intelligence decision support system for the management of type 1 diabetes", Nature Metabolism, Vol. 2, No. 7, pp. 612-619, 2020). El-Fathi et al. and Toffanin et al. developed run-to-run learning algorithms to adjust CR (A. El Fathi et al., “A Model-Based Insulin Dose Optimization Algorithm for People with Type 1 Diabetes on Multiple Daily Injections Therapy,” IEEE Transactions on Biomedical Engineering, Vol. 68, No. 4, pp. 1208-1219, 2020 and C. Toffanin et al., “Toward a run-to-run adaptive artificial pancreas: In silico results,” IEEE Transactions on Biomedical Engineering, Vol. 65, No. 3, pp. 479-488, 2017).El-Fathi et al.'s algorithm was tested in an 11-day pilot randomized parallel study of 21 young adults, comparing the algorithm's adjustments with those made by a physician (A. ElFathi et al., "A pilot non-inferiority randomized controlled trial to assess automatic adjustments of insulin doses in adolescents with type 1 diabetes on multiple daily injections therapy," Pediatric Diabetes, vol. 21, no. 6, pp. 950-959, 2020). The feasibility of Toffanin et al.'s algorithm was tested in a one-month study of 18 adults (M. Messori et al., "Individually adaptive artificial pancreas in subjects with type 1 diabetes: a one-month proof-of-concept trial in free-living conditions," Diabetes Technol. Ther, vol. 19, no. 10, pp. 560-571, 2017). Finally, Herrero et al. proposed a novel bolus calculator for CR adjustment based on a run-by-run approach (P. Herrero et al., “Method for automatic adjustment of aninsulin bolus calculator: In silico robustness evaluation under intra-day variability,” Computer Methods and Programs in Biomedicine, Vol. 119, No. 1, pp. 1-8, 2018). The bolus calculator proposed by Herrero et al. was tested using computer simulations.

[0033] CR adjustment according to these and other methods may be suboptimal because they do not take into account the correction dose before and after meals. In addition, these and other methods only adjust CR, while CF is calculated from the total daily insulin dose using the static 100 rule equation, which is proposed in "Guidelines for optimal bolus calculator settings in adults" by J.Walsh, R.Roberts and T.Bailey, Journal of Diabetes Science and Technology, Vol. 5, No. 1, pp. 129-135, 2011. Although the 100 rule is widely used for CF estimation, data from children and adolescents show that the actual required dose is on average stronger than those calculated by the 100 rule. In addition, CF may be affected by several factors including puberty, age and body mass index. Some studies have also reported that CF may be different between boys and girls during puberty. Additionally, due to the amplitude of changes in the secretion of different hormones, diurnal variation in insulin sensitivity occurs throughout the day. Therefore, tracking the changes in CF is also useful for calculating the optimal insulin dose.

[0034] In recent years, reinforcement learning has gained increasing popularity to solve a variety of problems, such as drug dosage, autonomous driving, and the board game "Go", by way of non-limiting examples.

[0035] Various embodiments disclosed herein include a novel learning method to simultaneously adjust CR and CF in PWD receiving MDI therapy. Q-learning methods can be used to adjust CR using novel states and reward functions that include the effects of correction doses before and after meals and the rate of change of postprandial glucose levels. To adjust CF, a nearest neighbor method combined with Q-learning can be used to achieve finite sample convergence because pure correction bolus data is scarce. To account for intra-day variability, a CF for daytime and a CF for nighttime can be used. The learning method according to the various embodiments disclosed herein can be compared with the algorithm of Herrero et al. using a validated simulator in a real-world scenario over an 8-week period.

[0036] Figure 11 is a flow chart illustrating a method 100 for determining a meal bolus insulin dose, a correction bolus dose, and / or a basal dose according to some embodiments. At operation 102, the method 100 includes tracking changes in CR using a Q-learning algorithm. At operation 104, the method 100 includes tracking changes in CF using a nearest neighbor Q-learning algorithm. At operation 106, the method 100 includes determining a meal bolus dose in response to the tracked CR and the tracked CF. At operation 108, the method 100 includes estimating a dose.

[0037] A. Bolus Calculation

[0038] The formula used to calculate the bolus dose is as follows:

[0039]

[0040] Among them G m is the blood glucose level (mmol / L), G T is the target glucose level (mmol / L), CHO is the amount of carbohydrate in the meal (g), and IOB is the insulin on board from the previous insulin dose. By way of non-limiting example, formula (1) can be used in Figure 1 Operation 106 of method 100 is at determining a prandial bolus dose.

[0041] B. Problem Modeling

[0042] Reinforcement Learning: The reinforcement learning framework can be assumed to be a discrete Markov decision process (S, A, P, r) with a state space S, an action space A, and transition dynamics P (s k+1 |s k , a k ) and reward r. The agent is in state s k Next take action a k and reaches state s k+1 After receiving the reward The agent's goal is to maximize long-term rewards:

[0043]

[0044] where γ∈[0,1) is a discount factor that weights the preference for immediate (small γ) over future (large γ) rewards.

[0045] The agent chooses its action based on the policy π: S→A. The value function under the policy π is can be described as being in state s k The expected return at time T in the future when following strategy π is:

[0046]

[0047] Using the backward recursion equation, The value can be rewritten as:

[0048]

[0049] in It is action a k When taken from state s k To state s k+1 The state transition probability.

[0050] Using the Bellman optimality equation, we can find the maximum function on the strategy π as:

[0051]

[0052] The optimal policy at time step k is defined as Optimal action-value function is defined as when followed by π * Expected return at time (s):

[0053]

[0054] In some embodiments, a model-free temporal difference method may be used to estimate In such an embodiment, the action-value function (5) can be rewritten as:

[0055]

[0056] where α∈[0, 1) is the learning rate and u is the index of the action in the action space A.

[0057] Nearest Neighbor Q-Learning: Let S be a compact state space and ρ be a metric in S. For every scalar h>0, a finite set of states can be found The state set uses v i Discretize S into the h-covering sphere centered at:

[0058] Make

[0059] ρ(s,v i )<h (8)

[0060] Assume that the state set v has been i , the Q value is estimated, expressed as Q={Q(v i , a), v i ∈S h , a∈A}, then the Q value of any state-action pair (s, a) can be obtained by using the nearest neighbor v iThe weighted average of the Q values ​​in S is estimated as follows:

[0061]

[0062] Where W is a weighting function that satisfies

[0063]

[0064] State-action pair (v i , the Q value of a) can be estimated as follows:

[0065] Q k+1 (v i , a)=(1-α k )Q k (v i , a)+α k JQ k (v i , a) (11)

[0066] where α k ∈[0, 1) is the learning rate, and JQ is the average value of each state-action pair (v i , the joint nearest neighbor Q value operator of a) is as follows:

[0067]

[0068] Unlike standard Q-learning, each state-action pair (v i , the Q value of a) is estimated using all states present in its neighborhood.

[0069] In order to learn the optimal policy, the agent should visit all state-action pairs and improve its current policy by choosing actions that have been tried in the past and that have contributed the most to the cumulative reward. One way to achieve this is to use an ε-greedy policy. For example, the ε-greedy policy was proposed in RS Sutton and AG Barto, "Reinforcement learning: An introduction", 2011, where the agent explores (trying new actions) with probability ε and exercises (using experience) with probability 1-ε. The agent uses exercise to exploit prior knowledge and exploration to identify new options. The agent chooses the optimal action to generate the maximum possible reward for a given state. During the learning period, ε starts at 0.9 and gradually decreases to a small value of 0.1.

[0070] C. Learning Methods for CR

[0071] In this section, states, actions, and rewards are defined for the CR learning algorithm.

[0072] 1) State Space

[0073] Status for each meal type (breakfast, lunch, and dinner) including (i) the rate of change of glucose in the period between t1 and t2 after the meal, and (ii) postprandial glucose error, defined as:

[0074] E k =G min (k)-G T (13)

[0075] Among them G T is the target glucose level and G min (k) is the lowest glucose level in the period between t3 and t4 after the meal. t1, t2, t3, and t4 are tuning parameters. In the case of a pre- or post-meal correction bolus, a predicted blood glucose profile is used, which is calculated after each correction bolus using a linear prediction model with parameters of the current glucose level and a weighted average of the rate of change of two consecutive glucose values ​​within the last 60-minute window.

[0076] value is the percentage of time spent below 4.0 mmol / L between 1 and 6 hours after a meal. Can be defined as the rate of change (ROC) range Error margin|E k |≤E T and percentage of time ROC T and E T This choice of state representation and target state allows the method to target strict and stable postprandial glucose levels.

[0077] 2) Action Space

[0078] Let A(s m ) is state s m The set of all possible actions under . For all states, the action space A can be defined as follows:

[0079] A(s m )={a m |1,-1,0} (14)

[0080] where 1, -1, and 0 represent an increase, decrease, and unchanged CR relative to the previous day's value, respectively.

[0081] 3) Reward Function

[0082] Taking action After that, the agent receives the following rewards

[0083]

[0084] Where n is a scaling multiplier that is equal to 1 if the glucose level in the postprandial period is not less than 4.0 mmol / L; otherwise, it is equal to 2. m is defined as:

[0085]

[0086] and Is to record state-action pairs A counter of how many times it has been accessed.

[0087] The reward function (15) encourages the learning method to take the approach of not increasing the postprandial glucose error E k+1 and glucose change rate If the action results in the desired state without hypoglycemia The agent receives a large positive reward. If the action does not result in a state change without hypoglycemia, the agent receives a constant reward of 1. In all other cases, the agent receives a negative reward. m is included to learn from the state-action pair during the initial learning phase by To the next state The offline average reward of gives an additional incentive to the learning agent to promote early and safe convergence.

[0088] The details of the learning method for CR estimation are given in Algorithm 1.

[0089]

[0090] k1=0.2, k2=0.15 and k h =(0.05E k -0.05) was chosen empirically using clinical guidelines and aimed to converge within 7 iterations.

[0091] Figure 2 is a flow chart illustrating a learning method 200 for CR according to some embodiments. By way of non-limiting example, the learning method 200 may be used to Figure 1 Algorithm 1 discussed above is an example of the learning method 200 .

[0092] At operation 202, the learning method 200 includes initializing a discount factor that weights the preference for immediate versus future rewards, a learning rate, and an ε-greedy policy probability. At operation 204, the learning method 200 includes initializing an action-value function using clinical guidelines. At operation 206, the learning method 200 includes evaluating the current state.

[0093] The learning method 200 repeats operations 208, 210, 212, 214, and 216. At operation 208, the learning method 200 includes selecting an action from the current state using an ε-greedy policy. At operation 210, the learning method 200 includes obtaining a CR value from the selected action. At operation 212, the learning method 200 includes applying the obtained CR value to observe the reward and the next state. At operation 214, the learning method 200 includes updating the action-value function. At operation 216, the learning method 200 includes evaluating the next state. The learning method 200 may return to operation 208 to use the next state as the current state.

[0094] D. Learning Methods for CF

[0095] In this section, a learning method for estimating CF is disclosed. A CF for daytime (7:00 AM to 12:00 AM) and a CF for nighttime (12:00 AM to 7:00 AM) are disclosed. The CF for daytime and nighttime are estimated using a nearest neighbor Q-learning algorithm. The nearest neighbors in conjunction with the Q-learning method can be chosen to achieve finite sample convergence because pure (i.e., not accompanied by a meal bolus) correction bolus data is scarce.

[0096] 1) State Space

[0097] The state used for the CF learning algorithm in interval I including (i) the glucose change rate at the time of the correction bolus and the period between t1 and t2 after the correction bolus, respectively and (ii) Glycemic Risk Index (GRI) k,c , calculated as a linear combination of the hypoglycemic and hyperglycemic components, and (iii) the corrected glucose error, defined as:

[0098] E′ k =G min,c (k)-G T (17)

[0099] Among them G T is the target glucose level and G min,c (k) is the lowest glucose level in the period between t3 and t4 after the correction bolus injection. t1, t2, t3 and t4 are tuning parameters.

[0100] value It can be the percentage of time spent below 4.0 mmol / L between 1 and 6 hours after the correction bolus. is defined as |E′ k |≤E T and ROC T and E T The choice of state representation and target state is the same as discussed above for the CR learning algorithm, with the addition of a glycemic risk index parameter to track the overall quality of glycemia in interval I.

[0101] 2) Action Space

[0102] Set A(s c ) is state s c The set of all possible actions under . For all states, the action space A can be defined as follows:

[0103] A(s c )={a c |1, -1, 0} (18)

[0104] where 1, -1, and 0 represent an increase, decrease, and unchanged CF relative to the previous day's value, respectively.

[0105] 3) Reward Function

[0106] In response to taking action Agents receive the following rewards

[0107]

[0108] Where n is a scaling multiplier that is equal to 1 if the glucose level during the corrected time period is not less than 4.0 mmol / L, and equal to 2 otherwise. c and r c is defined as:

[0109]

[0110] in Is to record state-action pairs A counter of how many times it has been accessed.

[0111] The reward function (19) encourages the learning method to consider the GRI in interval I. k,c At the same time, the correction of glucose error E′ is not increased. k+1 and glucose change rate If the action results in the desired state without hypoglycemia The agent receives a large positive reward. If the action does not result in a state change without hypoglycemia, the agent receives a constant reward of 1. In all other cases, the agent receives a negative reward. c is added to learn during the initial learning phase by generating a state-action pair for each successful action. To the next state The offline average reward of gives additional incentives to the learning agent to promote early and safe convergence.

[0112] The details of the learning method for CF estimation are given in Algorithm 2.

[0113]

[0114] k c1 =0.2, k c2 = 0.15 and k c3 =(0.05E k,c -0.05) was chosen empirically using clinical guidelines and aimed at convergence of the algorithm within 7 iterations. Is to record state-action pairs A counter of how many times it has been accessed.

[0115] Figure 3 is a flow chart illustrating a learning method 300 for CF according to some embodiments. By way of non-limiting example, the learning method 300 may be used to Figure 1 104 to track changes in CF. Algorithm 2 discussed above is an example of the learning method 300.

[0116] At operation 302, the learning method 300 includes initializing a discount factor that weights the preference for immediate versus future rewards and an ε-greedy policy probability. At operation 304, the learning method 300 includes constructing a finite state space set. At operation 306, the learning method 300 includes initializing an action-value function using clinical guidelines. At operation 308, the learning method 300 includes setting a counter value to zero for each state-action pair, the counter value indicating the number of times the corresponding state-action pair has been visited. At operation 310, the learning method 300 includes evaluating the current state.

[0117] The learning method 300 repeats operations 312, 314, 316, 318, and 320. At operation 312, the learning method 300 includes selecting an action from the current state using an ε-greedy policy. At operation 314, the learning method 300 includes obtaining a CF value from the selected action. At operation 316, the learning method 300 includes applying the obtained CF value to observe the reward and the next state. At operation 318, the learning method 300 includes determining a joint nearest neighbor Q-value operator for each state closest to the current state and incrementing the corresponding counter value by one. At operation 320, the learning method 300 includes determining an updated action-value function and a next learning rate for each state closest to the current state. The learning method 300 may return to operation 312 to use the next state as the current state.

[0118] E. Learning Algorithms for Basal

[0119] In this subsection, states, actions, and rewards are defined for the base learning algorithm.

[0120] 1) State Space

[0121] State of the art for base learning methods Includes: (i) Average glucose error U k , is calculated as the average glucose G in the time period between 2:00 am and 7:00 am mean Subtract target glucose level G T , (ii) the percentage of time spent below 4.0 mmol / L and above 10.0 mmol / L, respectively, between 2:00 am and 7:00 am and (iii) glucose rate of change during the time period between 4:00 AM and 7:00 AM, and (iv) hypoglycemia treatment during the time period between 10:00 PM and 2:00 AM. Can be defined as and

[0122] The choice of state representation and target state allows the algorithm to target strict and stable fasting glucose levels.

[0123] 2) Action Space

[0124] Set A(s b ) can be state s b The set of all possible actions under . For all states, the action space A can be defined as follows:

[0125] A(s b )={ab |1, -1, 0} (22)

[0126] where 1, -1, and 0 represent an increase, decrease, and unchanged basis, respectively, relative to the previous day's value.

[0127] 3) Reward Function

[0128] In response to taking action Agent receives rewards

[0129]

[0130] Where n is a scaling multiplier that is equal to 1 if the glucose level during the period between 2:00 AM and 7:00 AM is not less than 4.0 mmol / L, and is equal to 2 otherwise. b is defined as:

[0131]

[0132] and Is to record state-action pairs A counter of how many times it has been accessed.

[0133] The reward function (23) encourages the agent to take actions that do not increase the average glucose error ΔU k 、 and If the selected action leads to the target state The agent receives a large positive reward. If the chosen action does not result in a change of state, the agent receives a constant reward of 1. In all other cases, the agent receives a negative reward.

[0134] The details of the learning method for basis estimation are given in Algorithm 3.

[0135]

[0136] 3. Repetition

[0137] 4. Using the ε-greedy strategy from Select an action

[0138] 5. Use the following from action a k Get B k+1 :

[0139]

[0140] 6. Application B k+1 , observe the reward and the next state

[0141] 7. Update the action-value function (7)

[0142] 8.

[0143] k1 = 0.2, k2 = 0.15 and was chosen empirically using clinical guidelines and with the goal of convergence of the algorithm within 7 iterations.

[0144] To use the learning method safely and to ensure its robustness in abnormal days, updates of CR, CF, and baseline values ​​are limited to ±20% of the previous day's value.

[0145] Figure 4 is a flow chart illustrating a method 400 for estimating a basal dose according to some embodiments. By way of non-limiting example, the method 400 may be used to Figure 1 The method 400 is performed in operation 108 to estimate the dose for the bolus or basal dose. Algorithm 3 discussed above is an example of the method 400.

[0146] At operation 402, method 400 includes initializing a discount factor that weights the preference for immediate versus future rewards, a learning rate, and an ε-greedy policy probability. At operation 404, method 400 includes initializing an action-value function using clinical guidelines. At operation 406, method 400 includes evaluating the current state.

[0147] Method 400 repeats operations 408, 410, 412, 414, and 416. At operation 408, method 400 includes selecting an action from the current state using an ε-greedy strategy. At operation 410, method 400 includes obtaining a basal or bolus value from the selected action. At operation 412, method 400 includes applying the obtained basal or bolus value to observe the reward and the next state. At operation 414, method 400 includes updating the action-value function. At operation 416, method 400 includes evaluating the next state.

[0148] F. System Architecture of Bolus Calculator

[0149] Figure 55 is a block diagram of the architecture of a system 500 for an advanced bolus calculator 502 for a PWD with type 1 diabetes (e.g., individual 504) receiving MDI therapy, according to some embodiments. The methods discussed above (e.g., learning method 200, learning method 300, method 400) can use glucose, insulin, and meal data (e.g., at the end of each day) to calculate CR and CF (e.g., optimal CR and CF) for the next day using Algorithms 1 and 2 or using learning method 200 and learning method 300.

[0150] In addition to the advanced bolus calculator 502 and the individual 504, the system 500 includes a Q-learning method 506, a nearest neighbor Q-learning method 508, a state calculation 510, and a state calculation 512. By way of non-limiting example, the Q-learning method 506 may include Figure 2 Learning method 200. Also by way of non-limiting example, the Q learning method 506 may include Algorithm 1 discussed above. By way of non-limiting example, the nearest neighbor Q learning method 508 may include Figure 3 Learning method 300. Also by way of non-limiting example, the nearest neighbor Q-learning method 508 may include Algorithm 2 discussed above.

[0151] State calculation 510 determines a state in response to the received glucose history for individual 504 and data corresponding to meals and correction boluses. The state can be determined in response to one or more of the glucose data, insulin data, or meal data. State calculation 510 delivers the determined state to Q-learning method 506. Q-learning method 506 determines a CR in response to the determined state from state calculation 510 and the determined reward. Q-learning method 506 provides the CR to advanced bolus calculator 502.

[0152] State calculation 512 determines a state in response to the received glucose history and data corresponding to the meal and correction bolus for individual 504. State calculation 512 delivers the determined state to nearest neighbor Q-learning method 508. Nearest neighbor Q-learning method 508 determines a CF in response to the determined state from state calculation 512 and the determined reward. Nearest neighbor Q-learning method 508 provides the CF to advanced bolus calculator 502.

[0153] The advanced bolus calculator 502 determines a bolus dose in response to the CR, CF, a target glucose level (e.g., but not limited to, determined by a healthcare professional), and data corresponding to the meal. A bolus dose of insulin substantially corresponding to the determined bolus dose can be delivered to the individual 504 (e.g., but not limited to, using an injection pen, using an insulin pump).

[0154] Figure 66 is a block diagram of a system 600 according to some embodiments. The system 600 includes a treatment delivery system 602, a mobile device 604, one or more cloud servers 608, a medical care provider device 610, and in some embodiments, a glucose sensor 612. The glucose sensor may include an in vivo glucose sensor having a first portion configured to be placed subcutaneously and a second portion configured to be arranged on the skin. The second portion may be coupled to sensor electronics that include one or more of a power supply, a processor, and communication circuitry such as a transceiver for communicating with another component (e.g., mobile device 604). Communication may be via a wireless communication protocol. Figure 5 Portions of the system 500 may be performed by one or more of the therapy delivery system 602 , the mobile device 604 , the one or more cloud servers 608 , and the medical care provider device 610 .

[0155] By way of non-limiting example, the advanced bolus calculator 502 can be executed by the therapy delivery system 602. Also by way of non-limiting example, the therapy delivery system 602 can be executed by a mobile application 606 executed by a mobile device 604 or at one or more cloud servers 608 and then provided to the therapy delivery system 602 by the mobile device 604. As a specific non-limiting example, the one or more cloud servers 608 can execute the Q-learning method 506, the nearest neighbor Q-learning method 508, the state calculation 510, and the state calculation 512, and deliver the CR and CF to the therapy delivery system 602 via the mobile application 606. As another specific non-limiting example, if the therapy delivery system 602 has sufficient processing power, the therapy delivery system 602 can perform operations corresponding to the Q-learning method 506, the nearest neighbor Q-learning method 508, the state calculation 510, and the state calculation 512.

[0156] The therapeutic delivery system 602 may include one or more injection pens, insulin pumps, or other therapeutic delivery systems. By way of non-limiting example, the therapeutic delivery system 602 may include an injection pen including a cap including electronics that enable the therapeutic delivery system 602 to communicate with a glucose sensor 612 and a mobile application 606 executed by a mobile device 604. In some such examples, the glucose sensor 612 may detect an individual (e.g., but not limited to, Figure 5The device 604 can detect glucose levels in the individual 504 and provide the detected glucose levels with corresponding timestamps to the therapy delivery system 602 (e.g., via an RFID scanner of the therapy delivery system 602). The therapy delivery system 602 can interact with the mobile application 606 to deliver this glucose information to the mobile device 604. Alternatively, the glucose information can be obtained by the individual via a blood test and manually entered into an interface of the therapy delivery system 602 or the mobile application 606.

[0157] In some embodiments, the treatment delivery system 602 can also be configured to capture data corresponding to a meal. By way of non-limiting example, the treatment delivery system 602 can include an injection pen that includes a cap that includes electronics that enable the treatment delivery system 602 to detect the removal of the cap on the pen needle to detect that the pen is being used to deliver an injection before a meal. The electronics can detect the time the cap is removed and estimate the number of calories for the meal based on the time of day (e.g., if it is in the morning, a breakfast-sized dose; if it is near noon, a lunch-sized dose; if it is in the evening, a dinner-sized dose). Also by way of non-limiting example, the pen cap or mobile application 606 can include an interface that enables an individual to enter an estimate of how many calories they will consume in an upcoming meal. For example, options such as "small," "medium," and "large" can be provided, and the user can select one of the options, each corresponding to a number of carbohydrates. As another example, a list of selectable carbohydrate quantities can be presented to the individual, and the individual can select the carbohydrate quantity estimated to be most suitable for the upcoming meal.

[0158] In some embodiments, the healthcare provider device 610 can enable a healthcare professional to input information such as a target glucose level, which can be communicated by the medical treatment system 600 to whichever device will ultimately calculate the bolus. For example, if the therapy delivery system 602 will ultimately be used to calculate the bolus, the healthcare provider device 610 can be used to program the therapy delivery system 602 with the target glucose level.

[0159] G. Baseline Learning Method

[0160] The learning method discussed above is compared with the run-by-run algorithm and the algorithm of Herero et al. In the run-by-run algorithm, the CR of each main meal m (breakfast, lunch, and dinner) is adjusted as follows:

[0161]

[0162] Among them CR m is the initial CR of meal m, and is the percentage of time with hypoglycemia and hyperglycemia during the seven-hour period after a meal (or until the next meal if within seven hours), and G mean is the average post-meal glucose level over a seven-hour period (or until the next meal if within seven hours).

[0163] In the algorithm of Herero et al., the CR for each meal m is adjusted as follows:

[0164]

[0165] Where GHO is the amount of carbohydrate in the diet, G m is the glucose level at mealtime, G T is the target glucose level, W is the individual weight in kilograms (kg), B is the recommended insulin dose at mealtime, IOB represents the remaining insulin in the body, and B add is defined as:

[0166]

[0167] and It is the lowest glucose level between two and five hours after a meal.

[0168] For the baseline algorithm, CF is calculated using the static 100-rule equation.

[0169] H. Computer Simulation (In-Silico) Environment

[0170] The computer simulation environment included a PWD based on the Hovorka model of R. Hovorka et al. ("Partitioning glucose distribution / transport, disposal, and endogenous production during IVGTT," Am J Physiol Endocrinol Metab, Vol. 282, No. 5, pp. e992-e1007, 2002). The parameters in the model were sampled from a lognormal distribution, with the mean values ​​of the parameters obtained from Wilinska et al., "Simulation environment to evaluate closed-loop insulin delivery systems in type 1 diabetes," J Diabetes Sci Technol, Vol. 4, No. 1, pp. 132-44, 2010, and correlations from healthy individual data.

[0171] By making the model parameters of each individual oscillate periodically with random phase and frequency, and by adding daily random insulin and glucose fluxes, inter-day glucose variability was increased. By randomly varying the peak time of meal absorption, variability in meal absorption was achieved, which accounts for fast and slow carbohydrate meals (as discussed in "Stochastic Virtual Population of Subjects With Type 1 Diabetes for the Assessment of Closed-Loop Glucose Controllers" by A. Haidar et al., IEEE Trans Biomed Eng, Vol. 60, No. 12, pp. 3524-33, 2013). The simulated environment included noise in the glucose measurements, with a correlation of 80% and a coefficient of variation of 7% (see "Modeling the glucose sensor error" by A. Facchinetti et al., IEEE Trans Biomed Eng, Vol. 61, No. 3, pp. 620-9, 2014).

[0172] Each computer-simulated individual has a unique and optimal CR and long-acting basal insulin dose. A clinical dataset of 81 individuals was used to validate the simulator. The feasibility of the learning method disclosed herein was evaluated on 100 computer-simulated PWDs over an 8-week period. Each computer-simulated individual was initialized with non-optimal CR and CF values ​​by imposing a uniform random error of 30-70% on the optimal CR and CF values. Whenever the glucose level dropped below 4.0 mmol / L, each computer-simulated individual ate 15 g of carbohydrates, and this was repeated every 15 minutes until the glucose level increased to above 4.0 mmol / L.

[0173] The learning algorithm was tested in two scenarios: nominal and variance. In the nominal scenario, each computer-simulated individual consumed 40g, 60g, and 80g of carbohydrates for breakfast (7:00 AM), lunch (2:00 PM), and dinner (8:00 PM), respectively. In the variance scenario, random carbohydrate counting errors (with a mean of zero and a coefficient of variation of 40%) and random variation were added to the size and timing of the consumed meals, as shown in Table 1.

[0174]

[0175] Table I: Meal plan for variance scenario (uniform distribution)

[0176] A daytime correction bolus was generated if: (i) the glucose level was ≥10 mmol / L for at least 154 (90-205) minutes after a meal (e.g., but not limited to, due to carbohydrate counting error), similar to data reported from two studies encompassing 49,995 days, 296,685 meals (including snacks), and 61,654 insulin correction boluses, (ii) the glucose level was ≥10 mmol / L for at least 30 minutes after a snack, or (iii) the glucose level was ≥10 mmol / L due to missing 2.1 (0-9) boluses per week. Similarly, a nighttime correction bolus was generated if the glucose level was ≥10 mmol / L or if the glucose level was ≥10 mmol / L for at least 30 minutes after a snack. In real time, these nighttime correction boluses may occur due to suboptimal basal dosing, the dawn phenomenon, daytime physical activity, a nighttime snack, or a suboptimal dinner bolus.

[0177] III. Experimental Results

[0178] Figure 7 The percentage of time spent in the target range, below 4.0 mmol / L and above 10.0 mmol / L under the nominal scenario is shown (mean ± SD). Blood glucose results from the three algorithms (the disclosed algorithm, the run-by-run algorithm, and the algorithm of Herrero et al.) are reported for both standard and variance scenarios. Before starting the 60-day simulation, the first 7 days were simulated to calculate baseline results.

[0179] The tuning parameters are selected as t1 = t3 = 2.5 hours, t2 = t4 = 6 hours, and The selected value t i|i=1,2,3,4 and ROC T The intuition behind this is to account for the variability due to different profiles of meals and insulin absorption (e.g., low and high glycemic index foods). The threshold chosen The intuition behind this is to make the updates of the learning algorithm robust in the face of perturbations such as carbohydrate counting errors, delayed insulin boluses, and metabolic variability.

[0180] Nominal scenario: Figure 7The changes in blood glucose results for 100 computer-simulated individuals over 60 days using the methods disclosed herein and a baseline algorithm are shown. Comparing the final week of the 60-day simulation to the baseline, the average time spent in the target range increased from 55% to 68% using the embodiments disclosed herein, compared to 64% and 58% using the run-by-run and Herrero et al. algorithms, respectively. The average time spent below 4 mmol / L decreased from 8% to 0.8% using the embodiments disclosed herein, compared to 2.8% using the run-by-run algorithm and 0.5% using the Herrero et al. algorithm. The average time spent above 10 mmol / L decreased from 37% to 31% using the embodiments disclosed herein, compared to 33% using the run-by-run algorithm; the average time spent above 10 mmol / L increased to 41% using the Herrero et al. algorithm. Tables II and III report the blood glucose results, including overall, daytime, nighttime, post-meal, and post-correction time periods.

[0181] Figure 8 The percentage time spent in the target range, below 4.0 mmol / L and above 10.0 mmol / L under the variance scenario is shown (mean ± SD). Figure 8 The changes in blood glucose results for 100 computer-simulated individuals over 60 days using the embodiments disclosed herein and a baseline algorithm are shown. Comparing the final week of the 60-day simulation to the baseline, the average time spent in the target range increased from 51% to 64% using the embodiments disclosed herein, compared to 61% and 58% using the run-by-run and Herrero et al. algorithms, respectively. The average time spent below 4 mmol / L decreased from 9% to 1.9% using the embodiments disclosed herein, compared to 3.4% using the run-by-run algorithm and 2.3% using the Herrero et al. algorithm. The average time spent above 10 mmol / L decreased from 40% to 34% using the embodiments disclosed herein, compared to 35% using the run-by-run algorithm; the average time spent above 10 mmol / L remained unchanged using the Herrero et al. algorithm. Tables II and III report the blood glucose results, including overall, daytime, nighttime, post-meal, and post-correction time periods.

[0182]

[0183]

[0184]

[0185] Table II: Summary of blood glucose results under nominal and variance scenarios. Results are average

[0186] (SD).

[0187] IV. Discussion of Bolus Calculator

[0188] PWDs live with the lifelong burden of managing their glucose levels. With the rapid development of glucose sensor technology and smart insulin pens, it has become possible to automatically adjust PWD's therapy parameters (e.g., CR, CF, and basal dose) using algorithms. This paper discloses an advanced bolus calculator for individuals receiving MDI therapy that adjusts their CR and CF by analyzing their glucose and insulin data.

[0189] The disclosed bolus calculator is designed using a Q-learning approach to adjust the CR and a nearest neighbor Q-learning approach to adjust the CF. To account for diurnal variations in insulin sensitivity, separate CFs are disclosed for daytime and nighttime. To evaluate the performance of the disclosed bolus calculator, a computer simulation environment was used to test its performance by: (i) adding random carbohydrate counting errors with zero mean and a 40% coefficient of variation; and (ii) using random meal sizes and random meal intake times. Despite this, the disclosed bolus calculator improved time on target and reduced hypoglycemia.

[0190] Compared to the algorithm of Herrero et al., the disclosed method achieved a greater reduction in hypoglycemia during the day and the entire 24-hour period in both nominal and variance scenarios (Table II). In addition, the disclosed method achieved a greater time on target during the night and the 24-hour period in both nominal and variance scenarios (Table II). These benefits of using the disclosed method are explained by multiple factors: (i) the disclosed method adjusts the daytime and nighttime CFs separately, and (ii) the disclosed method uses novel state and reward functions designed to promote early and safe convergence.

[0191] Given the rising incidence of diabetes and the increasing shortage of specialized endocrinologists, the responsibility for managing PWD may increasingly shift to primary care physicians. Since some physicians may not be familiar with the use of glucose sensors and smart insulin pens, the disclosed bolus calculator can help make insulin dosage decisions efficiently in primary care settings. In addition, the disclosed algorithm can be utilized to make dosage adjustments more frequently than a physician (e.g., weekly versus every three to six months).

[0192] The disclosed method may have several limitations. First, the disclosed bolus calculator assumes that the meal insulin bolus is calculated based solely on carbohydrate content. Studies have shown that carbohydrate-matched, high-fat meals require higher insulin doses and can lead to prolonged hyperglycemia for up to five hours after the meal. Second, although the computer simulation studies discussed above show improvements in glycemic outcomes, simulations have their own shortcomings. Simulations tend to overestimate the benefits of interventions because they do not account for all the perturbations and uncertainties present in real-world settings.

[0193] V. Conclusion

[0194] Disclosed are methods for automatically adjusting the parameters of an insulin bolus calculator for people with pre-existing diabetes mellitus (PWD) receiving MDI therapy, as well as related systems and devices. The disclosed method is developed using reinforcement learning using a nearest neighbor approach applied to continuous glucose monitoring and insulin data. The disclosed method was tested on 100 computer-simulated individuals. Simulation results support the effectiveness of the method in improving glucose control.

[0195]

[0196]

[0197]

[0198] Table III: Summary of blood glucose results for the post-meal and post-correction bolus periods. Results are average

[0199] (SD).

[0200] Those skilled in the art will appreciate that the functional elements (eg, functions, operations, actions, processes, and / or methods) of the embodiments disclosed herein may be implemented in any suitable hardware, software, firmware, or a combination thereof. Figure 9 The diagrams illustrate non-limiting examples of implementations of the functional elements disclosed herein.In some embodiments, some or all of the functional elements disclosed herein may be executed by hardware specifically configured to perform these functional elements.

[0201] Figure 9900 is a block diagram of circuitry 900, which, in some embodiments, may be used to implement various functions, operations, actions, processes, and / or methods disclosed herein. Circuitry 900 includes one or more processors 902 (sometimes referred to herein as "processors 902") operatively coupled to one or more data storage devices (sometimes referred to herein as "storage devices 904"). Storage devices 904 include machine-executable code 906 stored thereon, and processors 902 include logic circuitry 908. Machine-executable code 906 includes information describing functional elements that may be implemented (e.g., executed) by logic circuitry 908. Logic circuitry 908 is adapted to implement (e.g., execute) the functional elements described by machine-executable code 906. When executing the functional elements described by machine-executable code 906, circuitry 900 should be considered to be dedicated hardware configured to perform the functional elements disclosed herein. In some embodiments, processor 902 may be configured to execute the functional elements described by machine-executable code 906 sequentially, concurrently (e.g., on one or more different hardware platforms), or in one or more parallel processing streams.

[0202] When implemented by the logic circuitry 908 of the processor 902, the machine executable code 906 is configured to adapt the processor 902 to perform the operations of the embodiments disclosed herein. For example, the machine executable code 906 may be configured to adapt the processor 902 to perform Figure 1 Method 100, Figure 2 Learning methods 200, Figure 3 Learning methods 300, Figure 4 Method 400, Figure 5 Q-learning method 506, Figure 5 Nearest neighbor Q learning method 508, Figure 5 State calculation 510, Figure 5 As another example, the machine executable code 906 may be configured to cause the processor 902 to execute at least a portion or all of the state calculation 512, Algorithm 1, Algorithm 2, and / or Algorithm 3. Figure 6 Therapeutic delivery system 602, Figure 6 Mobile applications 606, Figure 6 One or more cloud servers 608 and / or Figure 6 at least some or all of the operations discussed above for the medical care provider device 610 .

[0203] The processor 902 may include a general-purpose processor, a special-purpose processor, a central processing unit (CPU), a microcontroller, a programmable logic controller (PLC), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, other programmable devices or any combination thereof designed to perform the functions disclosed herein. A general-purpose computer including a processor is considered a special-purpose computer, while a general-purpose computer is configured to execute functional elements corresponding to the machine executable code 906 (e.g., software code, firmware code, hardware description) associated with the embodiments of the present disclosure. It is noted that the general-purpose processor (also referred to herein as a host processor or simply a host) may be a microprocessor, but in an alternative, the processor 902 may include any conventional processor, controller, microcontroller or state machine. The processor 902 may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in combination with a DSP core, or any other such configuration.

[0204] In some embodiments, the storage device 904 includes a volatile data storage device (e.g., random access memory (RAM)) and a non-volatile data storage device (e.g., flash memory, hard disk drive, solid-state drive, erasable programmable read-only memory (EPROM), etc.). In some embodiments, the processor 902 and the storage device 904 may be implemented as a single device (e.g., a semiconductor device product, a system on a chip (SOC), etc.). In some embodiments, the processor 902 and the storage device 904 may be implemented as separate devices.

[0205] In some embodiments, machine executable code 906 may include computer-readable instructions (e.g., software code, firmware code). By way of non-limiting example, the computer-readable instructions may be stored by storage device 904, directly accessed by processor 902, and executed by processor 902 using at least logic circuitry 908. Also by way of non-limiting example, the computer-readable instructions may be stored on storage device 904, transferred to a memory device (not shown) for execution, and executed by processor 902 using at least logic circuitry 908. Thus, in some embodiments, logic circuitry 908 includes electrically configurable logic circuitry 908.

[0206] In some embodiments, machine executable code 906 may describe hardware (e.g., circuitry) to be implemented in logic circuit 908 to perform a functional element. This hardware may be described at any of a variety of levels of abstraction, from low-level transistor layouts to high-level description languages. At a high level of abstraction, a hardware description language (HDL) such as an IEEE standard hardware description language (HDL) may be used. By way of non-limiting example, a VERILOG TM 、SYSTEMVERILOG TM or Very Large Scale Integration (VLSI) Hardware Description Language (VHDL TM ).

[0207] The HDL description can be converted to a description at any of a variety of other levels of abstraction as desired. As non-limiting examples, the high-level description can be converted to a logic-level description, such as a register transfer language (RTL), a gate-level (GL) description, a layout-level description, or a mask-level description. As non-limiting examples, the micro-operations to be performed by the hardware logic circuits (e.g., but not limited to, gates, flip-flops, registers) of logic circuit 908 can be described in RTL and then converted to a GL description by a synthesis tool, and the GL description can be converted to a layout-level description by a placement and routing tool, which corresponds to the physical layout of an integrated circuit of a programmable logic device, discrete gate or transistor logic, discrete hardware components, or a combination thereof. Therefore, in some embodiments, the machine executable code 906 may include an HDL, RTL, GL description, a mask-level description, other hardware descriptions, or any combination thereof.

[0208] In embodiments where the machine-executable code 906 includes a hardware description (at any level of abstraction), a system (not shown, but including the storage device 904) can be configured to implement the hardware description described by the machine-executable code 906. By way of non-limiting example, the processor 902 can include a programmable logic device (e.g., an FPGA or a PLC), and the logic circuit 908 can be electronically controlled to implement circuitry corresponding to the hardware description into the logic circuit 908. Also by way of non-limiting example, the logic circuit 908 can include hard-wired logic manufactured by a manufacturing system (not shown, but including the storage device 904) according to the hardware description of the machine-executable code 906.

[0209] Regardless of whether the machine executable code 906 includes computer-readable instructions or a hardware description, the logic circuit 908 is adapted to perform the functional elements described by the machine executable code 906 when implementing the functional elements of the machine executable code 906. Note that although a hardware description may not directly describe a functional element, the hardware description indirectly describes the functional elements that the hardware elements described by the hardware description are capable of performing.

[0210] As used in this disclosure, the term "module" or "component" may refer to a specific hardware implementation configured to perform the actions of the module or component, and / or a software object or software routine that may be stored on and / or executed by general-purpose hardware of a computing system (e.g., a computer-readable medium, a processing device, etc.). In some embodiments, the different components, modules, engines, and services described in this disclosure may be implemented as objects or processes that execute on a computing system (e.g., as separate threads). Although some of the systems and methods described in this disclosure are generally described as being implemented in software (stored on and / or executed by general-purpose hardware), specific hardware implementations or combinations of software and specific hardware implementations are also possible and contemplated.

[0211] As used in this disclosure, the term "combination" with respect to a plurality of elements may include all of the combinations of the elements, or any of various subcombinations of some of the elements. For example, the phrase "A, B, C, D, or a combination thereof" may refer to any one of A, B, C, or D; a combination of each of A, B, C, and D; and any subcombination of A, B, C, or D, such as A, B, and C; A, B, and D; A, C, and D; B, C, and D; A and B; A and C; A and D; B and C; B and D; or C and D.

[0212] The terms used in this disclosure and especially in the appended claims (e.g., the bodies of the appended claims) are generally intended to be “open” terms (e.g., the term “including” should be interpreted as “including, but not limited to,” the term “having” should be interpreted as “having at least,” the term “comprising” should be interpreted as “including, but not limited to,” etc.).

[0213] Additionally, if a specific number of claim recitations being introduced is intended, such intent will be expressly recited in the claim, and in the absence of such recitation, such intent is absent. For example, to aid understanding, the following appended claims may contain the use of the introductory phrases "at least one" and "one or more" to introduce claim recitations. However, the use of such phrases should not be construed to imply that the introduction of a claim recitation by the indefinite article "a" or "an" will limit any particular claim containing such introduced claim to embodiments containing only one such recitation, even when the same claim includes the introductory phrases "one or more" or "at least one" and an indefinite article such as "a" or "an" (e.g., "a" or "an" should be interpreted to mean "at least one" or "one or more"); the same applies to the use of definite articles used to introduce claims.

[0214] Furthermore, even if a specific number of claims is explicitly recited, those skilled in the art will recognize that such a recitation should be interpreted as meaning at least the recited number (e.g., the unmodified recitation of "two recitations" without other modifiers means at least two recitations, or two or more recitations). Furthermore, in those cases where a convention similar to "at least one of A, B, and C, etc." or "one or more of A, B, and C, etc." is used, such a construction is generally intended to include A alone, B alone, C alone, A and B together, A and C together, B and C together, or A, B, and C together, etc.

[0215] Furthermore, any disjunctive word or phrase indicating two or more alternative terms, whether in the specification, claims, or drawings, should be understood to include the possibility of one, either, or both of the terms. For example, the phrase "A or B" should be understood to include the possibility of "A" or "B" or "A and B."

[0216] Although the present disclosure has been described herein with respect to certain illustrated embodiments, those skilled in the art will recognize and appreciate that the invention is not limited thereto. Rather, many additions, deletions, and modifications may be made to the illustrated and described embodiments without departing from the scope of the invention as claimed below and its legal equivalents. Additionally, features from one embodiment may be combined with features from another embodiment while still falling within the scope of the invention as contemplated by the inventors.

[0217] Exemplary embodiments are set forth in the following numbered clauses:

[0218] 1. A system for determining an insulin dosage, the system comprising one or more computer-readable storage media having computer-readable instructions stored thereon, the computer-readable instructions configured to instruct the one or more processors to:

[0219] determining status based on the user's glucose, meal, and insulin information;

[0220] Based on the determined state, a Q-learning algorithm is used to track changes in carbohydrate ratio (CR);

[0221] Based on the determined state, tracking changes in a correction factor (CF) using a nearest neighbor Q-learning algorithm; and

[0222] In response to the tracked CR and the tracked CF, a prandial bolus, correction bolus, or basal dose is determined.

[0223] 2. A system according to claim 1, wherein the system is configured to instruct the one or more processors to track the change in the CR by: initializing an action-value function, selecting an action from a current state, obtaining a CR value based on the selected action, and determining a reward for the obtained CR value.

[0224] 3. The system of clause 1, wherein the computer-readable instructions are configured to instruct the one or more processors to track changes in the CR by:

[0225] Initialize the discount factor, learning rate, and ε-greedy policy probability that weight the preference for immediate versus future rewards;

[0226] Initialize the action-value function using clinical guidelines;

[0227] Assess the current status; and

[0228] Repeat the following:

[0229] Use the ε-greedy strategy to select an action from the current state;

[0230] Get the CR value from the selected action;

[0231] Apply the obtained CR value to observe the reward and next state;

[0232] Updating the action-value function; and

[0233] Evaluate the next state.

[0234] 4. The system of any preceding clause, wherein the computer-readable instructions are configured to instruct the one or more processors to track the change in CF using the nearest neighbor Q-learning algorithm by:

[0235] Initialize the discount factor that weights the preference for immediate versus future rewards and the ε-greedy policy probability;

[0236] Construct a finite state space set;

[0237] Initialize the action-value function using clinical guidelines;

[0238] setting a counter value to zero for each state-action pair, the counter value indicating the number of times the corresponding state-action pair has been visited;

[0239] Assess the current status; and

[0240] Repeat the following:

[0241] Use the ε-greedy strategy to select an action from the current state;

[0242] Get the CF value from the selected action;

[0243] Apply the obtained CF value to observe the reward and next state;

[0244] For each state closest to the current state, determine the joint nearest neighbor Q value operator and increment the corresponding counter value by one; and

[0245] For each state closest to the current state, an updated action-value function and the next learning rate are determined.

[0246] 5. A system according to any of the preceding clauses, wherein the computer readable instructions are configured to estimate a basal dose.

[0247] 6. The system of clause 5, wherein the computer-readable instructions are configured to estimate the basal dose by:

[0248] Initialize the discount factor, learning rate, and ε-greedy policy probability that weight the preference for immediate versus future rewards;

[0249] Initialize the action-value function using clinical guidelines;

[0250] Assess the current status; and

[0251] Repeat the following:

[0252] Use the ε-greedy strategy to select an action from the current state;

[0253] Obtain a base dose value from the selected action;

[0254] Apply the obtained basal dose value to observe the reward and next state;

[0255] Updating the action-value function; and

[0256] Evaluate the next state.

[0257] 7. A system according to any preceding clause, wherein the computer readable instructions are configured to:

[0258] A Q-learning algorithm is used to track changes in basal or bolus doses.

[0259] 8. A system according to any of the preceding clauses, wherein one or more of the Q-learning algorithm or the nearest neighbor Q-learning algorithm includes a state and a reward function.

[0260] 9. A system according to claim 8, wherein the state and reward function is based at least in part on a representation of one or more of: the effect of pre-meal and post-meal correction injections, the post-meal glucose error determined within a specified post-meal time window, or the rate of change of glucose level determined within a specified post-meal time window.

[0261] 10. The system of clause 9, wherein the designated postprandial time window is an optimal postprandial time window for measuring glucose concentration, wherein the optimal postprandial time window is optionally predetermined.

[0262] 11. A method for determining a prandial bolus insulin dose, the method comprising:

[0263] determining status based on the user's glucose, meal, and insulin information;

[0264] Based on the determined state, a Q-learning algorithm is used to track changes in carbohydrate ratio (CR);

[0265] Based on the determined state, tracking changes in a correction factor (CF) using a nearest neighbor Q-learning algorithm; and

[0266] In response to the tracked CR and the tracked CF, the dose of the meal bolus is determined.

[0267] 12. The method of claim 11, wherein tracking the change in the CR comprises initializing an action-value function, selecting an action from a current state, obtaining a CR value based on the selected action, and determining a reward for the obtained CR value.

[0268] 13. The method of clause 11, wherein tracking the changes in the CR comprises:

[0269] Initialize the discount factor, learning rate, and ε-greedy policy probability that weight the preference for immediate versus future rewards;

[0270] Initialize the action-value function using clinical guidelines;

[0271] Assess the current status; and

[0272] Repeat the following:

[0273] Use the ε-greedy strategy to select an action from the current state;

[0274] Get the CR value from the selected action;

[0275] Apply the obtained CR value to observe the reward and next state;

[0276] Updating the action-value function; and

[0277] Evaluate the next state.

[0278] 14. The method of any one of statements 11 to 13, wherein tracking the changes in the CF comprises:

[0279] Initialize the discount factor that weights the preference for immediate versus future rewards and the ε-greedy policy probability;

[0280] Construct a finite state space set;

[0281] Initialize the action-value function using clinical guidelines;

[0282] setting a counter value to zero for each state-action pair, the counter value indicating the number of times the corresponding state-action pair has been visited;

[0283] Assess the current status; and

[0284] Repeat the following:

[0285] Use the ε-greedy strategy to select an action from the current state;

[0286] Get the CF value from the selected action;

[0287] Apply the obtained CF value to observe the reward and next state;

[0288] For each state closest to the current state, determine the joint nearest neighbor Q value operator and increment the corresponding counter value by one; and

[0289] For each state closest to the current state, an updated action-value function and the next learning rate are determined.

[0290] 15. The method according to any one of clauses 10 to 14, further comprising estimating a basal dose or a bolus dose.

[0291] 16. The method of clause 15, wherein estimating the basal dose or bolus dose comprises:

[0292] Initialize the discount factor, learning rate, and ε-greedy policy probability that weight the preference for immediate versus future rewards;

[0293] Initialize the action-value function using clinical guidelines;

[0294] Assess the current status; and

[0295] Repeat the following:

[0296] Use the ε-greedy strategy to select an action from the current state;

[0297] Obtain a base dose value from the selected action;

[0298] Apply the obtained basal dose value to observe the reward and next state;

[0299] Updating the action-value function; and

[0300] Evaluate the next state.

[0301] 17. A method according to any one of clauses 10 to 16, comprising:

[0302] A Q-learning algorithm is used to track changes in basal or bolus doses.

[0303] 18. A method according to any of clauses 10 to 17, wherein one or more of the Q-learning algorithm or the nearest neighbour Q-learning algorithm comprises a state and a reward function.

[0304] 19. A method according to clause 18, wherein the state and reward function is based at least in part on a representation of one or more of: the effect of pre-meal and post-meal correction boluses, the post-meal glucose error determined within a specified post-meal time window, or the rate of change of glucose level determined within a specified post-meal time window.

[0305] 20. The method according to clause 19, wherein the specified postprandial time window is an optimal postprandial time window for measuring glucose concentration, wherein the optimal postprandial time window is optionally predetermined.

Claims

1. A system for determining an insulin dosage, the system comprising one or more processors and one or more computer-readable storage media having computer-readable instructions stored thereon, the computer-readable instructions configured to instruct the one or more processors to: determining status based on the user's glucose, meal, and insulin information; Based on the determined state, a Q-learning algorithm is used to track changes in carbohydrate ratio (CR); Based on the determined state, a nearest neighbor Q-learning algorithm is used to track changes in the correction factor (CF); as well as A prandial bolus, correction bolus, or basal dose is determined in response to the tracked changes in CR and the tracked changes in CF.

2. The system of claim 1 , wherein the system is configured to instruct the one or more processors to track the change in the CR by initializing an action-value function, selecting an action from a current state, obtaining a CR value based on the selected action, and determining a reward for the obtained CR value.

3. The system of claim 1 , wherein the system is configured to instruct the one or more processors to track the changes in the CR by: Initialize the discount factor, learning rate, and ε-greedy policy probability that weight the preference for immediate versus future rewards; Initialize the action-value function using clinical guidelines; Assess the current state; as well as Repeat the following: Select an action from the current state using an ε-greedy strategy; Get the CR value from the selected action; Apply the obtained CR value to observe the reward and next state; Updating the action-value function; and Evaluate the next state.

4. The system of claim 1 , wherein the computer-readable instructions are configured to instruct the one or more processors to utilize the nearest neighbor Q-learning algorithm to track the changes in the CF by: Initialize the discount factor that weights the preference for immediate versus future rewards and the ε-greedy policy probability; Construct a finite state space set; Initialize the action-value function using clinical guidelines; For each state-action pair, setting a counter value to zero, the counter value indicating the number of times the corresponding state-action pair has been visited; Assess the current state; as well as Repeat the following steps: Select an action from the current state using an ε-greedy strategy; Get the CF value from the selected action; Apply the obtained CF value to observe the reward and next state; For each state closest to the current state, determine a joint nearest neighbor Q value operator and increment the corresponding counter value by one; as well as For each state closest to the current state, an updated action-value function and next learning rate are determined. The system of claim 1 , wherein the computer readable instructions are configured to estimate a basal dose.

6. The system of claim 5, wherein the computer readable instructions are configured to estimate the basal dose by: Initialize the discount factor, learning rate, and ε-greedy policy probability that weight the preference for immediate versus future rewards; Initialize the action-value function using clinical guidelines; Assess the current state; as well as Repeat the following steps: Select an action from the current state using an ε-greedy strategy; Obtain a base dose value from the selected action; Apply the obtained basal dose value to observe the reward and next state; Updating the action-value function; and Evaluate the next state.

7. The system of claim 1 , wherein the computer-readable instructions are configured to: A Q-learning algorithm is used to track changes in basal or bolus doses.

8. The system of claim 1, wherein one or more of the Q-learning algorithm or the nearest neighbor Q-learning algorithm comprises a state and a reward function.

9. A system according to claim 8, wherein the state and reward function is based at least in part on a representation of one or more of: pre-meal and post-meal correction bolus effects, post-meal glucose error determined within a specified post-meal time window, or rate of change of glucose level determined within a specified post-meal time window. 10 . The system of claim 9 , wherein the designated postprandial time window is an optimal postprandial time window for measuring glucose concentration, wherein the optimal postprandial time window is predetermined.

11. A method for determining a prandial bolus insulin dose, the method comprising: determining status based on the user's glucose, meal, and insulin information; Based on the determined state, a Q-learning algorithm is used to track changes in carbohydrate ratio (CR); Based on the determined state, a nearest neighbor Q-learning algorithm is used to track changes in the correction factor (CF); as well as The dose of the prandial bolus is determined in response to the tracked changes in CR and the tracked changes in CF.

12. The method of claim 11 , wherein tracking the changes in the CR comprises: Initialize the action-value function, select an action from the current state, obtain a CR value based on the selected action, and determine a reward for the obtained CR value.

13. The method of claim 11 , wherein tracking the changes in the CR comprises: Initialize the discount factor, learning rate, and ε-greedy policy probability that weight the preference for immediate versus future rewards; Initialize the action-value function using clinical guidelines; Assess the current state; as well as Repeat the following steps: Select an action from the current state using an ε-greedy strategy; Get the CR value from the selected action; Apply the obtained CR value to observe the reward and next state; Updating the action-value function; and Evaluate the next state.

14. The method of claim 11 , wherein tracking the changes in the CF comprises: Initialize the discount factor that weights the preference for immediate versus future rewards and the ε-greedy policy probability; Construct a finite state space set; Initialize the action-value function using clinical guidelines; For each state-action pair, setting a counter value to zero, the counter value indicating the number of times the corresponding state-action pair has been visited; Assess the current state; as well as Repeat the following steps: Select an action from the current state using an ε-greedy strategy; Obtaining a CF value from the selected action; Apply the obtained CF value to observe the reward and next state; For each state closest to the current state, determine a joint nearest neighbor Q value operator and increment the corresponding counter value by one; as well as For each state closest to the current state, an updated action-value function and next learning rate are determined.

15. The method of claim 11, further comprising estimating a basal dose or a bolus dose.

16. The method of claim 15, wherein estimating the basal dose or bolus dose comprises: Initialize the discount factor, learning rate, and ε-greedy policy probability that weight the preference for immediate versus future rewards; Initialize the action-value function using clinical guidelines; Assess the current state; as well as Repeat the following steps: Select an action from the current state using an ε-greedy strategy; Obtain a base dose value from the selected action; Apply the obtained basal dose value to observe the reward and next state; Updating the action-value function; and Evaluate the next state.

17. The method according to claim 11, comprising: A Q-learning algorithm is used to track changes in basal or bolus doses.

18. The method of claim 11, wherein one or more of the Q-learning algorithm or the nearest neighbor Q-learning algorithm comprises a state and a reward function.

19. A method according to claim 18, wherein the state and reward function is based at least in part on a representation of one or more of: pre-meal and post-meal correction bolus effects, post-meal glucose error determined within a specified post-meal time window, or rate of change of glucose level determined within a specified post-meal time window.

20. The method of claim 19, wherein the designated postprandial time window is an optimal postprandial time window for measuring glucose concentration, wherein the optimal postprandial time window is predetermined.