Reinforced learning method for optimizing patient therapy delivery by a medical device

The reinforced learning method addresses the inefficiencies in current SCS therapy optimization by using a model-based approach to adaptively adjust therapy parameters in response to changing patient states and device performance, enhancing pain relief efficacy and reducing the need for manual recalibrations.

WO2025119668A1PCT designated stage expired Publication Date: 2025-06-12BIOTRONIK SE & CO KG
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/083253
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-08
Filing Date
2024-11-22
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Current methods for optimizing spinal cord stimulator (SCS) therapy parameters are inefficient, relying on manual adjustments that are infrequent and only responsive to long-term trends, leading to suboptimal pain relief and potential therapy ineffectiveness over time.

Method used

A computer-implemented reinforced learning method that dynamically optimizes SCS therapy parameters by using a model-based approach, incorporating real and simulated environments to adaptively adjust parameters in response to changing patient states and device performance.

Benefits of technology

This method enables personalized, adaptive optimization of SCS therapy parameters, improving pain relief efficacy by continuously updating therapy settings in response to real-time patient data and feedback, thus avoiding the need for frequent manual recalibrations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024083253_12062025_PF_FP_ABST
    Figure EP2024083253_12062025_PF_FP_ABST
Patent Text Reader

Abstract

A reinforced learning method is used to optimize patient therapy delivered by an implanted medical device. The real environment of the device and patient provides inputs to a reinforced learning algorithm which has an agent that uses them as states and rewards. The agent applies a policy to map states into actions and modify the real environment according to the actions. The reinforced learning algorithm further comprises a model that simulates the real environment and generates simulated states and rewards and outputs these to the agent. The agent thus receives both real and simulated states and rewards and maximizes rewards according to the value function responsive to both the real and simulated environments. The rewards are then used to affect the real environment, including through associating rewards with control parameters of the medical device to optimize therapy.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] REINFORCED LEARNING METHOD FOR OPTIMIZING PATIENT THERAPY DELIVERY BY A MEDICAL DEVICE

[0002] The invention relates to the application of reinforced learning for optimizing patient therapy for pain relief and other patient physiological, psychological or behavioral states with a spinal cord stimulator (SCS) device or another implantable medical device (IMD).

[0003] An SCS device is a type of IMD that is used for chronic pain therapy of a patient. SCS devices and other types of IMD are also used to modify and / or modulate physiological, psychological or behavioral states of a patient. An example physiological state is posture or gait. An example psychological state is depression or stress. An example behavioral state is sleep or pain.

[0004] SCS is a treatment for chronic pain that uses electrical current sent through electrodes on one or more leads implanted in the epidural space dorsal to the spinal cord. Stimulation parameters include at least: pulse width, frequency, amplitude and electrode selection. Additionally, duty cycling and burst stimulation can be introduced as further stimulation parameters to modulate therapy and better manage battery consumption through the introduction of alternating periods of carrier and envelope stimulation.

[0005] These SCS parameters are first set during initial therapy programming. However, given the large number of individual parameters, the set of possible parameter combinations is extensive. The choice of parameters needs to take into account several factors, such as manufacturers recommendations, therapy type (sub-perception or traditional), pain etiology, pain location, lead location, patient anatomy, stimulation perception threshold, previous experience with an SCS device, physician preference, and others. The parameter space is characterized by high dimensionality, and inevitable constraints and discontinuities make the therapeutically effective parameter space sparse. These characteristics of the parameter space and the nonlinear nature of the neurophysiology make it difficult to find the optimal set parameter set. First, multiple local optima may exist, making the search for the global optimum difficult. Secondly, nonlinearity including discontinuity in the parameter space may cause the local tuning to diverge, driving the solution farther away from the optimum. Hence, the analytical solution of the optimal set of parameters is usually unknown, and parameter optimization is typically achieved empirically, for each patient, over time. While initial tuning can be attained shortly after the implant, the real challenge begins as patients resumes performance of normal daily functions. Daily activity, as well as sleep and medication, may impact the effectiveness of the therapy and the pain state, thus altering the parameter space. When the parameter and pain state space changes over time, the combination of parameters that were previously optimal may become suboptimal and may even become substantially ineffective. Therefore, capturing the changes in the environment parameter and state space, preferably with minimal delay, is important for maintaining therapy efficacy over the long term.

[0006] In summary, there is a need to (1) individually model the therapy environment and state space including the patient, (2) update the therapy parameter space taking into account the present state of the patient and the history of the patient as well as device states, and to (3) adaptively adjust therapy parameters, or recommend such adjustments, to provide a dynamically changing optimal configuration of the therapy parameter space that updates continually within a relatively short time frame to respond to the patient’s instant pain state or reporting state. This will provide an improvement over adjustments that are infrequent and only responsive to long term trends and patient feedback to a clinician.

[0007] In the state of the art, for sub-perception therapy, parameter optimization is achieved manually by following a carefully developed workflow to test multiple parameter sets until a patient experiences satisfactory pain relief. Device optimization to a patient typically takes place in two steps. There is an initial set up during a trial period of approximately one week and a further set up after permanent implant of the device. SCS therapy is discontinued if sufficient pain relief is not achieved by the end of the trial or after permanent implant if attempts to re-program the device are unsuccessful. US 2021 / 0252288 Al utilizes a model that estimates the relationship between a plurality of pain states of a patient and the SCS stimulation parameters based on input from the patient who provides pain state information. Although the model generates time-varying therapy in response to the input changes, the model becomes quite resistant to change as more and more data is accumulated by the model over time, because an ever greater amount of data is required to alter the model. Moreover, error accumulation due to state drift may cause a set of therapy parameters that were previously effective to become less effective or substantially ineffective.

[0008] According to first aspect of the disclosure there is provided a computer-implemented (model-based) reinforced learning method for optimizing therapy and / or treatments delivered to a patient by a medical device and / or a practitioner, the method comprising: receiving states and rewards from a real environment; providing an agent comprising a policy and a value function, the agent being configured to apply the policy to map the states received as input into actions and output the actions to cause modification of the real environment; providing a model configured to simulate the real environment as a simulated environment and connected to receive as input the actions from the agent, the model being further configured to generate simulated states and simulated rewards and to output these to the agent, so that the agent receives real and simulated states and rewards; and iteratively optimizing the agent and / or the agent-actions and / or -interventions according to the value function to maximize rewards from the real and simulated environments through the agent’s receipt of the real and simulated states and awards.

[0009] A (model-based) reinforcement learning (RL) algorithm is applied to dynamically optimize therapy parameters by adjusting both the environment and a model of the environment over time, where the environment includes a medical device such as an SCS device and the patient. The (model-based) RL algorithm takes advantage of the knowledge of the dynamics of the environment to predict the states relevant to the pain experience, enabling proactive and timely optimization of the therapy. Personalized optimization of the control of SCS parameters can be achieved over time. The dynamics of a patient’s decision-making and behaviors can be included as inputs into the model. Moreover, patients and / or physicians may set priorities in the behaviors and conditions, which are reflected in the reward used to optimize the therapy parameters. The contributions to, and impacts of, pain can be included in the RL algorithm to predict the pain state and to optimize therapy for individual patients. For example, the model can determine an optimal set of therapy parameters to apply having regard to the instantaneous pain experienced by the patient and other behavioral / physiological states in order to move the patient away from undesired patient states into desired target patient states.

[0010] Certain embodiments of the invention are able to provide a personalized, i.e., patientspecific, model-based pain treatment intervention which is optimized according to outputs of a computer-automated process. The pain treatment intervention encompasses SCS parameter control and recommended action output to patients as may be provided by remote care via a medical device communication system that is in communication with a neuroservice data center or the like allowing input from personnel and automated systems.

[0011] An adaptive model of the environment can be provided where environment encompasses patient states and SCS states as well as their inter-related dynamics.

[0012] A reward can be provided that is personalized by being based on each patient’s priorities in states and actions.

[0013] The patient-specific optimization provided by the model automatically varies over time according to changes in patient behavior, patient physiology or patient psychology and changes in device performance, since all these can be included in the real environment and the simulated environment. The automatic updating of the optimization can therefore avoid the need for manual re-programming or other interventions and can provide a more effective therapy outcome for an SCS patient. The model runs a simulation of the real environment and is configured to continually update its estimates of the patient states even when the patient feedback from the real environment is absent, enabling intervention and changes to the intervention even without patient feedback from the real environment. The model is also updated and optimized adaptively in response to real states and real rewards from the real environment. Through these features, the model simulating the real environment can address long term drift in data. This can be important, since when a patient is under a long-term treatment regimen, for example using SCS, it is often observed that a patient’s baseline in respect of behavioral, physiological, psychological or other factors changes. This change is referred to as drift, since the changes are most often gradual and continuous. However, sometimes a jump, i.e., discontinuous change, is observed. The patient baseline set at the time the IMD was implanted through the original calibration is therefore no longer optimum. The model incorporates a mechanism that can track and adapt to such deviation, thus increasing the success probability and effective duration of a long-term treatment without manual intervention. A full recalibration in a clinical setting can thereby be avoided or at least deferred.

[0014] The RL method may further comprise: the model receiving as input the states and rewards from the real environment; and iteratively updating the simulated environment of the model responsive to the real states and rewards.

[0015] In some embodiments of the RL method, at least one other one of the rewards is associated with a care intervention for the patient.

[0016] In some embodiments of the RL method, certain ones of the rewards are associated with respective values of control parameters of a medical device that is delivering therapy to the patient.

[0017] According to a second aspect of the disclosure there is provided a computer readable medium storing instructions for performing the computer-implemented RL method of the first aspect.

[0018] According to a third aspect of the disclosure there is provided a medical device communication system, comprising: a medical device for delivering therapy to a patient, the medical device delivering the therapy according to a set of control parameters; and a remote computing resource arranged in data communication connection with the medical device, the remote computing resource comprising a processor and a computer readable medium storing instructions for performing the RL method of the first aspect on the processor, wherein the method uses sensor data received from the medical device as states and rewards, and wherein the remote computing resource is further configured to transmit the control parameter values determined by the agent to the medical device.

[0019] In some embodiments of the medical device communication system, the remote computing resource is further configured to output a care intervention message according to at least one other one of the rewards.

[0020] According to a fourth aspect of the disclosure there is provided a medical device for treating a patient, comprising: an electrical energy source; therapeutic components for patient treatment that are powered by the electrical energy source, the therapeutic components being driven according to a set of control parameters; a receiver arranged to receive control data with values of the control parameters via a data communication connection; and a memory and a processor, wherein the memory stores a computer program which when executed on the processor varies the control parameters based on control parameter values received by the receiver, said control parameter values having been determined by a remote computing resource according to the RL method of the first aspect.

[0021] In some embodiments, the medical device of the fourth aspect further comprises a transmitter arranged to transmit sensor data via the data communication connection to the remote computing resource, the sensor data providing states and rewards for the computer- implemented RL method performed by the remote computing resource. The states and rewards are input to the agent and optionally also to the model.

[0022] In some embodiments, the medical device of the fourth aspect further comprises a control circuit arranged to receive the control parameters and a drive circuit arranged to deliver electrical signals to at least one electrode according to the control parameters. In one example, the control circuit, drive circuit and the at least one electrode are configured to provide spinal cord stimulation to a patient.

[0023] According to a fifth aspect of the disclosure there is provided a method of delivering therapy to a patient via a medical device, the method comprising: providing the patient with a medical device having: an electrical energy source; therapeutic components for patient treatment that are powered by the electrical energy source, the therapeutic components being driven according to a set of control parameters; a receiver arranged to receive control data with values of the control parameters via a data communication connection from a remote computing resource; and a memory and a processor, storing a computer program in the memory of the medical device which when executed on the processor of the medical device varies the control parameters of the medical device based on control parameter values received by the receiver; determining the control parameter values at the remote computing resource according to a RL method, the RL method comprising: receiving states and rewards from a real environment; providing an agent comprising a policy and a value function, the agent being configured to apply the policy to map the states received as input into actions and output the actions to cause modification of the real environment; providing a model configured to simulate the real environment as a simulated environment and connected to receive as input the actions from the agent, the model being further configured to generate simulated states and simulated rewards and to output these to the agent, so that the agent receives real and simulated states and rewards; and iteratively optimizing the agent and / or the agent-actions and / or -interventions according to the value function to maximize rewards from the real and simulated environments through the agent’s receipt of the real and simulated states and awards, wherein certain ones of the rewards are associated with respective values of the control parameters.

[0024] In some embodiments, the method of treatment of the fifth aspect is performed such that the medical device further comprises a transmitter arranged to transmit sensor data via the data communication connection to the remote computing resource, and the sensor data providing states and rewards for the RL method performed by the remote computing resource.

[0025] In some embodiments, the method of treatment of the fifth aspect is performed such that the model receives as input the states and rewards from the real environment, and wherein the RL method iteratively updates the simulated environment of the model responsive to the real states and rewards.

[0026] In some embodiments of the system and / or methods, a / the medical device is configured to aid treatment of a patient, the medical device comprising sensors which sense the patient state or environment; a / the receiver is arranged to receive data from the medical device with values for the patient state or environment via a data communication connection, and a data processing system is configured process this data and provides input to the agent according to at least one of the reinforced learning methods (described above).

[0027] In some embodiments, the system further comprises a / the transmitter arranged to transmit sensor data via the data communication connection to the remote computing resource, and / or the sensor data providing states and rewards for the computer-implemented reinforced learning method performed by the remote computing resource.

[0028] In some embodiments, the states and / or rewards are input to the agent and / or input to the model.

[0029] This invention will now be further described, by way of example only, with reference to the accompanying drawings. Figure 1 is a schematic diagram of an IMD with an SCS stimulator as part of a medical device communication system with standard architecture.

[0030] Figure 2 is a schematic diagram showing the SCS stimulator attached to a lead arrangement of electrodes.

[0031] Figure 3 is a schematic diagram of a RL algorithm according to an embodiment of the invention.

[0032] Figure 4 is a schematic diagram shows the RL algorithm of Figure 3 in further detail.

[0033] Figure 5 is a flow diagram of an embodiment of the RL algorithm.

[0034] Figure 6 shows an example model of patient behavior which includes dynamic interactions between activity, pain, and sleep.

[0035] Figure 7 is a block diagram illustrating an example computing apparatus that may be used in connection with embodiments of the invention.

[0036] In the following detailed description, for purposes of explanation and not limitation, specific details are set forth in order to provide a better understanding of the present disclosure. It will be apparent to one skilled in the art that the present disclosure may be practiced in other embodiments that depart from these specific details.

[0037] Those skilled in the art will further appreciate that the services, functions and steps explained herein may be implemented using software, i.e. a computer program, stored in memory and functioning in conjunction with a programmed microprocessor, or using an Application Specific Integrated Circuit (ASIC), a Digital Signal Processor (DSP), a Programmable Logic Array (PLA), or a field programmable gate array (FPGA). As such references to a processor should include ASICs including artificial intelligence accelerator ASICs, DSPs, PLAs and FPGAs as well as central processor units (CPUs), graphics processor units (GPUs). References to memory in the following may refer to any one or more of: a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), and a static random access memory (SRAM).

[0038] References to computer program in the following refer to machine readable program instructions for carrying out operations and may be assembler instructions, instruction-set- architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages.

[0039] It will be understood that features and elements of a standard architecture described above with reference to Figure 1 will be incorporated in embodiments of the invention as appropriate and specific detail already described above may not be repeated in the following detailed description for corresponding features.

[0040] An IMD typically forms part of a medical device communication system (MDCS) that is configured to upload clinical data from the IMD on a continual basis via a MDCS to a remotely located data repository, which may be a data center. In an SCS device, clinical data includes both patient data relating to physiological monitoring of a patient's health, for example as collected by sensors of the IMD, and device data related to status and operation of the IMD, for example data logging each time an SCS device is used and at what stimulation level in terms of electrical pulse width and pulse frequency as well as control signals issued to the SCS device to optimize delivery of pain relief. These clinical data can then be accessed by health care staff accessing the data center using suitably designed software applications. Treatment plans and, if necessary, interventions can then be decided upon based on analysis of these clinical data.

[0041] Figure 1 shows a standard architecture of a MDCS for communication between an IMD implanted in a patient and a remotely located data center, which is accessible to health care staff using suitable application software. The IMD 10 includes therapeutic components for patient treatment, such as an SCS device 11 (or another IMD stimulator or IMD monitor) incorporating an implantable pulse generator (IPG), and therapeutic components for patient monitoring, such as an IMD sensor 13. The IMD 10 further comprises chipsets providing computing and telecommunications resource in the form of memory 17, a processor 18 and a wireless personal area network (WPAN) transceiver 12. Instead of or in addition to the WPAN transceiver 12, the IMD 10 may comprise a low-power wide-area network (LPWAN) transceiver 15. The IMD 10 is powered by an electrical energy source 14, in this example by a rechargeable battery 14. The IMD's rechargeable battery 14 of an implanted IMD can be charged in a contactless manner by a charger 20, for example a resonant inductive charger, which is placed on the patient's skin adjacent the implanted IMD 10 to charge its rechargeable battery 14. The charger 20 is itself provided with an energy source 24, here a rechargeable battery 24, as well as an external mains power connection 29, for example an external power jack, to power the charger 20 so that its rechargeable battery 24 can be recharged. Alternatively, the charger's battery 24 can be recharged wirelessly by placing on a mains-powered charging pad. The IMD WPAN transceiver 12 uses a suitable WPAN protocol such as Bluetooth Low Energy (BLE), Medical Implant Communication System (MICS) or Medical Device Radiocommunications Service (MedRadio), the latter two being almost identical protocols.

[0042] The patient is provided with a patient remote controller 30 (hereinafter also called “patient remote 30”) and / or a patient remote software, for example a smartphone with wireless transceivers 32, 34, 35, 36 respectively for WPAN, LPWAN, cellular (e.g., 4G / LTE / 5G) and WLAN communication. If a smartphone is used, this is a smartphone that is possessed, e.g., owned, by the patient in which the IMD 10 is implanted and on which a software application (‘app’) is installed. The app is then a so-called Software as a Medical Device (SaMD) which is defined by the United States Food and Drug Administration (FDA) as software intended to be used for one or more medical purposes that perform these purposes without being part of a hardware medical device.

[0043] The cellular transceiver 35 and WLAN transceiver 36 provide two different data communication paths for the patient remote 30 to upload clinical data from the IMD 10 via an internet connection 50 to a remotely located data center 60 (hereinafter also called “backend 60”) acting as repository for storage of clinical data and / or as a host for the services, for example a neuro-service data center. The remote’s WLAN transceiver 36 can upload data to the backend 60 via a router 38 and telephone line 40 using a wired internet connection 50. The internet connection 50 may provide access to one or more distributed networks and / or cloud services. The remote’s cellular transceiver 35 can upload data to the backend 60 via one or more cellular network base stations 42 (cellular towers), for example LPWAN-capable cellular network base station. In some cases, instead of a public communication network, a dedicated point-to-point transmission, e.g., via a dedicated telephone line, may be provided for uploading clinical data. In all these scenarios, uploading of clinical data from the IMD 10 to the backend 60 takes place via the intermediary of the patient remote 30, the latter thereby acting as a relay device.

[0044] Health care staff, such as health care professionals (HCPs), clinical specialists and representatives and remote care team members have access to the backend 60 via suitable portals 70 with the aid of a software application running on the backend 60 and / or the portal 70 to provide the necessary user interfacing, diagnostics and so forth. A portal 70 (or several portals 70) are provided by at least one workstation for health care staff to access a data center, e.g., the backend 60. Analysis software may also be run at the backend 60 to analyze clinical data from individual patients or groups of patients.

[0045] Figure 2 is a schematic diagram showing the SCS device 11 attached to a lead arrangement 80 comprising one or more leads 82 with each lead incorporating one or more electrodes 84. In the illustrated example, there is one electrode per lead and the electrodes 84 are N in number and labeled, El, E2, E3 .... EN. The SCS device 11 comprises an electrode drive circuit 19 for providing a suitable drive current to each of the electrodes 84 according to a treatment plan as controlled by a control circuit 16. The control circuit 16 delivers control signals to the drive circuit according to a treatment plan devised by a computer program running on the processor 18 (e.g., a microprocessor) of the IMD 10. The computer program is stored in the IMD memory 17. These hardware and software components operate collectively to provide an intelligent pulse generator to deliver electrical pulses to the electrodes conforming to a particular set of parameters, including amplitude, pulse width, frequency and / or duty cycle to provide stimulation therapy. For SCS, the electrode leads are implanted at or near a patient’s spinal cord to direct electrical signals into the patient’s tissue for spinal cord stimulation. The SCS device is implanted subcutaneously.

[0046] Embodiments of the invention use a RL algorithm to optimize delivery of therapy by a medical device, such as for chronic pain therapy or for other behavioral outcomes.

[0047] Figure 3 is a schematic diagram of operation of a RL algorithm 90 according to an embodiment of the invention and Figure 4 is schematic diagram of an example showing further detail. The overall purpose of the RL algorithm 90 is to optimize therapy for the patient. The patient and any medical device delivering therapy to that patient, e.g., an SCS device, forms the environment whose parameters the RL algorithm aims to optimize. The environment 100 encompasses the dynamic interactions among the patient’s physiological, psychological, and behavioral states 104, 106, 108 as well as the physical components of the SCS device 110. The environment 100 delivers states and rewards to the RL algorithm 90. The RL algorithm 90 has an agent 120 configured to act according to a policy 122 and a value function 124. The agent 120 is a decision-making entity which makes decisions according to policy 122. Policy 122 is a function that gives a probability associated with each of a plurality of possible actions 126 when the environment is in a particular state 112. The value function 124 reflects a perception of the “goodness” of a state and predicts the reward 114 of being in a particular state 112. Each environment experience induced by a decision is assigned a reward 114, which is then accumulated over time to estimate the longterm value. The optimal policy 122 is the one that maximizes the value function 124, i.e., provides the biggest cumulative reward 114.

[0048] Compared with a standard RL algorithm, the RL algorithm as described with reference to Figures 3 and 4 contains an additional component which is a model 100’ that mimics or simulates the environment 100. Elsewhere in this document, we therefore refer to the environment 100 as the real environment and the model 100’ as modeling the simulated environment. The objective of modeling the simulated environment is to estimate the states as accurately and precisely as possible to match the states of the real environment. The simulated environment 100’ includes the same dynamic interactions among the patient’s physiological, psychological, and behavioral states as in the real environment 100, these simulated counterparts in the model being labeled 104’, 106’, 108’ respectively. The physical components of the SCS device are also simulated as labeled 110’. The simulated environment 100’ of the model then delivers simulated states and rewards 112’, 114’ to the RL algorithm 90. In addition, the model 100’ is updated by the decisions and experiences (i.e., states and rewards) from the real environment 100 so that it tracks changes that occur in the real environment 100. The model 100’ simulates the experiences while varying simulated decisions to search for the best decision that would produce the best experience. By including a model which simulates the environment within the RL algorithm 90, the agent 120 can continue optimization in cooperation with the simulated environment 100’ even during periods of time when the real environment 100 is not supplying state and reward information to the agent 120. The inclusion of the model 100 thus serves to increase the optimization speed of the RL algorithm 90 in situations where state inputs from the real environment 100 are infrequent and / or are associated with long reaction times.

[0049] The RL algorithm 90 takes advantage of the knowledge of the dynamics and state-space of the real and simulated environments to predict the states relevant to the pain experience, enabling more rapid optimization of the therapy. Optimal behavior is learnt indirectly by the model by simulating actions and observing their outcomes such that the agent can optimize its policy and its value function using not only input from the real environment constituted by the SCS device and the patient but also from the simulated environment provided by the model. The value function is weighted by cumulative rewards that are updated by both real and simulated experience.

[0050] Fitting of the value function and policy optimization by the agent may utilize not only an RL algorithm but optionally also other computational components. For example, an evolutionary algorithm may be used such as a genetic algorithm, to incorporate a fitness function which allows optimization of the model based on experiences of other patients or more precisely their associated models. Deep learning with a neural network is another computational component that may be used in combination with the RL algorithm. Combining the RL algorithm with an evolutionary algorithm and / or a deep learning neural network may help account for otherwise difficult to optimize function space, e.g., for discontinuities in the function space.

[0051] The agent decisions are associated with patient care interventions. The agent includes (i) SCS programmer software and hardware, (ii) the persons (e.g., remote care representative, clinician), and (iii) the algorithm / software / hardware system using the SCS programmer.

[0052] In some embodiments, the actions are a set of SCS control parameters and remote care interventions. The SCS control parameter set may include any (permutation) of the following SCS control parameters:

[0053] 1. Stimulation amplitude

[0054] 2. Stimulation frequency

[0055] 3. Stimulus pulse width

[0056] 4. Duty cycling ratio (time when stimulation is on vs. off)

[0057] 5. Electrode configuration.

[0058] 6. Burst envelop frequency (burst interval)

[0059] 7. Burst duty cycle ratio (burst on vs. off)

[0060] Additional embodiments may include guiding actions by the physician related to treatment decisions and dosages, e.g. medication, steroid injection, facet joint injection / ablati on, etc.

[0061] Interventions may include suggestions or recommendations made to a patient to modify any one or more of the following (related to pain treatment):

[0062] 1. Patient behavior, e.g., suggestions to decrease or increase activity

[0063] 2. Patient physiology, e.g., a reminder to take medication

[0064] 3. Patient psychology, e.g., suggestions to talk with a caregiver or perform activities

[0065] 4. SCS Programmer contact to patient to provide emotional support

[0066] 5. SCS implantable pulse generator (IPG) HW status, e.g., recommendation to readjust lead location

[0067] 6. Adjust medication prescription or intake

[0068] 7. Suggestions to deliver focal injection treatments e.g. steroid 8. Suggestions of prescription and / or change of dose / type of anti-anxiety medications

[0069] 9. Suggestions of interventions such as facet joint injections, spinal mechanical surgery, etc.

[0070] Additional embodiments may comprise expanded output guidance to include options outside of SCS, considering that input data may come from an implantable bio monitor (loop recorder) or an wearable system / device.

[0071] The environment encompasses patient and SCS IPG / HW, and output states describe the status of the patient and SCS IPG / HW, and the model mimics, i.e., simulates the environment.

[0072] The patient environment may include, for example one or more of the following:

[0073] 1. Behavior, e.g., activity level, step count, patient locations inferred by GPS, sleep, pain level, etc.

[0074] 2. Physiology, e.g., posture, gait characteristics, body temperature, illness inferred by physiological measurements such as heart rate variability, breathing rate, etc.

[0075] 3. Psychology, e.g., stress level, depression, anxiety, kinesiophobia, etc.

[0076] SCS IPG / HW may include, for example one or more of the following:

[0077] 1. IPG body, e.g., sensors

[0078] 2. Lead, e.g., lead location, lead impedance.

[0079] The data collection method for fitting the patient environment model may include a patient app.

[0080] The model may be dynamic and update itself to better follow the environment based on inputs received from the environment. Divergence between the model state estimates and the real states as provided by the environment may be analyzed by the agent and based on the analysis it may be decided to update the model by changing one or more of the model parameters.

[0081] Divergence between the model state estimates and the real states as provided by the environment may be analyzed by the agent and based on the analysis it may be decided to trigger an alert to initiate an intervention such as a patient care intervention or SCS device intervention.

[0082] The policy function and / or the value function at the agent can be updated responsive to both real experience states provided by the environment and simulated experience states provided by the model.

[0083] The optimization may incorporate an explore mechanism incorporating a random element for searching the parameter space and a mechanism involving a local search to aid finding a global optimum and avoid becoming trapped in local optima.

[0084] A satisfaction index can be used across a patient population to update the reward and set initial values for new patients.

[0085] Patient goals are collected as part of an initialization process for a new patient, which can optionally be updated or repeated occasionally. Patient goals are applied to adjust rewards and value function weights so that behavioral and feedback optima are found which are preferred for individual patients. For example, some patients may have patient goals to be more active while tolerating higher pain levels while other patients may have minimization of pain levels as their sole patient goal.

[0086] The model may include a patient caregiver and interactions with the caregiver. For example, outputs for improving mental health and pain management may include recommendations on specific interactions with the caregiver or other persons in the patient environment. Moreover, inputs for updating the model states may be provided by the caregiver. Patients having similar models and / or similar rewards may be grouped for patient phenotyping to differentially evaluate the likelihood of success. In case of a low likelihood of success being predicted for a given patient phenotype, alternative rewards targeting different behavioral / therapy outcomes may be recommended to increase the likelihood of success.

[0087] Figure 5 is a flow diagram of an embodiment of the invention as applicable to SCS treatment managed through remote care. The process flow path involving the simulated environment, i.e., the model, is shown with dotted lines and the process flow path involving the real environment is shown with solid lines. The feedback provided by the model (dashed lines) is referred to as simulated optimization whereas the feedback provided by the environment (solid lines) is referred to as direct optimization.

[0088] In Step SI, priorities are set to define the benefits. This can be done by the patient, a health care professional or jointly. A patient may have different priorities among states (e.g., activity level, pain, taking medication, etc.) which may introduce conflicting interests.

[0089] In Step S2, preferred states and actions are set to initial values and thereafter iteratively optimized according to feedback of the environmental states. Initial parameter settings for any given patient can be made according to a generalized population model of a cohort of patients and / or from an individualized model of a particular patient, e.g., the same patient or another patient with high similarity who has been using an SCS device for some time and its parameter settings are thus optimized. The initial values can also take account of input provided by the patient and / or by a health care professional to take account of patient wishes and / or clinical preferences. Patient and clinicians may have different priorities for the action they prefer to take, which is reflected in the rewards, e.g., a preference to change SCS control parameters rather than taking opioid medication. The reward is defined in terms of states and actions prioritized by the patient / physician rt= f st, at~) where reward rtis a function of state stand action atat time t. The reward is individualized for each patient to account for different priorities patients may have for their preferred lifestyle. In Steps S3 and S4, the model updates and runs according to the parameters set in Steps SI and S2.

[0090] In Step S5, the agent evaluates and updates the benefits according both to model input from Step S4 and the environment input from Steps SI and S2.

[0091] In Steps S6 and S7, the set of therapy parameter values associated with the maximum benefit are determined and then these are output to provide feedback to both the real environment and the simulated environment of the model. For example, the therapy change actioned by Step S7 may change a parameter setting in the SCS device implanted in the patient (feedback to Step S2) and the same change of the same parameter in the simulated SCS device that forms part of the model (feedback to Step S4).

[0092] Figure 6 shows an example model of patient behavior which includes dynamic interactions between activity (A), pain (P), and sleep (S). Modulated values of activity and pain are shown with a prime, i.e., A’ and B’. The model may be augmented to include medication metabolization, physiological states / changes (e.g., due to illness), psychological states / changes, and / or device states that impact other states. In the illustrated example, the model updates take place in two levels, one optimizing the gains (G in the figure above), and one searching for dynamic parameters (H in the figure above). The model may use different filter types for fitting, e.g., a particle filter, Bayesian filter or Kalman filter, which may be an extended Kalman filter. The filter compares the estimated states (determined by the model) and the real states (observed in the environment) to calculate the error and then adjust the gains to reduce impact of noise and enhance the precision to which states are estimated. Non-stationary model dynamic parameters may be followed using augmented filter algorithms, for example batch machine learning with forward-backward smoothing. Dynamic model parameter optimization may be done through continuous operations or in batch operations.

[0093] Figure 7 is a block diagram illustrating an example computing apparatus 500 that may be used in connection with various embodiments described herein. For example, computing apparatus 500 may be used as the remote computing resource at the neuro-service data center 60 of Figure 1.

[0094] Computing apparatus 500 can be a server or any conventional personal computer, or any other processor-enabled device that is capable of wired or wireless data communication. Other computing apparatus, systems and / or architectures may be also used, including devices that are not capable of wired or wireless data communication, as will be clear to those skilled in the art.

[0095] Computing apparatus 500 preferably includes one or more processors, such as processor 510. The processor 510 may be for example a CPU, GPU, TPU or arrays or combinations thereof such as CPU and TPU combinations or CPU and GPU combinations. Additional processors may be provided, such as an auxiliary processor to manage input / output, an auxiliary processor to perform floating point mathematical operations (e.g. a TPU), a specialpurpose microprocessor having an architecture suitable for fast execution of signal processing algorithms (e.g., digital signal processor, image processor), a slave processor subordinate to the main processing system (e.g., back-end processor), an additional microprocessor or controller for dual or multiple processor systems, or a coprocessor. Such auxiliary processors may be discrete processors or may be integrated with the processor 510. Examples of CPUs which may be used with computing apparatus 500 are, the Pentium processor, Core i7 processor, and Xeon processor, all of which are available from Intel Corporation of Santa Clara, California. An example GPU which may be used with computing apparatus 500 is Tesla K80 GPU of Nvidia Corporation, Santa Clara, California.

[0096] Processor 510 is connected to a communication bus 505. Communication bus 505 may include a data channel for facilitating information transfer between storage and other peripheral components of computing apparatus 500. Communication bus 505 further may provide a set of signals used for communication with processor 510, including a data bus, address bus, and control bus (not shown). Communication bus 505 may comprise any standard or non-standard bus architecture such as, for example, bus architectures compliant with industry standard architecture (ISA), extended industry standard architecture (EISA), Micro Channel Architecture (MCA), peripheral component interconnect (PCI) local bus, or standards promulgated by the Institute of Electrical and Electronics Engineers (IEEE) including IEEE 488 general-purpose interface bus (GPIB), IEEE 696 / S-100, and the like.

[0097] Computing apparatus 500 preferably includes a main memory 515 and may also include a secondary memory 520. Main memory 515 provides storage of instructions and data for programs executing on processor 510, such as one or more of the functions and / or modules discussed above. It should be understood that computer readable program instructions stored in the memory and executed by processor 510 may be assembler instructions, instructionset-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuitry, or either source code or object code written in and / or compiled from any combination of one or more programming languages, including without limitation Smalltalk, C / C++, Java, JavaScript, Perl, Visual Basic, .NET, and the like. Main memory 515 is typically semiconductor-based memory such as dynamic random access memory (DRAM) and / or static random access memory (SRAM). Other semiconductor-based memory types include, for example, synchronous dynamic random access memory (SDRAM), Rambus dynamic random access memory (RDRAM), ferroelectric random access memory (FRAM), and the like, including read only memory (ROM).

[0098] The computer readable program instructions may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0099] Secondary memory 520 may optionally include an internal memory 525 and / or a removable medium 530. Removable medium 530 is read from and / or written to in any well-known manner. Removable storage medium 530 may be, for example, a magnetic tape drive, a compact disc (CD) drive, a digital versatile disc (DVD) drive, other optical drive, a flash memory drive, etc. Removable storage medium 530 is a non-transitory computer-readable medium having stored thereon computer-executable code (i.e., software) and / or data. The computer software or data stored on removable storage medium 530 is read into computing apparatus 500 for execution by processor 510.

[0100] The secondary memory 520 may include other similar elements for allowing computer programs or other data or instructions to be loaded into computing apparatus 500. Such means may include, for example, an external storage medium 545 and a communication interface 540, which allows software and data to be transferred from external storage medium 545 to computing apparatus 500. Examples of external storage medium 545 may include an external hard disk drive, an external optical drive, an external magneto- optical drive, etc. Other examples of secondary memory 520 may include semiconductor- based memory such as programmable read-only memory (PROM), erasable programmable readonly memory (EPROM), electrically erasable read-only memory (EEPROM), or flash memory (block-oriented memory similar to EEPROM).

[0101] As mentioned above, computing apparatus 500 may include a communication interface 540. Communication interface 540 allows software and data to be transferred between computing apparatus 500 and external devices (e.g., printers), networks, or other information sources. For example, computer software or executable code may be transferred to computing apparatus 500 from a network server via communication interface 540. Examples of communication interface 540 include a built-in network adapter, network interface card (NIC), Personal Computer Memory Card International Association (PCMCIA) network card, card bus network adapter, wireless network adapter, Universal Serial Bus (USB) network adapter, modem, a network interface card (NIC), a wireless data card, a communications port, an infrared interface, an IEEE 1394 fire-wire, or any other device capable of interfacing system with a network or another computing device. Communication interface 540 preferably implements industry-promulgated protocol standards, such as Ethernet IEEE 802 standards, Fiber Channel, digital subscriber line (DSL), asynchronous digital subscriber line (ADSL), frame relay, asynchronous transfer mode (ATM), integrated digital services network (ISDN), personal communications services (PCS), transmission control protocol / Intemet protocol (TCP / IP), serial line Internet protocol / point to point protocol (SLIP / PPP), and so on, but may also implement customized or non-standard interface protocols as well.

[0102] Software and data transferred via communication interface 540 are generally in the form of electrical communication signals 555. These signals 555 may be provided to communication interface 540 via a communication channel 550. In an embodiment, communication channel 550 may be a wired or wireless network, or any variety of other communication links. Communication channel 550 carries signals 555 and can be implemented using a variety of wired or wireless communication means including wire or cable, fiber optics, conventional phone line, cellular phone link, wireless data communication link, radio frequency ("RF") link, or infrared link, just to name a few.

[0103] Computer-executable code (i.e., computer programs or software) is stored in main memory 515 and / or the secondary memory 520. Computer programs can also be received via communication interface 540 and stored in main memory 515 and / or secondary memory 520. Such computer programs, when executed, enable computing apparatus 500 to perform the various functions of the disclosed embodiments as described elsewhere herein.

[0104] In this document, the term "computer-readable medium" is used to refer to any non-transitory computer-readable storage media used to provide computer-executable code (e.g., software and computer programs) to computing apparatus 500. Examples of such media include main memory 515, secondary memory 520 (including internal memory 525, removable medium 530, and external storage medium 545), and any peripheral device communicatively coupled with communication interface 540 (including a network information server or other network device). These non-transitory computer-readable media are means for providing executable code, programming instructions, and software to computing apparatus 500. In an embodiment that is implemented using software, the software may be stored on a computer- readable medium and loaded into computing apparatus 500 by way of removable medium 530, VO interface 535, or communication interface 540. In such an embodiment, the software is loaded into computing apparatus 500 in the form of electrical communication signals 555. The software, when executed by processor 510, preferably causes processor 510 to perform the features and functions described elsewhere herein.

[0105] I / O interface 535 provides an interface between one or more components of computing apparatus 500 and one or more input and / or output devices. Example input devices include, without limitation, keyboards, touch screens or other touch-sensitive devices, biometric sensing devices, computer mice, trackballs, pen-based pointing devices, and the like. Examples of output devices include, without limitation, cathode ray tubes (CRTs), plasma displays, light-emitting diode (LED) displays, liquid crystal displays (LCDs), printers, vacuum florescent displays (VFDs), surface-conduction electron-emitter displays (SEDs), field emission displays (FEDs), and the like.

[0106] Computing apparatus 500 also includes optional wireless communication components that facilitate wireless communication over a voice network and / or a data network. The wireless communication components comprise an antenna system 570, a radio system 565, and a baseband system 560. In computing apparatus 500, radio frequency (RF) signals are transmitted and received over the air by antenna system 570 under the management of radio system 565.

[0107] Antenna system 570 may comprise one or more antennae and one or more multiplexors (not shown) that perform a switching function to provide antenna system 570 with transmit and receive signal paths. In the receive path, received RF signals can be coupled from a multiplexor to a low noise amplifier (not shown) that amplifies the received RF signal and sends the amplified signal to radio system 565.

[0108] Radio system 565 may comprise one or more radios that are configured to communicate over various frequencies. In an embodiment, radio system 565 may combine a demodulator (not shown) and modulator (not shown) in one integrated circuit (IC). The demodulator and modulator can also be separate components. In the incoming path, the demodulator strips away the RF carrier signal leaving a baseband receive audio signal, which is sent from radio system 565 to baseband system 560. If the received signal contains audio information, then baseband system 560 decodes the signal and converts it to an analogue signal. Then the signal is amplified and sent to a speaker. Baseband system 560 also receives analogue audio signals from a microphone. These analogue audio signals are converted to digital signals and encoded by baseband system 560. Baseband system 560 also codes the digital signals for transmission and generates a baseband transmit audio signal that is routed to the modulator portion of radio system 565. The modulator mixes the baseband transmit audio signal with an RF carrier signal generating an RF transmit signal that is routed to antenna system 570 and may pass through a power amplifier (not shown). The power amplifier amplifies the RF transmit signal and routes it to antenna system 570 where the signal is switched to the antenna port for transmission.

[0109] Baseband system 560 is also communicatively coupled with processor 510, which may be a central processing unit (CPU). Processor 510 has access to data storage areas 515 and 520. Processor 510 is preferably configured to execute instructions (i.e., computer programs or software) that can be stored in main memory 515 or secondary memory 520. Computer programs can also be received from baseband processor 560 and stored in main memory 510 or in secondary memory 520 or executed upon receipt. Such computer programs, when executed, enable computing apparatus 500 to perform the various functions of the disclosed embodiments. For example, data storage areas 515 or 520 may include various software modules.

[0110] The computing apparatus further comprises a display 575 directly attached to the communication bus 505 which may be provided instead of or addition to any display connected to the VO interface 535 referred to above.

[0111] Various embodiments may also be implemented primarily in hardware using, for example, components such as application specific integrated circuits (ASICs), programmable logic arrays (PLA), or field programmable gate arrays (FPGAs). Implementation of a hardware state machine capable of performing the functions described herein will also be apparent to those skilled in the relevant art. Various embodiments may also be implemented using a combination of both hardware and software. Furthermore, those of skill in the art will appreciate that the various illustrative logical blocks, modules, circuits, and method steps described in connection with the abovedescribed figures and the embodiments disclosed herein can often be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled persons can implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the invention. In addition, the grouping of functions within a module, block, circuit, or step is for ease of description. Specific functions or steps can be moved from one module, block, or circuit to another without departing from the invention.

[0112] Moreover, the various illustrative logical blocks, modules, functions, and methods described in connection with the embodiments disclosed herein can be implemented or performed with a general-purpose processor, a digital signal processor (DSP), an ASIC, FPGA, or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general- purpose processor can be a microprocessor, but in the alternative, the processor can be any processor, controller, microcontroller, or state machine. A processor can also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.

[0113] Additionally, the steps of a method or algorithm described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module can reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium including a network storage medium. An exemplary storage medium can be coupled to the processor such that the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium can be integral to the processor. The processor and the storage medium can also reside in an ASIC.

[0114] A computer readable storage medium, as referred to herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.

[0115] Any of the software components described herein may take a variety of forms. For example, a component may be a stand-alone software package, or it may be a software package incorporated as a "tool" in a larger software product. It may be downloadable from a network, for example, a website, as a stand-alone product or as an add-in package for installation in an existing software application. It may also be available as a client- server software application, as a web-enabled software application, and / or as a mobile application.

[0116] Embodiments of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer readable program instructions.

[0117] The computer readable program instructions may be provided to a processor of a general- purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions may also be stored in a computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement aspects of the function / act specified in the flowchart and / or block diagram block or blocks.

[0118] The computer readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0119] The illustrated flowcharts and block diagrams illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the Figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustration, and combinations of blocks in the block diagrams and / or flowchart illustration, can be implemented by special purpose hardware-based systems that perform the specified functions or acts or carry out combinations of special purpose hardware and computer instructions.

[0120] Apparatus and methods embodying the invention are capable of being hosted in and delivered by a cloud-computing environment. Cloud computing is a model of service delivery for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal management effort or interaction with a provider of the service. REFERENCE NUMERALS

[0121] I Medical Device Communication System (MDCS)

[0122] 10 Implantable Medical Device (IMD)

[0123] I I IMD stimulator, e.g., SCS device with attached electrodes

[0124] 12 IMD wireless personal area network (WPAN) transceiver

[0125] 13 IMD sensor

[0126] 14 IMD electrical energy source

[0127] 15 IMD low-power wide-area network (LPWAN) transceiver

[0128] 16 IMD control circuit

[0129] 17 IMD memory

[0130] 18 IMD processor

[0131] 19 IMD electrode drive circuit

[0132] 20 charger for IMD

[0133] 24 charger energy source

[0134] 29 charger external mains power connection

[0135] 30 patient remote controller (patient remote)

[0136] 32 smartphone wireless personal area network (WPAN) transceiver (e.g., BLE)

[0137] 34 smartphone low-power wide-area network (LPWAN) transceiver

[0138] 35 smartphone cellular transceiver

[0139] 36 smartphone wireless local area network (WLAN) transceiver

[0140] 38 router

[0141] 40 telephone line / telephone network

[0142] 42 cellular network base station (cellular tower)

[0143] 50 internet connection

[0144] 60 remotely located data center (backend)

[0145] 70 portal

[0146] 80 lead arrangement

[0147] 82 leads

[0148] 84 electrodes 0 reinforced learning (RL) algorithm

[0149] 100 real environment

[0150] 100’ model (simulated environment)

[0151] 104 patient behavior

[0152] 104’ patient behavior

[0153] 106 patient physiology

[0154] 106’ patient physiology

[0155] 108 patient psychology

[0156] 108’ patient psychology

[0157] 110 medical device status (e.g., SCS IPG status)

[0158] 110’ medical device status (e.g., SCS IPG status)

[0159] 112 real state

[0160] 112’ simul ated state

[0161] 114 real reward

[0162] 114’ simulated reward

[0163] 120 agent

[0164] 122 policy

[0165] 124 value function

[0166] 126 action

[0167] 500 computing apparatus

[0168] 505 communication bus

[0169] 510 processor

[0170] 515 main memory

[0171] 520 secondary memory

[0172] 525 internal memory

[0173] 530 removable medium

[0174] 535 I / O interface

[0175] 540 communication interface

[0176] 545 external storage medium

[0177] 550 communication channel

[0178] 555 electrical communication signals baseband system radio system antenna system display

Claims

Claims1. A computer-implemented reinforced learning method (90) for optimizing a spinal cord stimulation (SCS) therapy delivered to a patient by a medical device (10), the method comprising: receiving states (112) and rewards (114) from a real environment (100); providing an agent (120) comprising a policy (122) and a value function (124), the agent being configured to apply the policy to map the states received as input into actions (126) and output the actions to cause modification of the real environment; providing a model (100’) configured to simulate the real environment as a simulated environment and connected to receive as input the actions from the agent, the model being further configured to generate simulated states (112’) and simulated rewards (114’) and to output these to the agent, so that the agent receives real and simulated states and rewards; and iteratively optimizing the agent actions and / or interventions according to the value function to maximize rewards from the real and simulated environments through the agent’s receipt of the real and simulated states and awards, wherein the agent includes at least one of a SCS programmer software, a SCS programmer hardware, a remote care representative, a clinician, and an algorithm, the software or hardware system using a SCS programmer.

2. The method of claim 1, further comprising the model receiving as input the states and rewards from the real environment; and iteratively updating the simulated environment of the model responsive to the real states and rewards.

3. The method of claim 1 or 2, wherein at least one other one of the rewards is associated with a care intervention for the patient.

4. The method of any one of claims 1 to 3, wherein certain ones of the rewards are associated with respective values of control parameters of a medical device (10) that is delivering therapy to the patient.

5. A computer readable medium (515, 520, 525, 530, 545) storing instructions for performing the reinforced learning method (90) of any one of claims 1 to 4.

6. A medical device communication system (1), comprising: a medical device (10) for delivering therapy to a patient, the medical device delivering the therapy according to a set of control parameters; and a remote computing resource (60) arranged in data communication connection with the medical device, the remote computing resource comprising a processor (510) and a computer readable medium (515, 520, 525, 530, 545) storing instructions for performing the reinforced learning method (90) of claim 4 on the processor, wherein the method uses sensor data received from the medical device as states (112) and rewards (114), and wherein the remote computing resource is further configured to transmit the control parameter values determined by the agent to the medical device.

7. The medical device communication system (1) of claim 6, wherein the remote computing resource is further configured to output a care intervention message according to at least one other one of the rewards.

8. A medical device (10) for treating a patient, comprising: an electrical energy source (14); therapeutic components (11) for patient treatment that are powered by the electrical energy source (14), the therapeutic components being driven according to a set of control parameters; a receiver (12) arranged to receive control data with values of the control parameters via a data communication connection; and a memory (17) and a processor (18), wherein the memory stores a computer program which when executed on the processor varies the control parameters based on control parameter values received by the receiver, said control parameter values having beendetermined by a remote computing resource (60) according to the reinforced learning method of claim 4.

9. The medical device (10) of claim 8, further comprising a transmitter arranged to transmit sensor data via the data communication connection to the remote computing resource, the sensor data providing states and rewards for the computer-implemented reinforced learning method performed by the remote computing resource.

10. The medical device (10) of claim 9, wherein the states and rewards are input to the agent.

11. The medical device (10) of claim 9 or 10, wherein the states and rewards are input to the model.

12. The medical device (10) of any one of claims 8 to 11, further comprising a control circuit (16) arranged to receive the control parameters and an electrode drive circuit arranged to deliver electrical signals to at least one electrode according to the control parameters.

13. A method of delivering therapy to a patient via a medical device (10), the method comprising: providing the patient with a medical device (10) comprising: an electrical energy source (14); therapeutic components for patient treatment that are powered by the electrical energy source, the therapeutic components being driven according to a set of control parameters; a receiver (12) arranged to receive control data with values of the control parameters via a data communication connection from a remote computing resource (60); and a memory (17) and a processor (18),storing a computer program in the memory of the medical device which when executed on the processor of the medical device varies the control parameters of the medical device based on control parameter values received by the receiver; determining the control parameter values at the remote computing resource according to a reinforced learning method (90), the reinforced learning method comprising: receiving states (112) and rewards (114) from a real environment (100); providing an agent (120) comprising a policy (122) and a value function (124), the agent being configured to apply the policy to map the states received as input into actions (126) and output the actions to cause modification of the real environment; providing a model (100’) configured to simulate the real environment as a simulated environment and connected to receive as input the actions from the agent, the model being further configured to generate simulated states (112’) and simulated rewards (114’) and to output these to the agent, so that the agent receives real and simulated states and rewards; and iteratively optimizing the agent actions and / or interventions according to the value function to maximize rewards from the real and simulated environments through the agent’s receipt of the real and simulated states and awards, wherein certain ones of the rewards are associated with respective values of the control parameters.

14. The method of claim 13, wherein the medical device (10) further comprises a transmitter (12) arranged to transmit sensor data via the data communication connection to the remote computing resource, and wherein the sensor data provides states and rewards for the reinforced learning method performed by the remote computing resource.

15. The method of claim 13 or 14, wherein the model receives as input the states and rewards from the real environment, and wherein the reinforced learning method iteratively updates the simulated environment of the model responsive to the real states and rewards.

Citation Information

Patent Citations

  • Adaptive electrical neurostimulation treatment to reduce pain perception

    US20210252288A1

  • Systems and methods for providing neurostimulation therapy according to machine learning operations

    US20220323766A1

  • Systems and methods based on deep reinforcement learning and planning for shaping neural activity, rewiring neural circuits, augmenting neural function and / or restoring neural function

    US20230137595A1