accurate tissue proximity
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-16
- Publication Date
- 2026-08-11
Smart Images

Figure CN114631803B_ABST
Abstract
Description
[0001] Relevant application information
[0002] This application claims the benefit of U.S. Provisional Patent Application Serial No. 63 / 126,152, filed on December 16, 2020, the disclosure of which is incorporated herein by reference. Technical Field
[0003] This invention relates to medical systems, and specifically, but not exclusively, to catheter devices. Background Technology
[0004] Numerous medical procedures involve placing probes, such as catheters, inside a patient's body. Position sensing systems have been developed to track these probes. Magnetic position sensing is one method known in the art. In magnetic position sensing, a magnetic field generator is typically placed at a known location outside the patient's body. A magnetic field sensor within the distal end of the probe generates electrical signals in response to these magnetic fields, and these signals are processed to determine the coordinate position of the distal end of the probe. These methods and systems are described in U.S. Patents 5,391,199, 6,690,963, 6,484,118, 6,239,724, 6,618,612, and 6,332,089; PCT International Patent Publication WO 1996 / 005768; and U.S. Patent Application Publications 2002 / 0065455, 2003 / 0120150, and 2004 / 0068178, the entire disclosure of which is incorporated herein by reference. Position can also be tracked using systems based on impedance or current.
[0005] Treatment of arrhythmias is a medical procedure in which these types of probes or catheters have proven extremely useful. Arrhythmias, and specifically atrial fibrillation, have always been a common and dangerous medical condition, especially in the elderly.
[0006] The diagnosis and treatment of cardiac arrhythmias involve mapping the electrical properties of cardiac tissue, particularly the endocardium and cardiac volume, and selectively ablating cardiac tissue by applying energy. Such ablation can stop or alter unwanted electrical signals propagating from one part of the heart to another. Ablation methods disrupt unwanted electrical pathways by creating a non-conductive ablation focus. Various forms of energy delivery for creating ablation focuses have been disclosed, including the use of microwaves, lasers, and more commonly, radiofrequency energy to create conduction blocks along the cardiac tissue walls. In a two-step procedure (mapping followed by ablation), electrical activity at various points within the heart is typically sensed and measured by advancing a catheter containing one or more electrical sensors into the heart and acquiring data at multiple points. This data is then used to select the target endocardial region to be ablated.
[0007] Electrode catheters have been widely used in medical practice for many years. They are used to stimulate and map electrical activity in the heart, as well as to ablate sites of abnormal electrical activity. In use, the electrode catheter is inserted into a major vein or artery, such as the femoral artery, and then guided to the ventricle of interest. A typical ablation procedure involves inserting a catheter with one or more electrodes at its distal end into the ventricle. A reference electrode can be provided, typically taped to the patient's skin, or a second catheter positioned in or near the heart can be used to provide the reference electrode. RF (radio frequency) current is applied to the tip electrode of the ablation catheter, and the current flows through the surrounding medium (i.e., blood and tissue) to the reference electrode. The current distribution depends on the amount of contact between the electrode surface and the tissue compared to blood, which has a higher conductivity. Due to the resistance of the tissue, heating of the tissue occurs. The tissue is heated sufficiently to destroy the cells in the heart tissue, resulting in the formation of a non-conductive ablation focus within the heart tissue.
[0008] Therefore, when placing an ablation catheter or other catheter in the body (especially near the endocardial tissue), it is desirable to have a distal catheter tip that directly contacts the tissue. Contact can be confirmed, for example, by measuring the contact between the distal tip and the body tissue. U.S. Patent Application Publications 2007 / 0100332, 2009 / 0093806, and 2009 / 0138007 (the disclosures of which are incorporated herein by reference) describe methods for sensing the contact pressure between the distal catheter tip and tissue in a body cavity using a force sensor embedded within the catheter.
[0009] Numerous references have reported methods for determining electrode-tissue contact, including U.S. Patents 5,935,079, 5,891,095, 5,836,990, 5,836,874, 5,673,704, 5,662,108, 5,469,857, 5,447,529, 5,341,807, 5,078,714, and Canadian Patent Application 2,285,342. Many of these references, such as U.S. Patents 5,935,079, 5,836,990, and 5,447,529, determine electrode-tissue contact by measuring the impedance between the tip electrode and the return electrode. As disclosed in '529, it is well known that the impedance through blood is generally lower than that through tissue. Therefore, tissue contact is detected by comparing the impedance values between a set of electrodes with pre-measured impedance values when the electrode is in contact with tissue and when the electrode is in contact with blood only.
[0010] U.S. Patent 9,168,004, granted to Gliner et al., describes the use of machine learning to determine catheter electrode contact, which is incorporated herein by reference. Patent '004 describes a cardiac catheterization procedure performed by: memorizing the assignment of the contact state between the probe's electrodes and the heart wall as either a contact state or a non-contact state; determining a series of impedance phase angles of the current passing through one electrode to another, thereby identifying the maximum and minimum phase angles in the series; and adaptively defining a binary classifier at the midpoint of the limiting values. Test values are compared to a classifier adjusted by a hysteresis factor, and changes in contact state are reported when the test value exceeds or falls below the adjusted classifier.
[0011] U.S. Patent Publication 2013 / 0085416 to Mest, which is incorporated herein by reference, describes a method for in vivo recalibration of force-sensing probes, such as electrophysiological catheters, to generate an automatically zeroed region. The distal end of the catheter or other probe is placed within the patient's body cavity. Electrocardiogram (ECG) or impedance data, fluoroscopy or other real-time imaging data, and / or an electroanatomical mapping system are used to confirm the absence of tissue contact. Once the absence of tissue contact is confirmed, the system recalibrates the signal derived from the force sensor to correspond to a zero-gram force reading, and uses this calibrated baseline reading to generate and display a force reading based on the force sensor data. Summary of the Invention
[0012] According to embodiments of this disclosure, a method for obtaining a tissue proximity indication is provided, the method comprising: inserting a catheter into a body part of a living subject such that electrodes of the catheter contact tissue at corresponding locations within the body part; receiving signals provided by the electrodes; selectively rewarding and penalizing a reinforcement learning agent during a reinforcement learning exploration phase in response to at least one of the received signals to learn at least one tissue proximity strategy; applying the reinforcement learning agent during a reinforcement learning utilization phase to acquire a corresponding tissue proximity action to be taken in response to the at least one tissue proximity strategy, the tissue proximity action maximizing a corresponding expected reward; and providing a corresponding derived tissue proximity indication of the proximity of a given electrode in the electrodes to the tissue in response to the acquired corresponding tissue proximity action.
[0013] Furthermore, according to the embodiments of this disclosure, the application includes applying the reinforcement learning agent in a corresponding reinforcement learning exploitation phase within the reinforcement learning exploitation phase, in response to the corresponding state of the reinforcement learning agent and the corresponding available organization proximity action set to be taken, to acquire the corresponding organization proximity action among the organization proximity actions to be taken, the corresponding organization proximity action maximizing the corresponding expected reward; and the selective reward and punishment includes selectively rewarding and punishing the reinforcement learning agent within the reinforcement learning exploration phase in response to data from the corresponding final reinforcement learning exploitation phase within the reinforcement learning exploitation phase, the data including: the corresponding state in the state, the corresponding acquired organization proximity action among the organization proximity actions to be taken, and the corresponding actual reward.
[0014] Furthermore, according to embodiments of this disclosure, each organization proximity action in the available organization proximity action group to be taken includes: changing a corresponding organization proximity indicator in the derived organization proximity indicator provided in the corresponding last reinforcement learning utilization phase of the reinforcement learning utilization phase; and not changing the corresponding organization proximity indicator in the derived organization proximity indicator provided in the corresponding last reinforcement learning utilization phase of the reinforcement learning utilization phase.
[0015] Additionally, according to embodiments of this disclosure, each of the corresponding states includes: a corresponding tissue proximity indication provided in the corresponding last reinforcement learning utilization phase within the reinforcement learning utilization phase; and a corresponding impedance value of the given electrode calculated for the current corresponding reinforcement learning utilization phase within the reinforcement learning utilization phase.
[0016] Furthermore, according to embodiments of this disclosure, the method includes calculating the corresponding actual reward in response to the corresponding acquired organization proximity action in the organization proximity action to be taken and a corresponding reference organization proximity indication calculated independently of the reinforcement learning agent.
[0017] Additionally, according to embodiments of this disclosure, the method includes: calculating a corresponding impedance value of the given electrode for each of the corresponding states in response to at least one of the received signals; and calculating the corresponding reference tissue proximity indication independently of applying the reinforcement learning agent in response to at least one of the received signals.
[0018] Furthermore, according to embodiments of this disclosure, calculating the corresponding reference tissue proximity indication includes calculating the corresponding reference tissue proximity indication in response to the corresponding impedance value of the given electrode.
[0019] In addition, according to the implementation scheme of this disclosure, the calculation of the corresponding actual reward includes the calculation of the sum of future discount rewards.
[0020] Furthermore, according to the implementation scheme of this disclosure, the reinforcement learning agent is a deep reinforcement learning agent.
[0021] Furthermore, according to the implementation scheme of this disclosure, the reinforcement learning agent is a Q-learning model-free reinforcement learning agent.
[0022] Furthermore, according to embodiments of this disclosure, the method includes electrically coupling a first set of electrodes of the catheter and the given electrode to a first signal processing unit, wherein the calculation of a corresponding reference tissue proximity indication is performed by the first signal processing unit in response to receiving at least one signal provided by the given electrode; and electrically coupling a second set of electrodes of the catheter and the given electrode to a second signal processing unit, the second set of electrodes being different from the first set of electrodes, wherein the calculation of a corresponding impedance value of the given electrode for each of the corresponding states is performed in the second signal processing unit in response to receiving the at least one signal provided by the given electrode.
[0023] Furthermore, according to an embodiment of this disclosure, the method includes: calculating a corresponding impedance value of a corresponding electrode in the second set of electrodes by the second signal processing unit; applying the reinforcement learning agent to acquire a corresponding tissue proximity action to be taken in response to the calculated corresponding impedance value of the corresponding electrode in the second set of electrodes, the corresponding tissue proximity action maximizing the corresponding expected reward of the corresponding electrode in the second set of electrodes; and providing a corresponding derived tissue proximity indication of the proximity between the corresponding electrode in the second set of electrodes and the tissue in response to the acquired corresponding tissue proximity action of the corresponding electrode in the second set of electrodes.
[0024] Furthermore, according to embodiments of this disclosure, the method includes: applying a surface electrode to the skin surface of a living subject; electrically coupling the surface electrode to the first signal processing unit; and calculating a corresponding impedance value between the given electrode and the surface electrode, wherein the calculation of the corresponding reference tissue proximity indication is performed by the first signal processing unit in response to the calculated corresponding impedance value.
[0025] According to another embodiment of this disclosure, a system for acquiring a tissue proximity indication is also provided, the system comprising: a catheter configured to be inserted into a body part of a living subject and including electrodes configured to contact tissue at a corresponding location within the body part; and processing circuitry configured to: receive signals provided by the electrodes; selectively reward and punish a reinforcement learning agent during a reinforcement learning exploration phase, in response to at least one of the received signals, to learn at least one tissue proximity policy; apply the reinforcement learning agent during a reinforcement learning utilization phase to acquire a corresponding tissue proximity action to be taken in response to the at least one tissue proximity policy, the corresponding tissue proximity action maximizing a corresponding expected reward; and provide a corresponding derived tissue proximity indication of the proximity of a given electrode in the electrodes to the tissue in response to the acquired corresponding tissue proximity action.
[0026] Furthermore, according to an embodiment of this disclosure, the processing circuit is configured to: apply the reinforcement learning agent in a corresponding reinforcement learning exploitation phase within the reinforcement learning exploitation phase, in response to the corresponding state of the reinforcement learning agent and the corresponding available organization proximity action set to be taken, to acquire a corresponding organization proximity action among the organization proximity actions to be taken, the corresponding organization proximity action maximizing the corresponding expected reward; and selectively reward and punish the reinforcement learning agent within the reinforcement learning exploration phase in response to data from the corresponding final reinforcement learning exploitation phase within the reinforcement learning exploitation phase, the data including: the corresponding state in the state, the corresponding acquired organization proximity action among the organization proximity actions to be taken, and the corresponding actual reward.
[0027] Furthermore, according to embodiments of this disclosure, each organization proximity action in the available organization proximity action group to be taken includes: changing a corresponding organization proximity indicator in the derived organization proximity indicator provided in the corresponding last reinforcement learning utilization phase of the reinforcement learning utilization phase; and not changing the corresponding organization proximity indicator in the derived organization proximity indicator provided in the corresponding last reinforcement learning utilization phase of the reinforcement learning utilization phase.
[0028] Additionally, according to embodiments of this disclosure, each of the corresponding states includes: a corresponding tissue proximity indication provided in the corresponding last reinforcement learning utilization phase within the reinforcement learning utilization phase; and a corresponding impedance value of the given electrode calculated for the current corresponding reinforcement learning utilization phase within the reinforcement learning utilization phase.
[0029] Furthermore, according to an embodiment of this disclosure, the processing circuit is configured to calculate the corresponding actual reward in response to the corresponding acquired organization proximity action to be taken and the corresponding reference organization proximity indication calculated independently of the reinforcement learning agent.
[0030] Furthermore, according to an embodiment of this disclosure, the processing circuit is configured to: calculate a corresponding impedance value of the given electrode for each of the corresponding states in response to at least one of the received signals; and calculate the corresponding reference tissue proximity indication independently of applying the reinforcement learning agent in response to at least one of the received signals.
[0031] Furthermore, according to embodiments of this disclosure, the processing circuit is configured to calculate the corresponding reference tissue proximity indication in response to the corresponding impedance value of the given electrode.
[0032] Furthermore, according to an embodiment of this disclosure, the processing circuit is configured to: calculate the sum of future discount rewards; and calculate the corresponding actual reward in response to the calculated sum of future discount rewards.
[0033] Furthermore, according to the implementation scheme of this disclosure, the reinforcement learning agent is a deep reinforcement learning agent.
[0034] Furthermore, according to the implementation scheme of this disclosure, the reinforcement learning agent is a Q-learning model-free reinforcement learning agent.
[0035] Furthermore, according to embodiments of this disclosure, the catheter includes a first set of electrodes and a second set of electrodes, the second set of electrodes being different from the first set of electrodes. The system further includes: a first signal processing unit configured to be electrically coupled to the first set of electrodes and to calculate the corresponding reference tissue proximity indication in response to receiving at least one signal provided by the given electrode; and a second signal processing unit configured to be electrically coupled to the second set of electrodes of the catheter and to calculate a corresponding impedance value of the given electrode for each of the corresponding states in response to receiving the at least one signal provided by the given electrode.
[0036] Furthermore, according to an embodiment of this disclosure, the second signal processing unit is configured to calculate the corresponding impedance value of the corresponding electrode in the second set of electrodes; the processing circuit is configured to: apply the reinforcement learning agent to obtain a corresponding tissue proximity action to be taken in response to the calculated corresponding impedance value of the corresponding electrode in the second set of electrodes, the corresponding tissue proximity action maximizing the corresponding expected reward of the corresponding electrode in the second set of electrodes; and provide a corresponding derived tissue proximity indication of the proximity between the corresponding electrode in the second set of electrodes and the tissue in response to the obtained corresponding tissue proximity action to be taken by the corresponding electrode in the second set of electrodes.
[0037] Furthermore, according to embodiments of this disclosure, the system includes a body surface electrode configured to be applied to the skin surface of a live subject, the first signal processing unit being configured to be electrically coupled to the body surface electrode, wherein the first signal processing unit is configured to: calculate a corresponding impedance value between the given electrode and the body surface electrode; and calculate the corresponding reference tissue proximity indication in response to the calculated corresponding impedance value.
[0038] According to another embodiment of this disclosure, a software product is also provided, comprising a non-transitory computer-readable medium storing program instructions that, when read by a central processing unit (CPU), cause the CPU to: receive signals provided by electrodes; selectively reward and punish a reinforcement learning agent during a reinforcement learning exploration phase, in response to at least one of the received signals, to learn at least one organization proximity policy; apply the reinforcement learning agent during a reinforcement learning exploitation phase to acquire a corresponding organization proximity action to be taken in response to the at least one organization proximity policy, the corresponding organization proximity action maximizing a corresponding expected reward; and, in response to the acquired corresponding organization proximity action, provide a corresponding derived organization proximity indication of the proximity of a given electrode to the organization. Attached Figure Description
[0039] The invention will be understood from the following detailed description taken in conjunction with the accompanying drawings, wherein:
[0040] Figure 1 A schematic diagram of a medical protocol system constructed and operated according to an exemplary embodiment of the present invention;
[0041] Figure 2 For use Figure 1 A schematic diagram of the conduits in the system;
[0042] Figure 3 To perform reinforcement learning Figure 1 A schematic diagram of the system's components;
[0043] Figure 4 To include execution for Figure 1 A flowchart of the steps in reinforcement learning methods in a system;
[0044] Figure 5 To show the use of Figure 1 A schematic diagram of the reinforcement learning utilization phase in the system;
[0045] Figure 6 To include Figure 5 A flowchart of the steps performed during the utilization phase;
[0046] Figure 7 To show the use of Figure 1 A schematic diagram of the exploration phase of reinforcement learning in the system; and
[0047] Figure 8 To include Figure 7 The flowchart shows the steps performed during the exploration phase. Detailed Implementation
[0048] Overview
[0049] For example, in an electrophysiology (EP) laboratory, one or more catheters and surface electrode patches can be connected to a signal processing console that processes signals from the catheters and surface electrodes to perform various tasks associated with mapping and ablation.
[0050] Such signal processing consoles typically include a limited number of available input connectors for attaching electrodes and sensors. Additionally, conduits are being manufactured with an ever-increasing number of electrodes (dozens or even hundreds). For example, the Octaray from Biosense Webster (Irvine, CA). TM The catheter includes 48 mapping electrodes, two orientation sensor electrodes, and a magnetic sensor, all of which need to be connected to a signal processing console.
[0051] One solution is to replace the signal processing console with a console having one or more connectors. Another, more acceptable solution is to provide an auxiliary signal processing unit connected to a subgroup of electrodes, while the other electrodes are connected to the signal processing console. The auxiliary signal processing unit processes the signals received from the connected subgroup of electrodes, in parallel with the processing performed by the signal processing console (hereinafter referred to as the "main signal processing unit"). The outputs of the main and auxiliary signal processing units can then be processed by another processing device that performs tasks such as mapping, ablation control, and user interface rendering.
[0052] One problem with the above solutions involves the calculation of tissue proximity indication (TPI), which provides a measure of the proximity of each catheter electrode to tissue in a body part (e.g., the ventricle). The main signal processing unit can process the TPI in a different manner than the auxiliary signal processing unit. For example, the TPI can be calculated based on an impedance value obtained from the impedance between the catheter electrode and a surface electrode. When the surface electrode is connected to the main signal processing unit but not to the auxiliary signal processing unit, the main signal processing unit can calculate the TPI based on the impedance between the catheter electrode and the surface electrode, while the auxiliary signal processing unit can calculate the TPI based on bipolar impedance measured between each pair of catheter electrodes. Even for the same electrode, the TPI calculated based on bipolar impedance can differ from the TPI calculated by the main signal processing unit and may provide poorer results. In some examples, the main signal processing unit and the auxiliary signal processing unit may include different processing hardware, which can also lead to different quality of the calculated TPI.
[0053] The embodiments of the present invention address the above-mentioned problems by using reinforcement learning (e.g., Q-learning, deep Q-learning) to derive the TPI of electrodes connected only to the auxiliary signal processing unit based on the TPI calculated in the main signal processing unit, thereby providing similar results for electrodes connected only to the auxiliary signal processing unit.
[0054] One or more catheter electrodes (e.g., electrode A) are connected to both the main signal processing unit and the auxiliary signal processing unit, and can be used to calibrate the TPI calculated for catheter electrodes connected only to the auxiliary signal processing unit. The auxiliary signal processing unit calculates the impedance value of electrode A (e.g., based on a bipolar signal), and the main signal processing unit calculates the TPI of electrode A (e.g., based on the impedance value between electrode A and a surface electrode). The processing device receives the impedance value calculated by the auxiliary signal processing unit and obtains the TPI of the main signal processing unit. The processing device applies a reinforcement learning agent during the exploitation phase and rewards or punishes the agent during the exploration phase, as described in more detail below. The processing device performs subsequent exploitation and exploration phases in response to new impedance values and TPIs, and the agent is continuously trained between exploitation phases.
[0055] Other electrodes (e.g., electrode B) are connected only to an auxiliary signal processing box, which calculates the impedance values of these electrodes. The processing device then applies an agent in the utilization phase to determine the TPI of those electrodes based on input states (e.g., the impedance value of electrode B received from the auxiliary signal processing unit, and the contact state of previously provided electrode B (e.g., contact or non-contact)) and actions (e.g., changing or not changing the TPI from a previously provided TPI), thereby generating a first expected reward for that action. An action with the same state but opposite is input to the agent to generate a second expected reward for that action. The action associated with the highest reward is selected and taken.
[0056] The agent can be trained (during the exploration phase) using impedance data obtained from the auxiliary signal processing unit and TPI results obtained from the main signal processing unit. The actual reward is calculated based on whether the last action taken matches the TPI received from electrode A from the main signal processing unit. The actual reward can be calculated as the sum of future discounted rewards. The agent is then trained by rewarding and penalizing the agent based on the actual rewards. If the last action taken matches the TPI calculated by the main signal processing unit, the agent is rewarded (e.g., the actual reward may be +1). If the last action taken does not match the TPI calculated by the main signal processing unit, the agent is penalized (e.g., the actual reward may be -1). In some cases, the reward function can be discounted.
[0057] The state and the last action taken during the final use of electrode A can be input into the agent to generate the expected reward (e.g., a value between -1 and +1). The agent's parameters can be iteratively changed until the expected reward (i.e., the agent's actual output) is sufficiently close to the actual reward (i.e., the agent's desired output). A suitable loss function can be used to perform the comparison between the agent's actual output and the desired output.
[0058] System Description
[0059] See now Figure 1 This is a schematic diagram of a medical protocol system 20 constructed and operated according to an embodiment of the present invention. See now. Figure 2 , it is used for Figure 1 A schematic diagram of the conduit 40 in system 20.
[0060] The medical surgical system 20 is used to determine the orientation of the catheter 40, such as in... Figure 1 In illustration 25 and in Figure 2 As seen in more detail below. The catheter 40 includes a shaft 22 and a plurality of flexible arms 54 (only some are labeled for simplicity) for insertion into a body part of a living subject. The flexible arms 54 have corresponding proximal ends connected to the distal end of the shaft 22.
[0061] The conduit 40 includes an orientation sensor 53, which is positioned on the shaft 22 in a predefined spatial relationship relative to the proximal end of the flexible arm 54. The orientation sensor 53 may include a magnetic sensor 50 and / or at least one axial electrode 52. The magnetic sensor 50 may include at least one coil, such as, but not limited to, a biaxial or triaxial coil arrangement, to provide positional and orientation (including yaw) data. The conduit 40 includes a plurality of electrodes 55 positioned at different corresponding locations along each flexible arm 54 (only some are labeled for simplicity). Figure 2 (in the middle). Typically, catheter 40 can be used to map electrical activity in the heart of a living subject using electrode 55, or to perform any other suitable function in a body part of a living subject.
[0062] The medical surgical system 20 can determine the orientation and orientation of the shaft 22 of the catheter 40 based on signals provided by the magnetic sensor 50 and / or shaft electrodes 52 (proximal electrode 52a and distal electrode 52b) mounted on the shaft 22 on either side of the magnetic sensor 50. At least some of the electrodes, including the proximal electrode 52a, distal electrode 52b, magnetic sensor 50, and electrode 55, are connected to various drive circuits in the console 24 via wires extending through the shaft 22 via one or more catheter connectors 35. In some embodiments, at least two of the electrodes 55, shaft electrodes 52, and magnetic sensors 50 of each flexible arm 54 are connected to the drive circuits in the console 24 via the catheter connectors 35. In some embodiments, the distal electrode 52b and / or the proximal electrode 52a may be omitted.
[0063] Figure 2 The illustrations shown are chosen solely for clarity of concept. Other configurations of the axial electrode 52 and electrode 55 are also possible. Additional functionality may be included in the orientation sensor 53. For clarity, elements irrelevant to the embodiments disclosed in this invention, such as flushing ports, have been omitted.
[0064] The physician 30 navigates the catheter 40 to a target location in the body part (e.g., heart 26) of the patient 28 by manipulating the axis 22 and / or flexing from the sheath 23 using a manipulator 32 located near the proximal end of the catheter 40. The catheter 40 is inserted through the sheath 23, where the flexible arms 54 converge, and the flexible arms 54 are able to unfold and return to their intended functional shape only after the catheter 40 has retracted from the sheath 23. By housing the flexible arms 54 together, the sheath 23 also serves to minimize vascular trauma along its path to the target location.
[0065] The console 24 includes processing circuitry 41 (typically a general-purpose computer) and suitable front-end and interface circuitry 44 for generating signals in and / or receiving signals from surface electrodes 49 attached to the chest and back of the patient 28, or any other suitable skin surface, via wires passing through cable 39.
[0066] The console 24 also includes a magnetic induction subsystem. The patient 28 is placed in a magnetic field generated by a pad comprising at least one magnetic field radiator 42, driven by a unit 43 disposed in the console 24. The magnetic field radiator 42 is configured to emit an alternating magnetic field toward the area where a body part (e.g., heart 26) is located. The magnetic field generated by the magnetic field radiator 42 generates a direction signal in a magnetic sensor 50. The magnetic sensor 50 is configured to detect at least a portion of the emitted alternating magnetic field and provides the direction signal as a corresponding electrical input to processing circuitry 41.
[0067] In some embodiments, processing circuitry 41 uses orientation signals received from axial electrode 52, magnetic sensor 50, and electrode 55 to estimate the orientation of catheter 40 within an organ, such as a heart chamber. In some embodiments, processing circuitry 41 correlates the orientation signals received from electrodes 52, 55 with previously acquired magnetic position-calibration orientation signals to estimate the orientation of catheter 40 within the heart chamber. The positional coordinates of axial electrode 52 and electrode 55 can be determined by processing circuitry 41 based on (among other inputs) the measured impedance or current distribution between electrodes 52, 55 and surface electrode 49. Console 24 drives display 27, which shows the distal end of catheter 40 within heart 26.
[0068] Methods utilizing current distribution measurements and / or orientation sensing of external magnetic fields are implemented in various medical applications, for example, in devices manufactured by Biosense Webster Inc. (Irvine, California). The system is implemented and described in detail in U.S. Patents 5,391,199, 6,690,963, 6,484,118, 6,239,724, 6,618,612, 6,332,089, 7,756,576, 7,869,865 and 7,848,787, PCT Patent Publication WO 96 / 05768, and U.S. Patent Application Publications 2002 / 0065455A1, 2003 / 0120150A1 and 2004 / 0068178A1.
[0069] The system employs a position tracking method based on active current positioning (ACL) impedance. In some embodiments, processing circuitry 41 is configured to use the ACL method to generate a mapping between an indication of impedance and the orientation in the magnetic coordinate system of magnetic field radiator 42 (e.g., a current-orientation matrix (CPM)). Processing circuitry 41 estimates the orientation of axial electrodes 52 and 55 by performing a lookup in the CPM.
[0070] The processing circuit 41 is typically programmed with software to perform the functions described herein. This software may be downloaded electronically to a computer via a network, or alternatively or additionally set and / or stored on a non-transitory tangible medium (such as magnetic storage, optical storage, or electronic storage).
[0071] For the sake of simplicity and clarity, Figure 1 Only elements relevant to the technology disclosed in this invention are shown. System 20 typically includes additional modules and elements that are not directly related to the technology disclosed in this invention, and therefore these additional modules and elements are derived from... Figure 1 The corresponding descriptions were intentionally omitted.
[0072] The catheter 40 described above includes eight flexible arms 54, each arm 54 having six electrodes. By way of example only, any suitable catheter may be used in place of catheter 40 to perform the functions described above and below, such as catheters with different numbers of flexible arms and / or electrodes on each arm, or catheters with different distal end types such as balloon catheters, basket catheters, or lasso catheters.
[0073] The medical surgical system 20 can also use any suitable catheter, such as catheter 40 or different catheters, and any suitable ablation method to perform ablation of cardiac tissue. The console 24 may include an RF signal generator (not shown) configured to generate RF power applied by one or more electrodes of the catheter connected to the console 24 and one or more surface electrodes of the surface electrodes 49 to ablate the myocardium of the heart 26. The console 24 may include a pump (not shown) that pumps flushing fluid into a flushing channel to the distal end of the catheter performing ablation. The catheter performing ablation may also include a temperature sensor (not shown) for measuring the temperature of the myocardium during ablation and adjusting the ablation power and / or the flushing rate of the flushing fluid pump based on the measured temperature.
[0074] See now Figure 3 and Figure 4 . Figure 3 To perform reinforcement learning Figure 1 A schematic diagram of the components of system 20. Figure 4 To include execution for Figure 1The flowchart 400 shows the steps in the reinforcement learning method in System 20.
[0075] The catheter 40 includes electrodes 55 in a first group 60 and a second group 62. The first group 60 is different from the second group 62. For example, the first group 60 includes electrodes 1 to 14, and the second group 62 includes electrodes 15 to 48. For simplicity, the electrodes 55 have been described as being divided into two groups to explain how the electrodes 55 can be connected to different processing devices, and these two groups do not necessarily reflect how the electrodes 55 are positioned on the catheter 40.
[0076] Processing circuitry 41 includes a main signal processing unit 64, an auxiliary signal processing unit 66, and a processing device 68. Each of the main signal processing unit 64 and the auxiliary signal processing unit 66 may include corresponding signal processing circuitry, such as an analog-to-digital converter, noise reduction circuitry, and other filtering circuitry. The main signal processing unit 64 is configured to be electrically coupled to a first set of 60 electrodes 55 and optional body surface electrodes 49 (block 402). The auxiliary signal processing unit 66 is configured to be electrically coupled to a second set of 62 electrodes 55 (block 404). For training the reinforcement learning agent 70, one or more electrodes of the first set of 60 electrodes 55 are also coupled to the auxiliary signal processing unit 66. For simplicity, in the description provided below, it will be assumed that electrode 1 of the first set of 60 electrodes 55 is also coupled to the auxiliary signal processing unit 66. However, in practice, more than one electrode of the first set of 60 electrodes 55 may be coupled to the auxiliary signal processing unit 66 and provide data for training the reinforcement learning agent 70. Therefore, when the description below refers to electrode 1, the same or similar processing may also be performed on one or more electrodes of electrode 55 in the first group 60.
[0077] The catheter 40 is configured to be inserted into a living subject (e.g., Figure 1 Patient 28) body parts (e.g., Figure 1 The electrode 55 is located within the heart (frame 406). The electrode 55 is configured to contact tissue at a corresponding location within the body part. The surface electrode 49 is configured to be applied to the skin surface of a living subject.
[0078] An auxiliary signal processing unit 66 is configured to receive signals from electrode 1 in the second set of 62 electrodes 55 and the first set of 60 electrodes 55 of the catheter 40 (block 408). The auxiliary signal processing unit 66 is configured to calculate a corresponding impedance value 72 of electrode 1 over time in response to receiving at least one signal provided by electrode 1 (block 410). The signal provided by the electrode (e.g., electrode 1) can be received directly from the electrode via a wiring in the catheter 40 connected to the control console 24, or via a signal emitted from the electrode (e.g., electrode 1) and detected by a surface electrode 49. The auxiliary signal processing unit 66 is also configured to calculate the corresponding impedance value of the respective electrode in the second set of 62 electrodes 55.
[0079] As previously mentioned, each exploration phase follows immediately after the exploitation phase. For example, reinforcement learning agent 70 is first applied to learn the actions to be taken in the exploitation phase, and then the reinforcement learning agent 70 is rewarded or penalized in the exploration phase based on the actions taken in the most recent exploitation phase. Therefore, the application of reinforcement learning agent 70 in the exploitation phase is described first, and the rewarding or penalizing of reinforcement learning agent 70 in the exploration phase is described then. First, the parameters of reinforcement learning agent 70 are initialized (e.g., randomized), and the parameters are changed in successive exploitation and exploration phases, while reinforcement learning agent 70 learns how to follow the TPI calculated by the main signal processing unit 64.
[0080] The processing device 68 is configured to receive the impedance values 72 of the corresponding electrodes 55 of the second group 62 from the auxiliary signal processing unit 66.
[0081] Processing device 68 is configured to apply reinforcement learning agent 70 (box 412) during the reinforcement learning exploitation phase to acquire a corresponding organization proximity action 74 to be taken in response to at least one organization proximity policy, the organization proximity action maximizing the corresponding expected reward of electrode 1. The steps of box 412 are referenced. Figure 5 and Figure 6 A more detailed description was provided.
[0082] Similarly, the processing device 68 is configured to apply a reinforcement learning agent 70 to acquire a corresponding tissue proximity action 74 to be taken in response to the calculated corresponding impedance value of the corresponding electrode 55 of the second group 62, which maximizes the corresponding expected reward of the corresponding electrode 55 of the second group 62.
[0083] The processing device 68 is configured to provide a corresponding derived tissue proximity indication (TPI) 76 (block 414) of the proximity of electrode 1 to tissue in response to a corresponding tissue proximity action 74 acquired by electrode 1. For example, if the tissue proximity action 74 to be taken is to change the TPI from a previously provided TPI, and the previous TPI was equal to "contact", then the current TPI will be "no contact". If the tissue proximity action 74 to be taken is not to change the TPI from a previously provided TPI, and the previous TPI was equal to "contact", then the current TPI will be "contact".
[0084] Similarly, the processing device 68 is configured to provide a corresponding derived tissue proximity indication 76 of the proximity of the corresponding electrode 55 of the second group 62 to the tissue in response to a corresponding tissue proximity action 74 acquired by the corresponding electrode 55 of the second group 62. The processing device 68 is configured to render the derived TPI 76 of the electrode 1 and the corresponding electrode 55 of the second group 62 to the display 27.
[0085] The main signal processing unit 64 is configured to receive signals from the first set 60 electrodes 55 and the surface electrodes 49 of the catheter 40 (block 416). The main signal processing unit 64 is configured to calculate, in response to the received signals, the corresponding impedance values between the electrodes 1 and the surface electrodes 49 over time (block 418). In other embodiments, any suitable method, such as based on bipolar signals, can be used to calculate the impedance values. The main signal processing unit 64 is typically configured to calculate the impedance values over time for the other electrodes 55 of the first set 60.
[0086] The main signal processing unit 64 is configured to calculate a corresponding reference tissue proximity indicator (TPI) 78 (block 420) in response to the calculated corresponding impedance value of electrode 1 (i.e., in response to receiving at least one signal provided by electrode 1). The reference TPI 78 is described as a “reference” TPI because it can be used as a quality reference during the training of the reinforcement learning agent 70. The main signal processing unit 64 is configured to calculate the corresponding reference tissue proximity indicator 78 independently of the application of the reinforcement learning agent 70 and in response to at least one of the received signals (which are used to calculate the impedance value of electrode 1). Similarly, the main signal processing unit 64 is configured to calculate the reference TPIs for the other electrodes in the first set of 60 electrodes 55. The steps in blocks 402 to 420 can be performed in any suitable order.
[0087] Processing device 68 is configured to selectively reward and punish reinforcement learning agent 70 (box 422) during the reinforcement learning exploration phase to learn organizational proximity policies (or multiple policies) in response to data including reference TPI 78, impedance value 72, and other data (e.g., state, actions taken, and actual rewards), such as reference TPI 78. Figure 7 and Figure 8 A more detailed description follows. In each reinforcement learning exploration phase, the processing device 68 is configured to reward or punish the reinforcement learning agent 70. As described above, data (including a reference TPI 78 and an impedance value 72) is calculated in response to at least one of the signals received from the electrode 55 and the body surface electrode 49. The steps of boxes 408 to 422 (arrow 424) are repeated for subsequent exploitation and exploration phases.
[0088] In some implementations, the reinforcement learning agent 70 is a deep reinforcement learning agent. Deep reinforcement learning uses deep neural networks that incorporate deep learning into the solution, allowing the agent to make decisions from unstructured input data without manually designing the state space. In other implementations, the reinforcement learning agent 70 is a Q-learning model-free reinforcement learning agent. Q-learning is model-free (does not require an environment model) and can solve problems involving stochastic transitions and rewards. Any suitable reinforcement learning algorithm can be used, such as Monte Carlo, SARSA, Q-learning-λ, SARSA-λ, DQN, DDPG, A3C, NAF, TRPO, PPO, TD3, or SAC.
[0089] In implementation, some or all of the functions of processing circuitry 41 may be combined in a single physical component, or alternatively, implemented using multiple physical components. These physical components may include hardwired or programmable devices, or a combination of both. In some embodiments, at least some of the functions of processing circuitry 41 may be implemented by a programmable processor under the control of suitable software. This software may be downloaded to the device electronically via, for example, a network. Alternatively or otherwise, the software may be stored in a tangible, non-transitory computer-readable storage medium, such as optical, magnetic, or electronic memory.
[0090] See now Figure 5 and Figure 6 . Figure 5 To show the use of Figure 1 A schematic diagram of the reinforcement learning utilization phase in System 20. Figure 6 To include Figure 5 The flowchart 600 shows the steps performed during the utilization phase.
[0091] Each stage in the utilization phase has an associated state 500, which, along with a set of different available tissue proximity actions 502 to be taken for that stage, serves as input to the reinforcement learning agent 70. Each corresponding state 500 includes: a corresponding derived tissue proximity indication 76 of electrode 1 provided in the corresponding final reinforcement learning utilization phase. Figure 3); and the corresponding impedance value 72 of electrode 1 calculated for the current corresponding reinforcement learning utilization phase. Figure 3 Each available tissue proximity action 502 to be taken includes: changing the derived tissue proximity indicator 76 of electrode 1 provided in the last reinforcement learning utilization phase (box 504); and not changing the derived tissue proximity indicator 76 of electrode 1 provided in the last reinforcement learning utilization phase (box 506). Processing device 68 is configured to update state 500 (box 602) in response to the current impedance value 72 of electrode 1 and the last derived TPI 76 provided by electrode 1 (in the previous utilization phase). In the initial utilization phase, the last derived TPI 76 provided by electrode 1 can be set to an initial value, for example, no contact.
[0092] The processing device 68 is configured to apply the reinforcement learning agent 70 (box 604) in the corresponding reinforcement learning utilization phase to acquire the corresponding tissue proximity action of the electrode 1 to be taken 74 in response to the corresponding state 500 of the reinforcement learning agent 70 of the electrode 1 and the corresponding available tissue proximity action 502 set of the electrode 1 to be taken, which maximizes (box 508) the corresponding expected reward (box 510).
[0093] In one utilization phase, the processing device 68 is configured to: apply a reinforcement learning agent 70 (box 606) having a state 500 and an action (box 504) of changing the derived TPI of electrode 1, thereby generating a anticipated reward 510-1; apply a reinforcement learning agent 70 (box 608) having a state 500 and an action (box 506) of not changing the derived TPI of electrode 1, thereby generating a anticipated reward 510-2; and maximize the anticipated reward 510 (boxes 508, 610) by selecting a tissue proximity action 502 to be taken by (electrode 1), thereby resulting in a maximum anticipated reward 510 (e.g., from anticipated rewards 510-1 and 510-2), thereby generating a tissue proximity action 74 to be taken by electrode 1 (e.g., changing the TPI or not changing the TPI). The processing device 68 is configured to derive the TPI 76 (box 612) of electrode 1 from the tissue proximity action 74 to be taken, and to provide the derived TPI 76 (box 614) of electrode 1 to physician 30 (e.g., by rendering the derived TPI 76 of electrode 1 onto display 27). Figure 1 )).
[0094] Processing device 68 is configured to connect to auxiliary signal processing unit 66 ( Figure 3 The other electrodes 55 (e.g., the second group 62 and optionally at least one other electrode 55 of the first group 60) perform the steps described above in blocks 602 to 614.
[0095] See now Figure 7 and Figure 8 . Figure 7 To show the use of Figure 1 A schematic diagram of the exploration phase of reinforcement learning in System 20. Figure 8 To include Figure 7 The flowchart 800 shows the steps performed during the exploration phase.
[0096] Processing device 68 is configured to respond to a corresponding tissue proximity action 74 to be taken. Figure 5 (i.e., the tissue proximity action 74 to be taken in the final utilization phase is used in the current exploration phase) and the corresponding actual reward 700 for electrode 1 in the respective exploration phase is calculated by the main signal processing unit 64 independently of the corresponding reference tissue proximity indication 78 of electrode 1 calculated by the applied reinforcement learning agent 70 (i.e., calculated based on impedance data from the final utilization phase) (box 802). In some embodiments, the processing device 68 is configured to: calculate the sum of future discount rewards; and calculate the corresponding actual reward 700 in response to the calculated sum of future discount rewards. The steps of box 802 may include the processing device 68 being configured to derive a TPI 76 for each respective exploration phase (i.e., the tissue proximity action 74 to be taken in the final utilization phase is used in the current exploration phase) and the corresponding actual reward 700 for electrode 1 in the respective exploration phase is calculated by the main signal processing unit 64 independently of the applied reinforcement learning agent 70. Figure 3 (Derived in the final utilization stage) and reference TPI 78 ( Figure 3 (Calculated based on impedance data from the final utilization phase) Compare (box 804). If the derived TPI 76 for electrode 1 in one of the exploration phases is the same as the reference TPI 78 for electrode 1 in that exploration phase, the actual reward can be equal to a high value, such as +1; and if the derived TPI 76 is different from the reference TPI 78 for electrode 1 in that exploration phase, the actual reward can be equal to a low value, such as -1.
[0097] Processing device 68 is configured to selectively reward and punish reinforcement learning agent 70 (box 806) during the reinforcement learning exploration phase in response to data from the corresponding last reinforcement learning exploitation phase within the reinforcement learning exploitation phase. In other words, the data used in each exploration phase comes from the last exploitation phase preceding that exploration phase. The data used for each corresponding exploitation phase includes: a corresponding state in state 500 (i.e., state 500 in the corresponding last exploitation phase), the corresponding acquired organizational proximity action 74 to be taken, and so on. Figure 5 And the corresponding actual rewards for the actions taken in the corresponding final utilization phase (i.e., the actions to be taken).
[0098] In one exploration phase of the exploration phase, processing device 68 is configured to: apply a reinforcement learning agent 70 with a final state 500 and an action 702 taken (i.e., the organizational proximity action 74 to be taken acquired in the final utilization phase) (box 808); compare the actual output (e.g., expected reward 704) of the reinforcement learning agent 70 with the desired output (e.g., the actual reward 700 calculated in the step of box 802 for this exploration phase) using any suitable loss function (boxes 706, 810); at decision box 812, determine whether the actual output is within a given threshold of the desired output, and if the actual output is not within the given threshold (branch 814), processing device 68 is configured to update the parameters (e.g., weights) of the reinforcement learning agent 70 (box 816) and continue with the steps of box 808; and if the actual output is within the given threshold (branch 818), continue with the steps of box 802 for the next exploration phase. The processing device 68 is configured to update the parameters using any suitable optimization algorithm (e.g., gradient descent algorithms such as the Adam optimization algorithm).
[0099] As used herein, the term “about” or “approximately” for any numerical value or range indicates a suitable dimensional tolerance that allows a collection of parts or components to achieve the intended purpose as described herein. More specifically, “about” or “approximately” may refer to a range of ±20% of the enumerated value; for example, “about 90%” may refer to a range of values from 72% to 108%.
[0100] For clarity, the various features of the invention described in the context of individual embodiments may also be provided in combination in a single embodiment. Conversely, for brevity, the various features of the invention described in the context of individual embodiments may also be provided individually or in any suitable sub-combination.
[0101] The above embodiments are cited by way of example, and the invention is not limited to the specific examples shown and described above. Rather, the scope of the invention includes combinations and sub-combinations of the various features described above, as well as variations and modifications thereof, which will occur to those skilled in the art upon reading the above description and are not disclosed in the prior art.
Claims
1. A computer-implemented method for acquiring tissue proximity indication, the method being performed when a catheter is inserted into a body part of a living subject such that electrodes of the catheter contact tissue at corresponding locations within the body part, the method comprising the following steps executed by a computer processor: Receive the signal provided by the electrodes; In response to at least one of the received signals, the reinforcement learning agent is selectively rewarded and punished during the reinforcement learning exploration phase to learn at least one organizational proximity policy. The reinforcement learning agent is applied in the reinforcement learning utilization phase to acquire a corresponding organization proximity action to be taken in response to the at least one organization proximity policy, the corresponding organization proximity action maximizing the corresponding expected reward. as well as In response to the acquired tissue proximity action, a corresponding derived tissue proximity indication is provided between a given electrode in the electrodes and the tissue. The selective reward and punishment includes responding to data from the corresponding final reinforcement learning exploitation phase within the reinforcement learning exploration phase, selectively rewarding and punishing the reinforcement learning agent during the phase. The data includes: the agent's corresponding state, the corresponding acquired organization proximity action to be taken, and the corresponding actual reward. The method further includes: In response to at least one of the received signals, independently of applying the reinforcement learning agent to calculate the corresponding reference tissue proximity indication, and The corresponding actual reward is calculated based on the corresponding reference organization proximity indicator.
2. The method according to claim 1, wherein: The application includes applying the reinforcement learning agent in a corresponding reinforcement learning exploitation phase within the reinforcement learning exploitation phase, in response to the corresponding state of the reinforcement learning agent and the corresponding available organizational proximity action set to be taken, to acquire the corresponding organizational proximity action among the organizational proximity actions to be taken, the corresponding organizational proximity action maximizing the corresponding expected reward.
3. The method according to claim 2, wherein, Each organization proximity action in the group of available organization proximity actions to be taken includes: changing a corresponding organization proximity indicator in the derived organization proximity indicator provided in the corresponding last reinforcement learning utilization phase in the reinforcement learning utilization phase; and not changing the corresponding organization proximity indicator in the derived organization proximity indicator provided in the corresponding last reinforcement learning utilization phase in the reinforcement learning utilization phase.
4. The method of claim 3, wherein each of the corresponding states comprises: The corresponding organization proximity indicator in the derived organization proximity indicator provided in the corresponding last reinforcement learning utilization phase in the reinforcement learning utilization phase; And the corresponding impedance value of the given electrode calculated for the current corresponding reinforcement learning utilization phase in the reinforcement learning utilization phase.
5. The method of claim 4, further comprising calculating the corresponding actual reward in response to the corresponding acquired organization proximity action in the organization proximity action to be taken and a corresponding reference organization proximity indication calculated independently of applying the reinforcement learning agent.
6. The method according to claim 5, wherein, Calculating the corresponding actual reward includes calculating the sum of future discount rewards.
7. The method according to claim 5, further comprising: The corresponding impedance value of the given electrode is calculated for each of the corresponding states in response to at least one of the received signals.
8. The method according to claim 7, wherein, The calculation of the corresponding reference tissue proximity indication includes calculating the corresponding reference tissue proximity indication in response to the corresponding impedance value of the given electrode.
9. The method according to claim 7, further comprising: The first set of electrodes of the catheter and the given electrode are electrically coupled to a first signal processing unit, wherein the calculation of the corresponding reference tissue proximity indication is performed by the first signal processing unit in response to receiving at least one signal provided by the given electrode; as well as The second set of electrodes of the catheter and the given electrode are electrically coupled to a second signal processing unit, the second set of electrodes being different from the first set of electrodes, wherein the calculation of the corresponding impedance value of the given electrode for each of the corresponding states is performed in the second signal processing unit in response to receiving the at least one signal provided by the given electrode.
10. The method of claim 9, further comprising: The second signal processing unit calculates the corresponding impedance value of the corresponding electrode in the second group of electrodes; The reinforcement learning agent is applied to obtain a corresponding tissue proximity action to be taken in response to the calculated corresponding impedance value of the corresponding electrode in the second set of electrodes, the corresponding tissue proximity action maximizing the corresponding expected reward of the corresponding electrode in the second set of electrodes; as well as In response to the corresponding tissue proximity action to be taken by the corresponding electrode in the second set of electrodes, a corresponding derived tissue proximity indication of the proximity between the corresponding electrode in the second set of electrodes and the tissue is provided.
11. The method of claim 9, further comprising: The body surface electrode is applied to the skin surface of the living subject; The body surface electrodes are electrically coupled to the first signal processing unit; as well as Calculate the corresponding impedance value between the given electrode and the body surface electrode, wherein the calculation of the corresponding reference tissue proximity indication is performed by the first signal processing unit in response to the calculated corresponding impedance value.
12. The method according to claim 1, wherein, The reinforcement learning agent is a deep reinforcement learning agent.
13. The method according to claim 1, wherein, The reinforcement learning agent is a Q-learning model-free reinforcement learning agent.
14. A system for obtaining an organization proximity indicator, comprising: A catheter configured to be inserted into a body part of a living subject, and including electrodes configured to contact tissue at corresponding locations within the body part; as well as Processing circuit, the processing circuit being configured to: Receive the signal provided by the electrodes; In response to at least one of the received signals, the reinforcement learning agent is selectively rewarded and punished during the reinforcement learning exploration phase to learn at least one organizational proximity policy. The reinforcement learning agent is applied in the reinforcement learning utilization phase to acquire a corresponding organization proximity action to be taken in response to the at least one organization proximity policy, the corresponding organization proximity action maximizing the corresponding expected reward. as well as In response to the acquired tissue proximity action, a corresponding derived tissue proximity indication is provided between a given electrode in the electrodes and the tissue. The selective reward and punishment includes responding to data from the corresponding final reinforcement learning exploitation phase within the reinforcement learning exploration phase, selectively rewarding and punishing the reinforcement learning agent during the phase. The data includes: the agent's corresponding state, the corresponding acquired organization proximity action to be taken, and the corresponding actual reward. The processing circuit is further configured as follows: In response to at least one of the received signals, independently of applying the reinforcement learning agent to calculate the corresponding reference tissue proximity indication, and The corresponding actual reward is calculated based on the corresponding reference organization proximity indicator.
15. The system according to claim 14, wherein, The processing circuit is configured to: The reinforcement learning agent is applied in the corresponding reinforcement learning utilization phase of the reinforcement learning utilization phase to acquire the corresponding organization proximity action among the organization proximity actions to be taken in response to the corresponding state of the reinforcement learning agent and the corresponding available organization proximity action set to be taken, the corresponding organization proximity action maximizing the corresponding expected reward.
16. The system according to claim 15, wherein, Each organization proximity action in the group of available organization proximity actions to be taken includes: changing a corresponding organization proximity indicator in the derived organization proximity indicator provided in the corresponding last reinforcement learning utilization phase in the reinforcement learning utilization phase; and not changing the corresponding organization proximity indicator in the derived organization proximity indicator provided in the corresponding last reinforcement learning utilization phase in the reinforcement learning utilization phase.
17. The system of claim 16, wherein each of the respective states comprises: The corresponding organization proximity indicator in the derived organization proximity indicator provided in the corresponding last reinforcement learning utilization phase in the reinforcement learning utilization phase; And the corresponding impedance value of the given electrode calculated for the current corresponding reinforcement learning utilization phase in the reinforcement learning utilization phase.
18. The system according to claim 17, wherein, The processing circuitry is configured to calculate the corresponding actual reward in response to the corresponding acquired organization proximity action to be taken and the corresponding reference organization proximity indication calculated independently of the reinforcement learning agent.
19. The system according to claim 18, wherein, The processing circuit is configured to: calculate the sum of future discount rewards; and calculate the corresponding actual reward in response to the calculated sum of future discount rewards.
20. The system according to claim 18, wherein, The processing circuit is configured to: The corresponding impedance value of the given electrode is calculated for each of the corresponding states in response to at least one of the received signals.
21. The system according to claim 20, wherein, The processing circuit is configured to calculate the corresponding reference tissue proximity indication in response to the corresponding impedance value of the given electrode.
22. The system according to claim 20, wherein, The catheter includes a first set of electrodes and a second set of electrodes, the second set of electrodes being different from the first set of electrodes. The system further includes: A first signal processing unit, configured to be electrically coupled to the first set of electrodes, and to calculate the corresponding reference tissue proximity indication in response to receiving at least one signal provided by the given electrodes; and A second signal processing unit is configured to be electrically coupled to the second set of electrodes of the conduit, and in response to receiving the at least one signal provided by the given electrode, calculates the corresponding impedance value of the given electrode for each of the corresponding states.
23. The system according to claim 22, wherein: The second signal processing unit is configured to calculate the corresponding impedance value of the corresponding electrode in the second group of electrodes; The processing circuit is configured to apply the reinforcement learning agent to obtain a corresponding tissue proximity action to be taken in response to a calculated corresponding impedance value of the corresponding electrode in the second set of electrodes, the corresponding tissue proximity action maximizing the corresponding expected reward of the corresponding electrode in the second set of electrodes. as well as In response to the corresponding tissue proximity action to be taken by the corresponding electrode in the second set of electrodes, a corresponding derived tissue proximity indication of the proximity between the corresponding electrode in the second set of electrodes and the tissue is provided.
24. The system of claim 22, further comprising a body surface electrode configured to be applied to the skin surface of the living subject, the first signal processing unit configured to be electrically coupled to the body surface electrode, wherein the first signal processing unit is configured to: calculate a corresponding impedance value between the given electrode and the body surface electrode; and calculate the corresponding reference tissue proximity indication in response to the calculated corresponding impedance value.
25. The system according to claim 14, wherein, The reinforcement learning agent is a deep reinforcement learning agent.
26. The system according to claim 14, wherein, The reinforcement learning agent is a Q-learning model-free reinforcement learning agent.
27. A software product comprising a non-transitory computer-readable medium storing program instructions in the non-transitory computer-readable medium, the instructions causing the CPU, when read by a central processing unit (CPU): Receive signals provided by the electrodes; In response to at least one of the received signals, the reinforcement learning agent is selectively rewarded and punished during the reinforcement learning exploration phase to learn at least one organizational proximity policy. The reinforcement learning agent is applied in the reinforcement learning utilization phase to acquire a corresponding organization proximity action to be taken in response to the at least one organization proximity policy, the corresponding organization proximity action maximizing the corresponding expected reward. as well as In response to the acquired tissue proximity action, a corresponding derived tissue proximity indication is provided between a given electrode in the electrodes and the tissue. The selective reward and punishment includes responding to data from the corresponding final reinforcement learning exploitation phase within the reinforcement learning exploration phase, selectively rewarding and punishing the reinforcement learning agent during the phase. The data includes: the agent's corresponding state, the corresponding acquired organization proximity action to be taken, and the corresponding actual reward. When the instruction is read by the CPU, it further causes the CPU to: In response to at least one of the received signals, independently of applying the reinforcement learning agent to calculate the corresponding reference tissue proximity indication, and The corresponding actual reward is calculated based on the corresponding reference organization proximity indicator.
Citation Information
Patent Citations
Medical diagnosis, treatment and imaging systems
US20020065455A1
Wireless position sensor
US20030120150A1
High-gradient recursive locating system
US20040068178A1
Systems and methods for electrode contact assessment
US20070100332A1
Catheter with pressure sensing
US20090093806A1