Magnetic capsule endoscope autonomous control method and system based on combination of tube model predictive control and reinforcement learning
By combining Tube model predictive control and reinforcement learning, the motion control accuracy and stability of capsule endoscopes have been improved, solving the problems of disturbance resistance and multi-task switching in complex environments of existing systems, and realizing efficient and safe capsule endoscope navigation and scanning.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
- Filing Date
- 2026-02-09
- Publication Date
- 2026-06-23
AI Technical Summary
Existing robotic capsule endoscope magnetic control systems are insufficient in resisting disturbances when facing the complex and ever-changing working environment inside the digestive tract. They are unable to achieve smooth switching between multiple motion modes and precise positioning, and their control precision is insufficient to meet the needs of fine diagnosis.
A method combining Tube model predictive control and reinforcement learning is adopted. The capsule endoscope is driven by a permanent magnet component at the end of the robotic arm. Robust control and predictive optimization under the Tube MPC control framework are combined with the compensation control quantity generated by the reinforcement learning agent. This enables online learning and compensation of unknown disturbances and nonlinear characteristics in the system, thereby improving the system's adaptability and control accuracy.
It significantly improves the trajectory tracking accuracy and attitude stability of capsule endoscopes, enabling efficient switching in different medical examination scenarios, enhancing the system's anti-disturbance capability and robustness, and ensuring the feasibility and safety of control commands.
Smart Images

Figure CN122250906A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to medical robots, and more particularly to an autonomous control method and system for a magnetically controlled capsule endoscope based on a combination of Tube model predictive control and reinforcement learning. Background Technology
[0002] Capsule endoscopy is a novel diagnostic device for digestive tract diseases. Patients insert the capsule orally, and as it moves through the digestive tract, it captures images and wirelessly transmits them to an external receiving device for diagnosis. Compared to traditional invasive endoscopes, capsule endoscopy offers advantages such as being painless, non-invasive, and eliminating the risk of cross-infection. It has been widely used in the examination of small bowel diseases and is gradually being expanded to the clinical diagnosis of other parts of the digestive tract, including the stomach, esophagus, and colon.
[0003] Early capsule endoscopes relied primarily on the peristalsis of the digestive tract for passive movement. This method suffered from problems such as uncontrollable movement, long examination times, and the potential to miss lesions. To address these issues, actively controlled capsule endoscopes have become a research hotspot. Currently, magnetically controlled actuation is the mainstream technology for achieving active movement of capsule endoscopes. Its basic principle involves embedding a permanent magnet inside the capsule. The interaction between the external magnetic field and the internal magnet generates magnetic force and torque, thereby driving the capsule to move controllably within the digestive tract.
[0004] Existing magnetic control methods mainly include two types: one is based on generating a controllable magnetic field using an array of electromagnetic coils, and driving the capsule by adjusting the magnetic field distribution through adjusting the current in each coil; the other uses a robotic arm device with a permanent magnet at its end, where the movement of the robotic arm drives the external permanent magnet to move, thereby controlling the position and orientation of the capsule. Compared with the electromagnetic coil solution, the robotic arm-based magnetic control system has advantages such as simple structure, high magnetic field strength, and lower cost, making it more suitable for clinical application.
[0005] However, existing robotic capsule endoscope magnetic control systems still have the following shortcomings: (1) Existing control methods mostly adopt traditional PID control or open-loop control strategies, which are difficult to cope with the complex and ever-changing working environment in the digestive tract. When the capsule is subjected to external disturbances such as intestinal peristalsis and changes in body position, the system lacks effective anti-disturbance capabilities, causing the movement trajectory to deviate from the expected target and affecting the examination results.
[0006] (2) Existing systems typically design control algorithms for a single task, making it difficult to achieve smooth switching and coordinated control of multiple motion modes under a unified control framework, which reduces inspection efficiency.
[0007] (3) Existing control methods suffer from a significant decrease in positioning accuracy when faced with system model uncertainties and external disturbances, making it difficult to meet the clinical needs of precise diagnosis. Summary of the Invention
[0008] In order to solve all or part of the above-mentioned problems in the prior art, the purpose of this disclosure is to propose a magnetic motion control method for capsule endoscopes. This method improves the trajectory tracking accuracy and posture stability during the navigation and scanning process of capsule endoscopes and can be applied to different medical examination scenarios.
[0009] A control method for a magnetically controlled capsule endoscope involves applying magnetic force and magnetic torque to the capsule endoscope via a permanent magnet assembly on the end effector of a robotic arm, thereby driving the capsule endoscope to move. The actuator control input consists of a Tube MPC control input and a compensation control quantity, which is generated using a trained reinforcement learning agent.
[0010] The beneficial technical effects of this solution are as follows: By introducing reinforcement learning into the Tube model predictive control framework, the system's anti-disturbance capability, robustness, and control accuracy are significantly improved due to the ability to perform real-time pose feedback, online correction, and continuous optimization through reinforcement learning. This enables precise positioning, stable motion, and efficient switching between multiple scenarios for the capsule endoscope. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a schematic diagram of a robot and capsule endoscope in one embodiment.
[0013] Figure 2 This is a schematic diagram of the Tube MPC control flow framework in one implementation.
[0014] Figure 3 This is a schematic diagram of the framework combining Tube MPC with reinforcement learning algorithms. Detailed Implementation
[0015] In existing magnetic control systems for capsule endoscopes driven by permanent magnets at the end effector of robotic arms, the control methods mostly employ traditional PID or open-loop / empirical parameter tuning strategies. These strategies struggle to cope with disturbances caused by peristalsis and changes in body position within the digestive tract, as well as uncertainties in the system model. This can easily lead to deviations in the capsule's trajectory and instability, thereby affecting navigation efficiency and the quality of lesion observation. Furthermore, existing control strategies are often designed for single tasks, making it difficult to achieve smooth switching and coordinated control between different medical scenarios, such as large-scale rapid navigation and localized fine scanning / fixed-point observation. In addition, traditional PID does not adequately consider constraints such as the joint angle and speed limits of the robotic arm and the boundaries of the workspace, which may result in control commands becoming unexecutable or exceeding limits, affecting system safety and clinical usability.
[0016] To overcome the aforementioned shortcomings, this invention aims to provide a magnetically controlled motion control method for a capsule endoscope driven by a single permanent magnet and a robotic arm. This method utilizes Tube model predictive control (TBMC) with robust control and predictive optimization to suppress disturbances and model uncertainties, thereby improving trajectory tracking accuracy and attitude stability during capsule navigation and scanning. Furthermore, reinforcement learning is introduced to superimpose a compensating control quantity generated by a reinforcement learning agent onto the control input of the TMC. This compensating agent is used to learn and compensate for nonlinear characteristics, unknown disturbances, and residual model errors that are difficult to model precisely in the system, thus further enhancing control accuracy and the system's adaptability. This method supports multiple operational scenarios, including navigation, scanning, and lesion localization, within a unified control framework, enabling rapid and smooth switching between different task modes. During the control solution process, constraints such as robotic arm motion and workspace are explicitly addressed to ensure the feasibility and safety of control commands.
[0017] The following description, in conjunction with the accompanying drawings, clearly and completely describes how the technical solution of this case is implemented. Obviously, the described embodiments are only a part of the embodiments of this case, and not all of them. Based on the embodiments in this case, all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of this application.
[0018] (a) Robot System The robotic system comprises the following components: small permanent magnets, medium permanent magnets, a robotic arm, an end effector gripper, a capsule shell, and a camera. The small and medium permanent magnets can be purchased custom-made. The capsule shell and end effector are manufactured using 3D printing with a precision of 10μm (= 0.01mm).
[0019] The movement and rotation of the capsule endoscope are controlled by an N-degree-of-freedom (N≥6) robotic arm holding an external permanent magnet. The robotic arm and permanent magnet are commercial equipment and will not be described in detail here.
[0020] See Figure 1 As shown, the experimental setup includes a robotic arm, an end-effector assembly (containing a medium-sized permanent magnet) and a capsule endoscope (containing a small permanent magnet); the robotic arm is used to carry the end-effector assembly to apply magnetic force and magnetic torque to the capsule endoscope.
[0021] This specification adopts the following notation conventions: scalars are represented by lowercase upright letters (e.g., c), vectors are represented by lowercase bold letters (e.g., x), matrices are represented by uppercase upright letters (e.g., M), and unit vectors are represented by the superscript "cap" symbol (e.g., ...). The nominal vector is indicated by a superscript horizontal line (e.g.) ).
[0022] The formula for calculating the magnetic field is: ,in B is the magnetic field. It is the permeability of free space. It is a 3×3 identity matrix. It is the magnetic moment of the permanent magnet, which is determined by the material of the permanent magnet installed on the end effector of the robotic arm.
[0023] The magnetic force of the permanent magnet on the end effector of the robotic arm on the capsule endoscope: Torque: ,in , It is the magnetic moment of the permanent magnet at the end of the robotic arm. It is the magnetic moment of the capsule endoscope. It is the bifurcation product.
[0024] In MPC, magnetic-related modeling is used for: dynamic prediction (calculating magnetic force / torque from the magnetic field and advancing the state), constraint evaluation (state / input constraints), and cost function calculation; real-time state information is used to update initial values online and solve the closed loop.
[0025] (ii) Tube Model Predictive Control The Tube Model Predictive Control (Tube MPC) method is used to control the motion of a capsule endoscope. First, a discrete state-space model of the system is established. .in, The state information at time k includes the pose and velocity information of the capsule endoscope and the center pose and velocity information of the permanent magnet at the end of the robotic arm. This represents the control input at time k (the increment of the position and orientation of the permanent magnet at the end of the robotic arm). This represents the external disturbance at time k. , Let A be a bounded set of disturbances, A be the state transition matrix, and B be the control input matrix.
[0026] Within each control cycle, the traditional MPC controller minimizes the cost function. The rolling optimization solution for the objective must satisfy the tightened state constraints during the optimization process. , Where Q, R, and P are positive semi-definite weighted matrices. The reference trajectory is N, and the prediction step size represents the number of forward prediction steps in one optimization.
[0027] Although standard model predictive control (MPC) provides an effective control framework, it does not explicitly consider disturbances, which may lead to violations of constraints in practical applications.
[0028] See the Tube MPC control flow. Figure 2 Tube MPC overcomes this limitation by ensuring robust constraint satisfaction under bounded perturbations. Its core idea is to decompose the control input into two parts. ,in It is the nominal control input, obtained through MPC online rolling optimization; This is the nominal state; the feedback gain matrix K is obtained by solving the linear quadratic optimal control (LQR) problem. Where P is a positive definite solution to the discrete algebraic Riccati equation, and the discrete algebraic Riccati equation is: Where A and B are the system state transition matrix and control input matrix, respectively, and Q and R are the semi-definite state weight matrix and positive definite control weight matrix, respectively. The deviation between the actual state and the nominal state is defined as... This deviation is calculated according to the dynamic equation. By evolving and rationally designing the Q and R weight matrices, the feedback gain matrix K is obtained, which makes the closed-loop system matrix (A + BK) satisfy the stability condition. This ensures that the deviation e(k) is always constrained within the robust positive invariant set Z. This invariant set Z constitutes a "bind" around the nominal trajectory, which is used to characterize the maximum deviation range of the actual state from the nominal state under bounded disturbances. This ensures that the system can maintain robust stability even in the presence of external disturbances and model uncertainties.
[0029] Compared to conventional proportional-integral-derivative (PID) control, open-loop control, or general model predictive control methods commonly used in existing technologies, this invention, based on the Tube Model Predictive Control (TubeMPC) framework, can explicitly introduce external disturbances (such as changes in friction or water flow) and model uncertainty constraints (such as inaccurate state information or discrepancies between the kinematic model and the real world). By constructing a robust tube bundle Z, this invention can effectively suppress the influence of external disturbances such as gastrointestinal peristalsis and changes in body position on capsule movement, thereby maintaining the stable movement and posture of the capsule in complex and dynamic environments and significantly improving the robustness of the system.
[0030] (III) Tube MPC based on reinforcement learning Although the Tube model predictive control method described above can effectively handle external disturbances and model uncertainties, in practical applications, the system may still have nonlinear characteristics that are difficult to model accurately, unknown time-varying disturbances, and residual model errors. These factors can affect the further improvement of control accuracy.
[0031] To address this issue, this invention further introduces reinforcement learning methods into the Tube MPC control framework, constructing a hybrid control architecture of "Tube MPC + Reinforcement Learning," see [link to relevant documentation]. Figure 3 .
[0032] In this hybrid control architecture, the actual control input consists of three parts: ,in To enhance the control compensation generated by the reinforcement learning agent, the agent continuously learns through interaction with the environment and generates compensation control variables online. It is used to compensate for nonlinear characteristics and unknown disturbances in the system that are difficult to describe by traditional modeling methods, thereby further improving trajectory tracking accuracy and attitude control stability.
[0033] The reinforcement learning agent is trained using the deep reinforcement learning algorithm PPO, with the goal of maximizing the cumulative reward function. The policy parameters are updated by maximizing the pruning substitution objective function. , , in , For the current policy in state Choose below The probability, The probability before the update. These are pruning parameters used to limit the magnitude of policy updates to improve training stability. This is the estimated value of the dominance function. , To assess the state value of the network output, system state information is combined with Tube MPC control input. The scalar value is obtained by performing a forward computation on the value assessment network. The final state space... It consists of four parts: the current position, attitude, and velocity of the capsule endoscope; the tracking error between the capsule endoscope's position and the reference trajectory; the joint angles and velocities of the robotic arm, and the pose and velocity information of the robotic arm's end effector; and the Tube MPC control input. The motion space is the compensation control quantity. The range of values for the reward function is defined as follows: the reward function is designed to be negatively correlated with the trajectory tracking error and the attitude deviation. That is, the smaller the tracking error and the smaller the attitude deviation, the higher the reward obtained by the agent, thereby guiding the agent to learn to generate a compensating control quantity that can reduce the tracking error.
[0034] The training process of reinforcement learning agents is divided into two stages: first, offline pre-training is carried out on a simulation platform, using a large amount of interactive data in the simulation environment to enable the agent to learn effective compensation strategies; then, the pre-trained agent is deployed to a real system for online fine-tuning, so that it can adapt to actual disturbances and model errors in the real environment, and further improve control performance.
[0035] By introducing reinforcement learning, the hybrid control architecture proposed in this invention can fully leverage the advantages of Tube MPC in constraint handling and robustness, while utilizing the adaptive learning capability of reinforcement learning to compensate for unmodeled dynamics and unknown disturbances in the system, thereby achieving higher precision and stronger robust motion control of the capsule endoscope.
[0036] In summary, this invention further introduces reinforcement learning into Tube MPC. By using a reinforcement learning agent to learn online and generate compensating control variables, it can adaptively compensate for residual model errors and unknown disturbances, thereby achieving higher trajectory tracking accuracy and attitude stability than simply using Tube MPC or traditional control methods. Furthermore, while the control parameters of traditional control methods are typically fixed after being determined in the design phase and are difficult to adapt to environmental changes during system operation, this invention, through a reinforcement learning agent, can continuously optimize the compensation strategy through continuous interaction with the environment, enabling the system to possess online adaptive learning capabilities and better cope with the complex and ever-changing environments in practical applications.
[0037] Traditional PID control cannot explicitly handle system constraints, and while standard MPC can handle constraints, it may lead to constraint violations under disturbances. This invention employs the Tube MPC framework to tighten state constraints and control input constraints during the optimization process, ensuring that the actual state and control input still satisfy the original constraints under worst-case disturbances, thus guaranteeing the feasibility of control commands and the safety of system operation. The control algorithm is first fully modeled, verified, and debugged on a simulation platform. By testing and optimizing the algorithm in the simulation environment, it is ensured that the algorithm's performance and stability meet the expected requirements before deployment to the real system, thereby effectively avoiding potential safety risks in practical applications. This "simulation-first, experimental verification" development process not only shortens the system development cycle but also provides an efficient and low-cost testing environment for subsequent optimization of the control algorithm, further ensuring the safety and reliability of the system in clinical applications.
[0038] The various embodiments of the present invention have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein. The scope of the invention is defined by the appended claims.
Claims
1. A method for controlling a magnetically controlled capsule endoscope, wherein a permanent magnet assembly on the end effector of a robotic arm applies magnetic force and magnetic torque to the capsule endoscope, thereby driving the capsule endoscope to move, characterized in that: The actuator control input consists of the Tube MPC control input and the compensation control quantity, which is generated using a trained reinforcement learning agent.
2. The method according to claim 1, characterized in that, Tube MPC control input ,in It is the nominal condition. The nominal control input and feedback gain matrix K are obtained by solving the linear quadratic optimal control (LQR) problem; The state information at time k includes the pose and velocity information of the capsule endoscope and the center pose and velocity information of the permanent magnet at the end of the robotic arm.
3. The method according to claim 1, characterized in that, During training, the observation space of a reinforcement learning agent consists of four parts: the current position, orientation, and velocity of the capsule endoscope; the error between the position of the capsule endoscope and the trajectory point; the joint angles and velocities of the robotic arm, the pose and velocity information of the robotic arm's end effector; and the Tube MPC control input.
4. The method according to claim 1, characterized in that, The reward function of a reinforcement learning agent during training is negatively correlated with trajectory tracking error and attitude deviation. That is, the smaller the tracking error and the smaller the attitude deviation, the higher the reward the agent receives, thereby guiding the agent to learn to generate compensating control variables that can reduce tracking error.
5. The method according to claim 1, characterized in that: The reinforcement learning agent is trained using the deep reinforcement learning algorithm PPO.
6. The method according to claim 1, characterized in that: The reinforcement learning agent is first pre-trained offline on a simulation platform, and then the pre-trained agent is deployed to a real system for online fine-tuning.
7. A control system for a magnetically controlled capsule endoscope, wherein a permanent magnet assembly on the end effector of a robotic arm applies magnetic force and magnetic torque to the capsule endoscope, driving the capsule endoscope to move, characterized in that: The actuator control input consists of the Tube MPC control input and the compensation control quantity, which is generated using a trained reinforcement learning agent.