Eye surgery robot autonomous intraocular foreign matter removal system and method based on dynamic calibration and imitation learning
Through dynamic calibration and imitation learning, the problem of ophthalmic surgical robot relies on artificial remote operation is solved, and high-precision autonomous intraocular foreign body removal in the scenarios of motion scaling variability and RCM offset is achieved, ensuring the stability and generalization of the operation.
Patent Information
- Application Number
- CN202510463601.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-08-15
AI Technical Summary
Existing ophthalmic surgical robots rely on manual remote operation, have a steep learning curve, low operation efficiency, and face kinematic uncertainty caused by variability in motion scaling and shift in remote motor centers, which affects the accuracy and generalization ability of the surgery.
The ophthalmic robot system based on dynamic calibration and imitation learning is adopted. By constructing a rotation matrix alignment local coordinate system to the global coordinate system, combining the ACT framework of RCM dynamic calibration and imitation learning, the encoder and decoder are optimized, the observation data are calibrated in real time and the future action sequence is predicted, and foreign matter removal is achieved intraocular.
Adaptive operation in complex intraocular environments is achieved, which eliminates spatial uncertainty of the base points of motion, and ensures submillimeter-level positioning accuracy, stability and generalization ability of surgical operations.
Smart Images

Figure CN120478040A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of ophthalmic surgical robots, and in particular relates to an ophthalmic surgical robot autonomous intraocular foreign body removal system and method based on dynamic calibration and imitation learning. Background Art
[0002] Intraocular foreign body removal surgery requires submillimeter precision to safely remove debris close to delicate retinal tissue while minimizing iatrogenic damage. Robotic-assisted technology shows great potential in enhancing control stability and reducing physiological tremor. However, existing systems still rely on manual teleoperation, which has a steep learning curve and low operational efficiency. Although automated removal can achieve continuous high-precision operation, key challenges include accurately identifying the optimal grasping point (determination of the starting posture of the movement) and real-time continuous assessment of the adhesion status of the debris (movement state assessment and adaptive path planning). These obstacles have hindered the clinical application of automated solutions to date.
[0003] Recent breakthroughs in large-scale imitation learning have significantly advanced the automation of robotic manipulation, such as in the paper DeRi-IGP: Learning to Manipulate Rigid Objects Using Deformable Linear Objects via Iterative Grasp-Pull. For example, imitation learning has been applied to the Franka robotic arm in the nail transfer task of laparoscopic surgery training, such as in the paper Robotic Constrained Imitation Learning for the Peg Transfer Task in Fundamentals of Laparoscopic Surgery. It has also achieved success in orthopedic surgical path planning, such as in the paper Motion Planning and Control of Active Robot in Orthopedic Surgery by CDMP-based Imitation Learning and Constrained Optimization. These advances demonstrate that imitation learning can help robotic systems learn the operations of expert surgeons, thereby improving the accuracy of complex surgeries and lowering the learning threshold. However, the application of imitation learning to microscale ophthalmic robotic systems remains largely unexplored. These systems are typically equipped with micromanipulators and advanced intraocular imaging devices, generating multimodal datasets consisting of stereoscopic microscopic video streams and instrument kinematic data. This data can be used to train autonomous control strategies for ophthalmic microsurgery.
[0004] The potential of imitation learning in ophthalmic surgery was explored through a representative and complex intraocular foreign body removal procedure. However, applying imitation learning to ophthalmic surgical robots faces two unique challenges: Variable motion scaling: Intraocular foreign body removal systems allow the surgeon to adjust the control-to-motion ratio to switch between coarse and fine adjustment modes. This variability leads to kinematic uncertainty—the same control input will produce different instrument motions due to differences in the scaling factor. Directly training the imitation learning strategy using raw control signals that do not account for scaling variations will affect generalization capabilities in actual surgery. Remote center of motion (RCM) variation: During ophthalmic robotic surgery, instruments enter the eye through a fixed entry point (RCM point) in the sclera. Maintaining precise instrument motion around this pivot point is crucial to avoid tissue damage. However, slight positional deviations of the RCM point can disrupt the spatial coordinate system, leading to instrument positioning errors. Due to anatomical constraints, traditional robotic system recalibration methods are difficult to implement in this scenario. Summary of the Invention
[0005] In view of the above, the purpose of the present invention is to provide an autonomous intraocular foreign body removal system and method for an ophthalmic surgical robot based on dynamic calibration and imitation learning, so as to solve the problem that existing ophthalmic surgical robots rely on manual remote operation, resulting in a steep learning curve and low operating efficiency, realize autonomous intraocular foreign body removal surgery, and also solve the kinematic uncertainty problems caused by motion scaling variability and remote center of motion (RCM) offset in ophthalmic robot surgery, thereby ensuring the generalization ability and positioning accuracy of surgical operations.
[0006] To achieve the above-mentioned purpose of the invention, an embodiment provides an autonomous intraocular foreign body removal system for an ophthalmic surgical robot based on dynamic calibration and imitation learning, characterized by comprising:
[0007] A rotation matrix construction module is used to construct a three-dimensional eye model for a real eyeball and construct a rotation matrix that aligns the local coordinate system to the global coordinate system based on preset fixed points on the three-dimensional eye model;
[0008] An imitation learning module, which combines the rotation matrix-based RCM dynamic calibration with the ACT framework of imitation learning. This allows for compensation for changes in the remote center of motion during task execution, and optimizes the encoder and decoder in the ACT framework through imitation learning.
[0009] The foreign body removal module is used to dynamically calibrate the observation data of the ophthalmic surgical robot in real time through a rotation matrix, use the encoder and decoder after imitation learning to predict future action sequences based on the dynamically calibrated observation data, and control the ophthalmic surgical robot to execute the predicted future actions to achieve intraocular foreign body removal.
[0010] Preferably, constructing a rotation matrix for aligning the local coordinate system to the global coordinate system based on preset fixed points on the three-dimensional eye model includes:
[0011] Obtaining the position coordinates of three preset fixed points on the three-dimensional eye model reached by the end effector of the ophthalmic surgical robot in the current local coordinate system and the position coordinates in the global coordinate system, where the three points are not coplanar;
[0012] By comparing the relative orientation of each fixed point in two different coordinate systems, the rotation matrix can be solved.
[0013] Preferably, in the imitation learning module, based on compensating for the change of the remote motion center during task execution, imitation learning is performed to optimize the encoder and decoder in the ACT framework, including:
[0014] Collect current observation data t , which includes the camera image of the eyeball, the position and direction of the end effector of the ophthalmic surgical robot, and the current posture of the robot arm. Other data except the camera image constitute the current observation data And the observation data is dynamically calibrated using the rotation matrix;
[0015] The encoder in the ACT framework is used to t:t+k and observation data after dynamic calibration Learn a latent vector z using the decoder based on the current observation data o after dynamic calibration t , latent variable z and global coordinate system Predicting future action sequences Where t represents the current moment and k represents the length of the future time step;
[0016] Build based on the current action sequence a t:t+k and future action sequences The reconstruction loss of , the regularization loss of the latent vector z, and the imitation learning of the ACT framework are used to optimize the encoder and decoder in the ACT framework.
[0017] Preferably, the reconstruction loss adopts MSE loss, which is used to measure the current action sequence a t:t+k and future action sequences The error between .
[0018] Preferably, the regularization loss is expressed as:
[0019]
[0020] in, Indicates that the encoder is based on the current action sequence a t:t+k and the observation sequence after dynamic calibration The probability distribution of the output potential vector z, φ represents the parameters of the encoder, represents the standard Gaussian distribution, D KL (·∥·) represents the KL divergence between the two distributions, through L reg Make sure the distribution of the latent variable z follows a Gaussian distribution.
[0021] Preferably, in the foreign body removal module, the observation data of the ophthalmic surgical robot is dynamically calibrated using a rotation matrix, and a decoder after imitation learning is used to predict future action sequences based on the dynamically calibrated observation data, including:
[0022] By rotating the matrix and the current observation data o t Multiply to transform the observation data into the global coordinate system to achieve dynamic calibration;
[0023] Dynamically calibrated observation data Together with the current action sequence a t:t+k The encoder predicts the potential vector z, and then uses the decoder based on the current observation data o after dynamic calibration t , latent variable z and global coordinate system Predicting future action sequences
[0024] Preferably, the foreign matter removal module also stores the future action sequence into a FIFO buffer for future prediction optimization.
[0025] Preferably, the foreign body removal module also performs a time-weighted average on the stored future action sequences. Specifically, at each time step, a weighted average calculation is performed based on the action sequences stored in the buffer, where the weight of the earliest action is w i According to exponential decay, the most recent action has a greater impact on the current decision, and the final execution action is obtained.
[0026] To achieve the above-mentioned object of the invention, an embodiment of the present invention further provides a method for autonomously removing intraocular foreign bodies using an ophthalmic surgical robot based on dynamic calibration and imitation learning. The method utilizes the above-mentioned system and includes the following steps:
[0027] A rotation matrix construction module is used to construct a 3D eye model for a real eyeball, and a rotation matrix is constructed based on preset fixed points on the 3D eye model to align the local coordinate system to the global coordinate system.
[0028] The imitation learning module combines the rotation matrix-based RCM dynamic calibration with the ACT framework of imitation learning. This allows for compensation of changes in the remote motion center during task execution, and then optimizes the encoder and decoder in the ACT framework through imitation learning.
[0029] The foreign body removal module is used to dynamically calibrate the observation data of the ophthalmic surgical robot in real time through the rotation matrix. The encoder and decoder after imitation learning are used to predict the future action sequence based on the dynamically calibrated observation data, and the ophthalmic surgical robot is controlled to execute the predicted future actions to achieve intraocular foreign body removal.
[0030] To achieve the above-mentioned purpose of the invention, an embodiment also provides a computing device, including a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the above-mentioned method for autonomous intraocular foreign body removal by an ophthalmic surgical robot based on dynamic calibration and imitation learning.
[0031] Compared with the prior art, the present invention has the following beneficial effects:
[0032] The system of the present invention is based on a rotation matrix construction module, an imitation learning module, and a foreign body removal module. By constructing a multimodal data-driven dynamic calibration mechanism, it realizes the adaptive operation of the ophthalmic surgical robot in the complex intraocular environment. By using stereoscopic vision images and instrument kinematic data, combined with a dynamic RCM coordinate alignment strategy, the local coordinate system drift caused by anatomical structure differences is uniformly mapped to the global reference system, thereby eliminating the spatial uncertainty of the motion base point. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0034] Figure 1 1 is a schematic structural diagram of an autonomous intraocular foreign body removal system of an ophthalmic surgical robot based on dynamic calibration and imitation learning provided in an embodiment;
[0035] Figure 2 This is the overall framework diagram of RCM-ACT provided by the embodiment;
[0036] Figure 3 1 is a schematic diagram of the deployment process of the ophthalmic surgical robot provided in the embodiment. DETAILED DESCRIPTION
[0037] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not limit the scope of protection of the present invention.
[0038] The inventive concept of the present invention is as follows: The embodiment of the present invention provides an autonomous intraocular foreign body removal system for an ophthalmic surgical robot based on dynamic calibration and imitation learning, which includes an imitation learning framework for autonomous intraocular foreign body removal surgery, and trains the control strategy using only instrument kinematics and stereo vision data through expert demonstration of foreign body grasping and placement on a biomimetic eye model. In view of the inherent motion scaling variability of microsurgery, the strategy of directly training the instrument actuator-level displacement (rather than the original control signal) is used to avoid kinematic ambiguity, ensure the generalization ability of the strategy at different magnifications, and maintain scalability to multiple platforms. In response to the challenge of RCM variation, a dynamic coordinate alignment strategy is developed: real-time instrument position recalibration based on three anatomical landmarks, and the variation-contaminated kinematic data is converted into a unified global coordinate system through iterative rotation matrix updates. This software-defined spatial anchoring method can alleviate the cumulative error caused by RCM deviation without hardware recalibration and maintain sub-millimeter positioning accuracy. The study achieved submillimeter accuracy in instrument-foreign body interaction (average grasping positioning error of 0.12 mm) without explicit depth perception, confirming the key feasibility of automated intraocular foreign body removal surgery under real-world kinematic uncertainty.
[0039] Based on the above invention concept, Figure 1 As shown, the embodiment provides an ophthalmic surgical robot autonomous intraocular foreign body removal system 10 based on dynamic calibration and imitation learning, including: a rotation matrix construction module 11, an imitation learning module 12, and a foreign body removal module 13.
[0040] In this embodiment, the rotation matrix construction module 11 is used to construct a three-dimensional eye model for a real eyeball, and to construct a rotation matrix that aligns the local coordinate system to the global coordinate system based on preset fixed points on the three-dimensional eye model.
[0041] Specifically, the 3D eye model is manipulated by the ophthalmic surgical robot arm, which contains the black ring R black and a larger orange ring R orange , the foreign body removal task requires the robot to automatically grab the black ring R black And place it with high precision on the orange ring R orange At each time step t, the robot receives observation data o t , where the camera image of the eyeball model M shows a black ring R black and Orange Ring R orange The positions pRblack[t] and pRorange[t], and the pose of the robot end effector pt=(x t ,y t ,z t ) and direction rt. In addition, the robot also receives proprioceptive information X representing the current posture of the robotic arm t, also as observation data o t Based on these observation data, the ophthalmic surgical robot calculates and executes action a t , to move the end effector to grab the black ring R black Place it on the orange ring R orange superior.
[0042] During each task demonstration, the position of the insertion point in the eye will shift during data collection, resulting in a positional deviation of the remote center of motion (RCM). This causes the RCM point to change during different data collection processes, resulting in inconsistent coordinate systems. In order to solve this problem, an embodiment of the present invention proposes a dynamic calibration method for RCM. The dynamic calibration method is implemented by allowing the robotic instrument to reach three fixed points p1, p2, and p3 in the workspace. At each time step, by obtaining the positions of these three fixed points in the current coordinate system, the corresponding rotation matrix R can be calculated. t , used to align the local coordinate system to the global coordinate system. The rotation matrix R t This enables data collected in different coordinate systems to be converted into a unified reference frame. The process can be expressed as follows:
[0043] At each time step t, the robot end effector reaches the three preset fixed points p1(t), p2(t), and p3(t). These three fixed points are in the current local coordinate system. Defined in, and not coplanar, and also get three fixed points in the global coordinate system The position coordinates of each point p i The position of (t) is expressed as:
[0044] p i (t) = [x i (t),y i (t),z i (t)] T ,i=1,2,3
[0045] Rotation matrix R t The purpose is to calculate the local coordinate system through these three fixed points. Align to global coordinate system Specifically, the rotation matrix R can be solved by comparing the relative directions of the fixed points in different coordinate systems. t :
[0046] p i (t) = R t ·p i (0),i=1,2,3
[0047] Among them, i is the fixed point index, p i (0) represents the fixed point pi In the global coordinate system The position coordinates in .
[0048] The calculated rotation matrix R t It is used to dynamically calibrate the collected sensing data. Specifically, the rotation matrix R t Applies to all data collected at time step t. Specifically, for the observation data o collected at time step t t , which can be obtained by applying the rotation matrix R t The inverse transformation recalibrates the data to the global coordinate system
[0049] p i (t)'=R t ·p i (t)
[0050] By using R t Recalibration is performed for each new observation point to ensure that all data are transformed into a consistent global coordinate system, thereby improving the reliability of task learning and enabling accurate performance evaluation across different demonstrations.
[0051] In an embodiment, the imitation learning module 12 is used to combine the RCM dynamic calibration based on the rotation matrix with the ACT framework of imitation learning, that is, to obtain the RCM-ACT model, so as to realize the optimization of the encoder and decoder in the ACT framework by imitation learning based on the compensation for the change of the remote motion center during the task execution.
[0052] To train the RCM-ACT model for a new task, data transmission and dataset recording are first performed to collect human demonstration data. The system input consists of binocular images from a stereo microscope, providing real-time visual feedback from two camera perspectives. In addition to the visual binocular images, the input also includes local perception data, representing the Cartesian coordinates of the current pose and the rotational encoder values of the gripper (from which the rotation direction can be determined). These proprioceptive and visual observation data are crucial for the task because they provide real-time perception information about the robot's state and environment.
[0053] After data collection is complete, the RCM-ACT model is trained to predict future actions based on the current observations. These predicted actions correspond to the target joint positions of the robot arms at the next time step. The ACT model aims to predict the next action sequence based on the current state, mimicking the actions of a human operator under the same observation conditions. Ultimately, these target joint positions are used to guide the robot arm to complete the desired operation.
[0054] The RCM-ACT model combines the traditional ACT framework with RCM dynamic calibration to ensure compensation for the change of motion base points during task execution. Specifically, the rotation matrix is used to dynamically calibrate the observation data to achieve compensation, and then the encoder in the ACT framework is used to calculate the current action sequence a. t:t+k and observation data after dynamic calibration ((in (does not contain binocular images) to learn a latent vector z, i.e. The decoder is used to dynamically calibrate the current observation data o t , latent variable z and global coordinate system Predicting future action sequences Right now Where t represents the current moment and k represents the length of the future time step;
[0055] Build based on the current action sequence a t:t+k and future action sequences The reconstruction loss of the latent vector z and the regularization loss of the latent vector z are used to optimize the encoder and decoder in the ACT framework by imitating the ACT framework. reconst MSE loss is used to measure the current action sequence a t:t+k and future action sequences The error between them is expressed as:
[0056]
[0057] The regularization loss is expressed as:
[0058]
[0059] in, Indicates that the encoder is based on the current action sequence a t:t+k and the observation sequence after dynamic calibration The probability distribution of the output potential vector z, φ represents the parameters of the encoder, represents the standard Gaussian distribution, D KL (·∥·) represents the KL divergence between the two distributions, through L reg Ensure that the distribution of the latent variable z follows a Gaussian distribution, thereby improving the generalization ability of the model.
[0060] During training, RCM dynamic calibration is used to recalibrate the robot's position and orientation in real time, ensuring that the collected data is always aligned with the global coordinate system. This calibration step is crucial because it compensates for changes in the motion base point and ensures the model's reliability during task execution. Once training is complete, the RCM-ACT model predicts future actions based on the current state, and the robot executes these predicted actions to complete the task.
[0061] In an embodiment, the foreign body removal module 13 is used to dynamically calibrate the observation data of the ophthalmic surgical robot in real time through a rotation matrix, use the encoder and decoder after imitation learning to predict future action sequences based on the observation data after dynamic calibration, and control the ophthalmic surgical robot to execute the predicted future actions to achieve the removal of intraocular foreign bodies.
[0062] First, the robot’s position p is updated based on the calibration data provided by the RCM sensor. t =(x t ,y t ,z t ) and towards R t This ensures that the robot operates in a precisely calibrated coordinate system, thus maintaining high accuracy during task execution. Then, using the rotation matrix R t Aligning data to the global coordinate system Ensure that all subsequent data are accurately positioned in the global reference frame.
[0063] When the calibration is completed, action prediction is performed and the observation data of dynamic calibration is obtained. Together with the current action sequence a t:t+k The encoder predicts the potential vector z, and then uses the decoder based on the current observation data o after dynamic calibration t , latent variable z and global coordinate system Predicting future action sequences The predicted actions are stored in a FIFO (first-in-first-out) buffer for future prediction optimization.
[0064] The foreign body removal module 13 also performs a time-weighted average on the stored future action sequences. Specifically, at each time step, a weighted average calculation is performed based on the action sequences stored in the buffer, where the weight of the earlier action is w i According to exponential decay, the most recent action has a greater impact on the current decision, and the final execution action is obtained.
[0065] During the inference process, data collected by the robot, such as proprioception data (pt, rt, Xt) and actions, is normalized using predefined statistics to ensure that the input data fits within the model's expected range, improving prediction accuracy. After the model predicts the action, it is denormalized back to its original scale to ensure that the robot's actions accurately correspond to the real-world environment. Finally, the calculated action is transmitted to the robot, enabling it to perform the task of picking up the black ring and placing it on the orange ring. This entire process combines real-time calibration with action prediction to ensure the robot can complete the task accurately and efficiently.
[0066] The system provided by this embodiment combines RCM dynamic calibration with the ACT model, which is key to achieving smooth, precise, and reliable robotic motion. This fusion solution ensures that the robotic system can perform tasks stably even when the motion base point may change or drift over time, adapting to complex surgical environments.
[0067] The embodiment also provides a specific RCM-ACT overall framework diagram, such as Figure 2 As shown in the figure, an automatic execution system based on an eye surgery robot is demonstrated, which combines a binocular microscope and a Transformer model for task execution and optimization. First, the system collects binocular microscope images to obtain real-time visual information from the surgical perspective, and collects 5-degree-of-freedom actuator data sequences as input. During the training process, the surgical demonstration data and corresponding labels are combined, and RCM calibration is used to ensure that the data is processed in a consistent coordinate system. Subsequently, the data is pre-processed by normalization and converted into a model-usable format through an image embedding module. The entire task is completed by the Transformer encoder and Transformer decoder in the ACT framework. The encoder learns data features, and the decoder outputs the actuator action sequence, which ultimately guides the actuator to control the surgical tool to achieve precise and stable operation.
[0068] like Figure 3The embodiment also provides a schematic diagram of the deployment process of an ophthalmic surgical robot, which includes the aforementioned autonomous intraocular foreign body removal system based on dynamic calibration and imitation learning. A specific dataset serves as the input source, providing surgical demonstration data and microscope images for training the RCM-ACT model. The trained model is deployed to a computer (PC), which accurately calculates the robot's operation trajectory by analyzing the microscope image input in real time and combining it with the RCM dynamic calibration mechanism. The PC continuously obtains the robot's real-time status through the "Robot_Status" topic of ROS communication. The model predicts the next action command based on the result, and then transmits the control command to the execution end via the "Model_Output" topic. After the robot completes the operation according to the command, a closed-loop control loop is formed through a state feedback mechanism to ensure the continuity and stability of task execution. With the PC as the core control unit, the system integrates high-precision visual input and real-time status data through the online reasoning capabilities of the RCM-ACT model. It utilizes the ROS communication architecture to achieve millimeter-level precision in the ophthalmic surgical robot, maintaining operational consistency and system reliability in a dynamic surgical environment.
[0069] The embodiment also provides a performance comparison between the deployment test results and the model training evaluation test, as shown in Table 1. The results in Table 1 reveal significant performance differences between the various methods in terms of deployment capabilities and evaluation indicators, highlighting the key role of the dynamic RCM calibration and resampling mechanism. In the deployment test, the baseline ACT method was partially successful, grasping the black ring twice and placing the orange ring three times in five trials, but failed to complete the full task sequence due to insufficient coordination between subtasks. The RCM-ACT variant without the resampling mechanism showed slightly improved performance, achieving one successful grasp and four placements, but still failed to achieve full task completion. In contrast, the complete RCM-ACT model demonstrated strong reliability, successfully grasping the black ring three times, placing the orange ring five times, and completing the full task sequence three times in five trials. The RCM-ACT variant equipped with an encoder could not be effectively deployed because the noise introduced by the encoder caused excessive mechanical jitter, which made it impossible to perform actual execution operations.
[0070] Table 1 Performance comparison between deployment test results and model training evaluation test
[0071]
[0072] The specific implementation methods described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. An autonomous intraocular foreign body removal system for ophthalmic surgical robots based on dynamic calibration and imitation learning, characterized in that: include: A rotation matrix construction module is used to construct a three-dimensional eye model for a real eyeball and construct a rotation matrix that aligns the local coordinate system to the global coordinate system based on preset fixed points on the three-dimensional eye model; An imitation learning module, which combines the rotation matrix-based RCM dynamic calibration with the ACT framework of imitation learning. This allows for compensation for changes in the remote center of motion during task execution, and optimizes the encoder and decoder in the ACT framework through imitation learning. The foreign body removal module is used to dynamically calibrate the observation data of the ophthalmic surgical robot in real time through a rotation matrix, use the encoder and decoder after imitation learning to predict future action sequences based on the dynamically calibrated observation data, and control the ophthalmic surgical robot to execute the predicted future actions to achieve intraocular foreign body removal.
2. The autonomous intraocular foreign body removal system based on dynamic calibration and imitation learning for ophthalmic surgery robots according to claim 1 is characterized in that: A rotation matrix is constructed based on preset fixed points on the 3D eye model to align the local coordinate system to the global coordinate system, including: Obtaining the position coordinates of three preset fixed points on the three-dimensional eye model reached by the end effector of the ophthalmic surgical robot in the current local coordinate system and the position coordinates in the global coordinate system, where the three points are not coplanar; By comparing the relative orientation of each fixed point in two different coordinate systems, the rotation matrix can be solved.
3. The ophthalmic surgical robot autonomous intraocular foreign body removal system based on dynamic calibration and imitation learning according to claim 1 is characterized in that: In the imitation learning module, based on compensating for changes in the remote motion center during task execution, imitation learning is performed to optimize the encoder and decoder in the ACT framework, including: Collect current observation data t , which includes the camera image of the eyeball, the position and direction of the end effector of the ophthalmic surgical robot, and the current posture of the robot arm. Other data except the camera image constitute the current observation data And the observation data is dynamically calibrated using the rotation matrix; The encoder in the ACT framework is used to t:t+k and observation data after dynamic calibration Learn a latent vector z using the decoder based on the current observation data o after dynamic calibration t , latent variable z and global coordinate system Predicting future action sequences Where t represents the current moment and k represents the length of the future time step; Build based on the current action sequence a t:t+k and future action sequences The reconstruction loss of , the regularization loss of the latent vector z, and the imitation learning of the ACT framework are used to optimize the encoder and decoder in the ACT framework.
4. The autonomous intraocular foreign body removal system of an ophthalmic surgical robot based on dynamic calibration and imitation learning according to claim 3 is characterized in that: The reconstruction loss adopts MSE loss, which is used to measure the current action sequence a t:t+k and future action sequences The error between .
5. The autonomous intraocular foreign body removal system of an ophthalmic surgical robot based on dynamic calibration and imitation learning according to claim 3 is characterized in that: The regularization loss is expressed as: in, Indicates that the encoder is based on the current action sequence a t:t+k and the observation sequence after dynamic calibration The probability distribution of the output potential vector z, φ represents the parameters of the encoder, represents the standard Gaussian distribution, D KL (·∥·) represents the KL divergence between the two distributions, through L reg Make sure the distribution of the latent variable z follows a Gaussian distribution.
6. The autonomous intraocular foreign body removal system of an ophthalmic surgical robot based on dynamic calibration and imitation learning according to claim 3, characterized in that: In the foreign body removal module, the ophthalmic surgical robot's observation data is dynamically calibrated using a rotation matrix. The decoder after imitation learning is used to predict future action sequences based on the dynamically calibrated observation data, including: By rotating the matrix and the current observation data o t Multiply to transform the observation data into the global coordinate system to achieve dynamic calibration; Dynamically calibrated observation data Together with the current action sequence a t:t+k The encoder predicts the potential vector z, and then uses the decoder based on the current observation data o after dynamic calibration t , latent variable z and global coordinate system Predicting future action sequences 7. The autonomous intraocular foreign body removal system based on dynamic calibration and imitation learning for ophthalmic surgery according to claim 1, characterized in that: The foreign object removal module also stores the future action sequence into the FIFO buffer for future prediction optimization.
8. The ophthalmic surgical robot autonomous intraocular foreign body removal system based on dynamic calibration and imitation learning according to claim 7, characterized in that: The foreign body removal module also performs a time-weighted average on the stored future action sequences. Specifically, at each time step, a weighted average calculation is performed based on the action sequences stored in the buffer, where the weight of the earlier action is w i According to exponential decay, the most recent action has a greater impact on the current decision, and the final execution action is obtained.
9. A method for autonomous intraocular foreign body removal by an ophthalmic surgical robot based on dynamic calibration and imitation learning, characterized in that: The method utilizes the system according to any one of claims 1 to 8, comprising the following steps: A rotation matrix construction module is used to construct a 3D eye model for a real eyeball, and a rotation matrix is constructed based on preset fixed points on the 3D eye model to align the local coordinate system to the global coordinate system. The imitation learning module combines the rotation matrix-based RCM dynamic calibration with the ACT framework of imitation learning. This allows for compensation of changes in the remote motion center during task execution, and then optimizes the encoder and decoder in the ACT framework through imitation learning. The foreign body removal module is used to dynamically calibrate the observation data of the ophthalmic surgical robot in real time through the rotation matrix. The encoder and decoder after imitation learning are used to predict the future action sequence based on the dynamically calibrated observation data, and the ophthalmic surgical robot is controlled to execute the predicted future actions to achieve intraocular foreign body removal.
10. A computing device comprising a memory and one or more processors, wherein the memory stores executable code, characterized in that: When the one or more processors execute the executable code, they are used to implement the method for autonomous intraocular foreign body removal by an ophthalmic surgical robot based on dynamic calibration and imitation learning as described in claim 9.