A method of performing cataract surgery autonomously by a robot

CN122805377APending Publication Date: 2026-09-25HARBIN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611003424.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-07
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0007]本发明提供一种由机器人自主开展白内障手术的方法,解决现有白内障手术智能程度不足,术后效果参差不齐的问题

Benefits of technology

1.本发明构建了一套融合显微镜图像、力觉信息、机器人本体状态以及患者生理参数的多模态信息采集框架,实现了对白内障手术全过程的全面感知。相比传统仅依赖视觉信息的控制方式,本发明能够同时获取组织形态变化、器械接触状态以及患者实时状态等多维度信息,使机器人对手术环境具有更全面、更准确的认知能力,为后续自主决策提供可靠的数据基础。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122805377A_ABST
    Figure CN122805377A_ABST
Patent Text Reader

Abstract

The application discloses a method for performing cataract surgery autonomously by a robot, and belongs to the technical field of ophthalmic surgery robots. Step one: a multi-modal information acquisition platform for cataract surgery is established, visual information, force sensation information, robot state information and patient physiological information are collected in real time, and multi-source data synchronization is completed; step two: the original data collected in step one is preprocessed; step three: an expert demonstration database covering each surgical stage is established, and an expert knowledge graph is constructed; step four: the operation rules and decision logic of expert doctors are learned by using an imitation learning method, and an initial strategy for autonomous control of the robot is formed; step five: the initial strategy is combined with a digital twin cataract surgery environment, and the control strategy is further optimized through reinforcement learning; step six: different patient samples are verified, an artificial takeover mechanism is introduced to correct abnormal operations, and continuous optimization and iterative updating of the model are realized. The application solves the problems of insufficient intelligence of existing cataract surgery and uneven postoperative effects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of ophthalmic surgical robot technology, specifically relating to a method for performing cataract surgery autonomously by a robot. Background Technology

[0002] Cataracts are the leading cause of blindness worldwide, and minimally invasive treatments such as femtosecond LASIK and phacoemulsification are widely used both domestically and internationally. Doctors create an incision in the cornea to cut, pulverize, and remove the cloudy lens, finally placing an artificial lens as a replacement. As China gradually enters an aging society, the increasing number of patients is placing a greater strain on the healthcare system, and the demand for cataract surgery is shifting from simply restoring sight to achieving higher-quality, more precise visual reconstruction.

[0003] However, the anterior segment space is extremely small and the tissue is fragile. Traditional phacoemulsification surgery heavily relies on the surgeon's experience and adaptability. Furthermore, factors such as the surgeon's physiological tremors, fatigue, and misjudgment can lead to improper pressure on the eyeball, inability to continuously tear the sac to create a perfectly round incision, and incorrect intraocular lens placement, all of which affect postoperative outcomes. While femtosecond laser technology can precisely create an incision and shatter the diseased lens using a laser, it is expensive and still requires additional surgical intervention. The training curve for a skilled cataract surgeon is as long as three years, and the experience and skills of experts are difficult to quantify and fully pass on. In remote areas, the adoption rate of phacoemulsification technology is less than 30%.

[0004] Compared to most surgical procedures, cataract surgery is characterized by its fixed procedures, standardized operations, clearly defined evaluation indicators, and high repeatability. It mainly includes standardized steps such as corneal incision creation, continuous circular capsulorhexis, phacoemulsification, cortical aspiration, and intraocular lens implantation. Each stage has relatively clear operational goals and quantifiable evaluation indicators. For example, the capsulorhexis process can be evaluated using indicators such as capsulorhexis roundness and diameter error, while the phacoemulsification process can be evaluated using indicators such as emulsification efficiency, degree of tissue damage, and operation time. This highly standardized surgical procedure not only reduces uncertainties during the operation but also provides a good foundation for robots to learn the operational patterns of expert surgeons.

[0005] Meanwhile, a large number of cataract surgery cases continuously generate rich and structurally consistent expert demonstration data, enabling robots to establish stable decision-making models through artificial intelligence methods such as imitation learning, reinforcement learning, and knowledge graphs. Therefore, cataract surgery is considered one of the most suitable ophthalmic surgeries for achieving autonomous robotic operation, providing an important application scenario and technological foundation for research on autonomous surgical robots based on multimodal motion generation models.

[0006] With the development of multimodal large-scale model technology, autonomous technology for complex long-term surgical tasks based on imitation learning and deep learning, using robots as platforms, has become an important direction for the development of cataract surgery. Cataract surgery is characterized by fixed operational procedures, clearly defined key steps, quantifiable evaluation indicators, and abundant clinical data, providing a solid foundation for artificial intelligence models to learn from expert experience and construct autonomous decision-making strategies. Simultaneously, the development of multimodal perception technology enables robots to acquire visual, force, and environmental status information in real time, providing reliable support for perception, decision-making, and control during autonomous surgery. The aforementioned methods can quantify standardized procedures into a model knowledge graph before surgery, collect multimodal information for cataract surgery scenarios, and use it for real-time robot motion generation, ultimately significantly improving the automation level of cataract surgery. Summary of the Invention

[0007] This invention provides a method for performing cataract surgery autonomously by a robot, which solves the problems of insufficient intelligence and inconsistent postoperative results in existing cataract surgery.

[0008] This invention is achieved through the following technical solution: A method for performing cataract surgery autonomously by a robot, the method comprising the following steps: Step 1: Establish a multimodal information acquisition platform for cataract surgery. Utilize a surgical microscope, OCT imaging system, six-dimensional force sensor, and robot controller to collect visual information, force information, robot status information, and patient physiological information in real time, and complete the synchronization of multi-source data. Step 2: Preprocess the raw data obtained in Step 1; Step 3: An ophthalmologist performs cataract surgery using a remotely operated robot, simultaneously recording the robot's status and control actions, establishing an expert demonstration database covering all stages of the surgery, and constructing an expert knowledge graph. Step 4: Using the imitation learning method, train the action generation model based on the expert demonstration database established in Step 3, learn the operating rules and decision-making logic of expert doctors, and form the initial strategy for autonomous control of the robot. Step 5: Combine the initial strategy from Step 4 with the digital twin cataract surgery environment, and further optimize the control strategy through reinforcement learning; Step Six: Validate the model using different patient samples and introduce a manual intervention mechanism to correct any abnormal operations, thereby achieving continuous optimization and iterative updates of the model.

[0009] Furthermore, step one specifically involves setting a time... The collected visual information, OCT information, force information, robot state information, and patient physiological information are respectively represented as follows: Then construct the time Multimodal state vector:

[0010] By using a unified timestamp mechanism to synchronize various types of data, the correspondence between different modalities of data on the same time axis can be achieved.

[0011] Furthermore, step two specifically involves assuming the input image is... The convolution kernel is The bias term is The output feature map of the convolutional layer is then:

[0012] in This represents the output feature map of the convolutional layer. This represents the convolution operation. Indicates the bias term. Indicates the activation function; After processing by a multi-layer convolutional network, high-dimensional semantic features of the image are obtained:

[0013] in Indicates time Microscopic images collected, This represents the corresponding visual feature vector; For force signals, the Kalman filter algorithm is used to eliminate random noise; First, establish the system state equations:

[0014] in Indicates time The true force signal state, Represents the state transition matrix. Represent process noise; establish observation equations:

[0015] in This indicates the force sensor measurement value. Represents the observation matrix. Indicate the measurement noise; calculate the Kalman gain based on the Kalman filtering principle. And update the state estimate:

[0016] After filtering, a smooth force signal is obtained; For robot motion data, a sliding window algorithm is used to identify and automatically delete long pauses, false triggers, and invalid operation segments. Let the window length be The average speed within the window is The first one in the window The speed of each sampling point is The average speed of the end effector within the window is:

[0017] Let the time threshold be When satisfied Furthermore, if the duration exceeds a preset time threshold, the system determines that the data segment is an invalid pause segment and automatically removes it; For image rotation enhancement, let the original image coordinates be... The rotation angle is Then the coordinates after rotation are:

[0018] For brightness enhancement, a linear transformation is used: ; in Represents the original image. This represents the enhanced image. Indicates the brightness scaling factor. This indicates the brightness offset.

[0019] Furthermore, random noise is first generated:

[0020] in This indicates that the mean is 0 and the variance is 0. Gaussian distribution; Then the enhanced image is obtained:

[0021] To enhance the perturbation of the robot's motion trajectory, let the original trajectory point be... The random perturbation vector is The enhanced trajectory points are:

[0022] After the above processing, the optimized multimodal state vector is constructed. , This represents a data optimization function that consists of feature extraction, filtering and noise reduction, outlier removal, and data augmentation.

[0023] Furthermore, step three specifically involves the system recording the robot's state vector in real time during the surgery. And the corresponding control actions of expert doctors ,in:

[0024] Indicates that the doctor is at all times Position and attitude control quantities applied to the robot's end effector; The expert physician's procedure sequence was obtained through continuous recording:

[0025] Where N is the total number of sampling frames.

[0026] Furthermore, based on the cataract surgery process, the expert data was divided into the capsulorhexis stage, phacoemulsification stage, cortical aspiration stage, and intraocular lens implantation stage, and corresponding stage labels were established. An expert knowledge graph is constructed based on the aforementioned expert demonstration data. The nodes in the knowledge graph include surgical stage nodes, tissue state nodes, instrument state nodes, and operation action nodes. Edges represent the causal and temporal relationships between different nodes. An expert decision-making knowledge base is formed by statistically analyzing the state transition patterns in a large number of expert operations.

[0027] Furthermore, step four specifically involves analyzing the operational logic of expert doctors using imitation learning methods; establishing an action generation model based on a Transformer network, with the input being a multimodal state sequence of the current and historical moments:

[0028] The output is the predicted action for the next time step:

[0029] The model is trained by minimizing the error between the predicted action and the expert's actual action:

[0030] in For expert action, To predict actions for the model.

[0031] Furthermore, step five specifically involves establishing a surgical state space. Action space and the reward function:

[0032] in This indicates an organizational safety reward. This indicates a reward for the quality of the surgery. This indicates a reward for operational efficiency. These are the corresponding weighting coefficients; Negative rewards are given when capsule tearing, tissue damage, or contact force exceeding the safety threshold occurs; positive rewards are given when the roundness of the capsule tear is improved, the emulsification efficiency is improved, or the operation time is shortened. The policy network is trained using either the Proximal Policy Optimization (PPO) algorithm or the Soft Actor Critic (SAC) algorithm to obtain the optimal policy function. .

[0033] Furthermore, in step six, different patient samples, different eye structures, and different lesion conditions are used to validate the trained model; During the robot's autonomous execution, the difference between the model's output actions and the expert reference actions is calculated in real time, and a risk assessment function is established:

[0034] When the risk value exceeds the preset threshold ,Right now Upon this, the system immediately triggers the manual takeover mechanism, suspending autonomous operation and switching to remote doctor operation mode; the doctor's corrected operation data is recorded as follows:

[0035] And re-add it to the training dataset; At the same time, a module credibility assessment mechanism is established based on the error type, and the activation weights of various sensors are dynamically adjusted to continuously optimize the performance of the autonomous decision-making model.

[0036] A control system for autonomous cataract surgery performed by a robot, the control system using the method for autonomous cataract surgery as described above, the control system comprising: Data acquisition unit: Establish a multimodal information acquisition platform for cataract surgery, using surgical microscope, OCT imaging system, six-dimensional force sensor and robot controller to collect visual information, force information, robot status information and patient physiological information in real time, and complete multi-source data synchronization; Data preprocessing unit: preprocesses the raw data acquired by the data acquisition unit; Expert database construction unit: Ophthalmology experts perform cataract surgery through remotely operated robots, synchronously recording the robot's status and control actions, establishing an expert demonstration database covering all stages of surgery, and constructing an expert knowledge graph; Initial strategy generation unit: The action generation model is trained using an expert demonstration database built on the expert database construction unit based on the imitation learning method. It learns the operation rules and decision-making logic of expert doctors and forms the initial strategy for autonomous control of the robot. Control strategy optimization unit: The initial strategy obtained by the initial strategy generation unit is combined with the digital twin cataract surgery environment, and the control strategy is further optimized through reinforcement learning; Correction and Validation Unit: Validates the model using different patient samples and introduces a manual intervention mechanism to correct abnormal operations, thereby achieving continuous optimization and iterative updates of the model.

[0037] The beneficial effects of this invention are: 1. This invention constructs a multimodal information acquisition framework that integrates microscope images, force information, robot body state, and patient physiological parameters, achieving comprehensive perception of the entire cataract surgery process. Compared to traditional control methods that rely solely on visual information, this invention can simultaneously acquire multi-dimensional information such as tissue morphology changes, instrument contact status, and real-time patient status, enabling the robot to have a more comprehensive and accurate cognitive ability of the surgical environment, providing a reliable data foundation for subsequent autonomous decision-making.

[0038] 2. This invention establishes a complete data optimization process from data acquisition, feature extraction, filtering and noise reduction, outlier removal to data augmentation. Key visual features are extracted using convolutional neural networks, and force signal noise is eliminated using Kalman filtering and wavelet denoising methods. Invalid data such as pauses and errors are automatically removed, effectively improving the quality of training data. Simultaneously, various data augmentation methods are employed to expand the sample space, reducing the model's dependence on specific surgical scenarios and individual doctor operating habits, thus improving the system's generalization ability and robustness.

[0039] 3. This invention enables doctors to remotely operate robots to complete surgeries while simultaneously collecting expert demonstration data, further constructing an expert knowledge graph and realizing the digital expression of expert doctors' operational experience and decision-making logic. The system not only records the operations performed by the doctor but also learns the state transition patterns and decision-making basis at different surgical stages. This transforms the clinical knowledge accumulated by experienced doctors over a long period into a computable and accessible knowledge base, providing prior knowledge support for the robot's autonomous decision-making.

[0040] 4. This invention employs imitation learning technology to model expert operational behaviors, enabling the robot to learn the movement patterns and control strategies of doctors during key steps such as capsulorhexis, phacoemulsification, and cortical aspiration. By establishing a motion generation model based on deep neural networks, it achieves a transformation from "observational expert" to "understanding expert," giving the robot operational capabilities and risk avoidance abilities similar to those of expert doctors, significantly improving the stability and consistency of autonomous operation.

[0041] 5. This invention further introduces a reinforcement learning optimization mechanism to continuously optimize the strategies learned through imitation in a digital twin cataract surgery environment. The system can autonomously explore better operating methods through extensive simulation training, improving capsulorhexis roundness, reducing the risk of capsular tearing, increasing emulsification efficiency, and shortening surgical time while ensuring tissue safety. This breaks through the performance limitations imposed by solely relying on expert demonstration data, enabling continuous evolution and optimization of surgical strategies.

[0042] 6. This invention constructs a complete human-machine collaborative closed-loop learning mechanism. During the robot's autonomous execution, the current operational status is monitored in real time through a risk assessment model. When a potential danger is detected, a manual intervention mechanism is automatically triggered, allowing a doctor to intervene and correct the operation in a timely manner. Simultaneously, the corrected data is re-integrated into the training database for continuous model updates and iterative optimization. This mechanism realizes a closed-loop learning process of "autonomous execution—risk monitoring—human intervention—model update," enabling the system to continuously improve its performance and safety in practical applications.

[0043] 7. This invention is the first to organically combine multimodal perception, expert knowledge graphs, imitation learning, reinforcement learning, and human-machine collaborative closed-loop optimization to form a complete autonomous decision-making technology system for cataract surgery. This system not only enables the robot to autonomously execute key aspects of cataract surgery, but also dynamically adjusts its operational strategies based on different patient conditions and changes in the intraoperative environment, thereby improving intraoperative operational accuracy, reducing the incidence of complications, and significantly enhancing the consistency and stability of surgical quality.

[0044] In summary, this invention, by constructing a complete technical chain encompassing multimodal perception, expert knowledge learning, action generation decision-making, strategy optimization, and human-machine collaborative feedback, enables the robot to autonomously perceive, make decisions, execute, and optimize key steps in cataract surgery. This significantly improves the intelligence level, operational precision, and safety of cataract surgery robots, providing a new technical solution for achieving high-level autonomous ophthalmic surgical robots, and has broad clinical application prospects and industrialization value. Attached Figure Description

[0045] Figure 1 This is a flowchart illustrating the multimodal information acquisition, state construction, and data processing and optimization process of the present invention.

[0046] Figure 2 The flowchart illustrates the model learning, optimization, and closed-loop control process of this invention. Detailed Implementation

[0047] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods are omitted so as not to obscure the description of this application with unnecessary detail.

[0048] It should be understood that, when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0049] It should also be understood that the terminology used in this application specification is for the purpose of describing particular embodiments only and is not intended to limit the application. As used in this application specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0050] The following is in conjunction with the appendix to this application specification. Figure 1-2 The technical solutions in the embodiments of this application are clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0051] Many specific details are set forth in the following description in order to provide a full understanding of this application. However, this application may also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the spirit of this application. Therefore, this application is not limited to the specific embodiments disclosed below.

[0052] Implementation Method 1 This invention provides a method for performing cataract surgery autonomously by a robot, the method comprising the following steps: Step 1: Establish a multimodal information acquisition platform for cataract surgery. Utilize a surgical microscope, OCT imaging system, six-dimensional force sensor, and robot controller to collect visual information, force information, robot status information, and patient physiological information in real time, and complete the synchronization of multi-source data. Step 2: Preprocess the raw data obtained in Step 1, including visual feature extraction, force signal filtering, outlier removal, and data augmentation, to improve data quality and model training effectiveness. Step 3: Experienced ophthalmologists perform cataract surgery using a remotely operated robot, simultaneously recording the robot's status and control actions, establishing an expert demonstration database covering all surgical stages, and constructing an expert knowledge graph. Step 4: Using the imitation learning method, train the action generation model based on the expert demonstration database established in Step 3, learn the operating rules and decision-making logic of expert doctors, and form the initial strategy for autonomous control of the robot. Step 5: Combining the initial strategy from Step 4 with the digital twin cataract surgery environment, the control strategy is further optimized through reinforcement learning to obtain a safer and more efficient autonomous operation plan. Step Six: Validate the model using different patient samples and introduce a manual intervention mechanism to correct any abnormal operations, thereby achieving continuous optimization and iterative updates of the model.

[0053] Furthermore, step one specifically involves establishing a multimodal information acquisition platform for cataract surgery to acquire visual information, force information, robot status information, and patient physiological information in real time during the surgical process. Visual information is acquired by a surgical microscope and OCT imaging system, including anterior segment images, lens images, and the relative positions of instruments and tissues; force information is acquired by a six-dimensional force sensor installed at the end of the surgical instrument, including the forces and torques generated during contact between the instrument and eye tissues; robot status information includes joint angles, joint velocities, joint torques, and end effector pose; patient physiological information includes parameters such as eye movements, heart rate, and blood pressure.

[0054] Set at time The collected visual information, OCT information, force information, robot state information, and patient physiological information are respectively represented as follows: Then construct the time Multimodal state vector:

[0055] By synchronizing various types of data through a unified timestamp mechanism, the correspondence between different modalities of data on the same timeline is realized, providing a unified data foundation for subsequent expert knowledge learning.

[0056] Furthermore, step two specifically involves preprocessing and optimizing the raw multimodal data acquired in step one; For visual data, an image feature extraction method based on convolutional neural network (CNN) is used to extract key features such as lens edge, capsule boundary, emulsification needle tip position and artificial lens position from microscope images. Let the input image be The convolution kernel is The bias term is The output feature map of the convolutional layer is then:

[0057] in This represents the output feature map of the convolutional layer. This represents the convolution operation. Indicates the bias term. Indicates the activation function; After processing by a multi-layer convolutional network, high-dimensional semantic features of the image are obtained:

[0058] in Indicates time Microscopic images collected, This represents the corresponding visual feature vector; For force signals, the Kalman filter algorithm is used to eliminate random noise; First, establish the system state equations:

[0059] in Indicates time The true force signal state, Represents the state transition matrix. Represent process noise; establish observation equations:

[0060] in This indicates the force sensor measurement value. Represents the observation matrix. Indicate the measurement noise; calculate the Kalman gain based on the Kalman filtering principle. And update the state estimate:

[0061] After filtering, a smooth force signal is obtained; For robot motion data, a sliding window algorithm is used to identify and automatically delete long pauses, false triggers, and invalid operation segments. Let the window length be The average speed within the window is The first one in the window The speed of each sampling point is The average speed of the end effector within the window is:

[0062] Let the time threshold be When satisfied Furthermore, if the duration exceeds a preset time threshold, the system determines that the data segment is an invalid pause segment and automatically removes it; Simultaneously, data augmentation methods are employed to expand the training samples, thereby improving the model's generalization ability; for image rotation enhancement, the original image coordinates are set as follows: The rotation angle is Then the coordinates after rotation are:

[0063] For brightness enhancement, a linear transformation is used: ; in Represents the original image. This represents the enhanced image. Indicates the brightness scaling factor. This indicates the brightness offset.

[0064] Furthermore, for Gaussian noise enhancement, random noise is first generated:

[0065] in This indicates that the mean is 0 and the variance is 0. Gaussian distribution; Then the enhanced image is obtained:

[0066] To enhance the perturbation of the robot's motion trajectory, let the original trajectory point be... The random perturbation vector is The enhanced trajectory points are:

[0067] After the above processing, the optimized multimodal state vector is constructed. , This represents a data optimization function that consists of feature extraction, filtering and noise reduction, outlier removal, and data augmentation.

[0068] Furthermore, step three specifically involves the system recording the robot's state vector in real time during the surgery. And the corresponding control actions of expert doctors ,in:

[0069] Indicates that the doctor is at all times Position and attitude control quantities applied to the robot's end effector; The expert physician's procedure sequence was obtained through continuous recording:

[0070] Where N is the total number of sampling frames.

[0071] Furthermore, based on the cataract surgery process, the expert data was divided into the capsulorhexis stage, phacoemulsification stage, cortical aspiration stage, and intraocular lens implantation stage, and corresponding stage labels were established. An expert knowledge graph is constructed based on the aforementioned expert demonstration data. Nodes in the knowledge graph include surgical stage nodes, tissue state nodes, instrument state nodes, and operational action nodes; edges represent causal and temporal relationships between different nodes. By statistically analyzing the state transition patterns during a large number of expert operations, an expert decision-making knowledge base is formed, providing prior knowledge constraints for subsequent autonomous decision-making models.

[0072] Furthermore, step four specifically involves analyzing the operational logic of expert doctors using imitation learning methods; establishing an action generation model based on a Transformer network, with the input being a multimodal state sequence of the current and historical moments:

[0073] The output is the predicted action for the next time step:

[0074] The model is trained by minimizing the error between the predicted action and the expert's actual action:

[0075] in For expert action, To predict actions for the model.

[0076] Through imitation learning, the model can learn the operating habits, decision-making logic, and risk avoidance strategies of expert doctors at different stages of surgery, thereby establishing an initial control strategy that is close to the level of an expert.

[0077] Furthermore, step five specifically involves constructing a reinforcement learning optimization framework based on the imitation learning model, and further optimizing the strategy through a digital twin cataract surgery environment.

[0078] Establish surgical state space Action space and the reward function:

[0079] in This indicates an organizational safety reward. This indicates a reward for the quality of the surgery. This indicates a reward for operational efficiency. These are the corresponding weighting coefficients; Negative rewards are given when capsule tearing, tissue damage, or contact force exceeding the safety threshold occurs; positive rewards are given when the roundness of the capsule tear is improved, the emulsification efficiency is improved, or the operation time is shortened. The policy network is trained using either the Proximal Policy Optimization (PPO) algorithm or the Soft Actor Critic (SAC) algorithm to obtain the optimal policy function. .

[0080] By continuously interacting with the digital twin environment, the model can overcome the limitations of expert demonstration data, discover safer and more efficient operating strategies, and achieve autonomous optimization of key steps such as capsulorhexis and phacoemulsification.

[0081] Furthermore, in step six, different patient samples, different eye structures, and different lesion conditions are used to validate the trained model; During the robot's autonomous execution, the difference between the model's output actions and the expert reference actions is calculated in real time, and a risk assessment function is established:

[0082] When the risk value exceeds the preset threshold ,Right now Upon this, the system immediately triggers the manual takeover mechanism, suspending autonomous operation and switching to remote doctor operation mode; the doctor's corrected operation data is recorded as follows:

[0083] And re-added to the training dataset.

[0084] At the same time, a module credibility assessment mechanism is established based on the error type, and the activation weights of various sensors are dynamically adjusted to continuously optimize the performance of the autonomous decision-making model. Through the aforementioned human-machine collaborative closed-loop learning mechanism, the model can be continuously iterated and its performance improved in practical applications, thereby enhancing the safety, stability, and generalization ability of autonomous cataract surgery.

[0085] Implementation Method 2 This invention provides a specific method for performing cataract surgery autonomously by a robot, obtained using the method of Embodiment 1.

[0086] Step 1: After preoperative examinations, the patient is secured to the surgical platform, and corresponding surgical instruments are mounted on the robot's end effector. A multimodal information acquisition platform is constructed using a surgical microscope, OCT system, six-dimensional force sensor, robot controller, and patient physiological monitoring equipment. Spatial transformation relationships between the coordinate systems of each device are obtained through system calibration, and microscope images are acquired in real time. OCT images Sense of force Robot status information and patient physiological information , building time Multimodal state vector Among them, microscopic images are used to identify the limbus, pupil, and surgical instrument positions; OCT images are used to acquire information on the anterior capsule and lens structure; and force information is used to monitor the contact status between instruments and tissues.

[0087] Step 2: Process the state vector obtained in Step 1 The process involves processing the images. Based on microscopic and OCT images, the corneal incision location, capsulorhexis center, and lens boundary are identified, and the surgical target area is generated. Kalman filtering is used to process the force signal. Denoising is performed to remove outlier and invalid data, resulting in an optimized state vector. in, This represents the data optimization function. Simultaneously, it updates the target area's location in real-time based on the patient's eye movement, providing a basis for subsequent autonomous robot operation.

[0088] Step 3: Establish an expert demonstration database. Ophthalmologists perform cataract surgery via a remotely operated robot, including corneal incision creation, continuous circular capsulotomy, phacoemulsification, and intraocular lens implantation. The optimized state vector is recorded synchronously during the procedure. and robot control actions Construct expert operation sequence Among them, actions This represents the pose adjustment of the robot's end effector at the current moment. Furthermore, an expert knowledge graph is built based on the surgical procedure to describe the relationships between surgical stages, tissue states, and surgical actions.

[0089] Step 4: The robot determines the current state vector. The robot performs autonomous operations. When the state vector represents the corneal incision stage, the robot controls the scalpel to create a clear corneal incision along the planned trajectory; when the state vector represents the capsulorhexis stage, the robot controls the capsulorhexis forceps to grasp the edge of the anterior capsular membrane and complete continuous circular capsulorhexis along the target trajectory; when the state vector represents the phacoemulsification stage, the robot controls the phacoemulsification handpiece to complete lens nucleus fragmentation and aspiration; when the state vector represents the intraocular lens implantation stage, the robot controls the implanter to deliver the intraocular lens into the capsular bag and complete its positioning. New state vectors are continuously acquired during the operation. It also corrects the motion trajectory in real time based on the predicted actions.

[0090] Step 5: Input the expert operation sequence (D) into the imitation learning network for training. The input state sequence is... Output predicted action By minimizing expert actions With predicting actions The error between these parameters allows the model to learn the operational patterns of expert surgeons at different stages of surgery. Subsequently, reinforcement learning optimization is performed in a digital twin cataract surgery environment to establish a reward function. To obtain the optimal policy function During operation, the robot operates based on its current state vector. Generate optimal action They completed the surgical procedure independently.

[0091] Step Six: Establish a risk assessment function The robot assesses operational risks in real time during execution. Risks are mitigated when the instrument contact force exceeds a safety threshold, the capsulorhexis trajectory deviates from the target area, the posterior capsule is too close, or abnormal eye movements occur. The system automatically pauses autonomous operation and switches to remote operation mode for the doctor. The doctor's corrected operation data is recorded as follows: The system then re-integrates the knowledge graph and strategy model into the training database, updating and optimizing them to form a closed-loop learning mechanism of "autonomous execution - risk assessment - manual intervention - model update".

[0092] Through the above implementation methods, the present invention realizes autonomous perception, autonomous decision-making and autonomous operation of cataract surgery robot based on multimodal motion generation model, thereby improving the level of surgical automation, operational accuracy and safety.

[0093] Implementation Method 3 This embodiment provides a control system for autonomously performing cataract surgery using a robot. The control system employs a method for autonomously performing cataract surgery as described in Embodiment 1. The control system includes: Data acquisition unit: Establish a multimodal information acquisition platform for cataract surgery, using surgical microscope, OCT imaging system, six-dimensional force sensor and robot controller to collect visual information, force information, robot status information and patient physiological information in real time, and complete multi-source data synchronization; Data preprocessing unit: preprocesses the raw data acquired by the data acquisition unit; Expert database construction unit: Ophthalmology experts perform cataract surgery through remotely operated robots, synchronously recording the robot's status and control actions, establishing an expert demonstration database covering all stages of surgery, and constructing an expert knowledge graph; Initial strategy generation unit: The action generation model is trained using an expert demonstration database built on the expert database construction unit based on the imitation learning method. It learns the operation rules and decision-making logic of expert doctors and forms the initial strategy for autonomous control of the robot. Control strategy optimization unit: The initial strategy obtained by the initial strategy generation unit is combined with the digital twin cataract surgery environment, and the control strategy is further optimized through reinforcement learning; Correction and Validation Unit: Validates the model using different patient samples and introduces a manual intervention mechanism to correct abnormal operations, thereby achieving continuous optimization and iterative updates of the model.

Claims

1. A method for performing cataract surgery autonomously by a robot, characterized in that, The method includes the following steps: Step 1: Establish a multimodal information acquisition platform for cataract surgery. Utilize a surgical microscope, OCT imaging system, six-dimensional force sensor, and robot controller to collect visual information, force information, robot status information, and patient physiological information in real time, and complete the synchronization of multi-source data. Step 2: Preprocess the raw data obtained in Step 1; Step 3: An ophthalmologist performs cataract surgery using a remotely operated robot, simultaneously recording the robot's status and control actions, establishing an expert demonstration database covering all stages of the surgery, and constructing an expert knowledge graph. Step 4: Using the imitation learning method, train the action generation model based on the expert demonstration database established in Step 3, learn the operating rules and decision-making logic of expert doctors, and form the initial strategy for autonomous control of the robot. Step 5: Combine the initial strategy from Step 4 with the digital twin cataract surgery environment, and further optimize the control strategy through reinforcement learning; Step Six: Validate the model using different patient samples and introduce a manual intervention mechanism to correct any abnormal operations, thereby achieving continuous optimization and iterative updates of the model.

2. The method according to claim 1, characterized in that, Specifically, step one involves setting a time... The collected visual information, OCT information, force information, robot state information, and patient physiological information are respectively represented as follows: Then construct the time Multimodal state vector: By using a unified timestamp mechanism to synchronize various types of data, the correspondence between different modalities of data on the same time axis can be achieved.

3. The method according to claim 1, characterized in that, Step two specifically involves assuming the input image is... The convolution kernel is The bias term is The output feature map of the convolutional layer is then: in This represents the output feature map of the convolutional layer. This represents the convolution operation. Indicates the bias term. Indicates the activation function; After processing by a multi-layer convolutional network, high-dimensional semantic features of the image are obtained: in Indicates time Microscopic images collected, This represents the corresponding visual feature vector; For force signals, the Kalman filter algorithm is used to eliminate random noise; First, establish the system state equations: in Indicates time The true force signal state, Represents the state transition matrix. Represent process noise; establish observation equations: in This indicates the force sensor measurement value. Represents the observation matrix. Indicates measurement noise; Calculate the Kalman gain based on the Kalman filtering principle. And update the state estimate: After filtering, a smooth force signal is obtained; For robot motion data, a sliding window algorithm is used to identify and automatically delete long pauses, false triggers, and invalid operation segments. Let the window length be The average speed within the window is The first one in the window The speed of each sampling point is The average speed of the end effector within the window is: Let the time threshold be When satisfied Furthermore, if the duration exceeds a preset time threshold, the system determines that the data segment is an invalid pause segment and automatically removes it; For image rotation enhancement, let the original image coordinates be... The rotation angle is Then the coordinates after rotation are: For brightness enhancement, a linear transformation is used: ; in Represents the original image. This represents the enhanced image. Indicates the brightness scaling factor. This indicates the brightness offset.

4. The method according to claim 3, characterized in that, First, generate random noise: in This indicates that the mean is 0 and the variance is 0. Gaussian distribution; Then the enhanced image is obtained: To enhance the perturbation of the robot's motion trajectory, let the original trajectory point be... The random perturbation vector is The enhanced trajectory points are: After the above processing, the optimized multimodal state vector is constructed. , This represents a data optimization function that consists of feature extraction, filtering and noise reduction, outlier removal, and data augmentation.

5. The method according to claim 2, characterized in that, Step three specifically involves the system recording the robot's state vector in real time during the surgery. And the corresponding control actions of expert doctors ,in: Indicates that the doctor is at all times Position and attitude control quantities applied to the robot's end effector; The expert physician's procedure sequence was obtained through continuous recording: Where N is the total number of sampling frames.

6. The method according to claim 5, characterized in that, Based on the cataract surgery process, expert data is divided into the capsulorhexis stage, phacoemulsification stage, cortical aspiration stage, and intraocular lens implantation stage, and corresponding stage labels are established. An expert knowledge graph is constructed based on the aforementioned expert demonstration data. The nodes in the knowledge graph include surgical stage nodes, tissue state nodes, instrument state nodes, and operation action nodes. Edges represent the causal and temporal relationships between different nodes. An expert decision-making knowledge base is formed by statistically analyzing the state transition patterns in a large number of expert operations.

7. The method according to claim 1, characterized in that, Step four specifically involves analyzing the operational logic of expert doctors using imitation learning methods; establishing an action generation model based on a Transformer network, with the input being a multimodal state sequence of the current and historical moments: The output is the predicted action for the next time step: The model is trained by minimizing the error between the predicted action and the expert's actual action: in For expert action, To predict actions for the model.

8. The method according to claim 2, characterized in that, Step five specifically involves establishing the surgical state space. Action space and the reward function: in This indicates an organizational safety reward. This indicates a reward for the quality of the surgery. This indicates a reward for operational efficiency. These are the corresponding weighting coefficients; Negative rewards are given when capsule tearing, tissue damage, or contact force exceeding the safety threshold occurs; positive rewards are given when the roundness of the capsule tear is improved, the emulsification efficiency is improved, or the operation time is shortened. The policy network is trained using either the Proximal Policy Optimization (PPO) algorithm or the Soft Actor Critic (SAC) algorithm to obtain the optimal policy function. .

9. The method according to claim 2, characterized in that, In step six, different patient samples, different eye structures, and different lesion conditions are used to validate the trained model; During the robot's autonomous execution, the difference between the model's output actions and the expert reference actions is calculated in real time, and a risk assessment function is established: When the risk value exceeds the preset threshold ,Right now Upon this, the system immediately triggers the manual takeover mechanism, suspending autonomous operation and switching to remote doctor operation mode; the doctor's corrected operation data is recorded as follows: And re-add it to the training dataset; At the same time, a module credibility assessment mechanism is established based on the error type, and the activation weights of various sensors are dynamically adjusted to continuously optimize the performance of the autonomous decision-making model.

10. A control system for a robot to autonomously perform cataract surgery, characterized in that, The control system uses a method for autonomous cataract surgery performed by a robot as described in any one of claims 1-9, the control system comprising: Data acquisition unit: Establish a multimodal information acquisition platform for cataract surgery, using surgical microscope, OCT imaging system, six-dimensional force sensor and robot controller to collect visual information, force information, robot status information and patient physiological information in real time, and complete multi-source data synchronization; Data preprocessing unit: preprocesses the raw data acquired by the data acquisition unit; Expert database construction unit: Ophthalmology experts perform cataract surgery through remotely operated robots, synchronously recording the robot's status and control actions, establishing an expert demonstration database covering all stages of surgery, and constructing an expert knowledge graph; Initial strategy generation unit: The action generation model is trained using an expert demonstration database built on the expert database construction unit based on the imitation learning method. It learns the operation rules and decision-making logic of expert doctors and forms the initial strategy for autonomous control of the robot. Control strategy optimization unit: The initial strategy obtained by the initial strategy generation unit is combined with the digital twin cataract surgery environment, and the control strategy is further optimized through reinforcement learning; Correction and Validation Unit: Validates the model using different patient samples and introduces a manual intervention mechanism to correct abnormal operations, thereby achieving continuous optimization and iterative updates of the model.