Intelligent control method and system for microsurgery robot

By fusing multimodal perception through binocular endoscope and wrist force sensor, a closed-loop intelligent decision-making system for microsurgical robot system in dynamic environments is realized, solving the problem of fragmented response logic in existing technologies and improving the continuity and efficiency of automated surgery.

CN121987351BActive Publication Date: 2026-07-10CHENGDU BORNS MEDICAL ROBOTICS INC
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHENGDU BORNS MEDICAL ROBOTICS INC
Filing Date
2026-04-09
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Existing surgical robot systems lack intelligent decision-making capabilities when facing dynamic disturbances, resulting in fragmented and passive response logic. They are unable to form an intelligent closed loop of perception-evaluation-decision-adaptive continuation, which affects the continuity and efficiency of automated execution.

Method used

By simultaneously acquiring visual and multidimensional force information through binocular endoscope and wrist force sensor, intelligent decision-making is made based on the semantics of surgical steps, generating correction strategies and executing transition trajectories under safe space constraints, thus realizing a closed loop of intelligent decision-making through multimodal perception fusion.

Benefits of technology

It improves the accuracy of anomaly detection and the rationality of decision-making, ensures the continuity and fault tolerance of automated surgical procedures, reduces interruptions in the surgical rhythm, and improves operational efficiency and task success rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121987351B_ABST
    Figure CN121987351B_ABST
Patent Text Reader

Abstract

This application provides an intelligent control method and system for a microsurgical robot, belonging to the field of medical robot control technology. The method includes: acquiring visual information and multi-dimensional force information for the surgical instruments when executing each surgical step based on a step-time library; calculating the visual deviation of the instrument tip based on the visual information and calculating force anomaly indicators based on the multi-dimensional force information; using the currently executed surgical step as the decision context, determining whether preset anomaly triggering conditions are met based on the visual deviation and force anomaly indicators; if so, pausing the currently executed surgical step, determining a correction strategy based on the depth confidence level, and determining the correction amount corresponding to the correction strategy under predefined safety space geometric constraints; generating a transition trajectory based on the correction amount, and controlling the surgical instruments to execute the transition trajectory. This application can realize an intelligent decision-making closed loop based on the semantics of the surgical steps and multi-modal perception fusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of medical robot control technology, and in particular to an intelligent control method and system for a microsurgical robot. Background Technology

[0002] As robot-assisted surgery technology expands into more precise intracavitary, transoral, and open surgeries, the demand for robots to automatically perform procedural and delicate operations such as suturing and knot tying is becoming increasingly urgent. However, current surgical robot systems generally exhibit a core contradiction when dealing with such highly continuous tasks: the systems lack the intelligence to uniformly understand and respond in real time to the pre-set, rigid procedural operating steps and the dynamic, uncertain, and complexly constrained physical surgical environment. When common dynamic disturbances occur during surgery, such as tissue traction displacement, instrument slippage, or loss of field of vision, existing systems can mostly only take simple emergency stops or trigger fixed alarms, forcing the surgical procedure to be interrupted and returning all responsibility for the treatment to the surgeon, severely restricting the continuity and overall efficiency of automated execution. Therefore, how to enable robot systems to actively and smoothly assist in adjustments and safely resume operations under dynamic disturbances is a key bottleneck in improving the level of autonomy of surgical robots.

[0003] Existing technologies have made various attempts to improve the system's response capabilities, mainly focusing on two levels: first, enhancing the capabilities of a single sensing dimension, such as developing more robust visual tracking algorithms to improve positioning accuracy, or using more sensitive force sensors to detect contact force anomalies earlier; second, strengthening the system's planning and constraint mechanisms, such as designing more complex obstacle avoidance paths or setting stricter virtual safety boundaries. However, these improvements still fundamentally suffer from an unbridgeable gap: there is a lack of an intelligent decision-making layer based on the current surgical context and the system's own sensing reliability between the sensed abnormal signal and the corrective action that should be taken. This results in a fragmented and passive response logic when the system faces anomalies—either reacting slowly due to conservative threshold settings, or frequently interrupting the process due to accidental triggering, ultimately often falling into an inefficient cycle of "sensing anomalies, emergency pause, manual reset, and restarting," rather than forming an intelligent closed loop of "perception-evaluation-decision-adaptive continuation."

[0004] It is evident that the deep-seated flaw in existing technical solutions lies in the mechanical and discontinuous nature of their response modes. Therefore, there is an urgent need to propose a new system control paradigm that can achieve an intelligent decision-making closed loop based on multimodal perception fusion, with surgical procedure semantics at its core. Summary of the Invention

[0005] The purpose of this application is to provide an intelligent control method and system for a microsurgical robot to solve the above-mentioned problems.

[0006] To achieve the above objectives, firstly, this application proposes an intelligent control method for a microsurgical robot, the method comprising:

[0007] When performing each surgical step based on a preset step sequence library, visual information and multi-dimensional force information for the surgical instruments are acquired simultaneously through a binocular endoscope and a wrist force sensor.

[0008] The visual deviation at the tip of the computing device is calculated based on the visual information, and the force anomaly index is calculated based on the multidimensional force information.

[0009] Using the currently performed surgical step as the decision context, and based on the visual deviation and the force abnormality index, it is determined whether the preset abnormality triggering conditions are met.

[0010] If so, the currently executed surgical step is paused, and the corresponding correction strategy is determined according to the depth confidence level of the binocular endoscope. Under the predefined safety space geometric constraints, the correction amount corresponding to the correction strategy for correcting the instrument pose or operating parameters is determined.

[0011] A transition trajectory is generated based on the correction amount, and the surgical instrument is controlled to execute the transition trajectory.

[0012] In some implementations, determining the corresponding correction strategy based on the depth confidence level of the binocular endoscope includes:

[0013] When the depth confidence level is high, the corresponding correction strategy is determined to be full-degree-of-freedom correction of the instrument pose based on the three-dimensional spatial pose deviation.

[0014] When the depth confidence level is medium, the corresponding correction strategy is determined to be to limit the correction amplitude in three-dimensional space and to perform hybrid correction by combining the pixel deviation of the two-dimensional image.

[0015] When the depth confidence level is low, the corresponding correction strategy is determined to be either triggering security maintenance or requesting manual intervention.

[0016] In some implementations, before determining whether a preset abnormality triggering condition is met based on the visual deviation and the force abnormality index, using the currently performed surgical step as the decision context, the process includes:

[0017] The binocular endoscope is subjected to binocular positioning and epipolar correction, and the parallax search range of stereo matching is limited to the parallax interval corresponding to the preset working distance range of the surgical robot;

[0018] Perform stereo matching within the disparity interval to acquire disparity data;

[0019] The depth information of the instrument tip is calculated based on the parallax data;

[0020] Depth confidence is calculated based on at least one of the following: left-right consistency in the stereo matching process, uniqueness of the matching result, texture features of the image region, and temporal consistency of depth information.

[0021] In some embodiments, the method further includes:

[0022] Identify and track the pixel coordinates of the instrument tip in the imaging image of the binocular endoscope;

[0023] Calculate the pixel offset between the pixel coordinates and the center of the image field of view;

[0024] If the pixel offset exceeds the preset offset value, the binocular endoscope is controlled to move along the planned path so that the instrument tip returns to center.

[0025] In some embodiments, controlling the movement of the binocular endoscope along a planned path includes:

[0026] Determine whether the binocular endoscope triggers a preset collision conflict or a preset occlusion conflict when it performs centering control.

[0027] If not, then control the binocular endoscope to move along the planned path based on centering control;

[0028] If so, the optimal observation pose of the binocular endoscope is solved by taking the maximization of the visibility of the instrument tip and / or the maximization of the image coverage of the key area as the optimization objectives, and taking collision avoidance constraints and predefined safety space geometric constraints as conditions. The binocular endoscope is then controlled to move to the optimal observation pose.

[0029] In some implementations, the step timing library includes step sequences, action primitives, transition conditions, and backoff strategies. Before simultaneously acquiring visual and multidimensional force information about the surgical instruments via a binocular endoscope and wrist force sensor when executing each surgical step based on the preset step timing library, the library further includes:

[0030] A standardized sequence of steps is constructed for each surgical subtask in the target surgical scenario, wherein the surgical subtask includes at least one of threading, suturing, and knotting.

[0031] For each surgical step in the sequence of steps, at least one action primitive is associated with the action primitive, which includes at least one of the following: a primitive for controlling the pose trajectory of the surgical instrument, a clamping primitive, a force control primitive, and a lens follow-up primitive for controlling the binocular endoscope.

[0032] For each surgical step, corresponding transition conditions are set, including at least one of instrument placement threshold, visual alignment threshold, force threshold, time threshold, and confidence threshold.

[0033] A corresponding retraction strategy is preset for each surgical step. The retraction strategy includes at least one of re-clamping, retracting to a safe position, re-performing the alignment action, and switching to manual control mode.

[0034] In some embodiments, the method further includes:

[0035] In response to a manual takeover command, the automated control output based on the step timing library is frozen, and control over the surgical instruments and the binocular endoscope is transferred to manual operation.

[0036] After manual operation is completed, the current system state vector is collected. The system state vector includes the instrument pose, clamping state, target point visibility, force / torque state, and task execution confidence.

[0037] The system state vector is matched with the expected state of each surgical step in the step timing library. The surgical step with the highest matching degree is determined as the resynchronization node, and the next surgical step of the resynchronization node is executed.

[0038] In some implementations, the expected state includes a reference state vector and a tolerance state vector. Matching the system state vector with the expected states of each surgical step in the step timing library to determine the surgical step with the highest matching degree as the resynchronization node includes:

[0039] Surgical steps whose state range, formed by the reference state vector and the tolerance state vector, overlaps with the components of the system state vector are selected as candidate step nodes.

[0040] Calculate the normalized deviation of the system state vector from the reference state vector of each candidate step node in each state dimension;

[0041] Based on the normalized deviation, the matching score of each candidate step node is calculated, and the candidate step node with the highest matching score is determined as the resynchronization node.

[0042] Secondly, to achieve the above objectives, this application also proposes an intelligent control system for a microsurgical robot, comprising:

[0043] The sensing module includes a binocular endoscope and a wrist force sensor;

[0044] The execution module includes the robotic actuator and surgical instruments;

[0045] The decision and control module, including a control computing unit, is used to execute the intelligent control method for the microsurgical robot as described above.

[0046] Thirdly, to achieve the above objectives, this application also proposes a computer storage medium storing executable instructions, which, when executed by a processor, cause the processor to perform the intelligent control method for the microsurgical robot described above.

[0047] Compared with the prior art, the beneficial effects of this application include:

[0048] Firstly, this application uses the currently performed surgical step as the decision context, enabling the judgment of visual deviation and force abnormality indicators to be deeply bound to the semantics of specific surgical actions. This overcomes the drawbacks of the existing technology, which uses fixed thresholds, resulting in sluggish response or frequent false triggers. It achieves a leap from isolated signal processing to contextualized intelligent judgment, significantly improving the accuracy of abnormality perception and the rationality of decision-making.

[0049] Secondly, by introducing a flexible decision-making mechanism based on deep confidence grading, quantitative assessment and response to uncertainty are achieved. This allows for the autonomous selection of the most appropriate correction strategy based on the reliability of the perceived information itself, rather than simple interruption, when common intraoperative disturbances such as tissue displacement or instrument slippage occur. This not only significantly improves the continuity and fault tolerance of the automated surgical procedure but also ensures that all autonomous corrective actions are performed within strict geometric boundaries through predefined safety space constraints, fundamentally avoiding secondary risks caused by automated intervention.

[0050] Thirdly, this application generates a smooth transition trajectory based on the correction amount and continues it to subsequent steps. After the system completes the correction, it can naturally and smoothly resume the automated process, effectively eliminating the interruption of surgical rhythm and operational redundancy caused by the traditional emergency stop and reset method. This enables the robot system to simulate the coherent assistance capabilities of a senior surgical assistant and maintain a high level of operational efficiency and overall task success rate in a dynamic environment. Attached Figure Description

[0051] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation on the scope of this application.

[0052] Figure 1 This is a schematic diagram of the overall architecture of the intelligent control system of the microsurgical robot in the embodiments of this application;

[0053] Figure 2 This is a schematic diagram of the coordinate system and geometric relationship of the intelligent control system of the microsurgical robot in the embodiments of this application;

[0054] Figure 3 This is a schematic diagram of coordinate system transformation for the intelligent control system of the microsurgical robot in this embodiment of the application;

[0055] Figure 4 This is a flowchart illustrating the intelligent control method for a microsurgical robot in one embodiment;

[0056] Figure 5 This is a schematic diagram of the timing library structure for one embodiment;

[0057] Figure 6 This is a schematic diagram of the safety space geometric constraints in one embodiment;

[0058] Figure 7 This is a flowchart illustrating the intelligent control method for a microsurgical robot in another embodiment;

[0059] Figure 8 This is a flowchart illustrating the intelligent control method for a microsurgical robot in yet another embodiment. Detailed Implementation

[0060] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0061] All terms used in this application (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0062] For example, the terms "first," "second," etc., used in this application may be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish one element from another. Furthermore, the terms "comprising," "including," etc., used in this application indicate the presence of features, steps, operations, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, or components.

[0063] As mentioned earlier, existing technologies have made various attempts to improve the intelligent control system capabilities of microsurgical robots, mainly focusing on two levels: first, enhancing the capabilities of a single perception dimension, such as developing more robust visual tracking algorithms to improve positioning accuracy, or using more sensitive force sensors to detect contact force anomalies earlier; second, strengthening the system's planning and constraint mechanisms, such as designing more complex obstacle avoidance paths or setting stricter virtual safety boundaries. However, these improvements still fundamentally suffer from an unbridgeable gap: there is a lack of an intelligent decision-making layer based on the current surgical context and the reliability of the system's own perception between the perceived abnormal signal and the corrective action that should be performed. This results in a fragmented and passive response logic when the system faces anomalies—either due to conservative threshold settings leading to a sluggish response, or due to frequent interruptions caused by accidental triggering, ultimately often falling into an inefficient cycle of "perceiving an anomaly, emergency stop, manual reset, and restarting," rather than forming an intelligent closed loop of "perception-evaluation-decision-adaptive continuation." Thus, the deep-seated flaw of existing technologies lies in the mechanical and discontinuous nature of their response modes. To this end, this application proposes an intelligent control method and system for a microsurgical robot, which can realize an intelligent decision-making closed loop based on the semantics of surgical steps and multimodal perception fusion.

[0064] In this embodiment, for ease of description, the following description uses the intelligent control system of a microsurgical robot as the executing entity. Figure 1As shown, the intelligent control system of the microsurgical robot provided in this application embodiment is an integrated closed-loop system with a control computing unit as the core, a multimodal perception and interaction device as the boundary, and structured data recording as the extension. From the system architecture level, the overall system can be divided into the following functional modules: (1) a perception module, including a binocular endoscope and a wrist force sensor. The binocular endoscope is installed above the surgical area or inserted through a natural cavity to collect left and right eye image streams and complete stereo matching and depth calculation in the control computing unit. At the same time, it outputs a depth confidence (Conf) flag to evaluate the reliability of the three-dimensional perception information. The wrist force sensor is integrated after the wrist drive unit of the surgical instrument and before the end effector. It measures the force components (Fx, Fy, Fz) of the instrument end in three translational directions and the torque components (Tx, Ty, Tz) in three rotational directions in real time. After gravity compensation and zero bias calibration, it outputs pure multidimensional force information. (2) An execution module, including a robotic actuator and surgical instruments, wherein the robotic actuator may be a multi-degree-of-freedom robotic arm or a dedicated surgical robot platform, receiving position, speed, and force control commands from the control computing unit, and driving the instruments to complete predetermined actions inside or on the patient's body. Surgical instruments include end effectors such as needle holders, scissors, and grippers, whose clamping, opening and closing, joint bending, and other actions are driven by the actuator, and are terminals that directly interact with tissues, sutures, and needles. (3) A decision and control module, including a control computing unit, used for the steps in the various method embodiments of this application. In addition, it may further include: (4) A human-computer interaction module, including a doctor's console (master / slave / joystick) and a foot pedal / voice module (optional), wherein the doctor can directly control the instruments and lens through the master-slave operation mode, or input takeover commands through the joystick at any time during automatic execution to trigger the manual operation mode. The foot pedal / voice module is used to provide non-manual interaction methods, such as foot pedal switch triggering pause / resume, voice command "lock field of view" temporarily freezing the lens and automatic following, etc., to meet the needs of different surgical procedures and doctor's habits. (5) Data recording module, including data recording module, used to timestamp and structure key information in the automated execution process.

[0065] In addition, such as Figure 2 and Figure 3As shown, the system establishes a corresponding Cartesian coordinate system for each entity with independent motion or spatial significance in the surgical scene, serving as the benchmark for position and orientation description. {B} is the robot base coordinate system, fixed to the robot's base, and serves as the global reference system for the entire system. All other coordinate systems can be transformed to this coordinate system for unified calculations using a homogeneous transformation matrix. This coordinate system typically has its origin at the robot's mounting surface, with the Z-axis perpendicular to the mounting plane and pointing upwards. {T} is the instrument end effector coordinate system (TCP), fixed to the tool center point of the surgical instrument's end effector (such as the needle holder tip or scissor blade). The origin of this coordinate system is the instrument tip point, whose pose directly reflects the actual position and orientation of the surgical operation, and is the core focus of visual tracking, force measurement, and motion control. {E} is the binocular endoscope coordinate system, fixed to the optical center of the binocular endoscope, typically with the optical center of the left eye camera as its origin, and the Z-axis pointing towards the depth of the scene along the optical axis. This coordinate system is the reference benchmark for visual perception; the three-dimensional coordinates of the instrument tip obtained from binocular depth calculation are located within this coordinate system. The {S} coordinate system is specifically designed for endovascular / peroral surgical scenarios and is fixed to the inlet plane or geometric center of the standard cavity stent. This coordinate system allows the system to quantify the stent, an anatomical substitute structure, into precise mathematical constraints, a prerequisite for safe control in confined spaces. It is understandable that the {S} coordinate system and working domain constraints can be selectively disabled or replaced with virtual boundaries in open surgical scenarios, while the {B}, {T}, and {E} coordinate systems are universal.

[0066] In the {S} coordinate system, the system defines a three-layer progressive geometric constraint on the operable space of the surgical instruments: the entry plane π0 is defined as a coordinate plane (usually the XY plane) in the support coordinate system {S}, and its equation is: ( - )=0, where This is the normal vector pointing towards the inside of the cavity. This constraint requires the instrument tip to be located inside the cavity at the inlet plane, i.e. ( - ≥0, fundamentally preventing unintended collisions or tissue abrasions between the instrument and the stent edge during instrument withdrawal from the cavity. The permissible working domain (cylindrical / truncated cone) is defined within the inlet plane, based on the physical dimensions and expansion morphology of the stent, defining the legal range of motion of the instrument tip. When the stent is cylindrical, a cylindrical constraint is used: r≤R, 0≤ ≤L, where R is the distance from the instrument tip to the Z-axis, R is the stent radius, and L is the effective working depth. When the stent is a gradually expanding type (such as a partial esophageal stent or rectal stent), a truncated cone constraint is used: r ≤ R0 + k· , 0≤ ≤L, where R0 is the inlet radius and k is the expansion coefficient, both of which together determine the taper shape of the stent as it expands with depth. This constraint ensures that the instrument tip always moves within the physical cavity enclosed by the inner wall of the stent, without excessively compressing the tube wall. The forbidden zone is defined as one or more closed subsets in the {S} coordinate system. It can be a preset geometry (sphere, polyhedron, cylinder) or a 3D risk area generated in real time by segmentation of endoscopic images combined with depth information. For example, locations of large blood vessels to be avoided during suturing, the base of polyps, or previously placed clips. The constraints are the set of instrument key points. This means that key points of the instrument (including but not limited to: the instrument tip, the center of the instrument wrist joint, discrete points uniformly sampled along the instrument axis, and the front end of the endoscope lens) must not enter the restricted area. The restricted area constraint uses absolute prohibition logic. Once a key point of the instrument is detected to have entered or is about to enter the restricted area, the system will: immediately suspend the current automated execution; trigger the highest priority rollback strategy; if it is a force feedback master-slave mode, generate a strong repulsive force; record the event and prompt manual intervention.

[0067] Figure 2 and Figure 3 The transformation relationships between the coordinate systems are also annotated in the form of homogeneous transformation matrices. Specifically, E^T_B represents the transformation from {B} to {E}, obtained through forward kinematics of the robot and hand-eye calibration. This allows the system to transform the planned instrument motion in the robot base coordinate system {B} to the endoscope coordinate system {E} for visual prediction of the expected position. S^T_B represents the transformation from {B} to {S}, obtained through preoperative image registration (such as CT / MRI registration with robot space) or intraoperative probe point cloud acquisition. This transformation maps the instrument pose in the robot base coordinate system {B} to the support coordinate system {S} in real time, thereby performing cycle-by-cycle verification of the aforementioned entry plane, allowed working area, and prohibited area constraints. E^T_T is the transformation from {T} to {E}, which is obtained by combining the forward kinematics solution of the robot with E^T_B and B^T_T, or by directly tracking the tip of the instrument with vision. This transformation directly reflects the correspondence between "where the instrument is" and "what the endoscope sees", and is the direct input for visual deviation calculation and automatic endoscopic following control.

[0068] like Figure 4 As shown in the figure, this application provides an intelligent control method for a microsurgical robot, the method comprising the following steps:

[0069] In step S10, when performing each surgical step based on a preset step sequence library, visual information and multi-dimensional force information for the surgical instruments are acquired simultaneously through a binocular endoscope and a wrist force sensor.

[0070] In this embodiment, the step sequence library is a predefined data structure used to decompose complex surgical procedures (such as threading, suturing, and knotting) into a series of standardized and ordered step sequences. Taking threading as an example, it may include S1 clamping the needle, S2 aligning the suture hole, S3 advancing the suture, S4 loosening the needle, and S5 clamping the suture; taking suturing as an example, it may include S1 aligning the needle hole, S2 inserting the needle, S3 exiting the suture, and S4 retracting the suture; taking knotting as an example, it may include S1 forming a loop, S2 crossing the knot, and S3 tightening the knot. Each step sequence is associated with specific action primitives (such as pose trajectory primitives, clamping primitives, force control primitives, and camera follow primitives), transition conditions (such as instrument positioning thresholds, visual alignment thresholds, force thresholds, time thresholds, and confidence thresholds), and retraction strategies (such as re-clamping, retracting to a safe pose, re-executing the alignment action, and switching to manual control mode).

[0071] In some implementations, a standardized sequence of steps can be constructed for each surgical subtask in the target surgical scenario, the surgical subtask including at least one of threading, suturing, and knotting; at least one action primitive can be associated with each surgical step in the sequence of steps, the action primitive including at least one of a primitive for controlling the pose trajectory of the surgical instrument, a clamping primitive, a force control primitive, and a lens follow-up primitive for controlling the binocular endoscope; corresponding transition conditions can be set for each surgical step, the transition conditions including at least one of an instrument positioning threshold, a visual alignment threshold, a force threshold, a time threshold, and a confidence threshold; and a corresponding backoff strategy can be preset for each surgical step, the backoff strategy including at least one of re-clamping, retracting to a safe pose, re-executing the alignment action, and switching to manual control mode, and a step timing library can be defined. In addition, corresponding parameter sets can be constructed for different surgical scenarios and surgical subtasks, as shown in Table (1), and the structure of the step timing library containing the parameter sets is as follows. Figure 5 As shown, the sequence is: State - Primitive - Profile - Guard - Fallback.

[0072] Table (1)

[0073]

[0074] Step S20: Calculate the force anomaly index based on the visual deviation at the tip of the visual information calculation device and based on the multidimensional force information.

[0075] In this embodiment, visual deviation refers to the difference between the actual observed position of the instrument tip (or target object, such as a pinhole) and its expected position defined in the step timing library. It can be divided into two-dimensional image pixel deviation (…). , ) and three-dimensional spatial pose deviation ( , , , The pixel deviation in the two-dimensional image can be obtained by calculating the difference between the pixel coordinates of the tip and the pixel coordinates of the desired target point (such as the center of a pinhole) in the image. The three-dimensional spatial pose deviation can be obtained by using the principle of binocular stereo vision, triangulating the matching tip points in the left and right images to obtain their three-dimensional coordinates in the endoscope coordinate system, and then comparing them with the desired three-dimensional pose to obtain the position deviation. , , ,) and attitude deviation ( ).

[0076] Force anomaly indices are mathematical quantities used to quantify the degree to which force perception signals deviate from normal or expected states, such as abrupt changes (calculated by the change in the norm of the force / torque vector between adjacent sampling periods). ), rate of change (calculate the first derivative of force / torque with respect to time: ), torque mutation ( Pattern indicators (e.g., during the suturing advancement step, a sudden drop in normal force Fz accompanied by small fluctuations in tangential force Fx / Fy may be defined as an indicator of "tissue penetration" or "slippage precursor"), etc.

[0077] Step S30: Using the currently executed surgical step as the decision context, determine whether the preset abnormal triggering conditions are met based on the visual deviation and the force abnormality index.

[0078] In this embodiment, the decision context refers to the semantic information contained in the currently executed sequence of steps (State). For example, the context of step "S3 Advance the suture" means that the system expects to see the needle tip moving in a straight line and feel gradually increasing resistance, and therefore is particularly sensitive to lateral visual deviations and abrupt changes in resistance.

[0079] The exception triggering conditions contain a set of logical judgment rules, whose thresholds or logical combinations are dynamically associated with the current step context.

[0080] For example, the abnormal triggering conditions include visual conditions (visual deviation in a certain dimension is greater than the corresponding deviation threshold) and / or force-sensory conditions (force abnormality index in a certain dimension is greater than the corresponding index threshold). The specific values ​​of the deviation threshold and index threshold are dynamically related to the context of the surgical step and the precision of the surgical instruments. This is because different surgical step nodes have different tolerances for deviations and different sensitivities to abnormality types. Taking the suturing task as an example, the deviation threshold and index threshold of each step node in the suturing task can be shown in Table (2):

[0081] Table (2):

[0082]

[0083] Understandably, in the S2 alignment step, due to the highest precision required for needle tip and needle hole alignment, the positional deviation threshold is tightened to 0.3mm, the pixel deviation to 8px, and the posture deviation to 2.0°. In the S3 advancement step, because tissue penetration requires overcoming significant resistance, the force normal threshold is relaxed to 5.0N, and the frictional force generated by suture sliding in the tissue needs to be allowed; therefore, the force tangential threshold is relaxed to 2.5N. Since resistance fluctuations during advancement are normal, the force change rate threshold is relaxed to 8.0N / s. In the S1 needle clamping step, considering that excessive tangential force during clamping may indicate a risk of slippage, the force tangential threshold is set to 1.5N.

[0084] Step S40: If yes, pause the currently executed surgical step, determine the corresponding correction strategy based on the depth confidence level of the binocular endoscope, and determine the correction amount corresponding to the correction strategy for correcting the instrument pose or operating parameters under the predefined safe space geometric constraints.

[0085] In this embodiment, the depth confidence level is a classification based on the depth confidence (Conf) score, which can be divided into high (Conf≥0.6), medium (0.35≤Conf<0.6), and low (Conf<0.35). Depth confidence is a quantified index between 0 and 1, used to evaluate the reliability of depth information calculated through binocular stereo matching. In some implementations, this can be achieved by performing binocular calibration and epipolar correction on the binocular endoscope, limiting the disparity search range of stereo matching to a disparity interval corresponding to a preset working distance range of the surgical robot; performing stereo matching within the disparity interval to acquire disparity data; calculating the depth information of the instrument tip based on the disparity data; and calculating the depth confidence based on at least one of the following: left-right consistency during stereo matching, uniqueness of the matching result, texture features of the image region, and temporal consistency of the depth information.

[0086] The correction strategy is a correction scheme formulated for different depth confidence levels. In some implementations, when the depth confidence level is high, the corresponding correction strategy is to perform full-degree-of-freedom correction of the instrument pose based on the three-dimensional spatial pose deviation; when the depth confidence level is medium, the corresponding correction strategy is to limit the three-dimensional spatial correction amplitude and combine it with two-dimensional image pixel deviation for hybrid correction; when the depth confidence level is low, the corresponding correction strategy is to trigger safety hold or request manual intervention.

[0087] like Figure 6As shown, the safety space geometric constraints are a set of geometric constraint systems designed for the narrow, rigid boundary physical environment in intracavitary / peroral surgical scenarios. These include entrance plane constraints, permissible working domain constraints, distance safety boundaries, and prohibited zone constraints. The entrance plane constraints, permissible working domain constraints, and prohibited zone constraints will not be elaborated upon in this embodiment.

[0088] Distance from safety boundary This addresses the need for microscopic avoidance of instruments from intracavitary obstacles and sensitive tissue areas. ,in, The boundary of the work area (inner wall of the support or virtual safety wall). This is the preset minimum safe distance threshold. When < At this time, the system triggers a progressive protection response: speed limiting (the velocity component of the device moving towards the boundary is gradually reduced); stopping (movement in this direction is completely prohibited as it approaches the boundary further); barrier force feedback (if the system has a force feedback master, in...). Approaching A virtual repulsive force is generated pointing towards the safe zone to alert doctors to the boundary.

[0089] Furthermore, to achieve progressive protection of the boundary, this application transforms the safety space geometric constraints into predefined barrier functions:

[0090] .

[0091] in, The shortest distance from the instrument's critical point p to the boundary of the allowable working domain; The shortest distance from the critical point p of the instrument to the boundary of the restricted area; This is the preset minimum safe distance threshold.

[0092] when That is, when the instrument is far from the boundary, the barrier function does not produce active inhibition; when B(p)→ When the device approaches the safety boundary, the system generates a continuous velocity suppression factor based on the reciprocal or exponential function of B(p), so that the motion component of the device towards the boundary is smoothly reduced to zero; when B(p) < 0, that is, the device has entered the danger zone (which should theoretically be prohibited by the pre-verification), the system forcibly triggers the highest priority safety retreat command, pulling the device back to the safety zone along the gradient ascent direction.

[0093] By transforming discrete safety boundaries into continuous control and suppression fields through barrier functions, motion smoothness and system robustness are significantly improved. Furthermore, the embodiments of this application can cover various surgical procedures, including open, intracavitary, and transoral procedures, using the same set of mathematical language (coordinate system + inequalities + barrier function). Scene migration can be achieved through parameter set configuration, demonstrating strong versatility and scalability.

[0094] In this embodiment, the correction amount is a correction vector, which may include position correction amounts (Δx, Δy, Δz) and attitude correction amounts (Δx, Δy, Δz). ), force control target correction amount ( In some implementations, a preliminary pose correction can be generated based on a correction strategy (such as full-degree-of-freedom correction of the instrument pose based on three-dimensional spatial pose deviation). This preliminary pose correction is then substituted into the safety space geometric constraint model for verification. For example, it verifies whether the new pose obtained through correction using the preliminary pose correction is still within the permissible working domain and whether it maintains a distance greater than [missing information] from the restricted area. The distance. If it does not meet the safe space geometric constraints, then an optimization algorithm (such as gradient descent, quadratic programming) is used to solve for a feasible target correction amount that is closest to the initial pose correction amount, while satisfying all space geometric constraints.

[0095] Step S50: Generate a transition trajectory based on the correction amount, and control the surgical instrument to execute the transition trajectory.

[0096] In this embodiment, the transition trajectory refers to a continuous and smooth time sequence (path) connecting the current position / state of the instrument to its position / state after correction. It can specify not only the position but also the velocity, acceleration, and even jerk. In some implementations, trajectory planning algorithms (such as polynomial interpolation or spline curves) can be used to generate a continuous trajectory in terms of velocity and acceleration, with the correction amount as the target and the current position / state as the starting point, while limiting the maximum velocity and acceleration to ensure smooth motion.

[0097] In some implementations, after step S50, the current step node is determined by querying the step timing library, and the next surgical step of the current step node is executed.

[0098] In the intelligent control method for microsurgical robots proposed in this application, in a first aspect, this application uses the currently executed surgical step as the decision context, so that the judgment of visual deviation and force abnormality indicators can be deeply bound to the semantics of specific surgical actions, thereby overcoming the drawbacks of the use of fixed thresholds in the prior art, which leads to sluggish response or frequent false triggering. It realizes the leap from isolated signal processing to contextualized intelligent judgment, and significantly improves the accuracy of abnormality perception and the rationality of decision-making.

[0099] Secondly, by introducing a flexible decision-making mechanism based on deep confidence grading, quantitative assessment and response to uncertainty are achieved. This allows for the autonomous selection of the most appropriate correction strategy based on the reliability of the perceived information itself, rather than simple interruption, when common intraoperative disturbances such as tissue displacement or instrument slippage occur. This not only significantly improves the continuity and fault tolerance of the automated surgical procedure but also ensures that all autonomous corrective actions are performed within strict geometric boundaries through predefined safety space constraints, fundamentally avoiding secondary risks caused by automated intervention.

[0100] Thirdly, this application generates a smooth transition trajectory based on the correction amount and continues it to subsequent steps. After the system completes the correction, it can naturally and smoothly resume the automated process, effectively eliminating the interruption of surgical rhythm and operational redundancy caused by the traditional emergency stop and reset method. This enables the robot system to simulate the coherent assistance capabilities of a senior surgical assistant and maintain a high level of operational efficiency and overall task success rate in a dynamic environment.

[0101] In one embodiment, such as Figure 7 As shown, the method further includes:

[0102] Step A10: Identify and track the pixel coordinates of the instrument tip in the imaging image of the binocular endoscope.

[0103] In this embodiment, the instrument tip refers to the point on the end effector of the surgical instrument that has the most operational semantic characteristics, typically the needle tip of a needle holder, the center of the jaws of a grasping forceps, or the intersection of the blades of scissors. This point is the location with the highest precision requirements in surgical operations and is also the controlled object of visual servo control. The imaging image of a binocular endoscope refers to the two-dimensional grayscale or color image output by either monocular (usually the left) of the binocular endoscope.

[0104] In some implementations, the movement of the instrument tip can be tracked in continuous imaging images by filtering (such as Kalman filtering), thereby mapping the continuous movement of the instrument tip in physical space into a series of discrete, computable pixel coordinate sequences on the image plane, enabling the system to perceive the position of the instrument visually.

[0105] Step A20: Calculate the pixel offset between the pixel coordinates and the center of the image field of view.

[0106] In this embodiment, the image field of view center refers to the geometric center point of the image plane, which typically corresponds to the intersection of the endoscope's optical axis and the image plane, and is the position visually directly facing the lens. Pixel offset refers to the Euclidean distance or component difference between the pixel coordinates of the instrument tip and the image field of view center.

[0107] In some implementations, it is considered that the center of the image field of view may not be the ideal target for retrieval in certain surgical scenarios. For example, during suturing, surgeons often hold the instrument tip one-third of the way below the center of the field of view to allow more space above for observing the needle path. In this case, a weighted center offset can be applied to the image field of view so that the offset image field of view becomes the target point for retrieval, and the pixel coordinates are calculated relative to the pixel offset of the target point.

[0108] It is understandable that if the pixel offset does not exceed the preset offset value, the pose of the binocular endoscope will not be adjusted.

[0109] Step A30: If the pixel offset exceeds the preset offset value, control the binocular endoscope to move along the planned path so that the instrument tip returns to center.

[0110] In this embodiment, the preset offset value is a pre-set threshold parameter, but it is not a fixed constant. Instead, it is dynamically related to the currently performed surgical step. For example, in the suture threading step, which requires extremely high visual accuracy, the threshold can be set to 30px; while in the suture tightening step, which requires a higher global visual field, the threshold can be relaxed to 150px.

[0111] In some implementations, if the pixel offset exceeds a preset offset value, it is further determined whether the binocular endoscope triggers a preset collision conflict or a preset occlusion conflict when performing centering control; if not, the binocular endoscope is controlled to move along the planned path based on the centering control; if so, the optimal observation pose of the binocular endoscope is solved with the optimization goal of maximizing the visibility of the instrument tip and / or maximizing the image coverage of the key area, and with collision avoidance constraints and predefined safe space geometric constraints as conditions, and the binocular endoscope is controlled to move to the optimal observation pose.

[0112] Specifically, centering control refers to controlling the movement of the binocular endoscope to ensure that the instrument tip is positioned at the target centering point within the endoscope's field of view. Preset collision conflict refers to the anticipated physical spatial interference between the endoscope body (including the lens, sheath, and connecting arm) and the following objects if the movement follows the planned path of centering control: surgical instruments (especially end effectors and wrist joints); other cooperating robotic arms; patient anatomical structures (such as body walls and cavity walls); stent entry edges; or other intraoperative fixation devices. Spatial intersection can be performed between the binocular endoscope envelope, the instrument envelope, the stent geometry, and the prohibited zone set; and it can be determined whether, in the entire time domain of the binocular endoscope's centering control trajectory, the minimum distance between any two envelopes is less than a preset minimum safe distance threshold. If so, then a preset collision conflict is triggered.

[0113] Preset occlusion conflict refers to situations where, although no physical collision occurs when the endoscope moves to the target pose controlled by centering, the following visual functions are impaired: the instrument tip is obscured by other parts of the instrument (such as the sheath or light source line); the instrument tip is obscured by tissue, blood, lens fogging, etc.; critical operating areas (such as pinholes, suture loops, suture paths) fall into the edge of the field of view or blind spots; depth confidence decreases significantly due to the deterioration of the observation angle. The endoscope target pose, instrument tip pose, and a known 3D surface model of the tissue (which can be obtained from binocular depth reconstruction) are input into the ray projection engine; the rays emanating from the optical center of the endoscope along the optical axis and the field cone are simulated to determine whether the line of sight connecting the instrument tip and the optical center is truncated by other geometric objects; at the same time, the projection coordinates of the instrument tip under the target pose of the endoscope are calculated to determine whether the projection coordinates are within the effective area of ​​the image and whether they are in the area with high depth confidence; if the line of sight is truncated, the tip projection is located in the 10% area of ​​the image edge, or the estimated depth confidence is lower than the threshold, it is determined that an occlusion conflict has been triggered.

[0114] Maximizing the visibility of the instrument tip aims to ensure that the projection of the instrument tip onto the imaging plane meets the following conditions: it is located within the effective image area (avoiding going out of the frame); it is as close as possible to the center of the image field of view or the target point; it is located in an area with high depth confidence (rich texture, suitable lighting, and no occlusion); and the viewing angle is conducive to subsequent operations (e.g., observing the needle tip along the needle axis). Image coverage of key regions is used to quantify the regional interest associated with the currently performed surgical step. Key regions are dynamically associated with the currently performed surgical step. For example: suturing step: the triangular area formed by the needle hole, needle entry point, and needle exit point; threading step: the complete outline of the suture loop and the center of the needle hole; knotting step: the knot body and the two suture tails. Coverage is used to measure whether these key regions are completely within the field of view, whether they are in a clear area, and whether they are sufficiently magnified.

[0115] Define an evaluation function: This allows us to determine an optimal observation pose that minimizes the weighted sum of all penalty terms (the smaller the value, the better).

[0116] in, Vis is a visibility penalty term, representing the visibility score of the instrument tip, with a value range of [0, 1], where 1 indicates complete visibility and optimal viewing angle, and 0 indicates complete invisibility. In some implementations, Vis can be composed of binary visibility (whether the instrument tip is within the current field of view), center proximity (normalized distance of the tip projection from the center of the image field of view or the target point at the retracement), depth confidence (stereo matching confidence of the area where the tip is located), and viewing angle quality (the angle between the endoscope optical axis and the surface normal vector of the instrument tip; a smaller angle is more conducive to observing fine operations).

[0117] The coverage penalty term is Cov, which is the image coverage score of the key region. The value range is [0, 1], where 1 means that all key regions are completely within the field of view at the best resolution, and 0 means that they are completely missing.

[0118] Mot represents the motion cost penalty, a non-negative real number where a smaller value indicates a gentler, smoother motion. In some implementations, Mot can consist of pose change (the weighted distance between the current endoscope pose and the candidate pose), velocity and acceleration constraints (if the candidate pose differs too much from the current position, causing the planned motion trajectory to exceed preset velocity and acceleration limits, an additional penalty is imposed), and image abrupt change suppression (predicting the degree of image disturbance caused by lens motion using optical flow and introducing corresponding penalty terms). The motion cost penalty term encourages the optimizer to select the feasible pose closest to the current position with the smoothest motion, avoiding frequent and violent lens movements in pursuit of a perfect field of view, which can cause visual dizziness and surgeon discomfort.

[0119] The Occ term is an occlusion penalty term, where Occ is the occlusion penalty value, a non-negative real number. The smaller the value, the less severe the occlusion. Occlusion can originate from: instrument self-occlusion, i.e., the endoscope lens is obstructed by the instrument itself (such as the endoscope arm or light source line); tissue occlusion, i.e., the tip or critical area is obstructed by blood vessels, mucosal flaps, blood stains, etc.; and optical occlusion, i.e., lens fogging or blood staining leading to decreased image quality. The Occ term directly penalizes observation poses where the tip can be "seen" but is partially obstructed by other objects, driving the system to find the observation angle with no occlusion or minimal occlusion.

[0120] As a safety barrier penalty item, This is the safety barrier function, a non-negative real number. The smaller the value, the further away from the safety boundary it is. When the constraint is violated, this term tends to infinity. This term softens hard constraints into strong penalties in the objective function. When a candidate pose violates the safety boundary... A rapid increase in this value drastically degrades the overall score of the pose, leading to its exclusion by the optimizer. When the pose is far from the boundary, this value approaches zero and its impact on the overall objective function is negligible.

[0121] During the solution process, prior constraints can be applied based on the surgical scenario. For example, for an intracavitary scenario, the optimization variables can be reduced from six degrees of freedom to four degrees of freedom, optimizing only the endoscope's advance and retreat along the stent axis, its rotation angle around the axis, and its yaw angles in two orthogonal directions. For an open scenario, a finite search range (e.g., translation ±20mm, rotation ±15°) can be defined centered on the current endoscope pose, reducing global optimization to local optimization.

[0122] In the intelligent control method for the microsurgical robot proposed in this application, firstly, by real-time identification of the pixel coordinates of the instrument tip, calculation of the offset from the center of the field of view, and active control of the endoscope to return to center when the offset exceeds a threshold, the endoscope of the surgical robot possesses autonomous following capability. This mechanism enables the surgeon to ensure that the instrument tip remains in the central area of ​​the field of view without frequently adjusting the lens angle via foot pedals, voice commands, or manual cranks during high-precision and high-continuity operations such as suturing, knotting, and threading.

[0123] Secondly, by introducing a feasibility prediction of centering control, automatic vision tracking is no longer a blind, reckless mechanical pursuit, but a controllable behavior that is always carried out under the constraints of safety boundaries and visual functions, fundamentally eliminating the secondary risks caused by automation intervention.

[0124] Thirdly, when the centering control cannot be executed due to collision or occlusion conflict, the field of view adjustment problem is upgraded to a constrained multi-objective optimization problem, so that the system can still actively find the optimal observation pose that is most beneficial to the surgical operation under the premise of safety when facing the real dilemma of crowded physical space, limited viewpoint and unavoidable conflict.

[0125] In one embodiment, such as Figure 8 As shown, the method further includes:

[0126] Step B10: In response to the manual takeover command, freeze the automated control output based on the step timing library, and transfer control of the surgical instruments and the binocular endoscope to manual operation.

[0127] In this embodiment, the manual takeover command refers to a trigger signal issued by the doctor through the human-computer interaction interface, requesting the system to immediately stop automated execution and transfer control to manual operation. This command can originate from multiple channels: master / slave joystick (the doctor holds the master joystick and applies force or displacement exceeding a threshold, which the system automatically recognizes as an intention to take over); foot switch (the doctor presses a preset "manual priority" foot switch, generating discrete trigger signals); voice commands (the doctor issues preset voice commands such as "pause" or "I'll do it"); touchscreen / control panel (the doctor clicks the "manual takeover" button in the graphical interface); and emergency stop (an independent hard-wired emergency stop circuit is triggered, at which point not only is takeover initiated, but the system also enters a safety lockout state).

[0128] Freezing the automated control output means that the system immediately stops issuing pre-programmed motion commands, force control commands, camera follow commands, etc., generated based on the step timing library to the robot actuator. The trajectory being executed is truncated in the current cycle, and the robot enters a position-holding mode or a compliant control mode, rather than a free-movement state.

[0129] Control transfer refers to the system switching the input source of robot motion commands from the automated planning module to the master-slave mapping module. Afterward, the robot's actuators move entirely according to the doctor's input on the control panel; the system retains only necessary safety boundary constraints and force feedback mappings, and no longer interferes with the intended motion.

[0130] For example, during an automated suture threading operation, the system mistakenly identified the reflective mucus near the suture as the edge of the suture hole and was about to advance the needle tip in the wrong direction. The doctor instantly noticed the abnormality on the monitor and pressed the left foot pedal. From the time the foot pedal signal was issued to the complete freeze of the instrument movement, the measured delay was <50ms, and the needle tip moved only 0.3mm, causing no tissue damage or suture dislocation. The doctor then guided the needle tip to the correct suture hole using the main joystick. After releasing the foot pedal, the system automatically recognized the end of the operation and executed step B20.

[0131] Step B20: After the manual operation is completed, collect the current system state vector.

[0132] In this embodiment, the end of manual operation refers to the moment when the doctor completes the manual intervention and releases control. For example: the doctor releases the main joystick and presses the "Resume Auto" button; the doctor lifts the foot switch; the system detects that there is no operation input on the main joystick within a preset time (e.g., 2 seconds), and automatically determines that the manual operation is over; the doctor issues the voice command "Resume".

[0133] The system state vector is a structured, multi-dimensional dataset used to comprehensively describe the combined state of the surgical robot system, surgical instruments, and surgical environment over a time cross-section. This includes instrument pose, gripping state, target point visibility, force / torque state, and task execution confidence. It is denoted as: = [P, G, V, T, C].

[0134] Wherein, P represents the device pose state, which may include position, posture, wrist joint angle, device opening and closing angle, etc.; G represents the clamping state, including opening and closing amount (mm or percentage), clamping force (N) and clamping effectiveness indicator (binary, a Boolean value based on clamping force, opening amount, and force anomaly index, indicating whether the device effectively clamps the target object); V represents the target point visibility state, including target visibility indicator, pixel confidence and occlusion rate; T represents the force / torque state, including torque vectors of each joint, anomaly indicator and rate of change; C represents the task execution confidence conf, with a value range of [0, 1].

[0135] Step B30: Match the system state vector with the expected state of each surgical step in the step timing library, determine the surgical step with the highest matching degree as the resynchronization node, and execute the next surgical step of the resynchronization node.

[0136] In this embodiment, the expected state is a collective term for the reference state vector and tolerance state vector defined for each surgical step node in the step timing library.

[0137] In the step-series library, the i-th step node can be defined as:

[0138] .

[0139] in, This represents the reference state vector for this step, i.e. the ideal state when performing this step. For example, the reference state for the "S2 Align with the pinhole" step is: the instrument tip is 1mm directly above the pinhole, the angle between the optical axis and the pinhole axis is <5°, the clamping force is stable at 2N, the target point visibility is 1, and the depth confidence is ≥0.7. This represents the tolerance state vector for this step, i.e., the maximum acceptable deviation range of each state component when executing this step, such as position deviation ≤ 0.5mm, attitude deviation ≤ 3°, clamping force fluctuation ≤ 0.3N, etc. TaskPhase indicates the task stage to which the step belongs. Pre / Post indicates the constraint relationship before and after the step.

[0140] By calculating the similarity between the system state vector and the expected state of each step node, a matching score between 0 and 1 can be output for each step node. The step node with the highest matching score is the resynchronization node. The system will resume automated execution based on this resynchronization node, specifically by executing the next surgical step of the resynchronization node.

[0141] Example 1: When the doctor takes over, the system is in the "S3 Suture Advance" step, and the needle tip has partially penetrated the tissue. The doctor manually pulls the needle tip out completely and slightly adjusts the suture tail position. After the manual operation, the system collects the state vector: the instrument tip is 2mm above the needle exit point, the clamping state is normal, the needle hole is visible, the torque is normal, and the depth confidence score is 0.8. Matching calculations show that the current system state highly matches the reference state of the "S4 Suture Retrieval" step (Score=0.91), but deviates significantly from the reference state of the "S3 Advance" step (Score=0.52). The system determines that the doctor has completed the S3 step and directly executes the S4 (resynchronization node) suture retrieval action primitive.

[0142] Example 2: When the doctor takes over, the system is in the "S2 Aligning with Needle Hole" step. The doctor attempts to align manually but finds the needle tip slightly blunted, failing to insert successfully after multiple attempts. The doctor manually moves the instrument to a safe area and releases the lever. After the manual operation, the system collects the state vector: the instrument position is far from the needle hole, the clamping state is TRUE, but the clamping validity flag is set to FALSE due to "suspected slippage" triggered by multiple attempts. Matching calculations show that the scores of all steps requiring "effective clamping" (S2, S3, S4) are severely lowered; while the "S1 Needle Clamping" step has the highest matching score (Score=0.78). The system determines that the current needle tip may be damaged or the clamping is unstable, and suggests reverting to re-clamping the needle (resynchronization node). The interface prompts the doctor, and after confirmation, the system executes the "Release Needle - Re-clamp" sub-process of step S1.

[0143] In some implementations, surgical steps whose state range, formed by the reference state vector and the tolerance state vector, overlaps with each component of the system state vector can be used as candidate step nodes; the normalized deviation between the system state vector and the reference state vector of each candidate step node is calculated in each state dimension; based on the normalized deviation, the matching score of each candidate step node is calculated, and the candidate step node with the highest matching score is determined as the resynchronization node.

[0144] The state range formed by the reference state vector and the tolerance state vector is a high-dimensional hyperrectangular region. If the system state vector falls within this hyperrectangle, it means that the system state vector satisfies the conditions of the step node in that dimension. Overlap means that at least one component of the system state vector falls within the acceptable state interval of the corresponding dimension of the step node.

[0145] Normalized bias is a dimensionless real number used to quantify the degree to which the system state vector deviates from the reference state of the step node in the k-th dimension, and uses the tolerance of this dimension as a metric. Its mathematical definition is: . The larger the value, the greater the degree of mismatch in that dimension.

[0146] For each candidate step node ( ), calculate the matching score, its mathematical definition is: .in The weights for the k-th state dimension satisfy the following conditions: =1; The exponential kernel function normalizes the bias. Mapped to the interval (0,1). The larger, The smaller.

[0147] In the intelligent control method for microsurgical robots proposed in this application, in the first aspect, the current system state vector is systematically collected and structurally expressed at the end of the manual operation. Heterogeneous, asynchronous, and heterogeneous multimodal perception information such as instrument pose, clamping state, target point visibility, force / torque state, and depth confidence is unified into a high-dimensional state space expression that is time-aligned, dimensionally complete, and semantically rich. The unpredictable and difficult-to-model manual operation results of doctors are transformed into a mathematical object that can be quantitatively compared with a predefined step library.

[0148] Secondly, by using a hyperrectangle defined by both the reference state vector and the tolerance state vector as the geometric boundary, hard constraints are used to quickly eliminate step nodes that are physically incompatible with the current state. Then, the original state deviations, which have different dimensions and physical meanings across various dimensions, are unified and normalized into dimensionless, comparable normalized deviation values. Finally, a weighted exponential function is used to aggregate the multidimensional deviation vectors into a single, intuitive matching score, which is used to determine the resynchronization node that best matches the current state. This enables the system to intelligently distinguish between two fundamentally different scenarios: "the doctor has completed the current step and entered the next state" and "the doctor encountered difficulties and returned to a safe position," and to make differentiated responses accordingly. This capability makes human-machine handover no longer an interruption and reset of the process, but a seamless inheritance and intelligent continuation of the task context.

[0149] In one embodiment, a computer storage medium is provided that stores executable instructions that, when executed by a processor, cause the processor to perform the steps in the above method embodiments.

[0150] In one embodiment, an intelligent control system for a microsurgical robot is also provided, the system comprising: a sensing module including a binocular endoscope and a wrist force sensor; an execution module including a robot actuator and surgical instruments; and a decision and control module including a control computing unit for executing the intelligent control method for the microsurgical robot as described in any embodiment of this application.

[0151] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

[0152] Furthermore, those skilled in the art will understand that although some embodiments herein include certain features included in other embodiments but not others, combinations of features from different embodiments are intended to be within the scope of this application and form different embodiments. For example, any of the embodiments or implementations claimed above can be used in any combination. The information disclosed in this background section is intended only to enhance the understanding of the general background of this application and should not be construed as an admission or in any way implying that such information constitutes prior art known to those skilled in the art.

Claims

1. An intelligent control system for a microsurgical robot, characterized in that, include: The sensing module includes a binocular endoscope and a wrist force sensor; The execution module includes the robotic actuator and surgical instruments; The decision and control module includes a control computing unit for executing an intelligent control method for a microsurgical robot, the method comprising: A standardized sequence of steps is constructed for each surgical subtask in the target surgical scenario. At least one action primitive is associated with each surgical step in the sequence of steps. Corresponding transition conditions and preset fallback strategies are set for each surgical step. When performing each surgical step based on a preset step timing library, visual information and multidimensional force information for the surgical instruments are acquired simultaneously through a binocular endoscope and a wrist force sensor. The step timing library includes step sequences, action primitives, transition conditions, and rollback strategies. The visual deviation at the tip of the computing device is calculated based on the visual information, and the force anomaly index is calculated based on the multidimensional force information. The binocular endoscope is subjected to binocular positioning and epipolar correction, and the parallax search range of stereo matching is limited to the parallax interval corresponding to the preset working distance range of the surgical robot; Perform stereo matching within the disparity interval to acquire disparity data; The depth information of the instrument tip is calculated based on the parallax data; Depth confidence is calculated based on at least one of the following: left-right consistency in the stereo matching process, uniqueness of the matching result, texture features of the image region, and temporal consistency of depth information. Using the currently performed surgical step as the decision context, and based on the visual deviation and the force abnormality index, it is determined whether the preset abnormality triggering conditions are met. If so, the currently executed surgical step is paused, and the corresponding correction strategy is determined according to the depth confidence level of the binocular endoscope. Under the predefined safety space geometric constraints, the correction amount corresponding to the correction strategy for correcting the instrument pose or operating parameters is determined. A transition trajectory is generated based on the correction amount, and the surgical instrument is controlled to execute the transition trajectory; The step of determining the corresponding correction strategy based on the depth confidence level of the binocular endoscope includes: when the depth confidence level is high, determining the corresponding correction strategy as full-degree-of-freedom correction of the instrument pose based on the three-dimensional spatial pose deviation; when the depth confidence level is medium, determining the corresponding correction strategy as limiting the three-dimensional spatial correction amplitude and combining it with the two-dimensional image pixel deviation for hybrid correction; and when the depth confidence level is low, determining the corresponding correction strategy as triggering safety hold or requesting manual intervention.

2. The intelligent control system of the microsurgical robot according to claim 1, characterized in that, The method further includes: Identify and track the pixel coordinates of the instrument tip in the imaging image of the binocular endoscope; Calculate the pixel offset between the pixel coordinates and the center of the image field of view; If the pixel offset exceeds the preset offset value, the binocular endoscope is controlled to move along the planned path so that the instrument tip returns to center.

3. The intelligent control system of the microsurgical robot according to claim 2, characterized in that, Controlling the movement of the binocular endoscope along a planned path includes: Determine whether the binocular endoscope triggers a preset collision conflict or a preset occlusion conflict when it performs centering control. If not, then control the binocular endoscope to move along the planned path based on centering control; If so, the optimal observation pose of the binocular endoscope is solved by taking the maximization of the visibility of the instrument tip and / or the maximization of the image coverage of the key area as the optimization objectives, and taking collision avoidance constraints and predefined safety space geometric constraints as conditions. The binocular endoscope is then controlled to move to the optimal observation pose.

4. The intelligent control system of the microsurgical robot according to claim 1, characterized in that, The surgical sub-tasks include at least one of threading, suturing, and knotting; The action primitives include at least one of the following: a pose trajectory primitive for controlling the surgical instrument, a clamping primitive, a force control primitive, and a lens follow-up primitive for controlling the binocular endoscope. The transition conditions include at least one of the following: instrument positioning threshold, visual alignment threshold, force threshold, time threshold, and confidence threshold; The retraction strategy includes at least one of re-clamping, retracting to a safe position, re-performing the alignment action, and switching to manual control mode.

5. The intelligent control system of the microsurgical robot according to claim 1, characterized in that, The method further includes: In response to a manual takeover command, the automated control output based on the step timing library is frozen, and control over the surgical instruments and the binocular endoscope is transferred to manual operation. After manual operation is completed, the current system state vector is collected. The system state vector includes the instrument pose, clamping state, target point visibility, force / torque state, and task execution confidence. The system state vector is matched with the expected state of each surgical step in the step timing library. The surgical step with the highest matching degree is determined as the resynchronization node, and the next surgical step of the resynchronization node is executed.

6. The intelligent control system of the microsurgical robot according to claim 5, characterized in that, The expected state includes a reference state vector and a tolerance state vector. Matching the system state vector with the expected states of each surgical step in the step timing library to determine the surgical step with the highest matching degree as the resynchronization node includes: Surgical steps whose state range, formed by the reference state vector and the tolerance state vector, overlaps with the components of the system state vector are selected as candidate step nodes. Calculate the normalized deviation of the system state vector from the reference state vector of each candidate step node in each state dimension; Based on the normalized deviation, the matching score of each candidate step node is calculated, and the candidate step node with the highest matching score is determined as the resynchronization node.

7. A computer storage medium, characterized in that, The storage medium stores executable instructions, which, when executed by a processor, cause the processor to perform an intelligent control method for a microsurgical robot, the method comprising: A standardized sequence of steps is constructed for each surgical subtask in the target surgical scenario. At least one action primitive is associated with each surgical step in the sequence of steps. Corresponding transition conditions and preset fallback strategies are set for each surgical step. When performing each surgical step based on a preset step timing library, visual information and multidimensional force information for the surgical instruments are acquired simultaneously through a binocular endoscope and a wrist force sensor. The step timing library includes step sequences, action primitives, transition conditions, and rollback strategies. The visual deviation at the tip of the computing device is calculated based on the visual information, and the force anomaly index is calculated based on the multidimensional force information. The binocular endoscope is subjected to binocular positioning and epipolar correction, and the parallax search range of stereo matching is limited to the parallax interval corresponding to the preset working distance range of the surgical robot; Perform stereo matching within the disparity interval to acquire disparity data; The depth information of the instrument tip is calculated based on the parallax data; Depth confidence is calculated based on at least one of the following: left-right consistency in the stereo matching process, uniqueness of the matching result, texture features of the image region, and temporal consistency of depth information. Using the currently performed surgical step as the decision context, and based on the visual deviation and the force abnormality index, it is determined whether the preset abnormality triggering conditions are met. If so, the currently executed surgical step is paused, and the corresponding correction strategy is determined according to the depth confidence level of the binocular endoscope. Under the predefined safety space geometric constraints, the correction amount corresponding to the correction strategy for correcting the instrument pose or operating parameters is determined. A transition trajectory is generated based on the correction amount, and the surgical instrument is controlled to execute the transition trajectory; The step of determining the corresponding correction strategy based on the depth confidence level of the binocular endoscope includes: when the depth confidence level is high, determining the corresponding correction strategy as full-degree-of-freedom correction of the instrument pose based on the three-dimensional spatial pose deviation; when the depth confidence level is medium, determining the corresponding correction strategy as limiting the three-dimensional spatial correction amplitude and combining it with the two-dimensional image pixel deviation for hybrid correction; and when the depth confidence level is low, determining the corresponding correction strategy as triggering safety hold or requesting manual intervention.

8. The computer storage medium according to claim 7, characterized in that, The method further includes: Identify and track the pixel coordinates of the instrument tip in the imaging image of the binocular endoscope; Calculate the pixel offset between the pixel coordinates and the center of the image field of view; If the pixel offset exceeds the preset offset value, the binocular endoscope is controlled to move along the planned path so that the instrument tip returns to center.

9. The intelligent control system of the microsurgical robot according to claim 8, characterized in that, Controlling the movement of the binocular endoscope along a planned path includes: Determine whether the binocular endoscope triggers a preset collision conflict or a preset occlusion conflict when it performs centering control. If not, then control the binocular endoscope to move along the planned path based on centering control; If so, the optimal observation pose of the binocular endoscope is solved by taking the maximization of the visibility of the instrument tip and / or the maximization of the image coverage of the key area as the optimization objectives, and taking collision avoidance constraints and predefined safety space geometric constraints as conditions. The binocular endoscope is then controlled to move to the optimal observation pose.

10. The computer storage medium according to claim 7, characterized in that, The surgical subtask includes at least one of threading, suturing, and knotting; the action primitives include at least one of primitives for controlling the pose trajectory of the surgical instruments, clamping primitives, force control primitives, and lens follow-up primitives for controlling the binocular endoscope; the transition conditions include at least one of instrument positioning threshold, visual alignment threshold, force threshold, time threshold, and confidence threshold; the retraction strategy includes at least one of re-clamping, retracting to a safe position, re-performing the alignment action, and switching to manual control mode; the method further includes: In response to a manual takeover command, the automated control output based on the step timing library is frozen, and control over the surgical instruments and the binocular endoscope is transferred to manual operation. After manual operation is completed, the current system state vector is collected. The system state vector includes the instrument pose, clamping state, target point visibility, force / torque state, and task execution confidence. The system state vector is matched with the expected state of each surgical step in the step timing library. The surgical step with the highest matching degree is determined as the resynchronization node, and the next surgical step of the resynchronization node is executed.

Citation Information

Patent Citations

  • Trocar penetration depth indicator and guide tube positioning device

    CA2022069A1