A surgical robot suture trajectory planning method based on a world model and reinforcement learning

CN122805374APending Publication Date: 2026-09-25CHONGQING UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202611316345.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-28
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

现有方案通常在检测到组织状态变化后再修正轨迹,属于事后补偿,难以在动作执行前预判后续状态,因此容易出现轨迹调整滞后的问题

Benefits of technology

本发明通过构建融合缝合视觉信息、组织形变信息与机器人运动状态的统一缝合状态特征,并利用世界模型对未来多步的视觉特征变化、组织形变状态、缝合阶段变化以及视觉伺服映射修正量进行预测,使缝合轨迹规划能够预先感知软组织形变等未来状态变化,解决了现有技术难以预测软组织未来变化的问题,提高了缝合轨迹与组织实际状态的适配性;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122805374A_ABST
    Figure CN122805374A_ABST
Patent Text Reader

Abstract

The present application relates to a kind of surgical robot suture trajectory planning method based on world model and reinforcement learning, belong to robot autonomous operation and trajectory planning technical field.The method includes: S1, according to endoscope image and robot motion state, construct unified suture state feature in combination with suture visual information and tissue deformation information;S2, based on the state feature and candidate end pose increment sequence, predict future multi-step visual feature, tissue deformation and suture stage change using world model;S3, by reinforcement learning strategy, generate target end pose based on prediction result;S4, according to target end pose and visual servo mapping correction amount, adaptively correct visual servo mapping and closed loop execution;S5, update observation state and prediction result online.The present application predicts the future state change of soft tissue through world model, realizes trajectory planning and visual servo cooperation, can compensate deformation and occlusion mapping deviation, adapts six-stage suture process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of robot autonomous operation and trajectory planning technology, and relates to a surgical robot suture trajectory planning method based on world model and reinforcement learning. Background Technology

[0002] Suturing is a crucial step in surgical procedures. With the development of surgical robotics, utilizing robots for autonomous or semi-autonomous suturing has become an important direction for improving suturing accuracy, stability, and safety. In robotic surgery, the robotic arms need to perform actions such as needle insertion, withdrawal, traction, clamping, and knotting within a confined surgical space. Their motion trajectory must not only ensure that the instrument tip reaches the target position but also meet constraints such as reachability and telecentric motion, avoiding interference between robotic arms, preventing collisions between instrument arms and non-target tissues, and minimizing the impact of instrument movement on tissue edges and suture condition. In multi-arm suturing scenarios, trajectory planning also needs to coordinate the relative motion relationships between the needle-holding instruments, auxiliary traction instruments, sutures, needles, and target tissue. Therefore, the safety of trajectory planning directly affects the surgical robot's operational accuracy, tissue damage risk, and suturing efficiency.

[0003] In existing surgical robot suture trajectory planning schemes, the spatial state of the surgical area is typically determined first using intraoperative images, 3D image models, or preset suture task information. Then, the motion path of the robotic arm is generated based on the suture target and instrument motion constraints. For suture tasks, existing schemes generally identify the relative positions of the suture object, the suture needle, and the needle holder, calculate the instrument trajectory required for the suture needle to enter and exit the tissue, and for subsequent traction processes, and complete trajectory tracking through visual feedback or a robot control system. For example, Chinese patent CN115624387A discloses a suture control method, control system, readable storage medium, and robot system. This scheme first acquires the planned suture trajectory, then drives the first needle holder to clamp the suture needle and insert it into the suture object. It also acquires tissue morphology information based on real-time images of the suture area to update the suture trajectory, allowing the suture needle to exit from a predetermined position. Chinese patent CN113499166A discloses an autonomous stereoscopic vision navigation method and system for a corneal transplantation surgical robot. This method uses a stereomicroscope camera to acquire reference and target images, and extracts visual features through a convolutional neural network to obtain real-time target 3D information and surgical tool pose, thereby guiding needle insertion and suture knotting. Chinese patent CN114129263B discloses a surgical robot path planning method, system, device and storage medium. It formulates the operation path based on the three-dimensional image model of the target area, and determines whether the robot can complete the path planning after registration during the operation. If the path cannot be planned, the object to be operated on is adjusted and the calculation is recalculated.

[0004] On the other hand, research on incorporating world models and reinforcement learning into surgical robot control to handle dynamic environments is also gradually emerging. For example, Chinese patent CN120791749A discloses a surgical robot control method based on world models and reinforcement learning, which uses world models to predict changes in the state of the surgical environment and inputs the prediction results along with the current state parameters into a reinforcement learning-based control strategy optimization model to output an action selection strategy; Chinese patent CN122473841A discloses a robot visual servo control method and system driven by a world action model, which uses a world model to predict the desired visual state based on the current observed image sequence, task commands, and historical action sequences to drive visual servoing.

[0005] However, the aforementioned existing technologies still have the following shortcomings when applied to autonomous soft tissue suturing scenarios: First, the ability to predict future changes in soft tissue is insufficient. After the suture needle contacts the tissue, the tissue will undergo changes such as compression, slippage, and local bulging, and suture traction will also change the shape of the wound edges. Current methods usually correct the trajectory after detecting changes in tissue state, which is a post-event compensation and makes it difficult to predict subsequent states before the action is performed, thus easily leading to the problem of trajectory adjustment lag.

[0006] Second, the representation of the suturing state is incomplete. The suturing process involves not only the positional change of the suture needle relative to the target tissue, but also the morphology of the tissue edges, the function of the suture, the robot's end-effector pose, and the current stage of the suturing action. Existing solutions often only utilize local target information in the image, lacking a mechanism to jointly represent the visual state, tissue state, and robot motion state into a unified suturing state. This makes it difficult for the robot to fully understand the relationship between the current action and subsequent tissue changes when planning its trajectory.

[0007] Third, there is insufficient adaptability to changes in the suturing stage. The suturing process consists of continuous actions such as approaching the tissue, adjusting posture, inserting the needle into the tissue, advancing the needle, exiting the tissue, and pulling the suture. The movement objectives and safety constraints are different at different stages. If only a preset procedure or fixed trajectory is relied upon, it is easy to cause unstable transitions between stages, affecting the continuity and accuracy of the suturing action.

[0008] Fourth, there is insufficient coordination between trajectory planning and visual servoing. Visual servoing can make local corrections based on current image errors, but the mapping relationship between visual features and robot end effector motion is affected by tissue deformation, viewpoint changes, and instrument occlusion. If this mapping relationship cannot be adjusted according to changes in tissue state, it will reduce needle tip alignment accuracy and affect the stable execution of subsequent suturing trajectories.

[0009] In view of the above, how to enable the surgical robot to generate a safe, stable and adaptive suturing trajectory under the circumstances of continuous changes in tissue morphology and unstable local visual information during the autonomous suturing of soft tissue has become a technical problem that urgently needs to be solved in this field. Summary of the Invention

[0010] The present invention aims to solve the problem of how to predict future changes in soft tissue state, such as deformation, during suturing operations, and to achieve coordination between suturing trajectory planning and visual servo execution based on the prediction results, thereby improving the accuracy of surgical robot suturing trajectory planning and its adaptability to different suturing stages.

[0011] To address the aforementioned technical problems, this invention provides a surgical robot suture trajectory planning method based on a world model and reinforcement learning, comprising the following steps: S1. Obtain the current suturing status: Based on the endoscopic images and the robot's motion status, combined with suturing visual information and tissue deformation information, a unified suturing status feature is constructed; S2. World model prediction: Based on the unified suture state features and candidate end pose increment sequence, the world model is used to predict the visual feature changes, tissue deformation state, suture stage changes and visual servo mapping corrections in the future multiple steps, and the current potential state of the world model is obtained. S3. Reinforcement Learning Planning: The prediction results of the world model are used as the input of the future state. The reinforcement learning policy network generates the target end pose at the next moment and outputs the policy intervention weights. S4. Visual servoing execution: Based on the target end pose and the visual servoing mapping correction amount predicted by the world model, the basic visual servoing mapping matrix is ​​adaptively corrected, and closed-loop execution is completed. S5. Online Update: Update the observation state, the potential state of the world model, and the prediction results corresponding to the incremental sequence of each candidate terminal pose in each control cycle.

[0012] Furthermore, the method for constructing the unified suture state features described in S1 includes: inputting the endoscopic image, robot joint position, robot joint velocity, end effector pose, suture visual features, tissue visual features, suture target features, current suture stage, and visual servo execution error of the previous control cycle into a multimodal encoder to obtain the suture visual encoding vector, tissue deformation encoding vector, and robot-target conditional encoding vector, and then fusing the three into a unified suture state feature.

[0013] Further, the world model prediction in S2 includes: inputting the unified stitching state features into the latent dynamics module, combining the hidden state of the previous time step and the actual executed action of the previous time step to obtain the current world model latent state; generating multiple sets of candidate end pose increment sequences by the candidate action generation part in the reinforcement learning policy network based on the unified stitching state features, the current world model latent state, and the visual servoing execution error of the previous control cycle; using the current world model latent state and each set of candidate end pose increment sequences as conditions, performing rolling prediction of the future latent state corresponding to each candidate end pose increment sequence through a local temporal prediction head; and decoding the future latent state into the key visual feature prediction results, tissue deformation state prediction results, stitching stage prediction results, prediction uncertainty, and visual servoing mapping correction amount corresponding to each candidate end pose increment sequence through a multi-head prediction decoder.

[0014] Furthermore, the multimodal encoder includes a suture visual feature module, a tissue deformation module, a robot-target condition module, and a multimodal fusion head; the suture visual feature module is implemented by a combination of convolutional neural network and visual Transformer, and is used to extract visual features of local areas of the suture needle, suture thread, and needle holder; the tissue deformation module is used to extract the local deformation state of flexible tissue from tissue visual features, continuous image changes, and actual robot actions; the robot-target condition module is used to encode the robot motion state and suture target information into conditional features; and the multimodal fusion head is used to fuse the suture visual encoding vector, tissue deformation encoding vector, and robot-target condition encoding vector into a unified suture state feature.

[0015] Furthermore, the multi-head prediction decoder includes a visual feature prediction head, a tissue deformation prediction head, a suturing stage prediction head, an uncertainty estimation head, and a visual servo mapping head. The visual feature prediction head predicts the changing trends of the needle tip, needle body direction, target tissue region, and local suture region in subsequent images. The tissue deformation prediction head predicts tissue surface displacement and local morphological changes. The suturing stage prediction head determines whether the subsequent state belongs to the tissue approach, posture adjustment, tissue puncture, needle advancement, needle tip exit, or suture traction stage. The uncertainty estimation head indicates the reliability of the current prediction result. The visual servo mapping head describes the impact of flexible tissue deformation on the visual servo execution relationship.

[0016] Furthermore, the reinforcement learning policy network described in S3 is trained using a soft actor-critic algorithm, which includes two sequentially executed parts: candidate action generation and candidate action evaluation. The candidate action generation part generates the multiple sets of candidate end pose increment sequences. The candidate action evaluation part evaluates the candidate actions and determines the target end pose increment to be used in the current control cycle by combining the future prediction results of the world model corresponding to each candidate action.

[0017] Furthermore, the candidate action evaluation section calculates the evaluation value of each candidate action based on the prediction uncertainty in the future prediction results corresponding to each candidate action and the predicted organizational risk obtained from the predicted organizational deformation state. The first end pose increment in the candidate end pose increment sequence with the highest evaluation value is taken as the target end pose increment of the current control cycle, and the strategy intervention weight is output. The strategy intervention weight is used to adjust the degree of influence of the visual servo mapping correction amount predicted by the world model on the basic visual servo mapping matrix.

[0018] Further, the visual servoing execution in S4 includes: mapping the target end effector pose to target visual features in image space; fusing the currently detected visual features with the visual features predicted by the world model in the previous control cycle based on the visibility of key visual features, and calculating the current visual servoing execution error; updating the basic visual servoing mapping matrix based on the visual servoing mapping correction amount predicted by the world model and the policy intervention weight, to obtain an updated visual servoing mapping matrix to compensate for the visual-motion mapping deviation caused by soft tissue deformation, viewpoint changes, and instrument occlusion; and generating an end effector speed command based on the updated visual servoing mapping matrix and the current visual servoing execution error.

[0019] Furthermore, the suturing stage is divided into six stages: approaching the tissue, posture adjustment, tissue puncture, needle advancement, needle tip exit, and suture traction. When the distance between the needle tip and the target insertion point is less than a preset distance threshold, the posture adjustment stage is entered. When both the needle tip position error and posture error are less than the corresponding thresholds, the tissue puncture stage is entered. When the needle tip penetration depth or visual penetration amount reaches a preset threshold, the needle advancement stage is entered. When the needle body rotation angle or arc-shaped advancement amount reaches a preset threshold, the needle tip exit stage is entered. When the needle tip becomes visible again in the target exit area and the position error is less than a preset position error threshold, the suture traction stage is entered. When the suture displacement, wound edge spacing, or traction state reaches preset conditions, the current suturing is completed.

[0020] This invention also provides a system for implementing the above-mentioned surgical robot suture trajectory planning method, comprising: a state representation module, used to construct a unified suture state feature based on endoscopic images and robot motion state, combined with suture visual information and tissue deformation information; a world model module, used to predict changes in visual features, tissue deformation state, suture stage changes, and visual servo mapping corrections for future multiple steps based on the unified suture state feature and candidate end-effector pose increment sequences; a reinforcement learning strategy module, used to generate the target end-effector pose for the next moment based on the prediction results of the world model module; a visual servo execution module, used to adaptively correct the basic visual servo mapping matrix and complete closed-loop execution based on the target end-effector pose and the visual servo mapping corrections predicted by the world model module; and an online update module, used to update the observation state, the potential state of the world model, and the prediction results corresponding to each candidate end-effector pose increment sequence in each control cycle.

[0021] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention constructs a unified suturing state feature that integrates suturing visual information, tissue deformation information, and robot motion state. It also uses a world model to predict changes in visual features, tissue deformation state, suturing stage changes, and visual servo mapping corrections in the future, enabling suturing trajectory planning to anticipate future changes in soft tissue deformation and other future state changes. This solves the problem that existing technologies have difficulty predicting future changes in soft tissue and improves the adaptability of suturing trajectory to the actual state of the tissue. This invention uses the prediction results of the world model as the input of the future state to the reinforcement learning policy network for trajectory planning, and uses the visual servo mapping correction amount predicted by the world model to adaptively correct the basic visual servo mapping matrix. This achieves the synergy between trajectory planning and visual servo execution, and can compensate for the visual-motion mapping deviation caused by soft tissue deformation, viewpoint changes and instrument occlusion, thereby improving the accuracy of suture trajectory planning. This invention uses a suture stage prediction head to predict and judge six stages: approaching tissue, posture adjustment, tissue puncture, needle advance, needle tip exit, and suture traction. This enables the suture trajectory planning to adapt to the state changes of different suture stages, thus improving its adaptability to different suture stages. This invention improves the safety of suturing operations by introducing predictive uncertainty and predictive tissue risk into the candidate action evaluation mechanism through predictive uncertainty estimation and candidate action evaluation.

[0022] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description

[0023] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein: Figure 1 This is a flowchart of the surgical robot suture trajectory planning method based on world model and reinforcement learning according to the present invention; Figure 2 This is a general framework diagram of the surgical robot suture trajectory planning method based on world model and reinforcement learning of the present invention. Detailed Implementation

[0024] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0025] Example 1 like Figure 2 As shown, this embodiment proposes a surgical robot suture trajectory planning method based on a world model and reinforcement learning. This method addresses scenarios such as continuous changes in tissue state during soft tissue suture, shifts in the target needle entry and exit points with contact, potential occlusion of the needle tip in local stages, and further alteration of the wound edge position due to suture traction. To solve these problems, this invention does not employ a simple geometric path planning approach, nor does it use a method of directly outputting low-level control commands from reinforcement learning. Instead, it constructs a hierarchical trajectory planning method composed of a world model, a reinforcement learning strategy, and a visual servoing execution unit.

[0026] In this method, the current observation state is first constructed based on endoscopic images, robot joint states, end effector poses, suturing target features, and the current suturing stage. This is then combined with historical observation states within a certain time range and the robot's actual actions to form temporal state information. A multimodal encoder extracts suturing visual information, tissue deformation information, and conditional information between the robot and the suturing target from this temporal state information, and fuses them into a unified suturing state feature. The latent dynamics module obtains the current world model latent state based on the unified suturing state feature, the previous hidden state, and the previous actual action.

[0027] Subsequently, the candidate action generation part of the reinforcement learning strategy generates several sets of candidate end-effector pose increment sequences, hereinafter referred to as candidate action sequences, based on the current unified stitching state features and the current world model latent state. The world model, using each set of candidate end-effector pose increment sequences as conditions, predicts key visual features, tissue deformation state, stitching stage, prediction uncertainty, and visual servo mapping corrections within several future time steps. The candidate action evaluation part of the reinforcement learning strategy determines the optimal candidate action sequence based on each candidate action sequence and its corresponding future prediction results, and uses the first end-effector pose increment in this optimal candidate action sequence as the target end-effector pose increment for the current control cycle.

[0028] Finally, the visual servoing unit constructs target visual features based on the target end effector pose, and uses the tissue deformation prediction results corresponding to the selected candidate action sequence and the visual servoing mapping correction amount to correct the mapping relationship between the current visual feedback and the end effector motion, generating end effector speed commands and end effector pose control commands. After the robot completes the action of the current control cycle, it re-acquires endoscopic images and robot state, calculates the actual executed actions, updates the observed state, the potential state of the world model, and historical state information, and continues planning, switching stages, or ending the complete suturing process based on the completion status of the current suturing stage.

[0029] The suturing process in this invention is uniformly divided into six stages: approaching the tissue, adjusting the posture, puncturing the tissue, advancing the needle, withdrawing the needle tip, and pulling the suture.

[0030] I. Acquisition and Characterization of Suture Status Data During the suturing task performed by the surgical robot, the following information is acquired at the current moment: endoscopic image, robot joint state, robot end effector pose, suturing visual features, tissue visual features, suturing target features, visual servo execution error of the previous control cycle, and visibility information of key visual features. The set of observation states at the current moment can be represented as:

[0031] in, This represents the set of observed states at the current moment. Indicates the current time step. This represents the endoscopic image at the current moment. This represents the robot's joint position variable at the current moment. This represents the robot's joint velocity variable at the current moment. This indicates the current pose of the robot's end effector. This represents the visual characteristics of suturing at the current moment, including a comprehensive visual representation of the local areas of the suture needle, suture thread, and needle holder. It represents the visual characteristics of the organization at the current moment, which includes a comprehensive visual representation of the organization's surface texture, edges, and local morphology. This indicates the current suture target features, which include the current suture entry point and the current suture exit point. Represents variables at the suture stage. This indicates the visual servo execution error of the previous control cycle. A visibility mask representing the current key visual features.

[0032] The pose of the robot's end effector is represented by a homogeneous transformation matrix:

[0033] in, This indicates the current pose of the end effector. This represents the attitude rotation matrix of the end effector relative to the robot's base coordinate system at the current moment. This represents the position vector of the end effector relative to the robot's base coordinate system at the current moment. In this way, the position and orientation of the end effector are uniformly incorporated into subsequent world model predictions, reinforcement learning policy outputs, and visual servo control execution.

[0034] II. World Model Construction The world model is used to map the current observed state to the stitched potential state and predict the future stitched state. It consists of four parts: a multimodal encoder, a latent dynamics module, a local temporal prediction head, and a multi-head prediction decoder.

[0035] 1. Multimodal encoder The multimodal encoder consists of a stitching visual feature module, a tissue deformation module, a robot-target condition module, and a multimodal fusion head.

[0036] The suture visual feature module is implemented using a combination of a convolutional neural network and a visual Transformer. The convolutional network is used to extract local texture and edge information, while the visual Transformer is used to establish spatial relationships between local regions of the suture needle, suture thread, and needle holder. The input to this module is a continuous image sequence and its corresponding suture visual features; the output is a suture visual encoding vector.

[0037] in, This represents the stitching visual encoding vector at the current moment. This indicates the visual feature module for stitching; Indicates from Time's up A sequence of continuous endoscopic images at different times. Indicates from Time's up A continuous stitched visual feature sequence at any given moment.

[0038] The tissue deformation module extracts the local deformation state of flexible tissues from tissue visual features, continuous image changes, and actual robot actions. This module takes as input a sequence of tissue visual features, a sequence of endoscopic images, and a historical sequence of actual pose increments of the end effector, and outputs a tissue deformation encoding vector.

[0039] in, This represents the organizational deformation encoding vector at the current moment. This indicates the tissue deformation module. This represents the sequence of tissue visual features from time tL to time t. Indicates from t L time to t The sequence of end-effector pose increments actually executed by the robot at time 1.

[0040] The robot-target conditional module encodes the robot's motion state and stitching target information into conditional features. These features constrain the future prediction direction of the world model, ensuring that the prediction results align with the current robot pose and the stitching task objective. Its encoding vector is represented as:

[0041] in, This represents the robot-target conditional encoding vector at the current moment. This represents the robot-target condition coding submodule.

[0042] The multimodal fusion head integrates the suture visual encoding vector, tissue deformation encoding vector, and robot-target conditional encoding vector into a unified suture state feature, represented as:

[0043] in, This represents the characteristics of the uniform stitching state at the current moment. This indicates a multimodal fusion module. This indicates a feature concatenation operation. This represents the stitching visual encoding vector. This represents the organization deformation encoding vector. This represents the robot-target conditional encoding vector.

[0044] To balance image information representation capabilities with online inference overhead, in a set of feasible network configurations, the stitching visual encoding vector is set to 128 dimensions, the tissue deformation encoding vector is set to 128 dimensions, and the robot-target conditional encoding vector is set to 64 dimensions. These three are concatenated to form a 320-dimensional intermediate feature, which is then mapped to a 256-dimensional unified stitching state feature by a multimodal fusion head. The hidden states of the latent dynamics module and the latent states of the world model are both set to 256 dimensions. The above dimensions are a set of implementation configurations for this stitching task, and can be adjusted according to image resolution, computational resources, and training data size.

[0045] 2. Potential Dynamics Module The latent dynamics module describes the changes in the stitching state over time. Its inputs are the current unified stitching state features, the previous hidden state, and the robot's actual actions performed in the previous control cycle. The outputs are the current hidden state and the current latent state of the world model. The current hidden state is updated as follows:

[0046] in, Indicates the current hidden state. This represents the state update function in the potential dynamics submodule. This indicates the hidden state in the previous moment. This represents the actual end-effector pose increment executed by the robot in the previous control cycle. The current world model latent state is represented as:

[0047] in, This represents the potential state of the current world model. This represents the potential state generation function.

[0048] 3. Candidate Action Generation and Local Temporal Prediction To avoid a circular dependency between world model prediction and reinforcement learning action selection, this invention first generates M sets of candidate end-point pose increment sequences based on the current unified stitching state features, the current world model latent state, and the visual servoing execution error of the previous control cycle, using the candidate action generation part of the reinforcement learning policy network. The m-th candidate action sequence is represented as:

[0049] in, Indicates at time The generated first Group of candidate action sequences, Indicates at time Planning, intended for future use The candidate end pose increment, Indicates the number of time steps to predict the future. This indicates the number of candidate action sequences.

[0050] Since this method needs to sequentially complete candidate action generation, world model rolling prediction, and candidate evaluation within an online planning cycle, the prediction range H and the number of candidate sequences M are preferably determined according to the available inference time per cycle. In a feasible configuration, H=10 and M=16 can be chosen, that is, 16 sets of candidate end pose increment sequences are batch rolled to predict the next 10 state update steps. If the single state update cycle is... The corresponding predicted coverage duration is If the average time for a single-step potential prediction is... The available time budget for world model predictions is Then, during serial estimation, the following conditions are met. In parallel batch processing, the actual batch processing latency is used as a constraint. When computational resources are limited, M can be reduced; when organizational deformation or occlusion increases prediction uncertainty but still meets real-time requirements, M can be increased.

[0051] The candidate action generation process is represented as follows:

[0052] in, This represents a reinforcement learning policy model with parameter θ. The candidate action generation process does not use future predictions that have not yet been obtained, thus ensuring a clear sequential relationship between candidate action generation and world model predictions.

[0053] The local temporal prediction head, based on the current potential state of the world model and the candidate action sequences of each group, performs rolling predictions of future potential states. The initial prediction state and recursive process corresponding to the m-th candidate action sequence are expressed as follows:

[0054]

[0055] in, This represents the future predicted based on the m-th candidate action sequence at the current time. The potential state at any given moment This represents a local time series prediction function.

[0056] 4. Multi-head prediction decoder The multi-head prediction decoder decodes the future potential states corresponding to each group of candidate actions into prediction results that can be used for reinforcement learning and visual servoing. It includes a visual feature prediction head, a tissue deformation prediction head, a suture stage prediction head, an uncertainty estimation head, and a visual servoing mapping head.

[0057] The system includes several key components: a visual feature prediction head to predict the changing trends of the needle tip, needle body direction, target tissue area, and local suture area in subsequent images; a tissue deformation prediction head to predict tissue surface displacement and local morphological changes; a suturing stage prediction head to determine whether the subsequent state is approaching tissue, adjusting posture, puncturing tissue, advancing the needle, withdrawing the needle tip, or pulling the suture; an uncertainty estimation head to indicate the reliability of the current prediction results; and a visual servo mapping head to describe the impact of flexible tissue deformation on the visual servo execution relationship, enabling the visual servo execution unit to update the control mapping based on changes in tissue state.

[0058] For the m-th candidate action sequence, its future prediction result is expressed as:

[0059] in, The world model is represented in the 1st century. The output future prediction results under the condition of a group of candidate action sequences This represents the key visual features obtained from the prediction. This indicates the predicted state of tissue deformation. This indicates the predicted suture stage. This indicates the uncertainty of the prediction result. This represents the predicted visual servo mapping correction amount. Indicates the index of the future prediction time step. Indicates the time range for the forecast.

[0060] The world model is trained using historical suturing videos, robot state records, and sequences of actual robot actions. During training, the sequences of actual actions in the historical records are used as action conditions, while the visual features, tissue deformation states, and suturing stages observed after the actions are performed are used as supervision information.

[0061] During online planning, the trained world model is applied to each group of candidate action sequences to obtain future state predictions before the candidate actions are actually executed. The world model training loss is expressed as:

[0062] in, This represents the training loss of the world model. Represents the weights of the visual feature prediction loss. Represents true visual features. Represents the predicted visual features, This indicates the weighting of the predicted loss due to tissue deformation. It represents the actual state of tissue deformation. This indicates the predicted state of tissue deformation. This indicates the weight of the predicted loss during the stitching phase. Represents the cross-entropy loss function. Indicates the actual suturing stage. Indicates the predicted suture stage. This indicates the uncertainty of the forecast. This represents the weight of the loss due to prediction uncertainty. and These represent the actual and predicted visual servo mapping corrections, respectively. Indicates the visual servo mapping loss weights. Indicates the time range for the forecast.

[0063] In a preferred training configuration, a world model training sample consists of a data window spanning 10 consecutive world model time points. This window includes endoscopic images, robot state, actual actions performed, tissue state supervision information, and the next time-to-be-observed time point for five consecutive historical time points and five subsequent predicted time points. The total number of effective training windows after time synchronization, outlier frame removal, and annotation quality checks is set to 30,000, derived from at least 300 independent suture trajectories. Each of the six stages—tissue approach, posture adjustment, tissue puncture, needle advancement, needle tip exit, and suture traction—contains at least 5,000 effective windows. The training, validation, and test sets are divided in an 8:1:1 ratio according to independent suture trajectories, tissue samples, or experimental batches, containing 24,000, 3,000, and 3,000 windows respectively. Adjacent windows of the same continuous trajectory must not be divided across sets.

[0064] III. Reinforcement Learning Strategy Network This invention employs the Soft Actor-Critic (SAC) algorithm to train a reinforcement learning policy model. The reinforcement learning policy comprises two sequentially executed parts: candidate action generation and candidate action evaluation. The candidate action generation part generates multiple sets of continuous action sequences based on the current state; the candidate action evaluation part combines the future prediction results of the world model corresponding to each candidate action to evaluate the candidate actions and determine the action to be used in the current control cycle.

[0065] For the m-th candidate action sequence, the reinforcement learning evaluation input is represented as:

[0066] in, Indicates the first The reinforcement learning evaluation status of the candidate actions. The evaluation value of the candidate action. The parameter is Value assessment network Indicates the first Uncertainty statistics of candidate actions within the prediction time range. This represents the predicted tissue risk obtained based on the predicted tissue deformation state and candidate needle tip trajectories. and These represent the evaluation weights for predicting uncertainty and predicting organizational risk, respectively.

[0067] The target end-effector pose increment used in the current control cycle is directly taken from the optimal action of the selected candidate action sequence:

[0068] in, This represents the target end-effector pose increment. Indicates the position increment of the end effector. This represents the attitude increment of the end effector. Simultaneously, it outputs the policy intervention weights based on the selected candidate actions and their future predictions. This is used to adjust the degree to which the visual servo mapping correction amount predicted by the world model affects the underlying visual servo mapping matrix.

[0069] To ensure that the policy output directly corresponds to the robot's executable end-effector pose and avoids the output remaining at an abstract action category, this invention generates the target end-effector pose for the next time step by using the current end-effector pose and the pose increment of the policy output:

[0070] in, Indicates the target end effector pose at the next moment. This indicates the current pose of the end effector. Indicates the increment of pose The constructed Lie algebra matrix, This represents a matrix exponential mapping used to convert pose increments in Lie algebra form into rigid body pose transformations.

[0071] Instead of being sent directly to the robot control system as the final control command, it serves as the planning target for the visual servoing execution unit. The visual servoing execution unit performs closed-loop correction of the target pose based on the current visual feedback, and finally generates the end effector pose command to be sent to the robot control system.

[0072] During reinforcement learning training, the robot performs the current action and calculates the reward after receiving the next observation. The reward function is:

[0073] in, This indicates the reward for the current control period. Indicates the visual error weight. This represents the squared L2 norm of the visual servoing execution error obtained after the current action is executed. Indicates the smoothing weight of the action. Indicates organizational risk weights. This represents the organizational risk assessment quantity observed after the action is performed, used to characterize the risk of excessive compression, excessive stretching, or abnormal deformation. Indicates the weight of the reward upon completion of the stage. This indicates the contribution of the current action to achieving the goal of the stitching phase. This represents the penalty weight for prediction uncertainty. It represents the statistical measure of the uncertainty of the world model's prediction of the future of candidate actions.

[0074] In a set of preferred training configurations, the SAC policy model is trained for 5000 rounds, with 200 world model control moments executed per round, corresponding to no more than 1,000,000 environmental interactions.

[0075] The experience replay buffer capacity is set to 1,000,000, the random exploration warm-up steps are set to 20,000, the batch size is set to 256, and the discount factor is set. Set to 0.99, target network soft update coefficient The learning rate was set to 0.005. Both the actor network and the critic network were two-layer fully connected networks with 256 units per layer, and the learning rate was set to [value missing]. An evaluation is performed every 100 rounds on 20 independent validation trajectories.

[0076] Training is stopped when the task completion rate improvement is less than 1 percentage point in 5 consecutive evaluations, the closed-loop position error is no more than 1.5 mm, and no safety constraint violations occur. The checkpoint with the highest verification task completion rate is then selected.

[0077] Each reward sub-item is normalized to the nearest whole number before weighting. The weights are set as visual error weights. Motion smoothing weights Organizational risk weights Stage completion reward weight Prediction uncertainty penalty weight During the tissue puncture and needle advancement phases, Increased to 4.0. When the normalized organizational risk is not less than 0.70, an additional termination penalty of -5.0 will be imposed and the current candidate action will be rejected.

[0078] The policy network optimization objective is expressed as:

[0079] in, This represents the optimization objective of the policy network. This represents the expectation of the state input and policy output. Represents the entropy temperature coefficient. This represents the candidate action sequence sampled by the policy model. This represents the future prediction result obtained by the world model under the conditions of this action sequence.

[0080] IV. Constructing Visual Servo Targets and Visual Servo Execution To convert the target end-effector pose in robot space into target visual features in image space, so that the reinforcement learning policy output can be used by the visual servoing execution unit, the target end-effector pose corresponding to the selected candidate action is mapped to the visual servoing target:

[0081] in, Indicates the visual features of the target at the next moment. This represents the projection mapping function from the robot's end-effector pose to the visual features of the image. Indicates camera intrinsic and extrinsic parameters. This indicates the organizational deformation state at the next moment corresponding to the selected candidate action.

[0082] When key visual features in the current image are visible, the current visual features obtained by the visual detection module are used; when the needle tip, needle body, or target area is partially occluded, the predicted visual features of the current moment are supplemented by the world model of the previous control cycle.

[0083] The current visual features used for visual servoing are represented as follows:

[0084] in, This represents the key visual features detected in the current image. This represents the prediction result of the world model from the previous control cycle for the visual features at the current moment. Indicates a mask for the visibility of key visual features. This indicates element-wise multiplication.

[0085] Visual servo execution error Represented as:

[0086] The visual servoing actuator establishes a mapping relationship between changes in visual features and the motion of the end effector:

[0087] in, Indicates the rate at which visual features change over time. This represents the visual servo mapping matrix at the current moment. This indicates the current speed command of the end effector.

[0088] Update the visual servo mapping matrix based on the visual servo mapping correction amount corresponding to the selected candidate action and the intervention weights of the reinforcement learning policy:

[0089] in, This represents the updated visual servo mapping matrix. This represents the basic visual servoing mapping matrix obtained based on the camera model and the robot's kinematics model. This represents the visual servo mapping correction for the next time step predicted by the world model. This correction is used to compensate for visual-motion mapping biases caused by soft tissue deformation, viewpoint changes, and local occlusion.

[0090] Based on visual servoing execution errors, generate end effector speed commands:

[0091] in, This indicates the current end effector speed command. Indicates the visual servo convergence gain. Represents the visual servo mapping matrix The pseudo-inverse matrix. When the current visual feature deviates from the target visual feature, the visual servoing execution unit calculates the end-effector velocity based on the updated visual servoing mapping relationship, thereby gradually reducing the visual error.

[0092] Generate the next moment's end effector pose command based on the end effector velocity command:

[0093] in, This indicates the pose command of the end effector to be sent to the robot control system at the next moment. , Indicates speed command The constructed Lie algebra matrix, Indicates the control cycle. This represents the matrix exponential mapping.

[0094] V. Action Execution and Online Updates The robot control system controls the movement of the end effector according to the end position command at the next moment, so that the needle holder can drive the suture needle to complete the current stage of action, including the needle tip approaching the target insertion point, the needle tip posture adjustment, tissue puncture, needle advance, needle tip exit, and suture traction.

[0095] After the robot executes the current end-effector pose command, it re-acquires endoscopic images, robot joint states, end-effector pose, suture visual features, tissue visual features, and suture target features, and forms the next set of observation states:

[0096] in, Represents the set of observed states at the next moment. This indicates the endoscopic image at the next moment. This represents the robot's joint position variable at the next moment. This represents the robot's joint velocity variable at the next moment. This indicates the pose of the robot's end effector at the next moment. This indicates the visual features of the stitching at the next moment. Indicates the visual characteristics of the organization at the next moment. This indicates the target feature to be stitched at the next moment. Indicates the next moment in the suturing phase. This indicates the visual servo execution error at the current moment. A visibility mask representing the current key visual features.

[0097] Next time observation state set The data is written to the historical state buffer, and the earliest state exceeding the historical window length is deleted to form an updated historical state buffer. Subsequently, the stitching visual encoding, tissue deformation encoding, robot-target conditional encoding, unified stitching state features, and world model latent states are recalculated based on the updated historical states.

[0098] During the approach to tissue stage, when the distance between the needle tip and the target insertion point is less than a preset distance threshold, the posture adjustment stage is entered. During the posture adjustment phase, when both the needle tip position error and posture error are less than the corresponding threshold, the tissue puncture phase begins. For the tissue puncture stage, when the needle tip penetration depth or visual penetration amount reaches the preset threshold, the needle advancement stage begins. During the needle advance stage, when the needle body rotation angle or arc advance amount reaches the preset threshold, the needle tip exit stage begins. For the needle tip exit stage, when the needle tip becomes visible again in the target exit area and the position error is less than the preset position error threshold, the suture traction stage begins. During the suture traction phase, the current stitch is completed when the suture displacement, wound edge spacing, or traction status reaches the preset conditions.

[0099] If the goal of the current stitching stage has not been achieved, the trajectory planning for the current stage continues; if the goal of the current stitching stage has been achieved, the process switches to the next stitching stage; if the complete stitching sequence has been completed, the autonomous stitching trajectory planning process ends.

[0100] When the organizational risk exceeds the safety threshold, visual features are not visible for an extended period, or the robot's control state is abnormal, at least one of the following safety operations is performed: deceleration, pausing, or returning to the previous safe pose.

[0101] Example 2 like Figure 1 As shown, in one specific embodiment, the suture trajectory planning method of the present invention includes the following steps: Step 1: Obtain the endoscope image, robot joint position, robot joint velocity, end effector pose, suturing visual features, tissue visual features, suturing target features, current suturing stage, and visual servo execution error of the previous control cycle at the current moment, construct the current observation state set, and update the historical state buffer.

[0102] Step 2: Input the current observation state set into the multimodal encoder to obtain the suture visual encoding vector, tissue deformation encoding vector, and robot-target conditional encoding vector, and fuse them to obtain a unified suture state feature.

[0103] Step 3: Input the unified stitching state features, the hidden state of the previous moment, and the actual actions performed by the robot in the previous moment into the latent dynamics module to obtain the current hidden state and the current world model latent state.

[0104] Step 4: Input the current unified stitching state features, the current world model latent state, and the visual servoing execution error of the previous control cycle into the candidate action generation part of the reinforcement learning strategy to generate multiple sets of candidate end pose increment sequences.

[0105] Step 5: Input the incremental pose sequences of each candidate end pose into the local temporal prediction head of the current world model potential state, and predict the future potential state sequence corresponding to each candidate action sequence.

[0106] Step 6: Input the future potential state sequence into the multi-head prediction decoder to obtain the key visual features, tissue deformation state, suturing stage, uncertainty and visual servo mapping correction amount corresponding to each candidate action sequence.

[0107] Step 7: Input the current unified stitching state features, the current potential state of the world model, each candidate action sequence, the future prediction results corresponding to each candidate action, and the visual servo execution error of the previous control cycle into the candidate action evaluation part, calculate the candidate action evaluation value, determine the optimal candidate action sequence number, and output the strategy intervention weight.

[0108] Step 8: Take the first end-effector pose increment in the optimal candidate action sequence as the target end-effector pose increment for the current control cycle, and generate the next time-phase target end-effector pose based on the current end-effector pose.

[0109] Step 9: Map the planned target end pose and the tissue deformation prediction results corresponding to the selected candidate actions to target visual features in the image space; based on the visibility of key visual features, fuse the currently detected visual features with the visual features predicted by the world model at the current moment in the previous control cycle, and calculate the current visual servo execution error.

[0110] Step 10: Update the visual servo mapping matrix according to the visual servo mapping correction amount and policy intervention weight corresponding to the selected candidate action, and generate the end effector speed command according to the visual servo execution error.

[0111] Step 11: Apply robot workspace, telecentric motion, instrument collision and organizational safety constraints to the end effector speed command, generate the end effector pose control command for the next moment and send it to the robot control system for execution.

[0112] Step 12: After the robot executes the action, it re-acquires the observation state at the next moment, calculates the current actual action based on the actual end pose before and after execution, updates the historical state buffer, and determines whether to continue the current suturing stage, switch to the next suturing stage, end the complete suturing sequence, or enter a safe pause state.

[0113] Example 3 This embodiment provides a surgical robot suture trajectory planning system based on world model and reinforcement learning, including a state representation module, a world model module, a reinforcement learning strategy module, a visual servo execution module, and an online update module.

[0114] The data flow relationships between the modules are as follows: The state representation module outputs the unified stitching state features obtained by fusion to the world model module and the reinforcement learning strategy module respectively; the world model module outputs the future prediction results corresponding to each candidate action sequence to the reinforcement learning strategy module, and outputs the visual servo mapping correction amount and tissue deformation state prediction results to the visual servo execution module; the reinforcement learning strategy module outputs the target end pose and policy intervention weights to the visual servo execution module; the visual servo execution module outputs the end pose control command to the robot control system for execution; the online update module feeds back the updated observation state, historical state information and world model potential state to the state representation module and the world model module, thereby forming a data closed loop of state representation, future prediction, action planning, closed-loop execution and online update in each control cycle.

[0115] The state representation module is used to construct a unified suturing state feature based on endoscopic images and robot motion state, combined with suturing visual information and tissue deformation information. Its input signals are the current endoscopic image, robot joint position, robot joint velocity, end effector pose, suturing visual features, tissue visual features, suturing target features, current suturing stage, and visual servo execution error of the previous control cycle. Its output signal is the unified suturing state feature, which is transmitted to the world model module as a condition for future prediction and to the reinforcement learning strategy module as a state input for candidate action generation. Specifically, it implements the suturing state data acquisition and representation process described in the first part, as well as the encoding and fusion process of the multimodal encoder described in the second part.

[0116] The world model module is used to predict changes in visual features, tissue deformation state, suturing stage changes, and visual servo mapping corrections in the future based on the unified suturing state features and candidate end-effector pose increment sequences. Its input signals are the unified suturing state features output by the state representation module, the hidden state at the previous moment, the robot's actual actions at the previous moment, and multiple sets of candidate end-effector pose increment sequences generated by the reinforcement learning strategy module. Its output signals are the key visual feature prediction results, tissue deformation state prediction results, suturing stage prediction results, prediction uncertainties, and visual servo mapping corrections corresponding to each candidate action sequence. The future prediction results are transmitted to the reinforcement learning strategy module for candidate action evaluation, and the visual servo mapping corrections and tissue deformation state prediction results are transmitted to the visual servo execution module for visual servo mapping correction. Specifically, it implements the functions of the latent dynamics module, the local temporal prediction head, and the multi-head prediction decoder.

[0117] The reinforcement learning strategy module is used to generate the target end effector pose for the next moment based on the prediction results of the world model module. Its input signals are the unified stitching state features, the current world model latent state, the future prediction results corresponding to each candidate action sequence, and the visual servoing execution error of the previous control cycle. Its output signals are the target end effector pose increment and the strategy intervention weight for the current control cycle. After the target end effector pose increment and the current end effector pose are synthesized to obtain the target end effector pose for the next moment, it is transmitted to the visual servoing execution module along with the strategy intervention weight. This module specifically implements the candidate action generation and candidate action evaluation process.

[0118] The visual servoing execution module is used to adaptively correct the basic visual servoing mapping matrix and complete closed-loop execution based on the target end-effector pose and the visual servoing mapping correction amount predicted by the world model module. Its input signals are the target end-effector pose and policy intervention weights output by the reinforcement learning policy module, the visual servoing mapping correction amount and tissue deformation state prediction results output by the world model module, and the key visual features detected in the current endoscopic image. Its output signals are the end-effector speed command and the end-effector pose control command generated by it for the next moment. The pose control command is sent to the robot control system to drive the needle holder to move the suture needle to complete the current suturing stage action. Specifically, it realizes the construction of the visual servoing target and the visual servoing execution process.

[0119] The online update module is used to update the observation state, the potential state of the world model, and the prediction results corresponding to each candidate end-effector pose increment sequence in each control cycle. Its input signals are the observation state at the next moment after the robot executes the current pose control command, and the actual execution action calculated based on the actual end-effector pose before and after execution. Its output signals are the updated historical state buffer, unified stitching state features, and the potential state of the world model, which are fed back to the state representation module and the world model module for state representation and future prediction in the next control cycle. At the same time, it outputs a stitching stage switching signal to indicate whether to continue the current stitching stage, switch to the next stitching stage, end the complete stitching sequence, or enter a safe pause state. Specifically, it implements the action execution and online update process.

[0120] Example 4 The specific network structure of the world model can be alternatively implemented. In other embodiments, the multimodal encoder can be implemented using convolutional neural networks, visual Transformers, temporal convolutional networks, recurrent neural networks, state-space models, or combinations of the above networks; the latent dynamics module can be implemented using recurrent neural networks, gated recurrent units, long short-term memory networks, Transformer temporal models, state-space models, or diffusion prediction models; the multi-head predictive decoder can also be replaced with a shared decoder plus a structure with multiple task branches. The adopted world model can conditionally predict the future stitching state corresponding to each candidate action sequence based on the current stitching state, historical actual actions, and candidate end pose increment sequences.

[0121] Example 5 Alternative methods can be used to characterize the soft tissue deformation state. In other implementations, methods such as tissue key point displacement fields, dense optical flow, three-dimensional point cloud deformation fields, finite element approximation parameters, implicit neural field parameters, tissue edge curve changes, or probability distribution-based deformation state representation methods can be used. The tissue deformation state characterization adopted can reflect the state changes of soft tissue during suture needle contact, passage, and suture traction, and can be used as input for subsequent trajectory planning or visual servo correction.

[0122] Example 6 Alternative approaches can be used to execute visual servoing. In other implementations, image-based visual servoing, position-based visual servoing, hybrid visual servoing, learning-based visual servoing, or model predictive control visual servoing can be employed; the visual servoing mapping can also be obtained from a camera model, a robot kinematics model, an online estimation model, or a neural network model.

[0123] Example 7 Alternative methods can be used to acquire suture status information and extract features. In other implementations, suture status can be obtained using monocular endoscopic images, binocular endoscopic images, structured light images, depth images, intraoperative ultrasound images, force feedback information, or multi-sensor fusion information; suture visual features can be extracted using traditional image processing methods, deep learning segmentation models, key point detection models, target tracking models, or visual basic models.

[0124] Example 8 Alternative reinforcement learning policy algorithms can be employed. In other implementations, proximal policy optimization algorithms, deep deterministic policy gradient algorithms, double-delay deep deterministic policy gradient algorithms, maximum entropy reinforcement learning algorithms, model prediction reinforcement learning algorithms, offline reinforcement learning algorithms, or algorithms combining imitation learning and reinforcement learning can be used.

[0125] Candidate action generation and evaluation can be performed jointly by the same policy network, or separately by the candidate action generation network and the candidate action evaluation network. When implemented separately, the candidate action generation network generates multiple sets of candidate end-effector pose increment sequences, while the candidate action evaluation network evaluates and selects each candidate action sequence based on the current stitching state, the candidate action sequences, and the corresponding world model prediction results. The reinforcement learning policy can directly output the optimal candidate sequence number, or it can output the evaluation value, selection probability, or weighting coefficient of each candidate sequence. After determining the optimal candidate action sequence, the first end-effector pose increment in the selected candidate sequence can be used as the target end-effector pose increment for the current control cycle, or a range-constrained residual correction amount can be output based on the selected action.

[0126] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A surgical robot suture trajectory planning method based on world model and reinforcement learning, characterized in that, Includes the following steps: S1. Obtain the current suturing status: Based on the endoscopic images and the robot's motion status, combined with suturing visual information and tissue deformation information, a unified suturing status feature is constructed; S2. World model prediction: Based on the unified suture state features and candidate end pose increment sequence, the world model is used to predict the visual feature changes, tissue deformation state, suture stage changes and visual servo mapping corrections in the future multiple steps, and the current potential state of the world model is obtained. S3. Reinforcement Learning Planning: The prediction results of the world model are used as the input of the future state. The reinforcement learning policy network generates the target end pose at the next moment and outputs the policy intervention weights. S4. Visual servoing execution: Based on the target end pose and the visual servoing mapping correction amount predicted by the world model, the basic visual servoing mapping matrix is ​​adaptively corrected, and closed-loop execution is completed. S5. Online Update: Update the observation state, the potential state of the world model, and the prediction results corresponding to the incremental sequence of each candidate terminal pose in each control cycle.

2. The surgical robot suture trajectory planning method according to claim 1, characterized in that, The method for constructing the unified suture state features described in S1 includes: The endoscopic image, robot joint position, robot joint velocity, end effector pose, suture visual features, tissue visual features, suture target features, current suture stage, and visual servo execution error of the previous control cycle are input into the multimodal encoder to obtain the suture visual encoding vector, tissue deformation encoding vector, and robot-target conditional encoding vector, and then the three are fused into a unified suture state feature.

3. The surgical robot suture trajectory planning method according to claim 2, characterized in that, The world model predictions described in S2 include: The unified stitching state features are input into the latent dynamics module, and combined with the hidden state of the previous moment and the actual action executed in the previous moment, the current world model latent state is obtained. The candidate action generation part in the reinforcement learning policy network generates multiple sets of candidate end pose increment sequences based on the unified stitching state features, the current world model latent state, and the visual servo execution error of the previous control cycle. Using the current potential state of the world model and the candidate end pose increment sequences as conditions, the future potential state corresponding to each candidate end pose increment sequence is predicted in a rolling manner through a local temporal prediction head; The future potential state is decoded by a multi-head predictive decoder into the prediction results of key visual features, tissue deformation state, suturing stage, prediction uncertainty, and visual servo mapping correction amount corresponding to each candidate end pose increment sequence.

4. The surgical robot suture trajectory planning method according to claim 2, characterized in that, The multimodal encoder includes a suture visual feature module, a tissue deformation module, a robot-target condition module, and a multimodal fusion head; The suture visual feature module is implemented by a combination of convolutional neural network and visual Transformer, and is used to extract visual features of local areas of suture needle, suture thread and needle holder; The tissue deformation module is used to extract the local deformation state of flexible tissue from tissue visual features, continuous image changes, and actual robot actions. The robot-target condition module is used to encode the robot's motion state and stitching target information into conditional features; The multimodal fusion head is used to fuse the suture visual encoding vector, tissue deformation encoding vector, and robot-target conditional encoding vector into a unified suture state feature.

5. The surgical robot suture trajectory planning method according to claim 3, characterized in that, The multi-head prediction decoder includes a visual feature prediction head, a tissue deformation prediction head, a suture stage prediction head, an uncertainty estimation head, and a visual servo mapping head. The visual feature prediction head is used to predict the changing trends of the needle tip, needle body direction, tissue target area, and suture local area in subsequent images. The tissue deformation prediction head is used to predict tissue surface displacement and local morphological changes; The suturing stage prediction head is used to determine whether the subsequent state belongs to the stage of approaching tissue, posture adjustment, tissue puncture, needle advancement, needle tip exit, or suture traction. The uncertainty estimation header is used to indicate the reliability of the current prediction result; The visual servo mapping head is used to describe the impact of flexible tissue deformation on the visual servo execution relationship.

6. The surgical robot suture trajectory planning method according to claim 3, characterized in that, The reinforcement learning policy network described in S3 is trained using the soft actor-critic algorithm, which includes two sequentially executed parts: candidate action generation and candidate action evaluation. The candidate action generation section generates the multiple sets of candidate end-effector pose increment sequences; The candidate action evaluation section combines the future prediction results of the world model corresponding to each candidate action to evaluate the candidate action and determine the target end pose increment to be used in the current control cycle.

7. The surgical robot suture trajectory planning method according to claim 6, characterized in that, The candidate action evaluation section calculates the evaluation value of each candidate action based on the prediction uncertainty in the future prediction results corresponding to each candidate action and the predicted organizational risk obtained from the predicted organizational deformation state. The first end pose increment in the candidate end pose increment sequence with the highest evaluation value is taken as the target end pose increment of the current control cycle, and the strategy intervention weight is output. The strategy intervention weights are used to adjust the degree to which the visual servo mapping correction predicted by the world model affects the underlying visual servo mapping matrix.

8. The surgical robot suture trajectory planning method according to claim 1, characterized in that, The visual servoing execution described in S4 includes: The target end pose is mapped to target visual features in image space; Based on the visibility of key visual features, the currently detected visual features are fused with the visual features predicted by the world model in the previous control cycle to calculate the current visual servo execution error. Based on the visual servo mapping correction amount and strategy intervention weight predicted by the world model, the basic visual servo mapping matrix is ​​updated to obtain the updated visual servo mapping matrix, in order to compensate for the visual-motion mapping deviation caused by soft tissue deformation, viewpoint changes and instrument occlusion. Based on the updated visual servo mapping matrix and the current visual servo execution error, generate end effector speed commands.

9. The surgical robot suture trajectory planning method according to claim 1, characterized in that, The suturing process is divided into six stages: approaching the tissue, adjusting the posture, puncturing the tissue, advancing the needle, withdrawing the needle tip, and pulling the suture. When the distance between the needle tip and the target insertion point is less than a preset distance threshold, the posture adjustment phase begins. When both the needle tip position error and the posture error are less than the corresponding threshold, the tissue puncture stage begins. When the needle tip penetration depth or visual penetration amount reaches the preset threshold, the needle advance stage begins. When the needle body rotation angle or arc advance amount reaches the preset threshold, the needle tip exit stage begins. When the needle tip becomes visible again in the target needle exit area and the positional error is less than the preset positional error threshold, the suture traction stage begins. When the suture displacement, wound edge spacing, or traction status reaches the preset conditions, the current stitch count is completed.

10. A system for implementing the surgical robot suture trajectory planning method according to any one of claims 1 to 9, characterized in that, include: The state characterization module is used to construct a unified suturing state feature based on endoscopic images and robot motion state, combined with suturing visual information and tissue deformation information; The world model module is used to predict changes in visual features, tissue deformation, suturing stage changes, and visual servo mapping corrections in the next few steps based on the unified suture state features and candidate end pose increment sequences. The reinforcement learning strategy module is used to generate the target end pose at the next moment based on the prediction results of the world model module; The visual servoing execution module is used to adaptively correct the basic visual servoing mapping matrix and complete closed-loop execution based on the target end pose and the visual servoing mapping correction amount predicted by the world model module. The online update module is used to update the observation state, the potential state of the world model, and the prediction results corresponding to the incremental sequence of each candidate terminal pose in each control cycle.

Citation Information

Patent Citations

  • Autonomous stereoscopic vision navigation method and system for corneal transplantation surgical robot

    CN113499166A

  • Surgical robot path planning methods, systems, equipment and storage media

    CN114129263B

  • Stitching control method, control system, readable storage medium and robot system

    CN115624387A

  • Surgical robot control method and system based on world model and reinforcement learning and robot system

    CN120791749A

  • Visual servo control method and system for robot driven by world action model

    CN122473841A