Surgical robot motion planning method based on diffusion model

Through the motion planning method of surgical robots based on diffusion model, combined with physical constraints and advanced diffusion model architecture, the complexity and stability of the motion planning of surgical robots in complex environments is solved, and efficient, stable and safe motion trajectory generation is achieved.

CN120038756AInactive Publication Date: 2025-05-27BEIJING ROSSUM ROBOT TECH CO LTD

Patent Information

Application Number
CN202510394590.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-05-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The motion planning of surgical robots in complex surgical environments faces complexity, collision risks, and high requirements for real-time and stability caused by high degrees of freedom. The efficiency and stability of existing methods in high-dimensional space still need to be improved, and physical constraints and environmental dynamic changes are ignored.

Method used

A surgical robot motion planning method based on diffusion model is proposed, combining physical constraints and advanced diffusion model architecture, through point cloud encoder, the transformer structure of encoder only and the comprehensive loss function, the efficiency and stability of motion planning are improved, and the smoothness and safety of the planning path are ensured.

Benefits of technology

It significantly improves the speed and quality of the motion planning of surgical robots, enhances the stability and safety of the planned path, and is suitable for complex and dynamic surgical environments and different surgical tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120038756A_ABST
    Figure CN120038756A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of surgical robot motion control, and discloses a surgical robot motion planning method based on a diffusion model, and the method comprises the steps: obtaining obstacle point cloud data, an initial posture and a target posture in a surgical environment; encoding the obstacle point cloud data into a potential space, and combining the encoded information with the initial attitude and the target attitude into a condition code; a Transform structure of an encoder is adopted to replace a U-Net structure in a traditional diffusion model, and the diffusion model is trained through forward diffusion and reverse denoising processes; in training, using a comprehensive loss function to optimize model performance, combining configuration space loss, geometric task space loss and collision loss, and introducing physical constraints to ensure that the trajectory conforms to kinematic characteristics and avoid collision; and in the operation task, generating a motion track based on the trained diffusion model. According to the invention, the motion planning efficiency and safety of the surgical robot in a complex environment can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of surgical robot motion control, and more specifically, to a surgical robot motion planning method based on a diffusion model. Background Art

[0002] Surgical robots have important applications in minimally invasive surgery due to their high precision and flexibility. However, the motion planning of surgical robots in complex surgical environments faces many challenges, including the complexity of motion planning caused by high degrees of freedom (DOFs), the risk of collision between surgical instruments and human tissue, and high requirements for real-time performance and stability. Although traditional motion planning methods, such as the sampling-based RRT* algorithm and deep learning-based motion planning networks, have solved these problems to a certain extent, their efficiency and stability in high-dimensional space still need to be improved. In addition, existing methods often ignore the physical constraints of surgical robot motion and the dynamic changes of the environment, resulting in insufficient feasibility and safety of planning results. Summary of the invention

[0003] The purpose of this invention is to propose a surgical robot motion planning method based on a diffusion model, which improves the motion planning efficiency and stability of the surgical robot in complex environments by combining physical constraints and an advanced diffusion model architecture, while ensuring the smoothness and safety of the planned path.

[0004] To achieve the above object, the present invention proposes a surgical robot motion planning method based on a diffusion model, comprising:

[0005] Obtain obstacle point cloud data in the surgical environment, the initial posture and target posture of the surgical robot;

[0006] A point cloud encoder is used to encode the obstacle point cloud data in the surgical environment and compress the obstacle point cloud data into a latent space to extract feature information related to motion planning.

[0007] Combining the encoded obstacle information, the initial posture and the target posture into a conditional code;

[0008] The encoder-only Transformer structure is used to replace the U-Net structure in the traditional diffusion model, forming a diffusion model with the encoder-only Transformer structure as the core;

[0009] The conditional code is input into the diffusion model and trained through forward diffusion and reverse denoising processes. During the training process, a comprehensive loss function combining configuration space loss, geometric task space loss and collision loss is used to optimize the model performance to improve the quality of the generated trajectory. Meanwhile, in the reverse denoising process, physical constraints are introduced to modify the prediction results of the intermediate trajectory to ensure that the generated motion trajectory conforms to the kinematic characteristics of the surgical robot and avoids collisions with obstacles in the surgical environment.

[0010] Under the condition of a given surgical task, the motion trajectory of the surgical robot is generated by completing the trained diffusion model.

[0011] Optionally, the encoder-only Transformer structure captures the temporal dependency of the surgical robot motion sequence by:

[0012] The input includes the time step t and the condition code C, where the condition code C contains the encoded obstacle information, initial posture and target posture;

[0013] Use multi-head self-attention mechanism and feed-forward network to process input data and gradually generate motion trajectories;

[0014] The formula is:

[0015] z tk =MLP(t)+Embed(C)

[0016] Where: z tk is the label input to the Transformer, which represents the embedded representation at time step t; MLP is a multi-layer perceptron, which is used to map the time step t to the model dimension; Embed is an embedding operation, which maps the conditional code C to the model dimension; C is the conditional code, which contains the encoded obstacle information, the initial posture and the target posture of the surgical robot.

[0017] Optionally, the point cloud encoder adopts a compressed autoencoder structure, and the point cloud encoder is trained by the following formula:

[0018]

[0019] Among them: AE is the loss function of the encoder, which is used to train the point cloud encoder; K is the number of points in the point cloud data I; i is a point in the original point cloud; is the corresponding point in the reconstructed point cloud; λ is the regularization coefficient, which is used to control the regularization term of the parameter; θ ij are the parameters in the encoder and decoder.

[0020] Optionally, the forward diffusion process is implemented by the following formula:

[0021]

[0022] Where: z t is the noise data at time step t; α t is the noise coefficient, which controls the amount of noise added at time step t; z t-1 is the noise data at time step t-1; ∈ t is the Gaussian noise added at time step t, which obeys the standard normal distribution N(0, I); I is the unit matrix, representing the covariance matrix of the noise.

[0023] Optionally, the reverse denoising process gradually restores the motion trajectory through the following formula:

[0024]

[0025] in: is the trajectory after denoising at time step t-1; α t is the noise coefficient, which controls the amount of noise removed at time step t; is the trajectory at time step t; ∈ t is the noise predicted at time step t.

[0026] Optionally, the comprehensive loss function optimizes model performance by the following formula:

[0027] L=λ joint ·L joint +λ point ·L point +λ collision ·L collision

[0028] Where: L is the comprehensive loss function used to optimize model performance; λ joint is the weight coefficient of the configuration space loss; point is the weight coefficient of the geometric task space loss; λ collision is the weight coefficient of collision loss; L joint It is the configuration space loss used to optimize the joint rotation error of the surgical robot. The calculation formula is:

[0029]

[0030] Where: N is the number of time steps of the trajectory; is the joint angle of the generated trajectory at time step i; is the joint angle of the target trajectory at time step i; L point is the geometric task space loss, which is used to optimize the cumulative error of the kinematic chain. Its calculation formula is:

[0031]

[0032] Where: FK is the forward kinematics function, which maps the joint angle to the position of the end effector; L collision is the collision loss, which is used to ensure that the generated trajectory avoids collision with obstacles. Its calculation formula is:

[0033]

[0034] Where: h is the collision detection function; I is the obstacle point cloud data.

[0035] Optionally, the physical constraint implements collision detection through the following formula:

[0036]

[0037] Where: h is the collision detection function, which is used to detect the distance between the surgical robot point cloud and the obstacle point cloud; w s is a point in the surgical robot point cloud; I is the obstacle point cloud data; D(w s , I) is the minimum distance between the surgical robot point cloud and the obstacle point cloud; S is the safety distance threshold. If D(w s ,I) is less than or equal to the safety distance threshold S, then h(ws,I) returns a positive value, indicating that there is a collision risk.

[0038] Optionally, the physical constraints ensure that the generated motion trajectory complies with the kinematic characteristics of the surgical robot by:

[0039] In the reverse denoising process, the prediction results of the intermediate trajectories are modified to ensure that the generated trajectories conform to the robot's kinematic model;

[0040] The formula is:

[0041]

[0042] in: is the generated motion trajectory; α is the interpolation coefficient, which is used to control the smoothness of the trajectory; x goal The target posture.

[0043] Optionally, the convergence judgment criterion in the diffusion model training process is:

[0044] When the value of the comprehensive loss function no longer decreases significantly over multiple consecutive training cycles, the model is considered to have converged and the training is completed.

[0045] Optionally, under the condition of a given surgical task, the motion trajectory of the surgical robot is generated by completing the trained diffusion model, including:

[0046] Based on the given surgical task conditions, obtain obstacle point cloud data from the surgical environment, obtain the current initial posture of the surgical robot, and determine the target posture of the surgical task;

[0047] Encode obstacle point cloud data into latent space using a point cloud encoder;

[0048] Combining the encoded obstacle information, initial posture, and target posture into a conditional code;

[0049] The conditional code is input into the trained diffusion model, and the diffusion model generates the motion trajectory of the surgical robot through a reverse denoising process;

[0050] The generated motion trajectory is sent to the surgical robot control system, and the surgical robot moves according to the motion trajectory to perform the surgical task.

[0051] The beneficial effects of the present invention are:

[0052] The present invention significantly improves the speed of motion planning of surgical robots through the rapid generation capability of diffusion models, and can generate high-quality motion trajectories in a short time; the present invention combines physical constraints and advanced model architectures to improve the stability and reliability of motion planning, avoiding the instability and collision risks that may occur in traditional methods. At the same time, the present invention adopts a Transformer structure with only encoders, which can generate smooth and coherent motion trajectories, improving the accuracy and safety of surgical operations. The present invention can improve the efficiency and stability of motion planning of surgical robots in complex environments, while ensuring the smoothness and safety of the planned path, and can work effectively in complex and dynamic surgical environments, adapting to different surgical tasks and environmental changes.

[0053] The system of the present invention has other characteristics and advantages, which will be apparent from the drawings incorporated herein and the following detailed description, or will be described in detail in the drawings incorporated herein and the following detailed description, which together serve to explain the specific principles of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] The above and other objects, features and advantages of the present invention will become more apparent through a more detailed description of exemplary embodiments of the present invention in conjunction with the accompanying drawings, in which like reference numerals generally represent like components.

[0055] Figure 1 A flowchart of a surgical robot motion planning method based on a diffusion model according to an embodiment of the present invention is shown.

[0056] Figure 2 The working principle diagram of the diffusion model in one embodiment of the present invention is shown. DETAILED DESCRIPTION

[0057] The present invention will be described in more detail below with reference to the accompanying drawings. Although preferred embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present invention more thorough and complete, and to fully convey the scope of the present invention to those skilled in the art.

[0058] like Figure 1 As shown, a surgical robot motion planning method based on a diffusion model according to the present invention comprises the following steps:

[0059] S1: Data preparation phase, including:

[0060] S101: Obtaining obstacle point cloud data in the surgical environment, the initial posture and target posture of the surgical robot;

[0061] S102: Encode the obstacle point cloud data in the surgical environment using a point cloud encoder, and compress the obstacle point cloud data into a latent space to extract feature information related to motion planning;

[0062] S103: Combining the encoded obstacle information, initial posture and target posture into a conditional code;

[0063] S2: Model training phase, including:

[0064] S201: The U-Net structure in the traditional diffusion model is replaced by the encoder-only Transformer structure to form a diffusion model with the encoder-only Transformer structure as the core, wherein the encoder-only Transformer structure is used to capture the temporal dependency of the surgical robot motion sequence;

[0065] S202: inputting the conditional code into a diffusion model for training. The model training includes a forward diffusion process and a reverse denoising process. The forward diffusion process is: in the latent space, gradually adding Gaussian noise to the given original motion trajectory of the surgical robot to generate noisy data; the reverse denoising process is: starting from the noisy data, gradually denoising to restore the original motion trajectory of the surgical robot;

[0066] S203: During the training process, a comprehensive loss function is used to optimize the model performance. The comprehensive loss function combines the configuration space loss, the geometric task space loss and the collision loss. Meanwhile, during the reverse denoising process, physical constraints are introduced to modify the prediction results of the intermediate trajectory to ensure that the generated motion trajectory conforms to the kinematic characteristics of the surgical robot and avoids collisions with obstacles in the surgical environment.

[0067] S204: Optimize the parameters of the model by minimizing the comprehensive loss function, and improve the quality of the generated trajectory until the model converges and the training is completed;

[0068] S3: In the practical application stage, under the condition of a given surgical task, the motion trajectory of the surgical robot is generated by completing the trained diffusion model, including:

[0069] S301: Based on given surgical task conditions, obtain obstacle point cloud data from the surgical environment, obtain the current initial posture of the surgical robot, and determine the target posture of the surgical task;

[0070] S302: Encode the obstacle point cloud data into a latent space using a point cloud encoder; combine the encoded obstacle information, initial posture, and target posture into a conditional code;

[0071] S303: inputting the conditional code into the trained diffusion model, and the diffusion model generates a motion trajectory of the surgical robot through a reverse denoising process;

[0072] S304: Send the generated motion trajectory to the surgical robot control system, and the surgical robot moves according to the motion trajectory to perform the surgical task.

[0073] The following is combined with Figure 2 The technical principle of the present invention is explained.

[0074] Figure 2 The working principle of the diffusion model in the present invention is demonstrated. Its working process mainly includes three steps: embedding, training and generation, as follows:

[0075] (1) Embedding: The obstacle point cloud data is embedded into the latent space z through the point cloud encoder, and the initial pose CS init , target pose CS goal Together they form condition code C.

[0076] (2) Training: The conditional code C is randomly masked for classifier-free learning and then projected to the input token z with t tk Given the noise and motion trajectory X, the encoder-only Transformer-based motion predictor predicts the original clean motion

[0077] (3) Given condition C, sample random noise X in the dimension of the desired motion T , and then iterates from step T to step 1. At each step t, the motion predictor predicts a clean sample Then add noise to get it back to X t-1 .

[0078] A more specific description is given below.

[0079] 1. Conditional Embedding

[0080] refer to Figure 2 The conditional embedding part is the basic input module of the entire diffusion model. Its function is to encode the task-related conditional information into a form that the model can process. The input data includes the obstacle point cloud, the initial posture and the target posture of the surgical robot. The obstacle point cloud needs to be encoded by the point cloud encoder, which is the output of the point cloud encoder. The initial posture and the target posture do not need to be encoded. The conditional embedding part specifically includes:

[0081] (1) Obstacle point cloud: The point cloud encoder uses a compressed autoencoder (CAE) structure to encode the obstacle point cloud data I into the latent space Z. The point cloud encoder is trained by minimizing the reconstruction error and regularization term, as follows:

[0082]

[0083] Among them, l AE is the loss function of the encoder, which is used to train the point cloud encoder; K is the number of points in the point cloud data I; i is a point in the original point cloud; is the corresponding point in the reconstructed point cloud; λ is the regularization coefficient, which is used to control the regularization term of the parameter; θ ij are the parameters in the encoder and decoder.

[0084] Specifically, there may be various obstacles in the surgical environment (such as human tissue, other instruments, etc.). Therefore, a point cloud encoder is used to encode the shape and position information of the obstacles into a concise representation (latent space Z). The input of the point cloud encoder is the point cloud data I of the obstacle, represented by I∈R K×3 (where R indicates that the coordinates of each point are real numbers, K is the number of points in the obstacle point cloud, and 3 represents the three-dimensional coordinates of each point, that is, the point cloud data I is a K×3 real number matrix, in which each row represents the three-dimensional coordinates of a point). The obstacle point cloud is compressed into the latent space Z through the point cloud encoder, irrelevant information (such as texture) is filtered out, and features related to motion planning are extracted; the output is the encoded latent space representation Z. The point cloud encoder can remove noise and irrelevant features in the point cloud and focus on the obstacle shape and position information related to motion planning. The encoded obstacle information Z is used as a conditional input to help the model consider the existence of obstacles when generating motion trajectories, thereby providing environmental constraints.

[0085] (2) Initial posture: used to provide the robot’s initial joint state as the starting point for motion planning. Specifically, the initial joint state CS init , denoted as CSinit ∈R D , where R represents each joint angle is a real number, D is the robot's degree of freedom (DOFs), that is, the initial joint state CS init is a D-dimensional real vector, where each element corresponds to the angle of a joint. The initial pose is directly used as part of the conditional input without additional encoding. The initial pose is the starting point of motion planning, and the model needs to generate a motion trajectory from this state.

[0086] (3) Target posture: used to provide the robot’s target joint state as the end point of motion planning. Specifically, the target joint state CS goal , denoted as CS goal ∈R D , where R represents each joint angle is a real number, D is the robot's degree of freedom (DOFs), that is, the target joint state CS goal It is a D-dimensional real vector, where each element corresponds to the angle of a joint. It is directly used as part of the conditional input without additional encoding. The target pose is the end point of motion planning, and the model needs to generate a feasible path from the initial pose to the target pose. Together with the initial pose and obstacle information, it defines the complete requirements of the motion planning task.

[0087] (4) Combined condition code: The specific process is:

[0088] Input: encoded obstacle information Z, initial posture CSinit and target posture CSgoal.

[0089] Processing: Combine this information into a condition code C. The condition code C is a structure containing all necessary information to guide the diffusion model to generate a motion trajectory that meets the task requirements.

[0090] Output: Condition code C.

[0091] The role of the conditional code C is to integrate the obstacle information in the surgical environment, the initial posture and the target posture of the robot, and provide comprehensive contextual information for the diffusion model. This information helps the model understand the background and requirements of the task. In addition, in the reverse denoising process of the diffusion model, the conditional code C is used as input to guide the model to gradually remove noise and generate motion trajectories that meet the requirements of the task. Specifically, the information contained in the conditional code C helps the model avoid collisions with obstacles, smoothly transition from the initial posture to the target posture, and generate trajectories that meet the kinematic characteristics of the robot. Therefore, by integrating the obstacle information, initial posture and target posture into the conditional code C, the model can generate safe, effective and task-compliant motion trajectories. This not only improves the efficiency of the surgical robot's motion planning, but also enhances the safety and accuracy of the surgery.

[0092] (5) Linear layer: used to embed the conditional input (obstacle information, initial pose, and target pose) into the latent space of the diffusion model so that it can be compatible with other parts of the model. Specifically, the input of the linear layer includes the conditional code C (including the encoded obstacle information Z, the initial pose CSinit, and the target pose CSgoal), and the conditional input is mapped to the latent space of the diffusion model through one or more linear layers (fully connected layers). The linear layer fuses the conditional information (obstacles, initial pose, target pose) from different sources into a unified representation to achieve feature fusion. At the same time, the dimension of the conditional input is adjusted to the latent space dimension required by the diffusion model, thereby providing a complete task context for the diffusion model and helping the model generate motion trajectories that meet the task requirements.

[0093] 2. Diffusion model main body

[0094] The main part of the diffusion model is used to generate motion trajectories through the diffusion process. The diffusion model gradually restores the noisy data to the real motion sequence by gradually adding noise and denoising.

[0095] The diffusion model mainly adopts the Transformer structure with only encoder (corresponding to Figure 2 The motion predictor in the image processing module is specially designed to process sequence data, capture the temporal dependency of motion sequences, and generate smooth and coherent motion trajectories. The main task of the motion predictor is to gradually remove noise from noisy data and restore the motion trajectory that meets the task requirements.

[0096] Specifically, the U-Net structure in the traditional diffusion model is replaced by an encoder-only Transformer structure, that is, the U-Net backbone network in the traditional diffusion model is replaced with a multi-layer Transformer encoder. The temporal dependency of the motion sequence is captured by a multi-layer Transformer encoder. Each layer of the Transformer encoder contains a self-attention mechanism (Self-Attention) and a feed-forward network (Feed-Forward Network, FFN). The number of layers of the Transformer encoder is unlimited and can be selected according to actual needs. The time step t and the conditional code C are embedded in the input of the Transformer, and position encoding is added to the input sequence to help the model capture the position information in the time series.

[0097] That is, the core of the diffusion model is an encoder-only Transformer structure. It gradually removes noise through multiple layers of processing while taking into account time dependencies (i.e., the smoothness of the path). Each step refers to the conditional code C to ensure that the generated path meets the requirements of the surgical environment.

[0098] For example, the structure of a pure encoder Transformer has L layers, and the structure of each layer is as follows:

[0099] 1) Input embedding layer, including:

[0100] Time step embedding: Maps a time step t into a vector that matches the model dimension.

[0101] Conditional embedding: The conditional code C (including obstacle information Z, initial pose and target pose) is embedded into the model dimension.

[0102] Positional encoding: Adds position information to the input sequence to help the model capture temporal dependencies.

[0103] 2) Transformer encoder layer (repeated L times):

[0104] Multi-Head Self-Attention (MHSA): Calculates self-attention weights to capture temporal dependencies and global context in a sequence.

[0105] Feed-Forward Network (FFN): Performs nonlinear transformation on the features at each position.

[0106] Residual connection and layer normalization: Add the output of the sub-layer back to the input and normalize it to enhance the training stability and performance of the model.

[0107] 3) Output layer:

[0108] Linear Layer: Maps the output of the encoder to the target dimension (such as joint angle or position).

[0109] Activation Function: Such as ReLU or Tanh, used for nonlinear transformation.

[0110] Furthermore, the Transformer structure captures the temporal dependencies in sequence data through a multi-head self-attention mechanism, which is particularly suitable for processing motion sequence data; the encoder-only Transformer structure simplifies the model structure and improves computational efficiency.

[0111] The encoder-only Transformer structure uses a multi-head self-attention mechanism and a feedforward network to process the input data and gradually generate motion trajectories. The formula is as follows:

[0112] z tk =MLP(t)+Embed(C)

[0113] Where: z tkis the token input to the Transformer, representing the embedded representation at time step t; MLP is a multi-layer perceptron, used to map time step t to the model dimension; Embed is an embedding operation, which maps the conditional code C to the model dimension; Detailed description of the formula: Multi-head self-attention mechanism: Capture different features in the input sequence through multiple attention heads to enhance the model's ability to capture time dependencies. Feedforward network: The output of each time step is further processed by a feedforward network to enhance the model's nonlinear expression ability. Input processing: Time step t is mapped to the model dimension through a multi-layer perceptron (MLP), and the conditional code C is mapped to the model dimension through an embedding operation (Embed).

[0114] Diffusion process: The diffusion model is trained by adding Gaussian noise to the latent space and denoising it step by step.

[0115] The forward diffusion process is as follows:

[0116]

[0117] Among them, z t is the noise data at time step t; α t is the noise coefficient, which controls the amount of noise added at time step t; z t-1 is the noise data at time step t-1; ∈ t is the Gaussian noise added at time step t, which obeys the standard normal distribution N(0, I); I is the unit matrix, representing the covariance matrix of the noise.

[0118] The reverse denoising process is as follows:

[0119]

[0120] in, is the trajectory after denoising at time step t-1, which is the motion trajectory after denoising at step t-1 and is used for subsequent denoising steps; α t is the noise coefficient, which controls the amount of noise removed at time step t. It is a value between 0 and 1, usually a decreasing sequence, indicating that as t increases, the amount of noise added gradually decreases; is the trajectory at time step t, which is the trajectory of the current step and is used to calculate the trajectory of the previous step; ∈ t is the noise predicted at time step t, which is the noise predicted by the model, used to Remove the noise and restore These parameters work together to ensure that the model can gradually recover smooth, safe and mission-compliant motion trajectories from noisy data.

[0121] exist Figure 2 In the training part:

[0122] The input is: the conditional code C from the conditional embedding part and the noise data z t (obtained through the forward diffusion process).

[0123] The processing process is: gradually remove the noise through the Transformer structure of the encoder only, and restore the motion trajectory

[0124] The output is: denoised motion trajectory And it is optimized through the loss function module.

[0125] 3. Motion guide generation

[0126] This part corresponds to Figure 2 In the generation part, physical constraints are introduced in the denoising process of the diffusion model. By modifying the prediction results of the intermediate trajectory, the generated motion trajectory is ensured to conform to the kinematic characteristics of the surgical robot and avoid collisions with obstacles in the surgical environment. Physical constraints include kinematic constraints, dynamic constraints, and collision constraints.

[0127] The dynamic constraints include the joint angle limits of the surgical robot. Through the kinematic constraints, the trajectory generated by the model can be actually executed by the robot without exceeding the physical limits of the robot. The dynamic constraints include the speed and acceleration limits of the robot's movement. Through the dynamic constraints, the trajectory generated by the model is not only smooth, but also does not cause excessive mechanical stress or unstable movement to the robot during actual execution. The collision constraints ensure that the generated motion trajectory avoids collisions with obstacles in the surgical environment. Specifically: In each step of the denoising process, collision detection is used to check whether the generated trajectory collides with obstacles. If a potential collision risk is detected, the trajectory is adjusted to avoid the collision.

[0128] The collision detection function is as follows:

[0129]

[0130] Among them, h is the collision detection function, which is used to detect the distance between the surgical robot point cloud and the obstacle point cloud; w s is a point in the surgical robot point cloud; I is the obstacle point cloud data; D(w s , I) is the minimum distance between the surgical robot point cloud and the obstacle point cloud; S is the safety distance threshold. If D(w s ,I) is less than or equal to the safety distance threshold S, then h(ws,I) returns a positive value, indicating that there is a collision risk.

[0131] At each denoising step, the model adjusts the generated trajectory according to physical constraints, ensuring that the trajectory complies with kinematic and dynamic properties and avoids collisions.

[0132] Furthermore, the initial and final postures of the motion sequence are controlled by the “noise interpolation” method to ensure that the generated trajectory is consistent with the goal of the surgical task. The formula is as follows:

[0133]

[0134] in, is the initial pose of the generated trajectory. After the reverse denoising process, the initial joint state of the trajectory generated by the model is expressed as a D-dimensional vector, where D is the robot's degrees of freedom (DOFs); α is the interpolation coefficient, a value between 0 and 1, which is used to control the smooth transition between the initial pose and the target pose of the generated trajectory. When α is close to 1, the generated trajectory tends to maintain the original generated pose. When α is close to 0, the generated trajectory is more inclined to the target posture xgoal; goal is the target posture, that is, the target joint state that the robot needs to reach, expressed as a D-dimensional vector, which defines the final goal of the surgical task.

[0135] Through noise interpolation, the model can smoothly transition to the target pose in the generated trajectory, ensuring that the generated trajectory is consistent with the goal of the surgical task. At the same time, the interpolation process helps to avoid mutations or discontinuities in the generated trajectory, thereby improving the smoothness and operability of the trajectory. In this way, the diffusion model can generate smooth motion trajectories that meet the requirements of the surgical task, ensuring that the robot can perform tasks safely and accurately during surgery.

[0136] 4. Loss Function

[0137] Configuration space loss: used to optimize the joint rotation error, the formula is as follows:

[0138]

[0139] Where N is the number of time steps of the trajectory; is the joint angle of the generated trajectory at time step i; is the joint angle of the target trajectory at time step i;

[0140] Geometric task space loss: used to optimize the accumulated error of the kinematic chain, the formula is as follows:

[0141]

[0142] Among them, FK is the forward kinematics function, which maps the joint angle to the position of the end effector;

[0143] Collision loss: Ensure that the robot point cloud maintains a safe distance from obstacles. The formula is as follows:

[0144]

[0145] Among them, h is the collision detection function; I is the obstacle point cloud data.

[0146] Comprehensive loss function: Combine the above three loss functions to optimize model performance. The formula is as follows:

[0147] L=λ joint ·L joint +λ pdint ·L point +λ collision ·L collision

[0148] Among them, λ joint , point and λ collision are the weight coefficients of configuration space loss, geometric task space loss and collision loss respectively.

[0149] 5. Motion trajectory generation

[0150] Under the condition of a given surgical task, the motion trajectory of the surgical robot is generated by the diffusion model. The diffusion model predicts the original data through the following formula:

[0151]

[0152] Among them, DiffusionModel represents the diffusion model, z T is the noise input (randomly generated by the model), T is the time step, and C is the conditional input.

[0153] The motion trajectory generated by the diffusion model is actually a sequence of joint angles that describe the smooth and coherent motion process of the robot from the initial posture to the target posture. Specifically, the final generated motion trajectory can be expressed as: in: represents the joint angle vector at the i-th time step; D is the robot's degrees of freedom (DOFs), that is, the number of joints, and N is the number of time steps of the trajectory, which represents the total length of the trajectory.

[0154] The specific contents of the motion trajectory include:

[0155] (1) Joint angle sequence:

[0156] Each is a D-dimensional vector representing the joint angles of the robot at the i-th time step. For example, for a 7-DOF robot, each It is a 7-dimensional vector representing the angles of the 7 joints.

[0157] (2) Time step:

[0158] N represents the total number of time steps in the trajectory, which is usually predefined according to the task requirements. Each time step i corresponds to a specific moment in time. The complete motion process of the robot from the initial posture to the target posture is described.

[0159] The properties of motion trajectories include:

[0160] (1) Smoothness:

[0161] Generated trajectories It is smooth and continuous, avoiding sudden jumps or unreasonable movements.

[0162] (2) Safety:

[0163] Through physical constraints and collision detection, the trajectory avoids collision with obstacles in the surgical environment, ensuring the safety of the movement process.

[0164] (3) Task Compliance:

[0165] The trajectory starts from the initial posture CS init Start, reach the target posture CS goal , meeting the requirements of surgical tasks.

[0166] Uses of motion tracks include:

[0167] Generated motion trajectory It can be directly used to control the movement of the surgical robot. The specific steps are as follows:

[0168] (1) Trajectory decomposition:

[0169] The generated trajectory Decomposed into joint angles at each time step

[0170] (2) Motion control:

[0171] The joint angles at each time step Sent to the robot's control system, controlling the robot to move step by step to the target posture.

[0172] In an example, suppose there is a 7-DOF surgical robot whose task is to move from an initial pose CS init Move to target pose CS goal , the generated motion trajectory is as follows:

[0173]

[0174] Among them, θ i,jrepresents the jth (j∈1…7) joint angle at the i-th (i∈1…N) time step.

[0175] In the specific implementation process, the training process of the diffusion model in this method includes the following steps:

[0176] 1. Data preparation:

[0177] Preparing point cloud data of obstacles in surgical environment I.

[0178] Prepare the robot's initial posture CS init and target posture CS goal .

[0179] Prepare motion trajectory data X from the initial posture to the target posture.

[0180] These data are usually generated from simulation environments or collected from actual surgical scenarios.

[0181] 2. Conditional Embedding:

[0182] The obstacle point cloud data I is encoded into the latent space Z using a point cloud encoder.

[0183] The encoded obstacle information Z and initial posture CS init and target posture CS goal Combined into condition code C.

[0184] 3. Forward diffusion process:

[0185] Starting from the original motion trajectory X, Gaussian noise is gradually added, and after T steps of calculation, the completely noisy data z is finally obtained. T ; The process of adding noise at each step is expressed as:

[0186]

[0187] 4. Reverse denoising process:

[0188] From the completely noisy data z T Start by removing noise step by step and restore the original motion trajectory ( is the original motion trajectory predicted by the model, close to the original motion trajectory X), and each step of denoising is calculated using the following formula:

[0189]

[0190] Loss function optimization:

[0191] Use a comprehensive loss function to optimize model performance:

[0192] L=λ joint ·Ljoint +λ point ·L point +λ collision ·L collision

[0193] Use an optimizer (such as Adam) to minimize the loss function and update the model parameters.

[0194] 5. Model Validation:

[0195] The performance of the model is evaluated on the validation set to ensure that the generated motion trajectory meets the task requirements, avoids collisions, and has good smoothness and coherence.

[0196] In practical applications, the diffusion model trained in this method uses the reverse denoising process of the model to automatically generate motion trajectories, including the following steps:

[0197] 1. Data collection:

[0198] Obtain obstacle point cloud data from the surgical environment I; (usually acquired in real time from sensors in the surgical environment (such as lidar, camera)

[0199] Get the current initial posture CS of the surgical robot init , and determine the target posture CS of the surgical task goal ; (These data are usually defined by the requirements of the surgical task and can come from the surgical planning system or the physician's input)

[0200] 2. Conditional Embedding:

[0201] Use the point cloud encoder to encode the obstacle point cloud data I into the latent space Z;

[0202] The encoded obstacle information Z and initial posture CS init and target posture CS goal Combined into condition code C.

[0203] 3. Motion trajectory generation:

[0204] The conditional code C is input into the trained diffusion model; a completely randomized Gaussian noise vector z is randomly generated inside the diffusion model. T , whose distribution is the standard normal distribution N(0, I), and then the model gradually removes the noise data z through the reverse denoising process. T Start to gradually restore the motion trajectory In each step t, the model uses the conditional code C (containing obstacle information, initial posture and target posture) and the current noise data as input, and predicts the denoised data through the encoder-only Transformer structure. In the reverse denoising process, the collision detection function h is used to detect whether the generated trajectory collides with the obstacle, and the trajectory is adjusted to avoid the collision. The kinematic adjustment ensures that the generated trajectory conforms to the robot's kinematic model, and the noise interpolation method is used to control the initial and final postures of the trajectory. After T steps of reverse denoising, the model generates the final motion trajectory The process is expressed as:

[0205] 4. Path execution:

[0206] The generated motion trajectory Sent to the robot control system.

[0207] The robot moves precisely along the planned trajectory to perform surgical tasks.

[0208] In summary, the present invention provides an efficient, stable and safe surgical robot motion planning method by combining advanced diffusion models and physical constraints. This method not only improves surgical efficiency, but also enhances surgical safety and accuracy, and is suitable for robot motion planning in complex surgical environments.

[0209] The embodiments of the present invention have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments.

Claims

1. A surgical robot motion planning method based on a diffusion model, characterized in that: include: Obtain obstacle point cloud data in the surgical environment, the initial posture and target posture of the surgical robot; A point cloud encoder is used to encode the obstacle point cloud data in the surgical environment and compress the obstacle point cloud data into a latent space to extract feature information related to motion planning. Combining the encoded obstacle information, the initial posture and the target posture into a conditional code; The encoder-only Transformer structure is used to replace the U-Net structure in the traditional diffusion model, forming a diffusion model with the encoder-only Transformer structure as the core; Inputting the conditional code into the diffusion model and training it through forward diffusion and reverse denoising processes; During the training process, a comprehensive loss function combining configuration space loss, geometric task space loss and collision loss is used to optimize the model performance to improve the quality of the generated trajectory. At the same time, in the reverse denoising process, physical constraints are introduced to modify the prediction results of the intermediate trajectory to ensure that the generated motion trajectory conforms to the kinematic characteristics of the surgical robot and avoids collisions with obstacles in the surgical environment. Under the condition of a given surgical task, the motion trajectory of the surgical robot is generated by completing the trained diffusion model.

2. The surgical robot motion planning method based on the diffusion model according to claim 1 is characterized in that: The encoder-only Transformer architecture captures the temporal dependencies of the surgical robot motion sequences in the following way: The input includes the time step t and the condition code C, where the condition code C contains the encoded obstacle information, initial posture and target posture; Use multi-head self-attention mechanism and feed-forward network to process input data and gradually generate motion trajectories; The formula is: z tk =MLP(t)+Embed(C) Where: z tk is the label input to the Transformer, which represents the embedded representation at time step t; MLP is a multi-layer perceptron, which is used to map the time step t to the model dimension; Embed is an embedding operation, which maps the conditional code C to the model dimension; C is the conditional code, which contains the encoded obstacle information, the initial posture and the target posture of the surgical robot.

3. The surgical robot motion planning method based on the diffusion model according to claim 2 is characterized in that: The point cloud encoder adopts a compressed autoencoder structure, and the point cloud encoder is trained by the following formula: Where: l AE is the loss function of the encoder, which is used to train the point cloud encoder; K is the number of points in the point cloud data I; i is a point in the original point cloud; is the corresponding point in the reconstructed point cloud; λ is the regularization coefficient, which is used to control the regularization term of the parameter; θ ij are the parameters in the encoder and decoder.

4. The surgical robot motion planning method based on the diffusion model according to claim 3 is characterized in that: The forward diffusion process is implemented by the following formula: Where: z t is the noise data at time step t; α t is the noise coefficient, which controls the amount of noise added at time step t; z t-1 is the noise data at time step t-1; ∈ t is the Gaussian noise added at time step t, which obeys the standard normal distribution N(0, I); I is the unit matrix, representing the covariance matrix of the noise.

5. The surgical robot motion planning method based on the diffusion model according to claim 4 is characterized in that: The reverse denoising process gradually restores the motion trajectory through the following formula: in: is the trajectory after denoising at time step t-1; α t is the noise coefficient, which controls the amount of noise removed at time step t; is the trajectory at time step t; ∈ t is the noise predicted at time step t.

6. The surgical robot motion planning method based on the diffusion model according to claim 5 is characterized in that: The comprehensive loss function optimizes the model performance through the following formula: L=λ joint ·L joint +λ point ·L point +λ collision ·L collision Where: L is the comprehensive loss function used to optimize model performance; λ joint is the weight coefficient of the configuration space loss; point is the weight coefficient of the geometric task space loss; λ collision is the weight coefficient of collision loss; L joint It is the configuration space loss used to optimize the joint rotation error of the surgical robot. The calculation formula is: Where: N is the number of time steps of the trajectory; is the joint angle of the generated trajectory at time step i; is the joint angle of the target trajectory at time step i; L point is the geometric task space loss, which is used to optimize the cumulative error of the kinematic chain. Its calculation formula is: Where: FK is the forward kinematics function, which maps the joint angle to the position of the end effector; L collision is the collision loss, which is used to ensure that the generated trajectory avoids collision with obstacles. Its calculation formula is: Where: h is the collision detection function; I is the obstacle point cloud data.

7. The surgical robot motion planning method based on the diffusion model according to claim 6, characterized in that: The physical constraints implement collision detection through the following formula: Where: h is the collision detection function, which is used to detect the distance between the surgical robot point cloud and the obstacle point cloud; w s is a point in the surgical robot point cloud; I is the obstacle point cloud data; D(w s , I) is the minimum distance between the surgical robot point cloud and the obstacle point cloud; S is the safety distance threshold. If D(w s ,I) is less than or equal to the safety distance threshold S, then h(ws,I) returns a positive value, indicating that there is a collision risk.

8. The surgical robot motion planning method based on the diffusion model according to claim 7, characterized in that: The physical constraints ensure that the generated motion trajectory complies with the kinematic characteristics of the surgical robot in the following ways: In the reverse denoising process, the prediction results of the intermediate trajectories are modified to ensure that the generated trajectories conform to the robot's kinematic model; The formula is: in: is the generated motion trajectory; α is the interpolation coefficient, which is used to control the smoothness of the trajectory; x goal The target posture.

9. The surgical robot motion planning method based on the diffusion model according to claim 8, characterized in that: The convergence judgment criteria in the diffusion model training process are: When the value of the comprehensive loss function no longer decreases significantly over multiple consecutive training cycles, the model is considered to have converged and the training is completed.

10. The surgical robot motion planning method based on the diffusion model according to claim 9, characterized in that: Under the condition of a given surgical task, the motion trajectory of the surgical robot is generated by completing the trained diffusion model, including: Based on the given surgical task conditions, obtain obstacle point cloud data from the surgical environment, obtain the current initial posture of the surgical robot, and determine the target posture of the surgical task; Encode obstacle point cloud data into latent space using a point cloud encoder; Combining the encoded obstacle information, initial posture, and target posture into a conditional code; The conditional code is input into the trained diffusion model, and the diffusion model generates the motion trajectory of the surgical robot through a reverse denoising process; The generated motion trajectory is sent to the surgical robot control system, and the surgical robot moves according to the motion trajectory to perform the surgical task.

Citation Information

Patent Citations

  • Multi-modal trajectory prediction method by paying attention to scene and state

    CN114707630A

  • Mechanical arm 6D grabbing attitude estimation method based on Lie group SE (3) diffusion model

    CN118700139A

  • Robot motion planning optimization method

    CN118809583A

  • Industrial robot motion planning method based on diffusion model

    CN119217373A

  • Method and device for optimizing operation performance of hinged object of dexterous arm hand robot and medium

    CN119347799A

Cited By

  • Robot skill migration method and system based on multi-view object trajectory prediction

    CN120552075A

  • Active man-machine cooperation method driven by process knowledge and interactive semantics

    CN120653970A

  • Mechanical arm path planning method based on Transform and diffusion model

    CN120921369A

  • A Robotic Arm Path Planning Method Based on Transformer and Diffusion Model

    CN120921369B

  • Grouping point cloud driven arm operation system

    CN120949936A