Surgical robot auxiliary task segmentation method based on transition state clustering
Through the surgical robot assisted task segmentation method based on transition state clustering, the GMM clustering algorithm is used to identify the transition state of the surgical task, which solves the problem of insufficient recognition accuracy in the prior art, realizes accurate segmentation of surgical tasks and robot assistance, and improves surgical efficiency and safety.
Patent Information
- Application Number
- CN202510444021.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-04-10
AI Technical Summary
The existing surgical robot task segmentation method lacks accuracy and reliability when identifying the transition state of surgical tasks, making it difficult to meet the requirements of surgical robot autonomy.
A surgical robot assisted task segmentation method based on transition state clustering is proposed. By collecting visual data and kinematic data in real time, feature extraction and dimensionality reduction processing is performed, GMM clustering algorithm is used to identify the transition states of macroscopic and microscopic task stages, and a predefined robot assist strategy is triggered.
The precise segmentation and identification of surgical tasks are achieved, the accuracy and reliability of task segmentation are improved, the risk of surgery is reduced, and the efficiency and quality of surgical operations are improved.
Smart Images

Figure CN119952733A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of surgical robots, and more specifically, to a surgical robot-assisted task segmentation method based on transition state clustering. Background Art
[0002] With the continuous development of minimally invasive surgery and robot-assisted surgery, surgical robots are being used more and more widely. These technologies have significantly improved the treatment and recovery speed of patients by reducing surgical trauma, improving surgical accuracy and operational flexibility. However, the operation of surgical robots still relies on the control of surgeons, and there is a risk of surgical errors due to human factors (such as information loss, decision-making errors, fatigue or lack of concentration). In order to improve the safety and efficiency of surgery, existing task segmentation methods help robots plan and execute tasks by decomposing surgical tasks into motion primitives and determining their start and end times. However, existing task segmentation methods lack accuracy and reliability in identifying the transition states of surgical tasks, and it is difficult to meet the requirements of surgical robot autonomy. Summary of the invention
[0003] The purpose of this invention is to propose a surgical robot-assisted task segmentation method based on transition state clustering, to achieve accurate segmentation and identification of surgical tasks, and on this basis to provide timely robot assistance, reduce surgical risks, and improve surgical efficiency and quality.
[0004] To achieve the above objectives, the present invention proposes a surgical robot-assisted task segmentation method based on transition state clustering, comprising: S1: Real-time synchronous acquisition of visual data and robot kinematic data during surgical operations; S2: Performing real-time feature extraction and dimensionality reduction processing on the visual data to generate a compressed visual feature vector; S3: Performing real-time feature extraction on the kinematic data to generate a kinematic feature vector; S4: performing a first-layer GMM clustering based on the compressed visual feature vector to generate multiple visual feature clusters, each visual feature cluster representing a macro task stage; S5: Perform the second-level GMM clustering on the kinematic feature vectors in each visual feature cluster to generate multiple kinematic feature sub-clusters, each of which represents a different micro-operation stage in the same visual stage; S6: detecting the transition state of the macro task phase switching in real time according to the boundary changes of the visual feature cluster, and detecting the transition state of the micro operation phase adjustment in real time according to the boundary changes of the kinematic sub-cluster within the same visual feature cluster; S7: Triggering a predefined robot assistance strategy based on the detected task transition state.
[0005] Optionally, step S1 specifically includes: During the operation, the visual image of the surgical area is captured in real time by a camera to obtain a series of video frames; At the same time, the robot system records its own kinematic data in real time, including the position, linear velocity, angular velocity, tool direction and angle of the surgical tool tip; The collected visual images are adjusted to a uniform size and normalized.
[0006] Optionally, step S2 specifically includes: A pre-trained ResNet18 neural network model is used to extract an original visual feature vector from a visual image, wherein the dimension of the original visual feature vector is 512; the original visual feature vector includes: the position of the surgical tool in the current visual scene, the position and posture of the target object, the background tissue state in the surgical scene, and the relative spatial relationship between the tool and the target; The 512-dimensional original visual feature vector is reduced to 128 dimensions by a pre-trained autoencoder to obtain a compressed visual feature vector, wherein the encoder of the autoencoder includes three fully connected layers and the activation function is ReLU.
[0007] Optionally, step S3 specifically includes: The robot kinematic data is extracted to organize the position, linear velocity, angular velocity, tool direction and angle information of the surgical tool tip into an n-dimensional kinematic feature vector.
[0008] Optionally, step S4 specifically includes: Input continuous samples in time series, each sample contains 128-dimensional compressed visual features and n-dimensional kinematic feature vectors at the same moment; Perform the first-layer GMM clustering on the compressed visual feature vector of the current sample, assign the compressed visual feature vector of the current sample to the visual feature cluster with the highest probability, and generate the visual feature cluster label of the current sample; The first-level GMM clustering uses the full covariance matrix to capture the complex correlations between visual features and determines the optimal number of clusters by maximizing the silhouette coefficient.
[0009] Optionally, step S5 specifically includes: Based on the visual feature cluster label of the current sample, within the same visual feature cluster, the kinematic feature vector of the current sample is clustered by the second-level GMM, the kinematic feature vector of the current sample is assigned to the kinematic feature sub-cluster with the highest probability, and the kinematic feature sub-cluster label of the current sample is generated; The second-level GMM clustering uses a diagonal covariance matrix to reduce computational complexity and sets the number of clusters according to the number of task stages.
[0010] Optionally, when performing the first layer GMM clustering and the second layer GMM clustering, the EM algorithm is used to optimize GMM parameters in real time, and the GMM parameters include cluster mean, covariance and mixing weight.
[0011] Optionally, step S6 specifically includes: If the visual feature cluster label of the current sample is different from the visual feature cluster label of the previous sample, the current moment is determined to be a transition state of switching between macro task stages; If the visual feature cluster label of the current sample is the same as the visual feature cluster label of the previous sample, but the kinematic feature subcluster label of the current sample is different from the kinematic feature subcluster label of the previous sample, then the current moment is determined to be a transition state of the micro-operation stage adjustment.
[0012] Optionally, it also includes: Setting a sliding detection window, wherein the size of the sliding detection window is N consecutive samples; If the visual or kinematic subcluster labels of more than M samples in the sliding detection window change consistently, it is considered a valid transition, where M≤N.
[0013] Optionally, step S7 specifically includes: According to the detected transition state of macro-task phase switching or micro-operation phase adjustment, the corresponding robot assistance strategy is activated, and the assistance strategy includes adjusting the posture and movement of the surgical tool based on a predefined tool posture library.
[0014] The beneficial effects of the present invention are: The present invention performs a first-layer GMM clustering on the visual feature vector to capture global scene changes (such as macro task stages such as the tool approaching the target and the start of the grasping action), and then performs a second-layer GMM clustering on the kinematic feature vectors in each visual feature cluster to refine the action mode in the same stage (such as micro operation stages such as speed adjustment and tool operation). Through the hierarchical clustering analysis method, combined with visual features and kinematic features, the key transition states in the surgical task can be accurately identified to improve the accuracy and reliability of task segmentation. Furthermore, by using a pre-trained neural network model to automatically extract visual features and using an unsupervised learning GMM clustering algorithm, the real-time nature of task segmentation and robot assistance is ensured, delays during surgery are reduced, and surgical efficiency is improved. By real-time detection of task conversion status and providing robot assistance, the surgeon's manual adjustments during the operation are reduced, and surgical errors caused by human factors are reduced.
[0015] The system of the present invention has other characteristics and advantages, which will be apparent from the drawings incorporated herein and the following detailed description, or will be described in detail in the drawings incorporated herein and the following detailed description, which together serve to explain the specific principles of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The above and other objects, features and advantages of the present invention will become more apparent through a more detailed description of exemplary embodiments of the present invention in conjunction with the accompanying drawings, in which like reference numerals generally represent like components.
[0017] Figure 1 A step diagram of a surgical robot-assisted task segmentation method based on transition state clustering according to the present invention is shown. DETAILED DESCRIPTION
[0018] The present invention will be described in more detail below with reference to the accompanying drawings. Although preferred embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present invention more thorough and complete, and to fully convey the scope of the present invention to those skilled in the art.
[0019] like Figure 1 As shown, this embodiment provides a surgical robot-assisted task segmentation method based on transition state clustering, comprising: S1: Real-time synchronous acquisition of visual data and robot kinematic data during surgical operations; Specifically, during the operation, the camera captures the visual image of the surgical area in real time to obtain a series of video frames. At the same time, the robot system records its own kinematic data in real time, including the position, linear velocity, angular velocity, tool direction and angle of the tip of the surgical tool. The collected visual images are adjusted to a uniform size and normalized.
[0020] In this embodiment, during the operation, the visual image of the surgical area is captured in real time by a camera to obtain a series of video frames with a frame rate of 30fps; at the same time, the robot system records its own kinematic data in real time, including the position of the tip of the surgical tool, linear velocity, angular velocity, tool direction, gripper angle and other information, with a sampling frequency of 100Hz; the collected visual images are preprocessed, including adjusting the images to a uniform size of 224x224 pixels and normalizing them to adapt to the pre-trained neural network model; preferably, the robot kinematic data is subjected to sliding average filtering with a window size of 5 to reduce noise in the data, and then down-sampling is performed with a down-sampling rate of 50% to reduce the amount of data.
[0021] S2: Performing real-time feature extraction and dimensionality reduction processing on the visual data to generate a compressed visual feature vector; In this step, a pre-trained ResNet18 neural network model is used to extract an original visual feature vector from the visual image, wherein the dimension of the original visual feature vector is 512; the original visual feature vector includes: the position of the surgical tool in the current visual scene, the position and posture of the target object, the background tissue state in the surgical scene, and the relative spatial relationship between the tool and the target; the 512-dimensional original visual feature vector is reduced to 128 dimensions by a pre-trained autoencoder to obtain a compressed visual feature vector, wherein the encoder of the autoencoder includes three fully connected layers, and the activation function is ReLU.
[0022] Specifically, using a pre-trained ResNet18 model, historical surgical data or an existing training data set can be used to complete pre-training, and the ResNet18 model removes the classification layer and outputs 512-dimensional features.
[0023] The encoder structure of the AutoEncoder is a 3-layer fully connected (512→256256→128) structure, with 512-dimensional features as input and ReLU as activation function; the decoder is a 3-layer fully connected (128→256→256→512) structure, with the output layer using the Sigmoid function. The RMSprop optimizer is used for training, and the loss function is the mean square error (MSE). During the training phase, the decoder and the encoder participate in the training together, and the feature extraction capability of the encoder is optimized through reconstruction loss to ensure that the low-dimensional features can effectively retain the semantic information of the original data, that is, to ensure that the 128-dimensional features after dimensionality reduction can retain the original visual features to the greatest extent. In actual use, only the encoder structure of the trained AutoEncoder is used for dimensionality reduction, that is, the 512-dimensional visual features output by the ResNet18 model are reduced to 128-dimensional compressed visual feature vectors for subsequent clustering, while the decoder is not involved.
[0024] By reducing the dimension of visual features through autoencoders, key visual information can be retained and the amount of computation can be reduced. The reduced-dimensional feature vector (128 dimensions) can significantly reduce clustering complexity while maintaining task-related semantics.
[0025] S3: Performing real-time feature extraction on the kinematic data to generate a kinematic feature vector; This step extracts features from the robot kinematic data and organizes the position, linear velocity, angular velocity, tool direction and angle information of the surgical tool tip into a kinematic feature vector.
[0026] Preferably, the extracted robot kinematic data is also standardized (such as using Z-score standardization) to eliminate the influence of different dimensions on clustering and improve the robustness of kinematic sub-clusters. After standardization, the feature distribution is more concentrated, which is convenient for accurate GMM modeling.
[0027] S4: performing a first-layer GMM clustering based on the compressed visual feature vector to generate multiple visual feature clusters, each visual feature cluster representing a macro task stage; This step inputs continuous samples in a time series, each sample contains a 128-dimensional compressed visual feature and an n-dimensional kinematic feature vector at the same moment; the first-level GMM clustering is performed on the compressed visual feature vector of the current sample, the compressed visual feature vector of the current sample is assigned to the visual feature cluster with the highest probability, and the visual feature cluster label of the current sample is generated; the first-level GMM clustering uses the full covariance matrix to capture the complex correlations between visual features, and determines the optimal number of clusters by maximizing the silhouette coefficient.
[0028] Specifically, GMM (Gussian Mixed Model) clustering is a probability-based unsupervised learning method that characterizes the overall distribution structure of the data through a linear mixture of multiple Gaussian distributions. The covariance matrix type for the first layer of GMM clustering of visual feature vectors is full covariance, allowing arbitrary cluster shapes. The initial clustering can generate a number of clusters based on the initial frames (samples), and then cluster each visual feature vector frame by frame. The visual feature vector of each frame is assigned to a cluster with the largest posterior probability, and the output is the cluster label corresponding to each frame. During the clustering process, the optimal number of clusters is selected by maximizing the silhouette coefficient. K (Search range 2~9) ensures that the samples within the visual cluster are compact and the separation between clusters is high, thus improving the accuracy of transition detection.
[0029] The generated visual feature cluster set can be expressed as: C v ={( m 1,Σ1),( m 2,Σ2),…,( m K ,Σ K )},in, C v represents the visual feature clustering result (i.e., the set of visual feature clusters), m i and Σ i Respectively represent i The cluster means and covariance matrices of Gaussian components, i∈[1,K], K Represents the number of categories for visual clustering.
[0030] Each visual feature cluster represents a macro task stage (e.g., tool approaching the target, grasping, moving to the target location). For example, when the surgical tool moves from a distance to near the target object, the visual features will change and thus be clustered into different visual feature clusters.
[0031] Through the first layer of GMM clustering, clustering based on visual features is performed to divide the task stages according to scene changes (such as tool approach, grasping, and movement).
[0032] Preferably, in this step, the EM (Expectation-Maximization) algorithm is also used to optimize the GMM parameters: In step E: calculate the posterior probability (i.e. expected value) of each data point belonging to each Gaussian component (cluster); In the M step: using the posterior probability calculated in the E step as the weight, update the parameters (mean, covariance matrix and weight) of each Gaussian component to maximize the log-likelihood function; The E step and the M step are executed alternately until the model parameters converge or the preset number of iterations is reached. Convergence condition: the change in log-likelihood is less than the set threshold (such as 1e -4 ) or reaches a maximum number of iterations (such as 100).
[0033] By using the EM algorithm to optimize the parameters of GMM, the key transition states in surgical tasks can be more accurately identified, and the accuracy and reliability of task segmentation can be improved.
[0034] S5: Perform the second-level GMM clustering on the kinematic feature vectors in each visual feature cluster to generate multiple kinematic feature sub-clusters, each of which represents a different micro-operation stage in the same visual stage; This step is based on the visual feature cluster label of the current sample. Within the same visual feature cluster, the kinematic feature vector of the current sample is clustered by the second-level GMM, the kinematic feature vector of the current sample is assigned to the kinematic feature subcluster with the highest probability, and the kinematic feature subcluster label of the current sample is generated. The second-level GMM clustering uses the diagonal covariance matrix to reduce the computational complexity and sets the number of clusters according to the number of task stages.
[0035] Specifically, on the basis of each visual cluster, the kinematic features (such as tool position, speed, gripper angle, etc.) are clustered in the second layer of GMM. Within each visual cluster, subclusters are further divided based on kinematic features. The kinematic feature vector (n-dimensional) within each visual cluster is input, and the covariance matrix type is a diagonal covariance matrix, which only models the independent variance of each dimension; the number of clusters is set according to the number of surgical task stages (for example, 9 stages correspond to the number of clusters L =9). The kinematic feature clustering result can be expressed as: Ci k ={( m i1 ,Σ i1 ),( m i2 ,Σ i2 ),...,( m iL ,Σ iL )},in C i k Indicates i The kinematic feature clustering results in visual feature clusters are m ij and Σ ij Respectively represent j The mean and covariance matrices of kinematic clusters, j∈[1,L], L Represents the number of categories for kinematic clustering.
[0036] Each kinematic feature subcluster represents a different micro-action mode within the same visual stage (e.g., fast approach, fine adjustment, stable grasping). For example, within the visual cluster of "tool approaching target", the kinematic subcluster can distinguish between high-speed tool movement and low-speed fine adjustment. The second-level clustering refines the action details (e.g., speed change, gripper action, etc.) by targeting the kinematic features of samples within the same visual cluster, and achieves the distinction between different action modes (e.g., fast movement and precise adjustment) within the same visual stage.
[0037] Preferably, this step also uses the EM algorithm to optimize the GMM parameters, refer to step S4 for details.
[0038] This method avoids the complexity of multimodal fusion by separately processing visual and kinematic features while retaining the advantages of hierarchical grading.
[0039] S6: detecting the transition state of the macro task phase switching in real time according to the boundary changes of the visual feature cluster, and detecting the transition state of the micro operation phase adjustment in real time according to the boundary changes of the kinematic sub-cluster within the same visual feature cluster; The detection method of the transition state in this step is: if the visual feature cluster label of the current sample is different from the visual feature cluster label of the previous sample, then the current moment is determined to be a transition state of switching between the macro task stage; if the visual feature cluster label of the current sample is the same as the visual feature cluster label of the previous sample, but the kinematic feature subcluster label of the current sample is different from the kinematic feature subcluster label of the previous sample, then the current moment is determined to be a transition state of adjustment in the micro operation stage.
[0040] Specifically, in the feature space, the boundary of the cluster is the dividing line between different clusters, indicating the transition area of the data point from one state to another. For example, when the tool enters the "grasping" stage from the "approaching the target" stage, both the visual features and the kinematic features will cross the cluster boundary. By analyzing the clustering results, the transition points from one cluster to another are identified. These transition points mark the change of the task stage, that is, the transition state. Specifically, the identification of the transition state is achieved by detecting the change of the cluster labels of adjacent samples. If the cluster labels of adjacent samples change, it means that the task has changed from one stage to another, and this point is the transition state.
[0041] For example, the specific logic of transition state detection is: The input data is: continuous samples in time series, each sample contains 128-dimensional visual features and n-dimensional kinematic feature vectors at the same time; The visual features are extracted and compressed into a 128-dimensional visual feature vector, and the robot kinematic feature vector is extracted. Only the 128-dimensional visual feature vector is used for the first-level GMM clustering to divide the macro task stages (such as "approaching the target" and "grasping"). Within each visual cluster, only the 7-dimensional kinematic features are used for the second-level GMM clustering to refine the action mode (such as "fast movement" and "fine adjustment").
[0042] The transition state detection rule is: the visual feature cluster labels of adjacent samples are different (such as C v1 → C v2 ), it is identified as a transition state of the macro task stage. Within the same visual feature cluster, the kinematic sub-cluster labels of adjacent samples are different (e.g. C 1 k1 → C 1 k2 ), it is identified as a transition state of adjustment in the micro-operation stage.
[0043] For example: The previous sample (time t-1 ): Visual feature cluster C v1 (representing approach to the target), kinematic feature sub-cluster C 1 k1 (represents fast movement); Current sample (time t ): Visual feature cluster C v1 (representing approach to the target), kinematic feature sub-cluster C 1 k2 (stands for fine adjustment); Then determine: trigger the transition to the micro-operation stage and adjust auxiliary parameters (such as reducing the movement speed).
[0044] Preferably, this step also includes: Set a sliding detection window, where the size of the sliding detection window is N consecutive samples (such as N=3); If the visual or kinematic subcluster labels of more than M (e.g., M=2) samples in the sliding detection window change consistently, it is considered a valid transition, where M≤N.
[0045] Multi-frame consistency verification through sliding detection window filtering can reduce the impact of instantaneous noise or single-frame false detection.
[0046] S7: Triggering a predefined robot assistance strategy based on the detected task transition state.
[0047] This step activates the corresponding robot assistance strategy according to the detected transition state of macro-task phase switching or micro-operation phase adjustment, and the assistance strategy includes adjusting the posture and movement of the surgical tool based on a predefined tool posture library.
[0048] Preferably, the triggering of the robot-assisted action is achieved by dynamically weighting the user input command and the robot-assisted command, which specifically includes the following steps: (a) Define the mixing weights w ∈[0,1], where w =1 means it is completely controlled by the user. w =0 means the robot is completely controlled autonomously; (b) After detecting the task transition state, T Adjust weights according to preset rules w , generate the final control instruction: t cmd = w · t user +(1- w )⋅ t assist in, t user Input instructions for users, t assist To assist the robot with instructions; (c) Sending the final control instruction to the robot actuator to achieve a smooth transition.
[0049] Specifically, in the initial state, when no transition is detected, the default w=1, which is completely controlled by the user; when a macro or micro transition state is detected, weight adjustment is initiated: in the transition time window T (e.g. 100ms), w Linearly decreases from 1 to 0; exponential decay or S-shaped curve can be used for dynamic adjustment according to task requirements w Among them, user instructions ( t user ): remote operation input from the operator (such as handle displacement or force feedback); auxiliary instructions ( t assist ): predefined robot target pose or virtual fixture constraint force; based on the above weighted hybrid formula, real-time adjustment w , achieving a smooth transition between user control and robot assistance. Optimize robot assistance instructions ( t assist ) has a higher priority when w <1, the user input is partially or completely overwritten; if a collision or over-limit operation is detected (such as joint torque exceeding the threshold), the forced setting w =0, switch to fully autonomous safety mode. The above dynamic hybrid weight control can avoid sudden changes in instructions, improve operational comfort, and adapt to the needs of autonomy in different surgical stages (such as reducing user weight during delicate operations). The safety of surgery can be guaranteed through priority and exception handling mechanisms.
[0050] Based on the above, the method of the present invention realizes the real-time accurate segmentation and transition state recognition of surgical robot-assisted tasks through unsupervised learning hierarchical GMM cluster analysis of visual features and kinematic features, and provides timely robot assistance, effectively improving surgical efficiency and quality and reducing surgical risks. In addition, this method is applicable to a variety of surgical robot-assisted scenarios, including: (1) Minimally invasive surgery, which usually involves a series of clear operation steps, such as incision, instrument insertion, instrument operation, suturing, etc. Each step has a clear sequence and characteristic changes during the operation, which is suitable for applying the present invention to perform task segmentation and transition state recognition.
[0051] (2) Robot-assisted surgery: The tasks of robot-assisted surgery can be clearly divided into multiple stages. Each stage has unique visual and kinematic features, which is suitable for applying the present invention to perform task segmentation and transition state recognition.
[0052] (3) In surgical training and simulation environments, novice surgeons need to be assessed and guided in their skills. These environments are usually designed as a series of clear operating steps to help novice surgeons learn and master surgical skills. Each step has a clear sequence and feature changes during the training process, which is suitable for the application of the present invention for task segmentation and transition state recognition.
[0053] (4) Some complex surgical tasks, such as cardiac surgery and neurosurgery, involve multiple subtasks, each of which has clear operating steps and feature changes. These tasks are usually highly complex and risky, requiring precise operations and real-time feedback, and are also suitable for the application of the present invention for task segmentation and transition state recognition.
[0054] The embodiments of the present invention have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments.
Claims
1. A surgical robot-assisted task segmentation method based on transition state clustering, characterized in that: include: S1: Real-time synchronous acquisition of visual data and robot kinematic data during surgical operations; S2: Performing real-time feature extraction and dimensionality reduction processing on the visual data to generate a compressed visual feature vector; S3: Performing real-time feature extraction on the kinematic data to generate a kinematic feature vector; S4: performing a first-layer GMM clustering based on the compressed visual feature vector to generate multiple visual feature clusters, each visual feature cluster representing a macro task stage; S5: Perform the second-level GMM clustering on the kinematic feature vectors in each visual feature cluster to generate multiple kinematic feature sub-clusters, each of which represents a different micro-operation stage in the same visual stage; S6: detecting the transition state of the macro task phase switching in real time according to the boundary changes of the visual feature cluster, and detecting the transition state of the micro operation phase adjustment in real time according to the boundary changes of the kinematic sub-cluster within the same visual feature cluster; S7: Triggering a predefined robot assistance strategy based on the detected task transition state.
2. The method according to claim 1, characterized in that Step S1 specifically includes: During the operation, the visual image of the surgical area is captured in real time by a camera to obtain a series of video frames; At the same time, the robot system records its own kinematic data in real time, including the position, linear velocity, angular velocity, tool direction and angle of the surgical tool tip; The collected visual images are adjusted to a uniform size and normalized.
3. The method according to claim 2, characterized in that Step S2 specifically includes: A pre-trained ResNet18 neural network model is used to extract an original visual feature vector from a visual image, wherein the dimension of the original visual feature vector is 512; the original visual feature vector includes: the position of the surgical tool in the current visual scene, the position and posture of the target object, the background tissue state in the surgical scene, and the relative spatial relationship between the tool and the target; The 512-dimensional original visual feature vector is reduced to 128 dimensions by a pre-trained autoencoder to obtain a compressed visual feature vector, wherein the encoder of the autoencoder includes three fully connected layers and the activation function is ReLU.
4. The method according to claim 3, characterized in that Step S3 specifically includes: The robot kinematic data is extracted to organize the position, linear velocity, angular velocity, tool direction and angle information of the surgical tool tip into an n-dimensional kinematic feature vector.
5. The method according to claim 4, characterized in that Step S4 specifically includes: Input continuous samples in time series, each sample contains 128-dimensional compressed visual features and n-dimensional kinematic feature vectors at the same moment; Perform the first-layer GMM clustering on the compressed visual feature vector of the current sample, assign the compressed visual feature vector of the current sample to the visual feature cluster with the highest probability, and generate the visual feature cluster label of the current sample; The first-level GMM clustering uses the full covariance matrix to capture the complex correlations between visual features and determines the optimal number of clusters by maximizing the silhouette coefficient.
6. The method according to claim 5, characterized in that Step S5 specifically includes: Based on the visual feature cluster label of the current sample, within the same visual feature cluster, the kinematic feature vector of the current sample is clustered by the second-level GMM, the kinematic feature vector of the current sample is assigned to the kinematic feature sub-cluster with the highest probability, and the kinematic feature sub-cluster label of the current sample is generated; The second-level GMM clustering uses a diagonal covariance matrix to reduce computational complexity and sets the number of clusters according to the number of task stages.
7. The method according to claim 6, characterized in that When performing the first-layer GMM clustering and the second-layer GMM clustering, the EM algorithm is used to optimize the GMM parameters in real time, and the GMM parameters include cluster mean, covariance and mixing weight.
8. The method according to claim 7, characterized in that Step S6 specifically includes: If the visual feature cluster label of the current sample is different from the visual feature cluster label of the previous sample, the current moment is determined to be a transition state of switching between macro task stages; If the visual feature cluster label of the current sample is the same as the visual feature cluster label of the previous sample, but the kinematic feature subcluster label of the current sample is different from the kinematic feature subcluster label of the previous sample, then the current moment is determined to be a transition state of the micro-operation stage adjustment.
9. The method according to claim 8, characterized in that Also includes: Setting a sliding detection window, wherein the size of the sliding detection window is N consecutive samples; If the visual or kinematic subcluster labels of more than M samples in the sliding detection window change consistently, it is considered a valid transition, where M≤N.
10. The method according to claim 9, characterized in that Step S7 specifically includes: According to the detected transition state of macro-task phase switching or micro-operation phase adjustment, the corresponding robot assistance strategy is activated, and the assistance strategy includes adjusting the posture and movement of the surgical tool based on a predefined tool posture library.
Citation Information
Patent Citations
A multimodal surgical trajectory fast segmentation method based on unsupervised deep learning
CN109165550A
Master-slave motion control method based on pose identification and surgical robot system
CN116492064A
Evaluation and enhancement system for open athletic skills
CN116830167A
Surgical system with training or assist functions
US20190090969A1
Error detection method and robot system based on a plurality of pose identifications
US20230219221A1