Surgical Robot-Assisted Task Segmentation Method Based on Transitional State Clustering
By collecting and analyzing visual and kinematic data in real time, using neural networks and GMM clustering algorithms to identify the transition state of surgical tasks, the problem of insufficient accuracy and reliability of surgical robot task segmentation in the existing technology is solved, and the precise segmentation of surgical tasks and robot assistance is realized, which improves surgical efficiency and safety.
Patent Information
- Application Number
- CN202510444021.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-10
AI Technical Summary
The existing surgical robot task segmentation method lacks accuracy and reliability when identifying the transition state of surgical tasks, making it difficult to meet the requirements of surgical robot autonomy.
By collecting visual data and robot kinematic data in real time, using pre-trained neural networks and GMM clustering algorithms, visual and kinematic feature clusters are generated, transition states of surgical tasks are identified, and robot assisted strategies are triggered.
The precise segmentation and identification of surgical tasks is achieved, the risk of surgery is reduced, the efficiency and quality of surgery is improved, and surgical errors caused by human factors are reduced.
Smart Images

Figure CN119952733B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of surgical robots, and more specifically, to a method for segmenting surgical robot-assisted tasks based on transition state clustering. Background Art
[0002] With the continuous development of minimally invasive surgery and robot-assisted surgery technologies, surgical robots are increasingly widely used. These technologies have significantly improved the treatment effect and recovery speed of patients by reducing surgical trauma, improving surgical precision, and operation flexibility. However, the operation of surgical robots still depends on the control of surgeons, and there is a risk of surgical errors caused by human factors (such as information loss, decision-making errors, fatigue, or inattention, etc.). To improve the safety and efficiency of surgery, existing task segmentation methods help robots plan and execute tasks by decomposing surgical tasks into motion primitives and determining their start and end times. However, the existing task segmentation methods have insufficient accuracy and reliability in identifying the transition states of surgical tasks, and it is difficult to meet the requirements of surgical robot autonomy. Summary of the Invention
[0003] The object of the present invention is to propose a method for segmenting surgical robot-assisted tasks based on transition state clustering, realizing accurate segmentation and identification of surgical tasks, and providing timely robot assistance on this basis, reducing surgical risks, and improving surgical efficiency and quality.
[0004] To achieve the above object, the present invention proposes a method for segmenting surgical robot-assisted tasks based on transition state clustering, including:
[0005] S1: Real-time synchronously collect visual data and robot kinematic data during surgical operations;
[0006] S2: Perform real-time feature extraction and dimensionality reduction processing on the visual data to generate a compressed visual feature vector;
[0007] S3: Perform real-time feature extraction on the kinematic data to generate a kinematic feature vector;
[0008] S4: Perform the first-layer GMM clustering based on the compressed visual feature vector to generate multiple visual feature clusters, and each visual feature cluster represents a macroscopic task stage;
[0009] S5: Perform the second-layer GMM clustering on the kinematic feature vectors within each visual feature cluster to generate multiple kinematic feature sub-clusters, and each kinematic sub-cluster represents different microscopic operation stages within the same visual stage;
[0010] S6: Detect the transitional state of the macroscopic task phase switching in real time according to the boundary changes of the visual feature clusters, and within the same visual feature cluster, detect the transitional state of the microscopic operation phase adjustment in real time according to the boundary changes of the kinematic sub-clusters;
[0011] S7: Trigger the predefined robot assistance strategy according to the detected task transitional state.
[0012] Optionally, step S1 specifically includes:
[0013] During the operation, use a camera to capture the visual images of the operation area in real time to obtain a series of video frames;
[0014] At the same time, use the robot system to record its own kinematic data in real time, including the position, linear velocity, angular velocity, tool direction and angle of the tip of the surgical tool;
[0015] Adjust the collected visual images to a unified size and normalize them.
[0016] Optionally, step S2 specifically includes:
[0017] Use a pre-trained ResNet18 neural network model to extract the original visual feature vectors from the visual images, and the dimension of the original visual feature vectors is 512; the original visual feature vectors include: the position of the surgical tool in the current visual scene, the position and pose of the target object, the background tissue state in the surgical scene, and the relative spatial relationship between the tool and the target;
[0018] Reduce the dimension of the 512-dimensional original visual feature vectors to 128 dimensions through a pre-trained autoencoder to obtain the compressed visual feature vectors. The encoder of the autoencoder contains three fully connected layers, and the activation function is ReLU.
[0019] Optionally, step S3 specifically includes:
[0020] Extract features from the robot kinematic data, and organize the position, linear velocity, angular velocity, tool direction and angle information of the tip of the surgical tool into an n-dimensional kinematic feature vector.
[0021] Optionally, step S4 specifically includes:
[0022] Input consecutive samples in the time series, and each sample contains the 128-dimensional compressed visual features and the n-dimensional kinematic feature vector at the same moment;
[0023] Perform the first-layer GMM clustering on the compressed visual feature vector of the current sample, assign the compressed visual feature vector of the current sample to the visual feature cluster with the highest probability, and generate the visual feature cluster label of the current sample;
[0024] The first - layer GMM clustering uses a full - covariance matrix to capture the complex correlations between visual features and determines the optimal number of clusters by maximizing the silhouette coefficient.
[0025] Optionally, step S5 specifically includes:
[0026] Based on the visual - feature cluster labels of the current sample, within the same visual - feature cluster, perform second - layer GMM clustering on the kinematic - feature vectors of the current sample, assign the kinematic - feature vectors of the current sample to the kinematic - feature sub - cluster with the highest probability, and generate kinematic - feature sub - cluster labels for the current sample.
[0027] The second - layer GMM clustering uses a diagonal covariance matrix to reduce the computational complexity and sets the number of clusters according to the number of task phases.
[0028] Optionally, when performing the first - layer GMM clustering and the second - layer GMM clustering, use the EM algorithm to optimize the GMM parameters in real - time. The GMM parameters include cluster means, covariances, and mixing weights.
[0029] Optionally, step S6 specifically includes:
[0030] If the visual - feature cluster label of the current sample is different from that of the previous sample, it is determined that the current moment is a transition state for the macro - task - phase switch.
[0031] If the visual - feature cluster label of the current sample is the same as that of the previous sample, but the kinematic - feature sub - cluster label of the current sample is different from that of the previous sample, it is determined that the current moment is a transition state for the micro - operation - phase adjustment.
[0032] Optionally, it further includes:
[0033] Set a sliding detection window, and the size of the sliding detection window is N consecutive samples.
[0034] If the visual or kinematic sub - cluster labels of more than M samples within the sliding detection window change consistently, it is determined as a valid transition, where M ≤ N.
[0035] Optionally, step S7 specifically includes:
[0036] According to the detected transition state of the macro - task - phase switch or the micro - operation - phase adjustment, activate the corresponding robot - assisted strategy. The assisted strategy includes adjusting the posture and actions of the surgical tool based on a predefined tool - posture library.
[0037] The beneficial effects of the present invention are as follows:
[0038] The present invention captures global scene changes (such as macroscopic task stages like tool approaching the target, grasping action starting, etc.) by performing the first - layer GMM clustering on visual feature vectors, and then performs the second - layer GMM clustering on the kinematic feature vectors within each visual feature cluster to refine the action patterns within the same stage (such as microscopic operation stages like speed adjustment, tool operation, etc.). By means of hierarchical clustering analysis, combining visual features and kinematic features, the key transition states in surgical tasks can be accurately identified, improving the accuracy and reliability of task segmentation. Further, by using a pre - trained neural network model to automatically extract visual features and the GMM clustering algorithm of unsupervised learning, the real - time performance of task segmentation and robot assistance is ensured, reducing the latency during the surgical process and improving surgical efficiency. By real - time detecting the task transition state and providing robot assistance, the manual adjustment by surgeons during the operation is reduced, and surgical errors caused by human factors are decreased.
[0039] The system of the present invention has other characteristics and advantages, which will be obvious from the accompanying drawings incorporated herein and the subsequent detailed description, or will be described in detail in the accompanying drawings incorporated herein and the subsequent detailed description. These accompanying drawings and detailed description are jointly used to explain the specific principles of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] By describing the exemplary embodiments of the present invention in more detail in conjunction with the accompanying drawings, the above - mentioned and other objects, features, and advantages of the present invention will become more obvious. In the exemplary embodiments of the present invention, the same reference numerals generally represent the same components.
[0041] Figure 1 A step diagram showing a surgical robot - assisted task segmentation method based on transition - state clustering according to the present invention is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] The present invention will be described in more detail below with reference to the accompanying drawings. Although the preferred embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present invention more thorough and complete, and to fully convey the scope of the present invention to those skilled in the art.
[0043] As Figure 1 shown, this embodiment provides a surgical robot - assisted task segmentation method based on transition - state clustering, including:
[0044] S1: Real - time synchronously collect visual data and robot kinematic data during surgical operations;
[0045] Specifically, during the operation, the visual images of the operation area are captured in real time by a camera to obtain a series of video frames. At the same time, the kinematic data of the robot system itself is recorded in real time, including the position, linear velocity, angular velocity, tool direction, and angle of the tip of the surgical tool. The collected visual images are adjusted to a unified size and normalized.
[0046] In this embodiment, during the operation, the visual images of the operation area are captured in real time by a camera to obtain a series of video frames with a frame rate of 30 fps. At the same time, the robot system records its own kinematic data in real time, including information such as the position, linear velocity, angular velocity, tool direction, and gripper angle of the tip of the surgical tool, with a sampling frequency of 100 Hz. The collected visual images are preprocessed, including adjusting the images to a unified size of 224x224 pixels and performing normalization processing to adapt to the pre-trained neural network model. Preferably, a moving average filter is applied to the robot kinematic data with a window size of 5 to reduce the noise in the data, and then downsampling processing is performed with a downsampling rate of 50% to reduce the data volume.
[0047] S2: Perform real-time feature extraction and dimensionality reduction processing on the visual data to generate a compressed visual feature vector.
[0048] In this step, a pre-trained ResNet18 neural network model is used to extract the original visual feature vector from the visual image. The dimension of the original visual feature vector is 512. The original visual feature vector includes: the position of the surgical tool in the current visual scene, the position and pose of the target object, the state of the background tissue in the surgical scene, and the relative spatial relationship between the tool and the target. The 512-dimensional original visual feature vector is reduced to 128 dimensions through a pre-trained autoencoder to obtain a compressed visual feature vector. The encoder of the autoencoder contains three fully connected layers, and the activation function is ReLU.
[0049] Specifically, using the pre-trained ResNet18 model, pre-training can be completed using historical surgical data or existing training datasets. After removing the classification layer of the ResNet18 model, 512-dimensional features are output.
[0050] The encoder structure of the AutoEncoder is a three-layer fully connected network (512→256→128), with an input of 512-dimensional features and the ReLU activation function; the decoder is a three-layer fully connected network (128→256→256→512), and the sigmoid function is used in the output layer. During training, the RMSprop optimizer is adopted, and the loss function is the mean squared error (MSE). During the training phase, the decoder and the encoder participate in the training together. The feature extraction ability of the encoder is optimized through the reconstruction loss to ensure that the low-dimensional features can effectively retain the semantic information of the original data, that is, to ensure that the 128-dimensional features after dimensionality reduction can retain the original visual features to the greatest extent. In actual use, only the encoder structure part of the trained autoencoder is used for dimensionality reduction, that is, the 512-dimensional visual features output by the ResNet18 model are reduced to 128-dimensional compressed visual feature vectors for subsequent clustering, and the decoder does not participate.
[0051] Dimensionality reduction of visual features by the autoencoder can retain key visual information and reduce the computational amount. The feature vector (128-dimensional) after dimensionality reduction can significantly reduce the clustering complexity while maintaining the task-related semantics.
[0052] S3: Perform real-time feature extraction on the kinematic data to generate a kinematic feature vector;
[0053] In this step, feature extraction is performed on the robot kinematic data, and the position, linear velocity, angular velocity, tool direction, and angle information of the surgical tool tip are organized into a kinematic feature vector.
[0054] Preferably, the extracted robot kinematic data is also standardized (such as using Z-score standardization) to eliminate the influence of different dimensions on clustering, improve the robustness of kinematic sub-clustering, and the feature distribution is more concentrated after standardization, which is convenient for accurate modeling by GMM.
[0055] S4: Perform the first layer of GMM clustering based on the compressed visual feature vector to generate multiple visual feature clusters, and each visual feature cluster represents a macroscopic task stage;
[0056] In this step, continuous samples in the time series are input, and each sample contains 128-dimensional compressed visual features and n-dimensional kinematic feature vectors at the same moment; the first layer of GMM clustering is performed on the compressed visual feature vector of the current sample, the compressed visual feature vector of the current sample is assigned to the visual feature cluster with the highest probability, and the visual feature cluster label of the current sample is generated; the first layer of GMM clustering uses the full covariance matrix to capture the complex correlations between visual features and determines the optimal number of clusters by maximizing the silhouette coefficient.
[0057] Specifically, GMM (Gussian Mixed Model) clustering is a probability-based unsupervised learning method that characterizes the overall distribution structure of data through the linear mixture of multiple Gaussian distributions. The covariance matrix type for the first-layer GMM clustering of visual feature vectors is full covariance, allowing arbitrary cluster shapes. The initial clustering can generate a certain number of clusters based on a starting number of frames (samples). Subsequently, each visual feature vector is clustered frame by frame, and each frame's visual feature vector is assigned to a cluster with the highest posterior probability. The output is the cluster label corresponding to each frame. During the clustering process, the optimal number of clusters is selected by maximizing the silhouette coefficient K (search range 2 - 9) to ensure that samples within visual clusters are close and the separation between clusters is high, improving the accuracy of transition detection.
[0058] The generated set of visual feature clusters can be represented as: C v = {( μ 1, Σ1), ( μ 2, Σ2), …, ( μ K , Σ K )}, where C v represents the visual feature clustering result (i.e., the set of visual feature clusters), μ i and Σ i respectively represent the cluster mean and covariance matrix of the i th Gaussian component, i ∈ [1, K], K represents the number of classes of visual clustering.
[0059] Each visual feature cluster represents a macroscopic task stage (e.g., tool approaching the target, grasping, moving to the target position). For example, when the surgical tool moves from a distance to near the target object, the visual features change and are thus clustered into different visual feature clusters.
[0060] Through the first-layer GMM clustering, clustering based on visual features realizes the division of task stages according to scene changes (such as tool approaching, grasping, moving).
[0061] Preferably, the EM (Expectation-Maximization) algorithm is also used in this step to optimize the GMM parameters:
[0062] In the E-step: Calculate the posterior probability (i.e., the expected value) of each data point belonging to each Gaussian component (cluster);
[0063] In the M-step: Use the posterior probability calculated in the E-step as weights to update the parameters (mean, covariance matrix, and weight) of each Gaussian component to maximize the log-likelihood function;
[0064] Alternately execute the E-step and the M-step until the model parameters converge or reach the preset number of iterations. Convergence condition: The change in log-likelihood is less than the set threshold (e.g., 1e -4 ).) or reach the maximum number of iterations (e.g., 100 times).
[0065] By using the EM algorithm to optimize the parameters of the GMM, the key transition states in the surgical task can be identified more accurately, improving the accuracy and reliability of task segmentation.
[0066] S5: Perform a second-layer GMM clustering on the kinematic feature vectors within each visual feature cluster to generate multiple kinematic feature sub-clusters, where each kinematic sub-cluster represents different microscopic operation stages within the same visual stage;
[0067] In this step, based on the visual feature cluster labels of the current samples, within the same visual feature cluster, perform a second-layer GMM clustering on the kinematic feature vectors of the current samples, assign the kinematic feature vectors of the current samples to the kinematic feature sub-cluster with the highest probability, and generate the kinematic feature sub-cluster labels of the current samples; The second-layer GMM clustering uses a diagonal covariance matrix to reduce the computational complexity and sets the number of clusters according to the number of task stages.
[0068] Specifically, based on each visual clustering, perform subsequent second-layer GMM clustering on the kinematic features (such as the position of the tool, speed, gripper angle, etc.), and within each visual cluster, further divide sub-clusters based on the kinematic features. Input the kinematic feature vectors (n-dimensional) within each visual cluster, the covariance matrix type is a diagonal covariance matrix, only modeling the independent variances of each dimension; the number of clusters is set according to the number of surgical task stages (e.g., 9 segments corresponds to the number of clusters L =9). The kinematic feature clustering result can be expressed as: C i k ={( μ i1 ,Σ i1 ),( μ i2 ,Σ i2 ),...,( μ iL ,Σ iL )}, where C i k represents the kinematic feature clustering result in the i th visual feature cluster, μ ij and Σ ij represent the mean and covariance matrix of the j th kinematic clustering respectively, j ∈ [1, L], LIndicates the number of categories for kinematic clustering.
[0069] Each kinematic feature sub-cluster represents different micro-action patterns within the same visual stage (such as rapid approach, fine adjustment, stable grasping). For example, within the visual cluster of "tool approaching the target", the kinematic sub-clusters can distinguish between the actions of the tool moving at high speed and making fine adjustments at low speed. Through the second-layer clustering of the kinematic features of samples within the same visual cluster, the action details (such as speed changes, gripper actions, etc.) are refined, enabling the differentiation of different action patterns (such as rapid movement and precise adjustment) within the same visual stage.
[0070] Preferably, this step also uses the EM algorithm to optimize the GMM parameters, specifically referring to step S4.
[0071] This method avoids the complexity of multi-modal fusion by separating the visual and kinematic features, while retaining the advantages of hierarchical classification.
[0072] S6: Detect the transition state of the macro-task stage switching in real time according to the boundary changes of the visual feature clusters, and within the same visual feature cluster, detect the transition state of the micro-operation stage adjustment in real time according to the boundary changes of the kinematic sub-clusters;
[0073] The detection method for the transition state in this step is as follows: If the visual feature cluster label of the current sample is different from that of the previous sample, it is determined that the current moment is the transition state of the macro-task stage switching; if the visual feature cluster label of the current sample is the same as that of the previous sample, but the kinematic feature sub-cluster label of the current sample is different from that of the previous sample, it is determined that the current moment is the transition state of the micro-operation stage adjustment.
[0074] Specifically, in the feature space, the boundary of a cluster is the dividing line between different clusters, representing the transition region of data points from one state to another. For example, when the tool enters the "grasping" stage from the "approaching the target" stage, both the visual features and kinematic features will cross the boundary of the cluster. By analyzing the clustering results, the conversion points from one cluster to another are identified. These conversion points mark the change of the task stage, that is, the transition state. Specifically, the identification of the transition state is achieved by detecting the change of the clustering labels of adjacent samples. If the clustering labels of adjacent samples change, it indicates that the task has switched from one stage to another, and this point is the transition state.
[0075] For example, the specific logic for transition state detection is as follows:
[0076] The input data is: consecutive samples in the time series, and each sample contains 128-dimensional visual features and n-dimensional kinematic feature vectors at the same moment;
[0077] Extract and compress visual features into a 128-dimensional visual feature vector, and extract the kinematic feature vector of the robot; only use the 128-dimensional visual feature vector for the first layer of GMM clustering to divide the macro task phases (such as "approaching the target", "grasping"), and within each visual cluster, only use the 7-dimensional kinematic features for the second layer of GMM clustering to refine the action patterns (such as "fast moving", "fine adjustment").
[0078] The transition state detection rule is: if the visual feature cluster labels of adjacent samples are different (such as C v1 → C v2 ), it is recognized as the transition state of the macro task phase. Within the same visual feature cluster, if the kinematic sub-cluster labels of adjacent samples are different (such as C 1 k1 → C 1 k2 ), it is recognized as the adjustment transition state of the micro operation phase.
[0079] For example:
[0080] The previous sample (at time t-1 ): visual feature cluster C v1 (representing approaching the target), kinematic feature sub-cluster C 1 k1 (representing fast moving);
[0081] The current sample (at time t ): visual feature cluster C v1 (representing approaching the target), kinematic feature sub-cluster C 1 k2 (representing fine adjustment);
[0082] Then it is determined that: the micro operation phase transition is triggered, and the auxiliary parameters are adjusted (such as reducing the moving speed).
[0083] Preferably, this step further includes:
[0084] Set a sliding detection window, and the size of the sliding detection window is N consecutive samples (such as N = 3);
[0085] If the visual or kinematic sub-cluster labels of more than M (such as M = 2) samples within the sliding detection window change consistently, it is determined as an effective transition, where M ≤ N.
[0086] Through multi-frame consistency verification by sliding detection window filtering, the influence of instantaneous noise or single-frame misdetection can be reduced.
[0087] S7: Trigger a predefined robot-assisted strategy according to the detected task transition state.
[0088] This step activates the corresponding robot-assisted strategy according to the detected transition state of the macro task phase switch or micro operation phase adjustment. The auxiliary strategy includes adjusting the posture and movement of the surgical tool based on a predefined tool posture library.
[0089] Preferably, the triggering of the robot-assisted action is achieved by dynamically mixing the user input command and the robot-assisted command, which specifically includes the following steps:
[0090] (a) Define the mixing weight w ∈ [0,1], where w = 1 indicates full user control, w = 0 indicates full autonomous control by the robot;
[0091] (b) After detecting the task transition state, adjust the weight T within the transition time window w according to a preset rule to generate the final control command:
[0092] τ cmd = w · τ user + (1 - w ) ⋅ τ assist
[0093] where, τ user is the user input command, τ assist is the robot-assisted command;
[0094] (c) Send the final control command to the robot actuator to achieve a smooth transition.
[0095] Specifically, in the initial state, when no transition is detected, by default w = 1, that is, full user control; when a macro or micro transition state is detected, weight adjustment is started: within the transition time window T (such as 100 ms), w linearly decreases from 1 to 0; according to the task requirements, exponential decay or S-shaped curve can be used to dynamically adjust w . Among them, the user command ( τ user ) comes from the teleoperation input of the operator (such as handle displacement or force feedback); the auxiliary command ( τ assist): predefined robot target poses or virtual fixture constraint forces; based on the above weight mixing formula, by adjusting in real time w , a smooth transition between user control and robot assistance is achieved. Preferably, the robot assistance instruction ( τ assist ) has a higher priority. When w <1, the user input is partially or completely overwritten; if a collision or out-of-limit operation (such as joint torque exceeding the threshold) is detected, forcefully set w = 0 and switch to the full-autonomous safety mode. Through the above dynamic mixed weight control, instruction mutations can be avoided, the operation comfort can be improved, and the requirements for autonomy in different surgical stages can be adapted (such as reducing the user weight during fine operations). The safety of the surgery can be guaranteed through the priority and exception handling mechanism.
[0096] Based on the above, the method of the present invention realizes the real-time accurate segmentation and transition state recognition of the surgical robot assistance task through unsupervised hierarchical GMM clustering analysis of visual features and kinematic features, and provides timely robot assistance, effectively improving the surgical efficiency and quality and reducing the surgical risk. In addition, this method is applicable to a variety of surgical robot assistance scenarios, including:
[0097] (1) Minimally invasive surgery, which usually involves a series of clear operation steps, such as incision, insertion of instruments, operation of instruments, suturing, etc. Each step has a clear sequence and characteristic changes during the surgery, and is suitable for applying the present invention for task segmentation and transition state recognition.
[0098] (2) Robot-assisted surgery, the tasks of robot-assisted surgery can be clearly divided into multiple stages, and each stage has unique visual and kinematic features, and is suitable for applying the present invention for task segmentation and transition state recognition.
[0099] (3) In surgical training and simulation environments, skill assessment and guidance are required for novice surgeons. These environments are usually designed as a series of clear operation steps to help novice surgeons learn and master surgical skills. Each step has a clear sequence and characteristic changes during the training, and is suitable for applying the present invention for task segmentation and transition state recognition.
[0100] (4) Some complex surgical tasks, such as cardiac surgery, neurosurgery, etc., involve multiple subtasks, and each subtask has clear operation steps and characteristic changes. These tasks usually have high complexity and risks, and require precise operation and real-time feedback, and are also suitable for applying the present invention for task segmentation and transition state recognition.
[0101] The embodiments of the present invention have been described above. The above description is exemplary and not exhaustive, and is also not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments.
Claims
1. A surgical robot-assisted task segmentation method based on transitional state clustering, characterized in that, Including: S1: Real-time synchronously collect visual data and robotic kinematic data during surgical operations; S2: Perform real-time feature extraction and dimensionality reduction processing on the visual data to generate compressed visual feature vectors; S3: Perform real-time feature extraction on the kinematic data to generate kinematic feature vectors; S4: Perform the first layer of GMM clustering based on the compressed visual feature vectors to generate multiple visual feature clusters, and each visual feature cluster represents a macroscopic task stage; This step specifically includes: Input consecutive samples in the time series, and each sample contains a 128-dimensional compressed visual feature and an n-dimensional kinematic feature vector at the same moment; Perform the first layer of GMM clustering on the compressed visual feature vector of the current sample, assign the compressed visual feature vector of the current sample to the visual feature cluster with the highest probability, and generate the visual feature cluster label of the current sample; S5: Perform the second layer of GMM clustering on the kinematic feature vectors within each visual feature cluster to generate multiple kinematic feature sub-clusters, and each kinematic sub-cluster represents different microscopic operation stages within the same visual stage; This step specifically includes: Based on the visual feature cluster label of the current sample, within the same visual feature cluster, perform the second layer of GMM clustering on the kinematic feature vector of the current sample, assign the kinematic feature vector of the current sample to the kinematic feature sub-cluster with the highest probability, and generate the kinematic feature sub-cluster label of the current sample; S6: Real-time detect the transition state of macroscopic task stage switching according to the boundary change of the visual feature cluster, and within the same visual feature cluster, real-time detect the transition state of microscopic operation stage adjustment according to the boundary change of the kinematic feature sub-cluster; This step specifically includes: If the visual feature cluster label of the current sample is different from that of the previous sample, it is determined that the current moment is the transition state of macroscopic task stage switching; If the visual feature cluster label of the current sample is the same as that of the previous sample, but the kinematic feature sub-cluster label of the current sample is different from that of the previous sample, it is determined that the current moment is the transition state of microscopic operation stage adjustment; S7: Trigger predefined robotic assistance strategies according to the detected task transition state.
2. The method according to claim 1, wherein Step S1 specifically includes: During the operation, use a camera to capture visual images of the surgical area in real time to obtain a series of video frames; At the same time, use the robotic system to record its own kinematic data in real time, including the position, linear velocity, angular velocity, tool direction and angle of the tip of the surgical tool; Adjust the collected visual images to a unified size and perform normalization.
3. The method according to claim 2, wherein Step S2 specifically includes: Use a pre-trained ResNet18 neural network model to extract the original visual feature vectors from the visual images, and the dimension of the original visual feature vectors is 512; The original visual feature vectors include: the position of the surgical tool in the current visual scene, the position and pose of the target object, the background tissue state in the surgical scene, and the relative spatial relationship between the tool and the target; Reduce the 512-dimensional original visual feature vector to 128 dimensions through a pre-trained autoencoder to obtain a compressed visual feature vector. The encoder of the autoencoder includes three fully connected layers, and the activation function is ReLU.
4. The method according to claim 3, wherein Step S3 specifically includes: Extract features from the robot kinematic data, and organize the position, linear velocity, angular velocity, tool direction, and angle information of the surgical tool tip into an n-dimensional kinematic feature vector.
5. The method according to claim 4, wherein In step S4, The first layer of GMM clustering uses a full covariance matrix to capture the complex correlations between visual features and determines the optimal number of clusters by maximizing the silhouette coefficient.
6. The method according to claim 5, wherein In step S5, The second layer of GMM clustering uses a diagonal covariance matrix to reduce the computational complexity and sets the number of clusters according to the number of task stages.
7. The method according to claim 6, wherein When performing the first layer of GMM clustering and the second layer of GMM clustering, use the EM algorithm to optimize the GMM parameters in real time. The GMM parameters include cluster means, covariances, and mixing weights.
8. The method according to claim 1, wherein It also includes: Set a sliding detection window, and the size of the sliding detection window is N consecutive samples; If the visual or kinematic sub-cluster labels of more than M samples in the sliding detection window change consistently, it is determined as a valid transition, where M ≤ N.
9. The method according to claim 1, wherein Step S7 specifically includes: According to the detected transition state of the macro task stage switch or micro operation stage adjustment, activate the corresponding robot assistance strategy. The assistance strategy includes adjusting the posture and actions of the surgical tool based on a predefined tool posture library.
Citation Information
Patent Citations
Surgical system with training or assist functions
US20190090969A1