Semantic Segmentation and Pose Estimation Method and System for Surgical Robot Instruments under Endoscopic Images
By adopting an encoder-decoder architecture-based pose estimation model in laparoscopic surgical robotic instruments, combining semantic segmentation and timing information, the problems of high error rate and poor real-time performance in the prior art are solved, and high-precision pose estimation and control are achieved, which improves the accuracy and safety of the surgery.
Patent Information
- Application Number
- CN202411656886.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-19
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-11-19
AI Technical Summary
The semantic segmentation and position estimation methods of existing laparoscopic surgical robotic devices have large error rates, poor real-time and generalization, and rarely consider the constraints of timing information, making it difficult to meet the needs of precise surgical operations.
The pose estimation model of surgical robot instruments based on the encoder-decoder architecture is adopted, and the pose estimation model of a single encoder-dual decoder based on the encoder-decoder architecture is combined with the backbone network, feature encoder, semantic segmentation decoder, pose estimation decoder and its data cache queue and time series module are used to perform semantic segmentation and pose estimation accuracy, using timing information to improve estimation accuracy.
By combining semantic segmentation results and timing information, the calculation accuracy and stability of surgical instrument position estimation are improved, and higher control accuracy and safety are achieved. The binary segmentation mIoU index is not less than 92%, the part segmentation mIoU index is not less than 80%, and the frame rate is not less than 20 frames per second.
Smart Images

Figure CN119579894B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and particularly to a semantic segmentation and pose estimation method and system for surgical robot instruments under a laparoscope view. Background Art
[0002] Laparoscopic minimally invasive surgery has become one of the important directions in the modern surgical field. Compared with traditional open abdominal surgery, laparoscopic minimally invasive surgery has the advantages of less trauma, faster patient recovery, and fewer postoperative complications. In laparoscopic minimally invasive surgery, the precise operation of surgical instruments is crucial for the surgical effect. Accurately segmenting the surgical instrument mask during the operation, estimating and tracking the spatial positions and postures of the wrist and rod parts of the surgical instrument, and the opening and closing angle of the end clamp can provide strong support for doctors, improve the accuracy, safety, and efficiency of the surgery, and also provide strong support for pose correction and auxiliary decision-making during the surgical process.
[0003] Semantic segmentation is an image processing technology whose goal is to classify each pixel in an image into different semantic categories, identify different types of instruments, biological tissue, etc. Compared with traditional image segmentation technologies, its advantage lies in enabling the computer to better recognize and understand the content in the image, with more accurate segmentation results and faster algorithm running speed. Currently, the semantic segmentation algorithms for laparoscopic surgical robot instruments mainly use convolutional neural networks or vision Transformer networks as the main structure, and the accuracy is acceptable.
[0004] Currently, there are mainly the following two types of methods for surgical instrument pose estimation: The feature point-based method estimates the pose by extracting and matching key points in the image, but in a complex surgical scenario, the extraction and matching of key points are prone to errors; the template matching-based method requires pre-establishing a three-dimensional model library of the instrument, with poor real-time performance and generalization. However, the existing methods have a large error rate, or poor real-time performance and generalization, and rarely consider the constraint effect of temporal information. The existing patent document CN118236166A discloses an automatic tracking system and method for surgical instruments, which includes: a robotic arm; a visual sensing unit arranged at the end of the robotic arm for acquiring image data; a host computer connected to the visual sensing unit for analyzing the image data based on a neural network model to obtain the position information of the surgical instrument and the position information of the biological cavity tissue, and obtaining the movement information of the robotic arm according to the position information of the surgical instrument and the position information of the biological cavity tissue; a control unit connected to the host computer and the robotic arm for controlling the movement direction of the robotic arm according to the movement information of the robotic arm. This existing technology does not perform pose estimation in combination with semantic segmentation and does not mention the opening and closing angle of the end clamp.
[0005] In addition, in practical applications, the pose feedback system of surgical instruments usually relies on various sensors in the mechanical structure, such as joint encoders, force sensors, etc. However, due to problems such as the hysteresis characteristics, plastic deformation, and backlash of the mechanical transmission structure, and the problem that the elastic deformation caused by insufficient stiffness of the instrument cannot be accurately feedback by the sensor, the pose feedback relying solely on the sensor often has delays and errors, and it is difficult to meet the requirements of precise surgical operations.
[0006] Therefore, it is urgent to improve the control accuracy and safety of laparoscopic robot surgery. Summary of the Invention
[0007] The technical problem to be solved by the present invention is:
[0008] The present invention provides a semantic segmentation and pose estimation method and system for surgical robot instruments under laparoscopic images to solve the problems that the existing semantic segmentation and pose estimation methods for surgical robot instruments have a large error rate, poor real-time performance and generalization ability, and rarely consider the constraint effect of temporal information, and it is difficult to meet the requirements of precise surgical operations.
[0009] The technical solution adopted by the present invention to solve the above technical problems is:
[0010] A semantic segmentation and pose estimation method for surgical robot instruments under laparoscopic images, the implementation process of the method is as follows:
[0011] Based on the encoder-decoder architecture, a single encoder-double decoder pose estimation model for surgical robot instruments is established, including a backbone network, a feature encoder (feature encoding module), a semantic segmentation decoder, a pose estimation decoder, and its data cache queue and time series module;
[0012] The backbone network is used to extract multi-scale information of the original endoscopic image, and the feature encoder aggregates the encoded features from the features extracted by the backbone network, including local image feature information at different abstraction levels of the image;
[0013] The semantic segmentation decoder is used to generate a segmentation image from the image feature map, and the segmentation image provides implicit geometric constraints for the pose estimation decoder;
[0014] Based on the geometric constraints provided by the semantic segmentation result, the pose estimation decoder regresses the position and attitude parameters of the wrist and rod of the surgical instrument relative to the camera, and the opening and closing angle of the instrument end from the feature map;
[0015] The time series module is used to construct a time series data cache queue for continuous frame analysis, reduce frame jitter, and improve the accuracy of the pose estimation and end opening and closing angle estimation algorithms.
[0016] Furthermore, the backbone network extracts features by concatenating MiST blocks and patch embeddings. The backbone network uses a network based on MiST blocks (Mixed Swin-Transformer) to extract target multi-scale features. The sizes of the feature maps are H / 4×W / 4, H / 8×W / 8, H / 16×W / 16, and H / 32×W / 32 in sequence. Each extraction stage contains Patch (patch) embedding and MiSTBlock for feature extraction and transformation. Specifically:
[0017] (1) Among them, the MiST block is mainly based on Swin-Transformer. Normalization is used before the Swin-Transformer-T module and fed into Swin-Transformer-T for attention operation. Its output is concatenated with the features of the input MiST block, and the concatenated result is used as the output of the MiST block;
[0018] (2) For the PatchEmbedding patch embedding, the first embedded patch uses a 5x5 convolutional kernel and a patch embedding with a stride of 4 to extract features of 1 / 4 size. Subsequent embedded patches use a 3x3 convolutional kernel and a patch embedding with a stride of 2 to extract features of 1 / 8, 1 / 16, and 1 / 32 sizes respectively;
[0019] (3) In the backbone network, the patch embedding module is concatenated with Ni (i = 1 to 4) MiST blocks. The input image passes through patch embedding and N1 MiST blocks to obtain a matrix of Feature 1, whose dimension is H / 4x W / 4x C1. Feature 2 passes through patch embedding and N2 MiST blocks to obtain a Feature 2 matrix, and so on. The value of Ci (i = 1 to 4) represents the richness of the features. C1 to C4 take 32, 64, 128, 256, and N1 to N4 take 2, 2, 2, 2. Features 1 to 4 are low-level to high-level feature information contained in the image. Low-level features include information such as texture and edges, and high-level features include information such as abstract semantic-level information.
[0020] Further, the semantic segmentation decoder is as follows: The features aggregated in the encoder are mapped through an MLP, and the output is the result of semantic segmentation. The semantic segmentation result provides constraints for the pose estimation decoder. The pose estimation decoder is as follows: To obtain the pose of the surgical instrument, the pose estimation decoder uses three cascaded MLP blocks, accepts the input of the aggregated features and the semantic segmentation result, obtains the estimated cache of the instrument pose in the current frame, and writes it into the data cache queue of the pose estimation decoder. The time series module uses an xLSTM network to balance long-term and short-term accuracy. The time series module obtains the timing cache from the pose estimation decoder and outputs the pose parameters of the instrument shaft, wrist, and the opening and closing angle of the end effector. The instrument shaft and wrist each have 6 degrees of freedom, represented by x, y, z, pitch, yaw, and roll, where: x, y, z represent the displacement of the shaft or wrist coordinate system relative to the camera coordinate system; pitch, yaw, and roll represent the rotation angle of the shaft or wrist coordinate system relative to the camera coordinate system.
[0021] Further, for the surgical robot instrument pose estimation model with a single encoder - double decoder, its training and environment settings are as follows:
[0022] (1) Optimizer and learning rate scheduling: Use the AdamW optimizer, and the default learning rate is set to 2e - 4. The learning rate scheduling is as follows: In the first quarter of the training stage, gradually increase the learning rate of the encoder from 1 / 8 of the default value to 1 / 4. In the last three - quarters of the training stage, gradually decrease the learning rate of the encoder from 1 / 4 of the default value to 1 / 16. The learning rate of the decoder remains 1 times the default learning rate until the last quarter of the training stage, and then the learning rates of the two decoders are reduced to 1 / 3 of the default learning rate.
[0023] (2) Data augmentation of the dataset: Based on the mmseg library for data augmentation, perform the following cascades: Random size cropping and random position cropping with a probability of 0.5, horizontal and vertical random flipping with a probability of 0.5, random distortion of brightness, contrast, saturation, and hue with a probability of 0.5 and a range of ±25%. Use a data wrapper to perform different data augmentations on the original data and repeat 10 times to increase the generalization of the augmented data.
[0024] (3) Training environment: The important environment configuration components are as follows: numpy = 1.19.5, CUDA = 11.3, ptorch = 1.11.0, torchvision = 0.12.0, mmcv = 1.6.0, mmseg = 0.29.1. Use two RTX3090 graphics cards for training for 20000 iters.
[0025] (4) Forward and backward propagation during training: During training, load the pre-trained weights of Swin-T and initialize the weights of the remaining network modules using the kaiming initialization method.
[0026] Furthermore, compare the semantic segmentation prediction image output by the semantic segmentation decoder with the label image of the original image through the designed loss function. The value of the loss function is used to measure the difference between the two, and the optimizer uses the derivative of the loss function to update the network parameters.
[0027] Furthermore, the loss function mixes the following several functions with corresponding weights. Specifically: (1) Edge loss function 1 Boundary_loss: Extract the gradient edge mask image of the ground truth image through the laplacian operator. The mask image is obtained by performing edge extraction on the ground truth image using a 3*3 laplacian operator; calculate dice_loss and ce_loss under the gradient edge mask image and add them to the total loss;
[0028] (2) Edge loss function 2 InverseTransformLoss: A loss function for boundary-aware segmentation that focuses on the alignment between the target and the predicted boundary; use a distance metric based on transformation parameters to measure the difference between the prediction result and the true boundary, and obtain the distance between the two according to the transformation relationship; specifically, use a small neural network to predict the affine transformation matrix between the two boundaries and output a matrix with 6 parameters, representing the affine transformation between the prediction result and the true boundary; calculate InverseTransformLoss from the above 6 parameters and add it to the total loss;
[0029] (3) diceloss: As shown in Equation (1);
[0030] The loss function diceloss based on the Dice coefficient is often used to handle the class imbalance problem
[0031]
[0032] where: p i is the probability predicted by the model, g i is the ground truth, and ∈ is a smoothing term to prevent the denominator from being zero;
[0033] (4) focalloss: As shown in Equation (2), focalloss is an improved version of the cross-entropy loss,
[0034] FL(p t ) = -α t (1 - p t ) γlog(p t ) (2)
[0035] where:
[0036] p t is the predicted probability of class t;
[0037] α t is the class weight factor (used to balance positive and negative samples), taking 0.5;
[0038] γ is the focusing parameter, used to reduce the weight of easy-to-separate samples, taking 3;
[0039] y is the value of the true label of the image pixel. When y = 1, p t = p; when y = 0, p t = 1 - p;
[0040] (5) BCELoss: as described in Equation (3);
[0041]
[0042] where N is the number of pixels participating in the loss calculation; y i is the true label of pixel i; is the predicted probability of pixel i.
[0043] A semantic segmentation and pose estimation system for surgical robot instruments under endoscopic images. The system has program modules corresponding to the steps of the above technical solution and executes the steps in the semantic segmentation and pose estimation method for surgical robot instruments under endoscopic images when running.
[0044] A computer-readable storage medium stores a computer program, and the computer program is configured to implement the steps of the semantic segmentation and pose estimation method for surgical robot instruments under endoscopic images when called by a processor.
[0045] The present invention has the following beneficial technical effects:
[0046] The present invention unifies semantic segmentation, instrument pose estimation, and temporal constraints into a multi-task learning framework; the semantic segmentation results are added as constraints to the instrument pose estimation link to improve its calculation accuracy; temporal modeling enhances the stability of pose estimation. The present invention introduces a mixed loss function and adds semantic segmentation edge information to the loss calculation. The pose estimation method based on a visual neural network proposed by the present invention can effectively solve the above technical problems by combining semantic segmentation results with temporal information. This method can provide an intelligent tool for laparoscopic robotic surgery, improving the control accuracy and safety of laparoscopic robotic surgery. The present invention can calculate the position and attitude parameters of the wrist and rod of the surgical instrument, as well as the opening and closing angle of the end clamp. It realizes the semantic segmentation of the surgical instrument, with the mIoU index of binary segmentation not less than 92%, the mIoU index of part segmentation not less than 80%, and the frame rate running on an RTX3090 graphics card not less than 20 frames per second. It realizes the learning of target features using the semantic segmentation task. Description of the Drawings
[0047] Figure 1 It is a flowchart of the method of the present invention.
[0048] Figure 2 It is a network structure block diagram of the surgical robot instrument pose estimation model in the present invention.
[0049] Figure 3 : Schematic diagram of the loss function and network parameter update in the training process (training process); the predicted image input to the semantic segmentation decoder is compared with the labeled image corresponding to the original image, and through the total loss function, the loss value (loss) is output, the loss function is differentiated, and the network is backpropagated to update the parameters of the network model.
[0050] Figure 4 : From left to right are the original image, the labeled image, and the semantic segmentation predicted image output by the semantic segmentation decoder.
[0051] Figure 5 It is a screenshot of the log output during the code training process.
[0052] Figure 6It is a coordinate system change diagram of six parameters x, y, z, pitch, yaw, and roll of the rod part and the wrist part. The origin O of the camera coordinate system, the origin O' of the rod / wrist coordinate system, (x, y, z) represents the position of O' in the O coordinate system, and (pitch, yaw, roll) represents the rotation angles of the rod / wrist coordinate system relative to the camera coordinate system, and the rotation directions are along the x / y / z axes of the camera coordinate system respectively; specifically, use a coordinate system, rotate by pitch angle along Xc, rotate by yaw angle along Yc, rotate by roll angle along Zc, and translate the origin to the (x, y, z) position of the camera coordinate system, then the rod / wrist coordinate system is obtained. Detailed implementation manners
[0053] Combined with Figure 1-6 , the implementation of the semantic segmentation and pose estimation method for the surgical robot instrument under the real endoscopic image of the present invention is elaborated as follows:
[0054] As Figure 1 shown, based on the encoder-decoder architecture, a single encoder-double decoder surgical robot instrument pose estimation model is established, including a backbone network, an encoder, a semantic segmentation decoder, a pose estimation decoder, and its data cache queue and time series module. The backbone network is used to extract multi-scale information of the original endoscopic image, and the feature encoding module aggregates the encoded features from the features extracted by the backbone network, including the local image feature information of the image at different abstraction levels; the semantic segmentation decoder is used to generate a segmentation image from the image feature map, and the segmentation image provides implicit geometric constraints for the pose estimation decoder; the pose estimation decoder, based on the geometric constraints provided by the semantic segmentation result, regresses the position and attitude parameters of the wrist and rod parts of the surgical instrument relative to the camera, and the opening and closing angles of the instrument end; the time series module is used to construct a time series data cache queue for continuous frame analysis, reduce inter-frame jitter, and improve the accuracy of the pose estimation and end opening and closing angle estimation algorithms.
[0055] Specifically, the network structure is as Figure 2 shown, and features are extracted through the backbone network shown in Figure 2 by connecting the MiST block and patch embedding in series. The backbone network uses a network based on the MiST block (Mixed Swin-Transformer) to extract target multi-scale features, and the feature map sizes are H / 4×W / 4, H / 8×W / 8, H / 16×W / 16, and H / 32×W / 32 in sequence. Each extraction stage includes Patch (patch) embedding and MiSTBlock for feature extraction and transformation.
[0056] (1) Among them, the MiST block is mainly based on Swin-Transformer. Normalization is used before the Swin-Transformer-T module, and the input is fed into Swin-Transformer-T for attention operation. Its output is concatenated with the features input to the MiST block, and the concatenated result is used as the output of the MiST block.
[0057] (2) For the PatchEmbedding patch embedding, the first embedded patch uses a 5x5 convolutional kernel and a patch embedding with a stride of 4 to extract features of 1 / 4 size. Subsequent embedded patches use a 3x3 convolutional kernel and a patch embedding with a stride of 2 to extract features of 1 / 8, 1 / 16, and 1 / 32 sizes respectively.
[0058] (3) Figure 2 In it, C1 to C4 take 32, 64, 128, 256, and N1 to N4 can take 2, 2, 2, 2. Thus, it is ensured that high-level abstract features have a sufficient number of channels to store diverse abstract feature representations. Through the above design, image features can be captured at different scales, hierarchical feature representations can be constructed, and rich feature information can be provided for downstream tasks of subsequent decoders.
[0059] Feature 1 to Feature 4 are low-level to high-level feature information contained in the image. Low-level features such as texture and edge information, and high-level features such as abstract semantic-level information. The matrix obtained from the MiST block for Feature 1 contains the feature information in the image. The dimension of Feature 1 is H / 4 x W / 4 x C1, and Features 2 to 4 are similar. The value of Cx represents the richness of the features. C1 to C4 take [32, 64, 128, 128], and N1 to N4 are the stacking times of the MiST block, taking [2, 2, 2, 2]. In the backbone network, the patch embedding module is connected in series with Ni MiST blocks. The input image passes through the patch embedding and N1 MiST blocks to obtain the matrix of Feature 1, whose dimension is H / 4 x W / 4 x C1. Feature 2 passes through the patch embedding and N2 MiST blocks to obtain the Feature 2 matrix, and so on.
[0060] The aggregated features are mapped through an MLP, and the output is the result of semantic segmentation. The semantic segmentation result provides constraints for the pose estimation decoder. The pose estimation decoder uses three MLP blocks connected in series, accepts the input of the aggregated features and the semantic segmentation result, obtains the estimated cache of the instrument pose in the current frame, and writes it into the data cache queue. The time series module uses an xLSTM network to balance long-term and short-term accuracy. The time series module obtains the timing cache from the pose estimation decoder and outputs the pose parameters (6 degrees of freedom, x, y, z, pitch, yaw, roll) of the instrument rod and wrist, as well as the opening and closing angle of the end effector.
[0061] Dataset
[0062] (1) Endovis2017 Segmentation Dataset: A robotic sub-dataset from the instrument segmentation and tracking sub-challenge of the Endoscopic Vision Challenge, consisting of 10 abdominal surgery sequences recorded by the da Vinci surgical robot system. It contains 3,000 images, each with a valid size of 1280 * 1024 pixels and contains 1 - 2 instruments.
[0063] (2) CholecInstanceSeg Segmentation Dataset: Released in June 2024, including 41,900 annotated frames and 64,400 tool instances, each with a semantic mask and instance ID. Each image contains a specific difficult case, including motion blur, smoke, reflection, soft tissue occlusion or attachment, overexposure or underexposure, etc.
[0064] (3) Surgical Video and Corresponding Instrument Pose Dataset: A private dataset with a total video duration of 6 hours, 30 frames per second for video data, and 2 frames per second for instrument pose data. For frames with instrument pose data, the frame has corresponding instrument semantic segmentation annotations.
[0065] Training and Environment Settings
[0066] (1) Optimizer and Learning Rate Scheduling: Use the AdamW optimizer with a default learning rate set to 2e - 4. The learning rate scheduling is as follows: In the first quarter of training, gradually increase the learning rate of the encoder from 1 / 8 of the default value to 1 / 4. In the last three - quarters of training, gradually decrease the learning rate of the encoder from 1 / 4 of the default value to 1 / 16. In this way, it takes into account the learning speed in the early stage of the model and the adjustment fineness in the later stage, helping the model to reach a better range optimal solution in the later stage. The learning rate of the decoder remains 1 times the default learning rate until the last quarter of training, and the learning rates of the two decoders are reduced to 1 / 3 of the default learning rate.
[0067] (2) Data augmentation for the dataset: Based on the mmseg library, data augmentation is performed as follows: concatenate random size cropping and random position cropping with a probability of 0.5, horizontal and vertical random flipping with a probability of 0.5, and random distortion of brightness, contrast, saturation, and hue with a probability of 0.5 and a range of ±25%. Use a data wrapper to perform different data augmentations on the original data and repeat 10 times to increase the generalization of the augmented data. In addition, before final deployment, in order to make full use of the existing data, the K-Fold Cross-Validation method is used to divide the validation set. Specifically, take K = 5, divide the dataset into 5 parts, take 1 part as the validation set each time, and use the rest as the training set, repeat 5 times, and the model evaluation metric is the average of the 5 validation results. This method can balance the random error under different data partitions, provide a more stable and reliable model performance evaluation, reduce the network's dependence on a specific part of the data during the final training, and ensure that overfitting does not occur.
[0068] (3) Training environment: The important environment configuration components are as follows: numpy = 1.19.5, CUDA = 11.3, ptorch = 1.11.0, torchvision = 0.12.0, mmcv = 1.6.0, mmseg = 0.29.1. Use two RTX3090 graphics cards for training for 20,000 iterations.
[0069] (4) Forward and backward propagation during training: During training, load the pre-trained weights of Swin-T and initialize the weights of the remaining network modules using the kaiming initialization method. Use the preprocessed image as the network input, perform the forward propagation of the network, obtain the semantic segmentation result from the semantic segmentation decoder, compare it with the ground truth image, obtain the loss of the semantic segmentation task through a specially designed loss function, and then perform backpropagation of the gradient to update the network parameters; compare the pose estimation parameters obtained from the pose estimation decoder with the ground truth and perform backpropagation; compare the optimized pose estimation parameters obtained through the data cache queue and time series module with the ground truth and perform backpropagation.
[0070] Loss function design
[0071] Use a mixed loss function and introduce the edge loss functions Boundary and InverseTransformLoss. Mix the following loss functions according to the corresponding weights. The weight of the edge loss function can be appropriately increased in the later stage of training.
[0072] (1) Edge loss function 1Boundary_loss: Extract the gradient edge mask image of the ground truth image through the laplacian operator. The mask image is obtained by performing edge extraction on the ground truth image using a laplacian operator of size 3*3. Calculate dice_loss and ce_loss under the gradient edge mask image and add them to the total loss.
[0073] (2) Edge loss function 2InverseTransformLoss: A loss function for boundary-aware segmentation that focuses on the alignment between the target and the predicted boundary. Use a distance metric based on transformation parameters to measure the difference between the prediction result and the true boundary, and obtain the distance between the two according to the transformation relationship. Specifically, use a small neural network to predict the affine transformation matrix between the two boundaries, and output a matrix with 6 parameters, representing the affine transformation between the prediction result and the true boundary. Calculate InverseTransformLoss from the above 6 parameters and add it to the total loss.
[0074] (3) diceloss: As shown in Equation (1);
[0075] The loss function diceloss based on the Dice coefficient is often used to handle class imbalance problems
[0076]
[0077] where: p i is the probability predicted by the model, g i is the ground truth label, and ∈ is a smoothing term to prevent the denominator from being zero.
[0078] (4) focalloss: As shown in Equation (2);
[0079] focalloss is an improved version of the cross-entropy loss,
[0080] FL(p t )=-α t (1-p t ) γ log(p t ) (2)
[0081] where:
[0082] p t is the predicted probability of class t;
[0083] α t is the class weight factor (used to balance positive and negative samples), taking 0.5;
[0084] γ is the focusing parameter, which is used to reduce the weight of easily separable samples and takes the value of 3;
[0085] y is the value of the true label of the image pixel. When y = 1, p t = p t ; when y = 0, p t = 1 - p t .
[0086] (5) BCELoss: as described in Equation (3).
[0087]
[0088] where N is the number of pixels participating in the loss calculation; y i is the true label of pixel i; is the predicted probability of pixel i.
[0089] The weighted weights of the above five loss functions are obtained through debugging, and the weight ratio used in the actual method is 1.5:1:0.2:1:1.
[0090] Evaluation metrics
[0091] include but are not limited to mIoU, as described in Equation (1-4). mIoU is used to measure the accuracy rate between the prediction P and the true image GT.
[0092]
[0093] Among them:
[0094] C is the total number of categories,
[0095] P i represents the predicted region of the i-th category,
[0096] GT i represents the region of the label of the i-th category,
[0097] |P i ∩GT i | and |P i ∪GT i | represent the area of the intersection and the area of the union of the two regions.
[0098] The present invention can calculate the position and attitude parameters of the wrist and rod parts of the surgical instrument and the opening and closing angle of the end clamp. It realizes the semantic segmentation of the surgical instrument, the mIoU index of binary segmentation is not less than 92%, the mIoU index of part segmentation is not less than 80%, and the frame rate when running on an RTX3090 graphics card is not less than 20 frames per second. It can realize learning the target features using the semantic segmentation task.
[0099] In summary, the method proposed by the present invention for pose estimation based on a visual neural network can effectively solve the technical problems proposed by the present invention by combining semantic segmentation results with temporal information. The method of the present invention can provide an intelligent tool for laparoscopic robotic surgery, improving the control accuracy and safety of laparoscopic robotic surgery. The method of the present invention has been verified by simulation experiments and practical applications, both of which have verified the claimed technical effects and practicality of the present invention. As Figure 5 shown, it is a screenshot of the log output during the code training process, and the code based on the present invention has been developed and actually applied.
[0100] The method of the present invention has been verified by simulation experiments and practical applications, both of which have verified the claimed technical effects of the present invention.
[0101] The algorithm (method) proposed by the present invention is the underlying technical core of the present invention, and various products can be derived based on the algorithm.
[0102] Based on the algorithm (method) proposed by the present invention, a semantic segmentation and pose estimation system for surgical robot instruments under endoscopic images is developed using a programming language. The system has program modules corresponding to the steps of the above technical solution, and executes the steps in the semantic segmentation and pose estimation method of the surgical robot instruments under endoscopic images when running.
[0103] The computer program of the developed system (software) is stored on a computer-readable storage medium, and the computer program is configured to implement the steps of the semantic segmentation and pose estimation method of the surgical robot instruments under endoscopic images when called by a processor. That is, the present invention is materialized on a carrier to become a computer program product.
[0104] The software developed based on the present invention is configured on surgical robot instruments to form a terminal product.
[0105] Various embodiments of the systems and technologies described in the present invention can be implemented in digital electronic circuit systems, integrated circuit systems, dedicated ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs, the one or more computer programs can be executed and / or interpreted on a programmable system including at least one programmable processor, the programmable processor can be a dedicated or general programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0106] The computing procedures (also referred to as programs, software, software applications, or code) in the present invention include machine instructions for a programmable processor, and these computing procedures can be implemented using high-level procedures and / or object-oriented programming languages, and / or assembly / machine languages. As used in the present invention, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., magnetic disks, optical disks, memories, programmable logic device PLD) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0107] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this application can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this application can be achieved, and all are within the protection scope of the present invention.
Claims
1. A method for semantic segmentation and pose estimation of surgical robot instruments in laparoscopic images, characterized in that: The implementation process of the method is: Based on the encoder-decoder architecture, a single encoder-dual decoder surgical robot instrument pose estimation model is established, including a backbone network, feature encoder, semantic segmentation decoder, pose estimation decoder and its data cache queue and time series module; The backbone network is used to extract multi-scale information of the original endoscope image. The feature encoder aggregates the encoded features from the features extracted by the backbone network, which contains the local feature information of the image at different abstraction levels. The semantic segmentation decoder is used to generate a segmentation image from the image feature map. The segmentation image provides implicit geometric constraints for the pose estimation decoder. Based on the geometric constraints provided by the semantic segmentation results, the pose estimation decoder regresses the position and pose parameters of the wrist and rod of the surgical instrument relative to the camera, as well as the opening and closing angle of the instrument end from the feature map; Use the time series module to build a time series data cache queue for continuous frame analysis, reduce inter-frame jitter, and improve the accuracy of the pose estimation and terminal opening and closing angle estimation algorithms; The semantic segmentation decoder is as follows: the features aggregated in the encoder are mapped through an MLP and the output is the result of semantic segmentation. The semantic segmentation result provides constraints for the pose estimation decoder. The pose estimation decoder is as follows: To obtain the pose of the surgical instrument, the pose estimation decoder uses three series-connected MLP blocks, accepts the input of aggregated features and semantic segmentation results, obtains the estimated cache of the instrument pose in the current frame, and writes it into the data cache queue of the pose estimation decoder; the time series module uses the xLSTM network to take into account both long-term and short-term accuracy; The time series module obtains the timing buffer from the pose estimation decoder and outputs the pose parameters of the instrument rod and wrist and the opening and closing angle of the end fixture; the instrument rod and wrist have 6 degrees of freedom respectively, represented by x, y, z, pitch, yaw, and roll, where: x, y, and z represent the displacement of the coordinate system of the rod or wrist relative to the camera coordinate system; pitch, yaw, and roll represent the rotation angle of the coordinate system of the rod or wrist relative to the camera coordinate system.
2. The method for semantic segmentation and pose estimation of surgical robot instruments under laparoscopic images according to claim 1, characterized in that: The backbone network extracts features by connecting MiST blocks and patch embedding in series. The backbone network uses a network based on MiST blocks (Mixed Swin-Transformer) to extract multi-scale features of the target. The feature map sizes are H / 4×W / 4, H / 8×W / 8, H / 16×W / 16, and H / 32×W / 32 respectively; each extraction stage contains Patch embedding and MiSTBlock for feature extraction and transformation, specifically: (1) Among them, The MiST block is mainly based on Swin-Transformer. Normalization is used before the Swin-Transformer-T module and input into Swin-Transformer-T for attention operation. Its output is concatenated with the features of the input MiST block and the concatenated output is used as the output of the MiST block. (2) For PatchEmbedding, the first embedded patch uses a 5x5 convolution kernel and a stride of 4 to extract features of 1 / 4 size. The subsequent embedded patches use a 3x3 convolution kernel and a stride of 2 to extract features of 1 / 8, 1 / 16, and 1 / 32 sizes, respectively. (3) In the backbone network, the patch embedding module is connected in series with Ni (i=1~4) MiST blocks. The input image undergoes patch embedding and N1 MiST blocks to obtain the matrix of feature 1, whose dimension is H / 4x W / 4x C1. Feature 2 undergoes patch embedding and N2 MiST blocks to obtain the matrix of feature 2, and so on. The value of Ci (i=1~4) represents the richness of the feature. C1~C4 takes 32, 64, 128, 256, and N1~N4 takes 2, 2, 2, 2. Feature 1~Feature 4 are the low-level to high-level feature information contained in the image. The low-level features include texture and edge information, and the high-level features include abstracted semantic level information.
3. The method for semantic segmentation and pose estimation of surgical robot instruments under laparoscopic images according to claim 1, characterized in that: For the single encoder-dual decoder surgical robot instrument pose estimation model, the training and environment settings are as follows: (1) Optimizer and learning rate scheduling: The AdamW optimizer is used, and the default learning rate is set to 2e-4. The learning rate scheduling is as follows: in the first quarter of the training, the encoder's learning rate is gradually increased from the default value of 1 / 8 to 1 / 4, and in the last three quarters of the training, the encoder's learning rate is gradually reduced from the default value of 1 / 4 to 1 / 16. The decoder learning rate is kept at 1 times the default learning rate until the last quarter of training, when the learning rates of both decoders are reduced to 1 / 3 of the default learning rate; (2) Dataset data enhancement: Data enhancement is performed based on the mmseg library, and the following are performed in series: random size cropping and random position cropping with a probability of 0.5, horizontal and vertical random flipping with a probability of 0.5, and random distortion of brightness, contrast, saturation, and hue with a probability of 0.5 and a range of ±25%. The original data is enhanced with different data enhancements using a data wrapper, and the data is repeated 10 times to increase the generalization of the enhanced data; (3) Training environment: The important environment configuration components are as follows: numpy = 1.19.5, CUDA = 11.3, ptorch = 1.11.0, torchvision = 0.12.0, mmcv = 1.6.0, mmseg = 0.29.1; two RTX3090 graphics cards are used for graphics card training, and 20,000 iter are trained; (4) Forward and backward propagation of training: During training, the pre-trained weights of Swin-T are loaded, and the weights of the remaining network modules are initialized using the kaiming initialization method.
4. The method for semantic segmentation and pose estimation of surgical robot instruments under laparoscopic images according to claim 3, characterized in that: The semantic segmentation prediction image output by the semantic segmentation decoder is compared with the label image of the original image through the designed loss function, and the value of the loss function is used to measure the difference between the two.
5. The method for semantic segmentation and pose estimation of surgical robot instruments under laparoscopic images according to claim 4, characterized in that: The loss function mixes the following functions according to the corresponding weights, specifically: (1) Edge loss function 1Boundary_loss: The gradient edge mask map of the true value image is extracted by the laplacian operator. The mask map is obtained by using the laplacian operator of size 3*3 to extract the edge of the true value image; dice_loss and ce_loss are calculated under the corresponding gradient edge mask map and added to the total loss; (2) Edge loss function 2InverseTransformLoss: A boundary-aware segmentation loss function that focuses on the alignment between the target and the predicted boundary. It uses a distance metric based on the transformation parameters to measure the difference between the predicted result and the true boundary, and obtains the distance between the two based on the transformation relationship. Specifically, a small neural network is used to predict the affine transformation matrix between the two boundaries, and outputs a matrix with 6 parameters, which represents the affine transformation between the predicted result and the true boundary. The InverseTransformLoss is calculated from the above 6 parameters and added to the total loss. (3) diceloss: as shown in formula (1); The loss function diceloss based on the Dice coefficient is often used to deal with class imbalance problems Where: p i is the probability predicted by the model, g i is the true label (ground truth), ∈ is a smoothing term to prevent the denominator from being 0; (4) focal loss: as shown in formula (2); FL(p t )=-a t (1-p t ) γ log(p t ) (2) in: p t is the predicted probability of class t; α t is the category weight factor, which is 0.5; γ is the focusing parameter, which is used to reduce the weight of easy-to-separate samples and is set to 3; y is the value of the true image pixel label. When y=1, p t =p; when y = 0, p t =1-p; (5) BCELoss: as described in formula (3); Where N is the number of pixels involved in the loss calculation; y i is the true label of pixel i; is the predicted probability of pixel i.
6. A semantic segmentation and pose estimation system for surgical robot instruments under laparoscopic images, characterized by: The system has a program module corresponding to the steps of any one of claims 1 to 5, and executes the steps in the method for semantic segmentation and posture estimation of surgical robot instruments under laparoscopic images when running.
7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is configured to implement the steps of the method for semantic segmentation and pose estimation of a surgical robot instrument under a laparoscopic image according to any one of claims 1 to 5 when called by a processor.
Citation Information
Patent Citations
Automatic tracking system and method for surgical instrument
CN118236166A
Method, device and equipment for calculating absolute and relative poses of endoscope in operation
CN117671012A
Monocular self-supervision depth estimation method and system for laparoscope video image
CN117876453A