Surgical instrument needle tip track prediction method

Through the generative adversarial network model, the 3D convolutional neural network and multi-scale discriminator are used to solve the shortcomings of the existing technology in three-dimensional spatial processing, and high-precision needle tip trajectory prediction of surgical instruments is achieved, improving the accuracy of surgical operations.

CN119963593AActive Publication Date: 2025-05-09CHANGCHUN UNIV OF SCI & TECH
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510101819.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-09
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

The existing surgical instrument needle tip trajectory prediction methods have poor results when dealing with three-dimensional space, resulting in prediction frame blur, artifact and pixel loss, which in turn affects prediction efficiency and accuracy.

Method used

The generative adversarial network model is adopted, and the 3D convolutional neural network is used as a generator and multi-scale discriminator as a discriminator to generate a high-precision needle tip motion trajectory prediction image sequence through adversarial training.

Benefits of technology

It improves the accuracy and detailed performance of the needle tip trajectory prediction of surgical instruments, achieves efficient and high-precision prediction results, and enhances the accuracy and reliability of surgical operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963593A_ABST
    Figure CN119963593A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of deep learning, and particularly discloses a surgical instrument needle tip trajectory prediction method, which comprises the following steps of: constructing a surgical instrument motion trajectory data set by using a surgical instrument key motion image sequence, the method comprises the following steps: dividing a key motion image sequence of a surgical instrument into a first image sequence and a second image sequence with the same frame number according to a time sequence, and constructing a generative adversarial network model by taking a 3D convolutional neural network as a generator and a multi-scale discriminator as a discriminator; and training the generative adversarial network model based on a mapping relation between the first image sequence and the second image sequence, and then fixing parameters of the generator and the discriminator. And obtaining a real-time motion track image sequence of the surgical instrument, inputting the real-time motion track image sequence into the generative adversarial network model to generate a prediction image sequence, and obtaining a prediction result of the tip track of the surgical instrument based on the prediction image sequence. According to the invention, high-precision prediction is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of deep learning and relates to a method for predicting the needle tip trajectory of a surgical instrument. Background Art

[0002] At present, the popularity of surgical navigation systems is gradually increasing worldwide, which means that the market prospects of surgical navigation systems are very optimistic. During high-precision surgical operations, surgical navigation systems need to track the position of surgical instrument needle tips in real time to prevent misoperation. As the complexity of surgical operations continues to increase, the precise control and positioning of surgical instrument needle tips has become one of the key factors in improving the success rate of surgery.

[0003] The currently used methods for predicting the trajectory of surgical instrument needle tips mainly include geometric modeling, physics-based simulation, and adaptive control strategies. Geometric modeling predicts the trajectory by establishing a geometric model of the interaction between the needle tip and the tissue, physics-based simulation uses physical laws to simulate the behavior of the needle tip in the tissue, and adaptive control strategies dynamically adjust the motion path of the needle tip through feedback mechanisms. The above-mentioned traditional methods have promoted the development of needle tip trajectory prediction to a certain extent, but the processing effect of three-dimensional space is poor, which will cause blurring, artifacts, and even pixel loss in the prediction frame, resulting in significant limitations in prediction efficiency and accuracy, which needs further improvement.

[0004] In summary, how to efficiently and accurately predict the needle tip trajectory of surgical instruments remains a difficult problem. Summary of the invention

[0005] The purpose of the present invention is to provide a method for predicting the trajectory of a surgical instrument needle tip, which improves the accuracy and detail of the prediction of the trajectory of a surgical instrument needle tip and achieves efficient and high-precision prediction.

[0006] To achieve the above purpose, the specific technical solutions provided by the present invention are as follows: A method for predicting the trajectory of a surgical instrument needle tip comprises the following steps: Acquire a historical motion trajectory of the surgical instrument, wherein the historical motion trajectory includes all motion forms of the surgical instrument, and extract a key motion image sequence of the surgical instrument in each motion trajectory to construct a surgical instrument motion trajectory dataset; Preprocessing a surgical instrument key motion image sequence in a surgical instrument motion trajectory data set, and dividing the preprocessed surgical instrument key motion image sequence into a first image sequence and a second image sequence with the same number of frames in chronological order; A generative adversarial network model is constructed using a 3D convolutional neural network as a generator and a multi-scale discriminator as a discriminator, and after training the generative adversarial network model based on a mapping relationship between the first image sequence and the second image sequence, the prediction accuracy of the generative adversarial network model is evaluated, and when the prediction accuracy of the generative adversarial network model reaches a preset accuracy, the parameters of the generator and the discriminator are fixed; A real-time motion trajectory image sequence of the surgical instrument is acquired, and the real-time motion trajectory image sequence is input into the generator to generate a predicted image sequence, and a predicted result of the surgical instrument needle tip trajectory is obtained based on the predicted image sequence.

[0007] Preferably, preprocessing the key motion image sequence of the surgical instrument in the surgical instrument motion trajectory data set comprises the following steps: Remove noise from each frame of a critical motion image sequence of surgical instruments; The pixel values ​​of each frame in the key motion image sequence of surgical instruments are normalized and mapped from [0-255] to [0-1].

[0008] Preferably, training the generative adversarial network model based on the mapping relationship between the first image sequence and the second image sequence comprises the following steps: The surgical instrument motion trajectory dataset is divided into a training set and a test set; Select a complete multi-frame image sequence consisting of a first image sequence and a second image sequence from the training set as an input of the discriminator; Send the first image sequence in the training set as input to the generator, generate a predicted image sequence, calculate the generator loss function, and update the generator; The predicted image sequence generated by the generator is input into the discriminator, the discriminator checks the temporal consistency between the generated predicted image sequence and the first image sequence, calculates the discriminator loss function, and updates the discriminator until the discriminator loss function is minimized.

[0009] Preferably, evaluating the prediction accuracy of the generative adversarial network model comprises the following steps: The first image sequence in the test set is obtained as the input of the generator, and the generator generates a predicted image sequence; A complete multi-frame image sequence consisting of the first image sequence and the second image sequence in the test set is obtained as the input of the discriminator to evaluate the difference between the predicted image and the real image, and to verify the error between the three-dimensional coordinates of the surgical instrument needle tip in the predicted image and the three-dimensional coordinates of the surgical instrument needle tip in the real image. The above results are used to comprehensively evaluate the prediction accuracy of the generative adversarial network model.

[0010] Preferably, the difference between the predicted image and the real image is evaluated using the following formula: , in, Represents the number of frames of the image in the predicted image sequence, For the The real image corresponding to the input image sequence is For the The predicted image corresponding to the input image sequence, MSE is the mean square error, which is used to evaluate the difference between the predicted image and the real image. is a natural number, The value range is 1 to ; The lower the mean square error (MSE) value, the higher the prediction accuracy.

[0011] Preferably, verifying the error between the three-dimensional coordinates of the surgical instrument needle tip in the predicted image and the three-dimensional coordinates of the surgical instrument needle tip in the real image comprises the following steps: respectively extracting the three-dimensional coordinates of the needle tip of the surgical instrument in the predicted image and the three-dimensional coordinates of the needle tip of the surgical instrument in the real image; The error between the three-dimensional coordinates of the surgical instrument needle tip in the predicted image and the three-dimensional coordinates of the surgical instrument needle tip in the real image is calculated using the following formula: , in, Indicates the number of frames of the image in the predicted image sequence, is the three-dimensional coordinate of the needle tip in the real image, To predict the three-dimensional coordinates of the needle tip in the image, is the error, is a natural number, The value range is 1 to ; When the error When it is less than 1, it indicates that the needle tip trajectory of the surgical instrument has achieved the predicted accuracy.

[0012] Preferably, obtaining the prediction result of the needle tip trajectory of the surgical instrument based on the prediction image sequence includes the following steps: Determine the three-dimensional coordinates of three markers on the surgical instrument based on the predicted image sequence; The three-dimensional coordinates of the needle tip of the surgical instrument are determined based on the three-dimensional coordinates of the three markers on the surgical instrument to obtain a prediction result of the trajectory of the needle tip of the surgical instrument.

[0013] Preferably, determining the three-dimensional coordinates of three markers on the surgical instrument based on the predicted image sequence comprises the following steps: Based on the predicted image sequence generated by the generative adversarial network model, the edge contours of the three landmarks on the surgical instrument are obtained using the Canny edge detection algorithm; Randomly select three pixel points on the edge contour of the marker and determine the two-dimensional center coordinates of the marker according to the following formula , , in, , and are the coordinates of three non-collinear points on the edge contour of the landmark, is the two-dimensional center coordinate of the marker; Matching the landmarks in the images acquired by the two cameras of the binocular vision measurement device through the regional similarity matching algorithm; The three-dimensional coordinates of the marker are determined as follows: , in, , is the center distance between the two cameras of the binocular vision measurement device, is the focal length of the binocular vision measurement device, , are the two-dimensional coordinates of the landmark in the images acquired by the two cameras, are the three-dimensional coordinates of the landmark.

[0014] Preferably, the three-dimensional coordinates of the surgical instrument needle tip are determined based on the following formula: , in, , , , is the coordinate of the surgical instrument needle tip in the surgical instrument coordinate system, is the three-dimensional coordinate of the surgical instrument needle tip, , and are the coordinates of marker 1, marker 2 and marker 3 in three-dimensional space, is the translation vector, is the rotation vector, They are the unit vectors in the x, y, and z directions of the surgical instrument coordinate system.

[0015] Compared with the prior art, the present invention provides a method for predicting the needle tip trajectory of a surgical instrument. By acquiring the historical motion trajectory of the surgical instrument, a key motion image sequence of the surgical instrument is extracted from each motion trajectory to construct a surgical instrument motion trajectory data set. The key motion image sequence of the surgical instrument is preprocessed, and the preprocessed key motion image sequence of the surgical instrument is divided into a first image sequence and a second image sequence with the same number of frames in chronological order. A generative adversarial network model is constructed using a 3D convolutional neural network as a generator and a multi-scale discriminator as a discriminator. The generative adversarial network model is trained based on the mapping relationship between the first image sequence and the second image sequence, and the prediction accuracy of the generative adversarial network model is evaluated. When the prediction accuracy of the generative adversarial network model reaches a preset accuracy, the parameters of the generator and the discriminator are fixed. The real-time motion trajectory image sequence of the surgical instrument is acquired, and the real-time motion trajectory image sequence is input into the generator to generate a predicted image sequence, and the prediction result of the needle tip trajectory of the surgical instrument is obtained based on the predicted image sequence. The present invention innovatively applies the generative adversarial network model to the prediction of surgical instrument needle tip trajectory, wherein the generator in the generative adversarial network model is a 3D convolutional neural network, and the discriminator is a multi-scale discriminator. The generator attempts to generate realistic predicted images to deceive the discriminator, while the discriminator attempts to distinguish between real images and predicted images. This adversarial process prompts the generator to continuously improve the authenticity of the generated data, further improves the accuracy and detail performance of surgical instrument needle tip trajectory prediction, and achieves efficient and high-precision prediction. It is highly practical and worthy of promotion. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a flow chart of the present invention.

[0017] Figure 2 This is a simulation scene diagram of the present invention.

[0018] Figure 3 It is a partial view of the present invention.

[0019] Figure 4 The generative adversarial network model used in the present invention.

[0020] Reference numerals: 1. Three-axis sliding guide test bench; 2. Surgical instruments; 3. Binocular vision measurement equipment; 4. Computer; 5. Marker 1; 6. Marker 2; 7. Marker 3; 8. Surgical instrument needle tip. DETAILED DESCRIPTION

[0021] In order to accurately predict the trajectory of the needle tip of a surgical instrument and help surgical robots or doctors perform surgical operations more accurately, the existing prediction methods have poor processing effects on three-dimensional space, which will cause blurring, artifacts and even pixel missing in the prediction frames, resulting in significant limitations in prediction efficiency and prediction accuracy. The present invention provides a method for predicting the trajectory of the needle tip of a surgical instrument to solve the above-mentioned technical problems.

[0022] In order to enable those skilled in the art to better understand the technical solution of the present invention and to implement it, the following will be combined with the attached Figure 1 To the attached Figure 4 , the technical solution in the present invention is described clearly and in detail.

[0023] In the description of the present invention, it is to be understood that the terms “center”, “longitudinal”, “lateral”, “length”, “width”, “thickness”, “up”, “down”, “front”, “back”, “left”, “right”, “vertical”, “horizontal”, “top”, “bottom”, “inside”, “outside”, “clockwise”, “counterclockwise”, “axial”, “radial”, “circumferential”, etc., indicating orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as limiting the present invention.

[0024] In addition, it should be further explained that, in the description of the embodiments of the present invention, unless otherwise specified, “ / ” means or, for example, A / B can mean A or B: “and / or” in the text is merely a description of the association relationship of associated objects, indicating that three relationships may exist, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present invention, “multiple” refers to two or more than two.

[0025] The following terms "first", "second", "third" and "fourth" are used for descriptive purposes only and should not be understood as suggesting or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, features defined as "first", "second", "third" and "fourth" may explicitly or implicitly include one or more of the features. In the description of the present invention, unless otherwise specified, "plurality" means two or more.

[0026] Example 1 The flowchart for predicting the trajectory of the surgical instrument needle tip is as follows Figure 1 As shown, the specific steps include: The historical motion trajectory of the surgical instrument 2 is obtained, and the historical motion trajectory includes all motion forms of the surgical instrument 2. A key motion image sequence of the surgical instrument is extracted from each motion trajectory to construct a surgical instrument motion trajectory dataset.

[0027] The key motion image sequence of the surgical instrument in the surgical instrument motion trajectory data set is preprocessed, and the preprocessed key motion image sequence of the surgical instrument is divided into a first image sequence and a second image sequence with the same number of frames in chronological order.

[0028] A generative adversarial network model is constructed using a 3D convolutional neural network as a generator and a multi-scale discriminator as a discriminator. After the generative adversarial network model is trained based on the mapping relationship between the first image sequence and the second image sequence, the prediction accuracy of the generative adversarial network model is evaluated. When the prediction accuracy of the generative adversarial network model reaches a preset accuracy, the parameters of the generator and the discriminator are fixed.

[0029] A real-time motion trajectory image sequence of the surgical instrument 2 is acquired, and the real-time motion trajectory image sequence is input into the generator to generate a predicted image sequence, and a predicted result of the surgical instrument needle tip trajectory is obtained based on the predicted image sequence.

[0030] Among them, the detailed implementation plan of the main multiple steps is as follows: 1. Construct a surgical instrument motion trajectory dataset Figure 2 It is a simulation scene diagram. The required devices in the scene diagram include surgical instruments 2, binocular vision measurement equipment 3, three-axis sliding guide test bench 1 and computer 4, wherein the surgical instruments 2, binocular vision measurement equipment 3 and three-axis sliding guide test bench 1 are all electrically connected to the computer 4.

[0031] In order to obtain the historical motion trajectory of the surgical instrument 2, the surgical instrument 2 is first fixed on the three-axis sliding guide test bench 1 so that it faces the binocular vision measurement device 3, such as Figure 2 Then, the three-axis sliding guide test bench 1 is controlled to slowly move the surgical instrument 2, and the binocular vision measurement device 3 is used to collect the video of the movement of the surgical instrument 2. These videos need to cover the movement of the surgical instrument 2 on the three-axis sliding guide test bench 1. , , All forms of motion in three directions. During the video acquisition process, special attention should be paid to the lighting conditions to ensure that the image is clear and without distortion. Finally, 1,000 videos were collected, and a continuous 14-frame sequence of key motion images of surgical instruments was extracted from each video to construct a surgical instrument motion trajectory dataset.

[0032] 2. Preprocessing of key motion image sequences of surgical instruments Due to the historical motion trajectory video sequence of surgical instrument 2 and the inevitable presence of noise in the video, in order to prevent the influence of noise on the generative adversarial network model, the key motion image sequence of the surgical instrument needs to be preprocessed to improve the image quality. Specifically, the Kalman filtering technology is used to denoise the image, so as to more accurately estimate the true state of each frame image and reduce noise interference.

[0033] Next, the image pixel values ​​are normalized and mapped from [0-255] to [0-1] to improve the stability and convergence speed of network model training.

[0034] 3. Building a Generative Adversarial Network Model Specifically, in practical application, the present invention predicts the trajectory of the surgical instrument in the next 7 frames by analyzing the historical motion data of the surgical instrument 2 during the operation, especially the trajectory of the surgical instrument in the first 7 frames. The core idea of ​​this method is to use the generation ability and adversarial learning mechanism of the generative adversarial network model, combined with the binocular vision measurement device 3, to accurately predict the future motion path of the surgical instrument needle tip 8.

[0035] Generative Adversarial Network (GAN) is a deep learning model that generates data through adversarial training. GAN consists of two main parts: the generator and the discriminator. The generator tries to generate realistic predicted images to deceive the discriminator, while the discriminator tries to distinguish between real images and predicted images. This adversarial process prompts the generator to continuously improve the accuracy of the predicted images. In the prediction of the motion trajectory of surgical instruments 2, GAN can be used to generate realistic motion trajectories of surgical instruments 2, which can then be used for simulation, training or auxiliary decision-making during surgery. Specifically, the design and training of the generator and discriminator are as follows:

[0036] The generator uses a 3D Convolutional Neural Network, 3D-CNN, which has the advantage of being able to capture the motion patterns of surgical instrument 2 in space and time. 3D-CNN can effectively process and generate complex spatiotemporal data by performing convolution operations in spatial and temporal dimensions. The main structure of the generator includes an input layer, a 3D convolution layer, an upsampling layer, and an output layer. The input layer is used to receive the motion trajectory of the surgical instrument 2 at the previous moment. The 3D convolution layer is used to extract spatiotemporal features. By using multiple 3D convolution layers, the complex dynamics of the movement of the surgical instrument 2 can be captured. The upsampling layer upsamples the low-dimensional feature map to the target dimension through methods such as deconvolution or interpolation to generate the motion trajectory of the surgical instrument 2. The output layer outputs the generated motion trajectory data of the surgical instrument 2.

[0037] The discriminator adopts the design of Multi-Scale Discriminator, which can evaluate the generated data at different scales, thereby improving the discriminator's sensitivity to details and grasp of the overall structure. The main structure of the multi-scale discriminator includes an input layer, a multi-scale convolutional layer, and a fully connected layer. The input layer is used to receive real or generated surgical instrument 2 motion trajectory data. The multi-scale convolutional layer includes multiple convolutional layers of different scales, such as different convolution kernel sizes and step settings, so as to extract features at different scales. Each scale of the convolutional layer is responsible for capturing motion features at different levels. The fully connected layer fuses and discriminates the features extracted by the convolutional layer to determine whether the output data is real or generated.

[0038] When constructing the generative adversarial network model, each frame of the preprocessed 14-frame continuous sequence of key motion images of surgical instruments is divided into two parts with the same number of frames in chronological order. The first 7 frames are used as the first image sequence for input, and the last 7 frames are used as the second image sequence for target output.

[0039] The idea of ​​the generative adversarial network model originates from game theory and consists of a generator and a discriminator. The model reaches a Nash equilibrium through continuous adversarial games, thereby generating high-quality samples. The generator is used to simulate the distribution of the second image sequence to generate realistic image samples, and the discriminator is used to distinguish between real trajectories and generated trajectories. The generator includes an input layer, a 3D convolution layer, an upsampling layer, and an output layer. The input layer is used to receive the input image sequence of the previous moment, and the 3D convolution layer is used to extract low-dimensional spatiotemporal features on the input image sequence. The upsampling layer upsamples the low-dimensional spatiotemporal features to the target dimension to obtain the output image sequence, and the output layer outputs the generated output image sequence. The discriminator uses a multi-scale discriminator to discriminate the output image sequence of the generator at different resolutions to obtain the probability of true or false judgment.

[0040] Generative adversarial network models such as Figure 4 As shown in Figure 2, the main function of the generator is to simulate the distribution of the second image sequence to generate realistic image samples. As input, the neural network generates the output image The generator continuously learns from real image data The features and distribution of the image are used to try to generate samples that are highly similar to the real image data, thereby deceiving the discriminator and misclassifying the generated images as real images. It is a binary classifier used to determine whether the input image is a real image or a predicted image. When Should be close to 1; when the input is the predicted image hour, The output value of should be close to 0. The two feed back each other through the loss function and continuously adjust the parameters. In adversarial learning, the generator and the discriminator improve each other and eventually tend to the Nash equilibrium state. At this time, the predicted image generated by the generator is almost indistinguishable from the real image, and the discriminator also finds it difficult to distinguish the difference between the two.

[0041] Since each frame in the key motion image sequence of surgical instruments contains time dimension information, in order to effectively extract the spatiotemporal correlation features between image frames, the neural network used by the generator is a 3D convolutional neural network. 3D convolution can capture the dynamic changes and spatiotemporal features of the image sequence by simultaneously performing convolution operations on the width and height of the spatial image and the relationship dimension between the time frame. When processing each frame in the first 7 frames of the key motion image sequence of surgical instruments, the dimension of the input tensor is [N, C, T=7, H, W], where N represents the batch size, that is, the number of samples input into the network at one time; C represents the number of channels; T represents the number of frames, and T=7 means that the number of frames is fixed to 7 frames; H and W represent the height and width of the image. First, the network uses a 3×3×3 3D convolution kernel to extract low-level spatiotemporal features of the image sequence. This convolution kernel can establish local correlations between time frames and spatial positions. Subsequently, by stacking multiple layers of 3D convolution layers, higher-level spatiotemporal features are gradually extracted. In order to enable the network to capture more complex features while reducing computational overhead and overfitting risks, an activation function Relu is added after each convolution layer. In addition, to alleviate the degradation and gradient vanishing problems that may occur in the deep structure of the network, a residual module is added every 3 layers of convolution to improve the training efficiency and performance of the model. Finally, in order to gradually restore the extracted spatiotemporal features to the resolution of the target output, the network upsamples the features through 3D transposed convolution, while gradually reducing the number of channels of the feature map. After this process, the dimension of the network's final output tensor is consistent with the target, that is, [N, C, T=7, H, W], thereby generating an image sequence that meets the requirements.

[0042] In order to better perceive the local and global information in the image sequence, the discriminator uses a multi-scale discriminator. The multi-scale discriminator can effectively capture the spatiotemporal features at different scales by discriminating images at different resolutions. By discriminating the output image sequence of the generator at different resolutions, the probability of true or false judgment is obtained, including the following steps:

[0043] The output image sequence of the generator is multi-resolution downsampled to generate image sequences with different resolutions, including original resolution, 1 / 2 resolution, 1 / 4 resolution, and 1 / 8 resolution.

[0044] Subsequently, for each resolution of the image sequence, a 3D convolutional neural network is used to extract its spatiotemporal features, and true and false discrimination is performed independently at each resolution, and the corresponding probability is output.

[0045] Finally, by taking the average of the output probabilities of all resolutions, an overall true or false judgment probability is obtained. If the overall true or false judgment probability of the output is close to 1, it is judged to be a real image; if the overall true or false judgment probability of the output is close to 0, it is judged to be a predicted image.

[0046] In this way, the multi-scale discriminator can more comprehensively analyze the features of image sequences at different scales, thereby improving the robustness and discrimination ability of the discriminator.

[0047] One of the keys to the generative adversarial network model is to reasonably design the loss function. In the present invention, both the generator and the discriminator use binary cross entropy loss as the optimization target. The generator loss function is defined as:

[0048] , in, represents the discriminator, represents a generator, is the input image, The image generated by the generator, Represents the probability that the discriminator outputs the predicted image as a real image. The generator continuously optimizes the quality of the generated image by minimizing the loss function, making it gradually approach the real image.

[0049] Ultimately, the goal of the generator is to make It approaches 1, making it difficult for the discriminator to distinguish between the predicted image and the real image, increasing the difficulty of the discriminator's judgment.

[0050] The discriminator loss function is defined as: , Among them, the discriminator loss function is divided into two parts. The first part is , the second image sequence is expected in this part is judged to be true, i.e. is 1; the second part is , the expected prediction image in this part It was judged to be forged, that is, is 0.

[0051] Therefore, the goal of the discriminator loss function is to maximize the predicted second image sequence output as much as possible to 1, and minimize the predicted image output as much as possible to 0. By minimizing the discriminator loss function , the discriminator can better learn how to distinguish between real images and predicted images.

[0052] After the generative adversarial network model is built, the model training begins.

[0053] First, the dataset is divided into a training set and a test set in a ratio of 8:2, with 80% used for training and 20% for testing. At the same time, in order to speed up the training process and improve the stability of the model, the Adam optimizer is used, which can adaptively adjust the learning rate, thereby effectively alleviating the gradient vanishing or gradient exploding problems that may occur during the training process.

[0054] In addition, the use of the Adam optimizer can improve the training efficiency and prediction performance of the model, making the model converge faster and achieve better results. In the generative adversarial network model, the generator and the discriminator are trained alternately. The specific training process is as follows:

[0055] First, the real 14 frames of images are input into the discriminator, and the discriminator is expected to output a probability close to 1 to accurately identify the real image.

[0056] Next, the first 7 frames of image sequence are fed into the generator as input, and the generator generates the next 7 frames of image sequence. Subsequently, the generator loss function is calculated and the generator is updated.

[0057] Subsequently, the last 7 frames of images generated by the generator are input into the discriminator, which evaluates the quality of the last 7 frames according to the complete real 14 frames, and checks the temporal consistency between the last 7 frames and the first 7 frames. The discriminator loss function is calculated and the discriminator is updated.

[0058] By continuously alternating the training of the discriminator and the generator, the generator gradually generates more realistic image sequences, while the discriminator gradually enhances its ability to distinguish the authenticity and temporal rationality of images.

[0059] When the training reaches the point where the discriminator output probability is close to 0.5, it indicates that the discriminator can no longer distinguish between real images and predicted images, and the training process ends.

[0060] The trained generative adversarial network model is tested using a test set to evaluate the prediction accuracy of the generative adversarial network model. When the prediction accuracy of the generative adversarial network model meets the requirements, the training is terminated.

[0061] Specifically, after the generative adversarial network model training is completed, a continuous sequence of 14 frames of key motion images of surgical instruments are randomly extracted from the test set.

[0062] First, the first 7 frames of image sequence are input into the generator, and the generator generates the next 7 frames of predicted image sequence.

[0063] At the same time, the complete 14-frame sequence of surgical instrument key motion images is input into the discriminator to evaluate the difference between the generated last 7 frames and the real images.

[0064] After the generator completes the 7-frame predicted image sequence, the prediction accuracy of the model is verified from two aspects.

[0065] On the one hand, in order to fully evaluate the quality of the predicted image, the mean square error (MSE) is used as an evaluation indicator of prediction accuracy. The mean square error (MSE) is a commonly used indicator used to measure the degree of difference between the predicted value and the true value. The mean square error (MSE) is calculated according to the following formula:

[0066] , in, represents the number of predicted image sequences, For the The real image corresponding to the input image sequence is For the The predicted image corresponding to the input image sequence, MSE is the mean square error, which is used to evaluate the difference between the predicted image and the real image.

[0067] when When is 7, the mean square error MSE is calculated as follows: , The mean square error (MSE) is used to evaluate the difference between the predicted image and the real image, and its value reflects the accuracy of the prediction result. The smaller the loss of pixel value, that is, the lower the mean square error (MSE) value, the closer the predicted image is to the real image, and the higher the prediction accuracy.

[0068] In the present invention, when the mean square error MSE is less than 50, it indicates that the quality of the generated image has met the expected requirements.

[0069] On the other hand, in order to obtain the trajectory of the surgical instrument needle tip 8 in the three-dimensional space and verify the prediction accuracy, the three-dimensional coordinates of the surgical instrument needle tip 8 in the generated last 7 frames of the image sequence and the three-dimensional coordinates of the surgical instrument needle tip 8 in the real last 7 frames of the image sequence are extracted by computer 4, and the prediction accuracy of the model is evaluated by calculating the error between the two sets of three-dimensional coordinates.

[0070] In order to accurately determine the three-dimensional coordinates of the surgical instrument needle tip 8, three markers are set on the surgical instrument 2, wherein the marker is an optical positioning marker ball that can reflect infrared light. The relative positions of the three markers on the surgical instrument 2 and the surgical instrument needle tip 8 are determined. Therefore, based on the predicted image sequence and the binocular vision measurement device 3, the three-dimensional coordinates of the three markers on the surgical instrument 2 can be determined. Based on the three-dimensional coordinates of the three markers on the surgical instrument 2, the three-dimensional coordinates of the surgical instrument needle tip 8 can be determined, and the predicted result of the surgical instrument needle tip trajectory can be obtained.

[0071] Specifically, the calculation process of the three-dimensional coordinates of the surgical instrument needle tip 8 is as follows: (1) Based on the predicted image sequence generated by the generative adversarial network model, the edge contours of the three landmarks on the surgical instrument 2 are obtained using the Canny edge detection algorithm.

[0072] Specifically, the markers are marker one 5, marker two 6 and marker three 7, all of which are arranged on the surgical instrument 2. Marker one 5, marker two 6 and marker three 7 are reflected in the image as a circle, and then the Canny algorithm can detect the circular outline of the marker.

[0073] (2) Calculate the two-dimensional center coordinates of the landmark based on the edge pixel points on the edge contour.

[0074] The specific method is: randomly select three pixel points on the edge contour of the marker, and set the coordinates of the three non-collinear points on the edge contour of the marker as , and , determine the two-dimensional center coordinates of the marker according to the following formula , , in, , and are the coordinates of three non-collinear points on the edge contour of the landmark, are the two-dimensional center coordinates of the marker.

[0075] (3) Accurately match the landmarks in the images acquired by the two cameras of the binocular vision measurement device 3 through a regional similarity matching algorithm.

[0076] (4) Calculate the three-dimensional coordinates of the landmark.

[0077] Assume that the two-dimensional coordinates of the marker in the images obtained by the two cameras are , , since the camera 1 and the camera 2 in the binocular vision measurement device 3 are parallel, then , the three-dimensional coordinates of the marker are , determine the three-dimensional coordinates of the marker according to the following formula, , in, is the center distance between the two cameras of the binocular vision measurement device 3, is the focal length of the binocular vision measurement device 3, , are the two-dimensional coordinates of the landmark in the images acquired by the two cameras, are the three-dimensional coordinates of the landmark.

[0078] (5) Figure 3 As shown, the coordinates of marker 1 5, marker 2 6 and marker 3 7 in three-dimensional space are respectively , and , then the translation vector of the surgical instrument 2 relative to the binocular vision measurement device 3 is: , The rotation matrix of the surgical instrument 2 relative to the binocular vision measurement device 3 is: , in, , , Coordinate system for surgical instruments The unit vector of the axis, the rules for establishing the surgical instrument coordinate system are as follows Figure 4 As shown, , Suppose the coordinates of the surgical instrument needle tip 8 obtained after calibration of the surgical instrument 2 in the surgical instrument coordinate system are , then the three-dimensional coordinates of the surgical instrument needle tip 8 are: , The error is calculated by extracting the three-dimensional coordinates of the surgical instrument needle tip 8 from the real 7 frames of images and the generated 7 frames of predicted images. The error calculation formula is: , in, represents the number of images in the predicted image sequence, is the three-dimensional coordinate of the surgical instrument needle tip 8 in the real image, To predict the three-dimensional coordinates of the surgical instrument needle tip 8 in the image, is the error, is a natural number, The value range is 1 to , when the error When it is less than 1, it indicates that the needle tip trajectory of the surgical instrument has achieved the predicted accuracy.

[0079] To predict the number of images in an image sequence Take 7 as an example, then .

[0080] In the present invention, when the error When it is less than 1, it indicates that the needle tip trajectory of the surgical instrument has achieved the prediction accuracy and the training of the generative adversarial network model is completed.

[0081] During the training process of the generative adversarial network model, the parameters are adjusted dynamically. After the training is completed, the final parameters are selected as the final parameters of the generative adversarial network model. At this time, the parameters are fixed, and the generator can be used directly for prediction in the later stage.

[0082] When in use, the real-time motion trajectory image sequence of the surgical instrument 2 is directly obtained, and the real-time motion trajectory image sequence is input into the generator to generate a predicted image sequence. The predicted result of the surgical instrument needle tip trajectory is obtained based on the predicted image sequence. At this time, the surgical instrument needle tip trajectory is drawn, and based on the predicted surgical instrument needle tip trajectory, feedback is given to the doctor or surgical robot to help guide the surgical verification operation.

[0083] In summary, the surgical instrument needle tip trajectory prediction method provided by the present invention has the following advantages: 1. The present invention uses a generative adversarial network model to predict the motion trajectory of the surgical instrument needle tip 8. The generative adversarial network model consists of a generator and a discriminator. Through adversarial training between the two, the generator can continuously optimize its output, thereby generating high-precision prediction results. This adversarial learning mechanism can effectively capture the complex motion patterns of the surgical instrument 2 and improve the accuracy of the surgical instrument needle tip trajectory prediction.

[0084] 2. In the generative adversarial network model used in the present invention, the 3D convolutional neural network used by the generator can perform convolution operations in both spatial and temporal dimensions simultaneously, effectively capturing the complex motion trajectory of the surgical instrument 2 in three-dimensional space, and accurately predicting the dynamic changes of the surgical instrument needle tip 8 during the operation.

[0085] 3. In the generative adversarial network model used in the present invention, the discriminator adopts a multi-scale discriminator, which can evaluate the generated motion trajectory at different scales and capture features at different levels, thereby improving the ability to judge the authenticity of the generated trajectory. This multi-scale feature extraction capability can effectively distinguish between real trajectories and forged trajectories, and improve the prediction accuracy of the overall system.

[0086] 4. The binocular vision measurement device 3 used in the present invention is a near-infrared binocular vision measurement device 3, which tracks and locates through a reflective marker ball and calculates the three-dimensional coordinates of the surgical instrument needle tip 8 to further verify the prediction accuracy of the generative adversarial network model, improve the accuracy and detail expression of the surgical instrument needle tip trajectory prediction, and ensure the high accuracy of the prediction.

[0087] The invention is applied to surgical navigation systems. By predicting the trajectory of surgical instrument needle tips, it helps doctors make more accurate decisions during surgery, reduces the risk of misoperation, and improves the success rate of surgery. In robot-assisted surgery, predicting the trajectory of surgical instrument needle tips helps improve the accuracy and reliability of surgical robot operations, making the surgical process smoother.

[0088] In addition, by recording and analyzing the needle tip trajectory of surgical instruments, a large amount of surgical data can be accumulated, which helps to improve surgical methods and instrument design.

[0089] It will be understood that the present invention is described through some embodiments, and those skilled in the art will appreciate that various changes or equivalent substitutions may be made to these features and embodiments without departing from the spirit and scope of the present invention.

[0090] In addition, under the teachings of the present invention, these features and embodiments may be modified to adapt to specific circumstances and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of the present invention are within the scope of protection of the present invention.

Claims

1. A method for predicting the trajectory of a surgical instrument needle tip, characterized in that: The following steps are involved: Acquiring a historical motion trajectory of the surgical instrument (2), the historical motion trajectory including all motion forms of the surgical instrument (2), extracting a key motion image sequence of the surgical instrument in each motion trajectory, and using this to construct a surgical instrument motion trajectory data set; Preprocessing a surgical instrument key motion image sequence in a surgical instrument motion trajectory data set, and dividing the preprocessed surgical instrument key motion image sequence into a first image sequence and a second image sequence with the same number of frames in chronological order; A generative adversarial network model is constructed using a 3D convolutional neural network as a generator and a multi-scale discriminator as a discriminator, and after training the generative adversarial network model based on a mapping relationship between the first image sequence and the second image sequence, the prediction accuracy of the generative adversarial network model is evaluated, and when the prediction accuracy of the generative adversarial network model reaches a preset accuracy, the parameters of the generator and the discriminator are fixed; A real-time motion trajectory image sequence of the surgical instrument (2) is acquired, and the real-time motion trajectory image sequence is input into the generator to generate a predicted image sequence, and a predicted result of the surgical instrument needle tip trajectory is obtained based on the predicted image sequence.

2. The surgical instrument needle tip trajectory prediction method according to claim 1, characterized in that: Preprocessing the key motion image sequences of surgical instruments in the surgical instrument motion trajectory dataset includes the following steps: Remove noise from each frame of a critical motion image sequence of surgical instruments; The pixel values ​​of each frame in the key motion image sequence of surgical instruments are normalized and mapped from [0-255] to [0-1].

3. The surgical instrument needle tip trajectory prediction method according to claim 1, characterized in that: The generative adversarial network model is trained based on the mapping relationship between the first image sequence and the second image sequence, comprising the following steps: The surgical instrument motion trajectory dataset is divided into a training set and a test set; Select a complete multi-frame image sequence consisting of a first image sequence and a second image sequence from the training set as an input of the discriminator; Send the first image sequence in the training set as input to the generator, generate a predicted image sequence, calculate the generator loss function, and update the generator; The predicted image sequence generated by the generator is input into the discriminator, the discriminator checks the temporal consistency between the generated predicted image sequence and the first image sequence, calculates the discriminator loss function, and updates the discriminator until the discriminator loss function is minimized.

4. The surgical instrument needle tip trajectory prediction method according to claim 3, characterized in that: Evaluating the prediction accuracy of the generative adversarial network model includes the following steps: The first image sequence in the test set is obtained as the input of the generator, and the generator generates a predicted image sequence; A complete multi-frame image sequence consisting of the first image sequence and the second image sequence in the test set is obtained as the input of the discriminator to evaluate the difference between the predicted image and the real image, and to verify the error between the three-dimensional coordinates of the surgical instrument needle tip (8) in the predicted image and the three-dimensional coordinates of the surgical instrument needle tip (8) in the real image. The above results are used to comprehensively evaluate the prediction accuracy of the generative adversarial network model.

5. The surgical instrument needle tip trajectory prediction method according to claim 4, characterized in that: The difference between the predicted image and the real image is evaluated using the following formula: , in, Represents the number of frames of the image in the predicted image sequence, For the The real image corresponding to the input image sequence is For the The predicted image corresponding to the input image sequence, MSE is the mean square error, which is used to evaluate the difference between the predicted image and the real image. is a natural number, The value range is 1 to ; The lower the mean square error (MSE) value, the higher the prediction accuracy.

6. The surgical instrument needle tip trajectory prediction method according to claim 4, characterized in that: Verifying the error between the three-dimensional coordinates of the surgical instrument needle tip (8) in the predicted image and the three-dimensional coordinates of the surgical instrument needle tip (8) in the real image comprises the following steps: Respectively extracting the three-dimensional coordinates of the surgical instrument needle tip (8) in the predicted image and the three-dimensional coordinates of the surgical instrument needle tip (8) in the real image; The error between the three-dimensional coordinates of the surgical instrument needle tip (8) in the predicted image and the three-dimensional coordinates of the surgical instrument needle tip (8) in the real image is calculated using the following formula: , in, Represents the number of frames of the image in the predicted image sequence, is the three-dimensional coordinate of the needle tip in the real image, To predict the three-dimensional coordinates of the needle tip in the image, is the error, is a natural number, The value range is 1 to ; When the error When it is less than 1, it indicates that the needle tip trajectory of the surgical instrument has achieved the predicted accuracy.

7. The surgical instrument needle tip trajectory prediction method according to claim 1, characterized in that: Obtaining a prediction result of a surgical instrument needle tip trajectory based on a prediction image sequence includes the following steps: Determining the three-dimensional coordinates of three markers on the surgical instrument (2) based on the predicted image sequence; The three-dimensional coordinates of the surgical instrument needle tip (8) are determined based on the three-dimensional coordinates of the three markers on the surgical instrument (2), and a prediction result of the surgical instrument needle tip trajectory is obtained.

8. The method for predicting the needle tip trajectory of a surgical instrument according to claim 7, characterized in that: Based on the predicted image sequence, the three-dimensional coordinates of three markers on the surgical instrument are determined, including the following steps: Based on the predicted image sequence generated by the generative adversarial network model, the edge contours of the three landmarks on the surgical instrument (2) are obtained by using the Canny edge detection algorithm; Randomly select three pixel points on the edge contour of the marker and determine the two-dimensional center coordinates of the marker according to the following formula , , in, , and are the coordinates of three non-collinear points on the edge contour of the landmark, is the two-dimensional center coordinate of the landmark; Matching the landmarks in the images acquired by two cameras of the binocular vision measurement device (3) by using a region similarity matching algorithm; The three-dimensional coordinates of the marker are determined as follows: , in, , is the center distance between the two cameras of the binocular vision measurement device (3), is the focal length of the binocular vision measurement device (3), , are the two-dimensional coordinates of the landmark in the images acquired by the two cameras, are the three-dimensional coordinates of the landmark.

9. The surgical instrument needle tip trajectory prediction method according to claim 7, characterized in that: The three-dimensional coordinates of the surgical instrument needle tip (8) are determined based on the following formula: , in, , , , is the coordinate of the surgical instrument needle tip (8) in the surgical instrument coordinate system, is the three-dimensional coordinate of the surgical instrument needle tip (8), , and are the coordinates of marker 1 (5), marker 2 (6) and marker 3 (7) in three-dimensional space, is the translation vector, is the rotation vector, They are the unit vectors in the x, y, and z directions of the surgical instrument coordinate system.

Citation Information

Patent Citations

  • Minimally invasive surgical instrument positioning method and system

    CN112734776A

  • Dynamic prediction method and system for use of surgical instrument

    CN114005022A

  • Track prediction method and device based on time attention convolutional network

    CN114116944A

  • Training method and device of trajectory prediction model and trajectory prediction method and device

    CN115861377A

  • Track prediction method based on attention neural network and generative adversarial network

    CN117874443A