Surgical instrument three-dimensional pose detection method and system based on monocular depth estimation

Through a monocular depth estimation method, combined with deep learning and image processing algorithms, the stability problem of three-dimensional posture estimation of surgical instruments in laparoscopic surgery is solved, and accurate and safe surgical navigation is achieved in complex environments, improving the accuracy and safety of laparoscopic surgery.

CN120496862APending Publication Date: 2025-08-15HUAZHONG UNIV OF SCI & TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510373258.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

Prior Art In laparoscopic surgery, the three-dimensional posture estimation method of surgical instruments relies on specific detection conditions, making it difficult to ensure stable reconstruction effect in complex and changeable surgical environments, affecting the accuracy and safety of the surgery.

Method used

Using a method based on monocular depth estimation, combined with deep learning technology and image processing algorithms, we construct laparoscopic surgical simulation scenarios, produce key points and monocular depth estimation data sets, train a monocular absolute depth estimation model and a two-dimensional key point detection model, couple two-dimensional key point detection and depth estimation results, eliminate abnormal depth values, and realize three-dimensional posture detection of surgical instruments.

Benefits of technology

It realizes real-time and accurate estimation of the three-dimensional posture information of the surgical instrument in laparoscopic surgery, improves the accuracy and safety of the surgery, reduces the cost of data set production, and enhances the stability and adaptability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496862A_ABST
    Figure CN120496862A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of surgical instrument control, and particularly discloses a surgical instrument three-dimensional pose detection method and system based on monocular depth estimation. The method comprises the following steps: constructing a laparoscopic surgery simulation scene, and making a surgical instrument key point and monocular depth estimation data set; training the monocular absolute depth estimation model by adopting a monocular depth estimation data set; based on the key points of the surgical instrument, adopting a boundary expansion and random cutting method to improve the detection stability of a two-dimensional key point detection model, and training the two-dimensional key point detection model; and coupling key point detection output by the two-dimensional key point detection model and a depth estimation result output by the monocular absolute depth estimation model, eliminating abnormal depth values of the key points, and restoring the three-dimensional pose of the surgical instrument. According to the method, the deep learning technology and the image processing algorithm are utilized, the three-dimensional pose information of the surgical instrument is accurately estimated from the laparoscopic image in real time, and the accuracy and safety of the laparoscopic surgery are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of surgical instrument control, and more specifically, relates to a method and system for detecting the three-dimensional posture of a surgical instrument based on monocular depth estimation. Background Art

[0002] With the continuous advancement of medical technology, laparoscopic surgery has been widely used in clinical treatment due to its advantages such as less trauma and faster recovery. However, the complexity and difficulty of laparoscopic surgery also place higher demands on the precise operation of surgical instruments. In laparoscopic surgery, the three-dimensional pose information of surgical instruments is crucial for the smooth progress of the operation, which directly affects the accuracy and safety of the operation. Current methods for estimating the three-dimensional pose of surgical instruments are often highly dependent on specific detection conditions, the addition of visual markers or the accuracy of feature matching in binocular vision, which makes it difficult to ensure stable reconstruction effects in complex and changing surgical environments. Therefore, there is an urgent need to develop a new method that is more efficient, stable and adaptable to overcome these limitations and improve the accuracy and safety of laparoscopic surgery. Summary of the Invention

[0003] To address the aforementioned deficiencies or improvements in the existing technology, the present invention provides a system and method for detecting the three-dimensional pose of surgical instruments based on monocular depth estimation. This method, incorporating the characteristics of the instrument's relative pose relative to the lens when the laparoscope is stationary at key frames, proposes a method for instrument pose estimation that couples depth estimation with key point detection. Monocular depth estimation is used to reconstruct the instrument's movement relative to the lens, while the three-dimensional positions of key points represent the instrument's position and pose. This method utilizes deep learning techniques and image processing algorithms to accurately and in real time estimate the three-dimensional pose information of surgical instruments from laparoscopic images, providing surgeons with intuitive and reliable surgical navigation and effectively improving the accuracy and safety of laparoscopic surgery.

[0004] To achieve the above objectives, according to one aspect of the present invention, a method for detecting the three-dimensional pose of a surgical instrument based on monocular depth estimation is proposed, comprising the following steps:

[0005] S100 builds a laparoscopic surgery simulation scene and uses it to create a dataset of surgical instrument key points and monocular depth estimation;

[0006] S200 builds a monocular absolute depth estimation model and uses a monocular depth estimation dataset to train the monocular absolute depth estimation model;

[0007] S300: constructing a two-dimensional key point detection model, and based on the surgical instrument key points, using boundary expansion and random cropping methods to improve the detection stability of the two-dimensional key point detection model, and training the two-dimensional key point detection model;

[0008] S400 couples the key point detection output by the two-dimensional key point detection model with the depth estimation results output by the monocular absolute depth estimation model to eliminate abnormal depth values of key points and restore the three-dimensional posture of surgical instruments.

[0009] As further preferred, step S100 includes the following steps:

[0010] S101 creates 3D models of surgical instruments and designs key points of each surgical instrument according to the characterization requirements of the surgical instruments;

[0011] S102 constructs a laparoscopic surgery simulation scene, randomly adjusts the surgical instrument posture, light source position and brightness, and background image style to achieve data generalization, saves the rendering result image, camera depth map, and surgical instrument key point position information, and extracts monocular depth estimation data to generate a data set.

[0012] As a further preferred embodiment, step S102 further includes:

[0013] A style transfer model is used to convert the style of the rendered image into the style of real-world images, achieving domain generalization of the dataset.

[0014] As a further preference, in step S200, DepthAnything is used as the monocular absolute depth estimation model, and SILogloss and L1loss are used as the loss functions of the monocular absolute depth estimation model to improve the smoothness of the prediction result.

[0015] As a further preference, the YOLOv8-pose model is used to construct a two-dimensional plane key point detection model.

[0016] As a further preferred embodiment, step S400 includes the following steps:

[0017] S401 adds depth information to each predicted key point according to the monocular depth estimation result to obtain the depth value of the key point;

[0018] S402 fills in the depth value of the missing key point;

[0019] S403 performs filtering on the depth values of the key points to remove abnormal values;

[0020] S404 calculates the position of the key points of the surgical instrument relative to the camera lens based on the camera parameters, and restores the three-dimensional posture of the surgical instrument.

[0021] As a further preferred embodiment, in step S401, for each key point p k =(x k ,y k ), its depth value Z k Get as:

[0022] Z k =D(x k ,y k )

[0023] Where D is the depth estimation result after boundary expansion.

[0024] As a further preferred embodiment, in step S402, the depth values of all pixels between the two key points are first extracted and arranged in order, and then the points with lost depth are eliminated and the least square method is used to fit the remaining data to fill in the lost depth values of the key points;

[0025] In step S403, based on the coordination of the overall movement of the instrument, a sudden change in the depth of the instrument tip is processed. When the difference between the instrument tip depth value and the previous moment is greater than a set threshold, the depth value is regarded as an abnormal value. The abnormal value is processed as follows:

[0026] depth err,j =depth err,j-1 +depth c,j -depth c,j-1

[0027] Among them, depth err,j is the correction value of the depth of the abnormal point, j is the current time series number, depth err,j-1 is the depth value of the abnormal point at the moment before, depth c,j is the depth value of the center point of the surgical instrument at the current moment, depth c,j-1 It is the depth value of the center point of the surgical instrument at the previous moment.

[0028] As a further preferred embodiment, in step S404, the position of the key point of the surgical instrument relative to the camera lens includes:

[0029]

[0030] Where, X c and Y c is the coordinate in the camera coordinate system, u and v are the coordinates in the image, c x and c y is the principal point coordinate of the camera, Z c is the depth value obtained by depth estimation, and f is the focal length of the camera.

[0031] According to another aspect of the present invention, a three-dimensional posture detection system for surgical instruments based on monocular depth estimation is provided, comprising:

[0032] The first main control module is used to build a laparoscopic surgery simulation scene and generate a dataset of surgical instrument key points and monocular depth estimation;

[0033] A second main control module is used to build a monocular absolute depth estimation model and train the monocular absolute depth estimation model using a monocular depth estimation dataset;

[0034] A third main control module is used to build a two-dimensional key point detection model, and based on the surgical instrument key points, use boundary expansion and random cropping methods to improve the detection stability of the two-dimensional key point detection model and train the two-dimensional key point detection model;

[0035] The fourth main control module is used to couple the key point detection output by the two-dimensional key point detection model and the depth estimation results output by the monocular absolute depth estimation model, eliminate abnormal depth values of the key points, and restore the three-dimensional posture of the surgical instrument.

[0036] In general, the above technical solutions conceived by the present invention have the following technical advantages compared with the existing technology:

[0037] 1. This invention uses monocular depth estimation to reconstruct the movement of instruments at near and far distances relative to the camera, characterizing the instrument's position and posture using the three-dimensional positions of key points. This method utilizes deep learning techniques and image processing algorithms to accurately estimate the three-dimensional position of surgical instruments in real time from laparoscopic images, providing surgeons with intuitive and reliable surgical navigation and effectively improving the accuracy and safety of laparoscopic surgery. The implementation of this invention will lay a solid foundation for the development of intelligent and precise laparoscopic surgery.

[0038] 2. This invention creates a dataset by building a simulation scene and uses style migration technology to improve the authenticity of the dataset, thereby reducing the production cost of the dataset and providing generalized data for deep learning models.

[0039] 3. The present invention achieves data enhancement through image boundary expansion and cropping, enabling the model to detect the positions of key points of instruments beyond the field of view, thereby improving the stability of model detection.

[0040] 4. The present invention uses the physical properties of the instrument itself to eliminate abnormal depth values of key points of the instrument through fitting straight lines and outlier detection, thereby improving the accuracy of three-dimensional posture detection of surgical instruments. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 This is a flow chart of a method for detecting three-dimensional posture of a surgical instrument based on monocular depth estimation according to an embodiment of the present invention;

[0042] Figure 2 is a flowchart of data set acquisition involved in an embodiment of the present invention;

[0043] Figure 3 This is an example diagram of key point settings of an instrument involved in an embodiment of the present invention;

[0044] Figure 4 Schematic diagram of the layout of a simulation data set scene according to an embodiment of the present invention;

[0045] Figure 5 2 is a schematic diagram of a key point detection image boundary expansion process according to an embodiment of the present invention;

[0046] Figure 6 This is a flow chart of the three-dimensional posture coupling of the device involved in the embodiment of the present invention;

[0047] Figure 7 Schematic diagram of filling missing depth values involved in an embodiment of the present invention;

[0048] Figure 8 Schematic diagram of random cropping of the tip area of the instrument in an embodiment of the present invention. DETAILED DESCRIPTION

[0049] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below may be combined with each other as long as they do not conflict with each other.

[0050] Example 1

[0051] In this embodiment, a method for detecting the three-dimensional pose of a surgical instrument based on monocular depth estimation is provided, comprising the following steps:

[0052] S100 builds a laparoscopic surgery simulation scene and uses it to create a dataset of surgical instrument key points and monocular depth estimation. This step specifically includes:

[0053] S101 creates 3D models of surgical instruments and designs key points of each surgical instrument according to the characterization requirements of the surgical instruments;

[0054] S102 constructs a laparoscopic surgery simulation scene, randomly adjusts the surgical instrument posture, light source position and brightness, and background image style to achieve data generalization, saves the rendering result image, camera depth map, and surgical instrument key point position information, and extracts monocular depth estimation data to generate a data set.

[0055] Specifically, step S102 further includes:

[0056] A style transfer model is used to convert the style of the rendered image into the style of real-world images, achieving domain generalization of the dataset.

[0057] S200 constructs a monocular absolute depth estimation model and uses a monocular depth estimation dataset to train the monocular absolute depth estimation model.

[0058] In this step, DepthAnything is used as the monocular absolute depth estimation model, and SILogloss and L1loss are used as the loss functions of the monocular absolute depth estimation model to improve the smoothness of the prediction results.

[0059] S300 constructs a two-dimensional key point detection model, and based on the surgical instrument key points, uses boundary expansion and random cropping methods to improve the detection stability of the two-dimensional key point detection model, and trains the two-dimensional key point detection model.

[0060] In this step, as a preference, the YOLOv8-pose model is used to construct a two-dimensional plane key point detection model.

[0061] As an alternative embodiment of the present invention, Figure 5 As shown, in this step, the geometric extrapolation method and the machine learning model prediction method are used to perform boundary expansion processing. Specifically:

[0062] First, identify the type of surgical instrument in the image;

[0063] The basic expansion distance is then calculated using geometric extrapolation:

[0064]

[0065] Where L is the length of the surgical instrument, obtained by measuring the Euclidean distance between the two endpoints of the surgical instrument in the image. R is the radius of curvature of the curved portion of the surgical instrument. S is the shape complexity of the surgical instrument. A is the area of the surgical instrument, obtained by calculating the number of pixels occupied by the instrument in the image. T is the type of surgical instrument, determined by a classification algorithm or manual annotation, with different weight coefficients corresponding to different types. α, β, γ, and δ are weight coefficients used to balance the impact of different parameters on the expansion distance and are adjusted based on actual data and experience.

[0066] Use a trained machine learning model (such as a convolutional neural network) to predict an adjustment coefficient k based on the type T, shape characteristics and other parameters of the surgical instrument, that is:

[0067] k=f(T, L, R, S, A)

[0068] Here, f is the prediction function of the machine learning model.

[0069] Adjust the expansion distance considering the image size constraint:

[0070]

[0071] Where W specifies the width of the image. H specifies the height of the image. min 、y min are the minimum coordinates of the initial boundary of the surgical instrument in the image. max 、y max are the maximum coordinates of the initial boundary of the surgical instrument in the image

[0072] Construct the final expansion distance calculation model:

[0073] D final =D base ×k

[0074] Where D final is the final expansion distance.

[0075] In addition, as preferred, Figure 8 As shown, in this step, the cropping process randomly crops an area of the instrument tip in the selected image, and then fills the background image in the area to achieve the effect that the instrument tip is blocked by the tissue.

[0076] S400 couples the key point detection output by the two-dimensional key point detection model with the depth estimation results output by the monocular absolute depth estimation model to eliminate abnormal depth values of key points and restore the three-dimensional posture of surgical instruments.

[0077] Preferably, this step includes:

[0078] S401 adds depth information to each predicted key point based on the monocular depth estimation result to obtain the depth value of the key point.

[0079] For each key point p k =(x k ,y k ), its depth value Z k Get as:

[0080] Z k =D(x k ,y k )

[0081] Where D is the depth estimation result after boundary expansion.

[0082] S402 fills in the depth value of the key point where the depth value is missing; specifically, first extract the depth values of all pixels between the two key points and arrange them in order, then remove the points where the depth is missing and use the least squares method to fit the remaining data to fill in the missing depth value of the key point.

[0083] S403 filters the depth values of the key points to remove outliers. Specifically, based on the coordination of the overall movement of the instrument, sudden changes in the depth of the instrument tip are processed. When the difference between the depth value of the instrument tip and the previous moment is greater than a set threshold, the depth value is considered an outlier. The outlier is processed as follows:

[0084] depth err,j =depth err,j-1 +depth c,j -depth c,j-1

[0085] Among them, depth err,j is the correction value of the depth of the abnormal point, j is the current time series number, depth err,j-1 is the depth value of the abnormal point at the moment before, depth c,j is the depth value of the center point of the surgical instrument at the current moment, depth c,j-1 It is the depth value of the center point of the surgical instrument at the previous moment.

[0086] S404 calculates the position of the key points of the surgical instrument relative to the camera lens based on the camera parameters, and restores the three-dimensional posture of the surgical instrument.

[0087] Preferably, the positions of the key points of the surgical instrument relative to the camera lens include:

[0088]

[0089] Where, X c and Y c is the coordinate in the camera coordinate system, u and v are the coordinates in the image, c x and c y is the principal point coordinate of the camera, Z c is the depth value obtained by depth estimation, and f is the focal length of the camera.

[0090] Based on any of the above embodiments or a combination of multiple embodiments, the present invention further provides a surgical instrument three-dimensional pose detection system based on monocular depth estimation, which is used to execute the method of any of the above embodiments or a combination of multiple embodiments, including:

[0091] The first main control module is used to build a laparoscopic surgery simulation scene and generate a dataset of surgical instrument key points and monocular depth estimation;

[0092] A second main control module is used to build a monocular absolute depth estimation model and train the monocular absolute depth estimation model using a monocular depth estimation dataset;

[0093] A third main control module is used to build a two-dimensional key point detection model, and based on the surgical instrument key points, use boundary expansion and random cropping methods to improve the detection stability of the two-dimensional key point detection model and train the two-dimensional key point detection model;

[0094] The fourth main control module is used to couple the key point detection output by the two-dimensional key point detection model and the depth estimation results output by the monocular absolute depth estimation model, eliminate abnormal depth values of the key points, and restore the three-dimensional posture of the surgical instrument.

[0095] Example 2

[0096] like Figure 1 As shown, an embodiment of the present invention provides a method for detecting the three-dimensional posture of a surgical instrument based on monocular depth estimation, comprising the following steps:

[0097] (1) Dataset acquisition. The cost of producing a dataset in a real scene is high, and due to hardware limitations, the accuracy of the dataset is difficult to guarantee. The present invention produces a dataset by building a simulation environment. The specific process is as follows: Figure 2 As shown, accurate depth information and key point location information can be obtained.

[0098] First, the surgical instruments must be accurately modeled, and representative key points that can reflect the position and posture of the instruments must be set for different surgical instruments, such as Figure 3 shown.

[0099] Use Blender modeling software to set up the simulation environment as follows Figure 4 As shown in the figure, the position and posture of the surgical instrument, the position and brightness of the light source, and the background image style are randomly adjusted to achieve the generalization of the dataset, and information such as the rendering result image, camera depth map, and instrument key point position are saved.

[0100] Use the style transfer model to convert the style of the simulated data rendered images into the style of real-world images, achieving domain generalization of the dataset and completing the data preparation.

[0101] 1. Instrument key point detection model training. During surgery, instruments are often located at the edge of the image. At this time, some key points may be out of view. In order to improve the stability of surgical instrument key point detection, this method performs boundary expansion processing on the input image. Figure 5 Data augmentation is achieved through boundary expansion, so that the key point model can learn to infer the position of key points of instruments outside the field of view based on the information within the field of view.

[0102] In addition, surgical instruments often need to separate organ tissues, and their tips may be obscured by them. This method uses cropping to randomly occlude some instrument tips in the dataset, allowing the key point detection model to handle occlusions.

[0103] After the above data enhancement, the two-dimensional plane key point detection model YOLOv8-pose is used for training to obtain the surgical instrument key point detection model.

[0104] 2. Monocular Depth Estimation Model Training. Since this method focuses solely on instrument depth estimation, after acquiring depth information captured by a camera in a simulated environment, the depth of the background plate is reset to zero. The model is then trained directly on the absolute depth of the surgical instrument relative to the camera to achieve absolute depth estimation. This method uses DepthAnything, pre-trained on a public dataset, for fine-tuning. SILogloss and L1loss are used as loss functions during training to improve the smoothness of the prediction results.

[0105] 3. Generating the 3D pose of the instrument. The key point detection model and depth estimation model mentioned above can be used to obtain the position of the instrument key points on the image and the depth information of the instrument in the field of view. The coupling process is as follows: Figure 5 shown.

[0106] First, use the monocular depth estimation result to add depth information to each predicted key point, that is, for each key point p k =(x k ,y k ), its depth value Z k Get as:

[0107] Z k =D(x k ,y k )(0.1)

[0108] Where D is the depth estimation result after boundary expansion.

[0109] Since the end of the surgical instrument handle and the tip of the instrument may not be visible, the depth value of the key point at this time does not exist. This study uses intestinal forceps as an example to deal with the problem. When the key points outside the connecting axis lose depth, the distortion of the camera imaging is ignored and the depth change from other instrument key points to the connecting axis key point is regarded as a continuous linear change. Figure 6 As shown in the figure, the depth values of all pixels between the two key points are first extracted and arranged in order, and then the points with lost depth are eliminated and the least squares method is used to fit the remaining data to fill in the lost depth values of the key points.

[0110] The above method can fill in the missing key point information of the depth value, but due to the complexity of the surgical scene, the depth value of the instrument tip may have abnormal mutations. This study handles the sudden change of the depth of the instrument tip based on the coordination of the overall movement of the instrument. When the difference between the depth value of the instrument tip and the previous moment is greater than the set threshold, the depth value is considered an abnormal value. The abnormal value is handled as follows:

[0111] depth err,j =depth err,j-1 +depth c,j -depth c,j-1 (0.2)

[0112] where depth err,j is the correction value of the depth of the abnormal point, j is the current time series number, depth err,j-1 is the depth value of the abnormal point at the moment before, depth c,j is the depth value of the center point of the surgical instrument at the current moment, depth c,j-1 It is the depth value of the center point of the surgical instrument at the previous moment.

[0113] Finally, the position of the instrument key points relative to the camera lens is calculated based on the camera parameters. The calculation formula is as follows:

[0114]

[0115] Among them, X c and Y c is the coordinate in the camera coordinate system, u and v are the coordinates in the image, c x and c y is the principal point coordinate of the camera, Z c is the depth value obtained by depth estimation, and f is the focal length of the camera.

[0116] At this point, the three-dimensional position of each key point of the instrument is restored, realizing the three-dimensional pose estimation of the surgical instrument.

[0117] It will be easily understood by those skilled in the art that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for detecting the three-dimensional pose of surgical instruments based on monocular depth estimation, characterized in that: The following steps are involved: S100 builds a laparoscopic surgery simulation scene and uses it to create a dataset of surgical instrument key points and monocular depth estimation; S200 builds a monocular absolute depth estimation model and uses a monocular depth estimation dataset to train the monocular absolute depth estimation model; S300: constructing a two-dimensional key point detection model, and based on the surgical instrument key points, using boundary expansion and random cropping methods to improve the detection stability of the two-dimensional key point detection model, and training the two-dimensional key point detection model; S400 couples the key point detection output by the two-dimensional key point detection model with the depth estimation results output by the monocular absolute depth estimation model to eliminate abnormal depth values of key points and restore the three-dimensional posture of surgical instruments.

2. A method for detecting three-dimensional posture of surgical instruments based on monocular depth estimation according to claim 1, characterized in that: Step S100 includes the following steps: S101 creates 3D models of surgical instruments and designs key points of each surgical instrument according to the characterization requirements of the surgical instruments; S102 constructs a laparoscopic surgery simulation scene, randomly adjusts the surgical instrument posture, light source position and brightness, and background image style to achieve data generalization, saves the rendering result image, camera depth map, and surgical instrument key point position information, and extracts monocular depth estimation data to generate a data set.

3. The method for detecting three-dimensional posture of surgical instruments based on monocular depth estimation according to claim 2, characterized in that: Step S102 further includes: A style transfer model is used to convert the style of the rendered image into the style of real-world images, achieving domain generalization of the dataset.

4. The method for detecting three-dimensional posture of surgical instruments based on monocular depth estimation according to claim 1, characterized in that: In step S200, DepthAnything is used as the monocular absolute depth estimation model, and SILogloss and L1loss are used as the loss functions of the monocular absolute depth estimation model to improve the smoothness of the prediction results.

5. The method for detecting three-dimensional posture of surgical instruments based on monocular depth estimation according to claim 1, characterized in that: The YOLOv8-pose model is used to build a two-dimensional plane key point detection model.

6. The method for detecting three-dimensional posture of surgical instruments based on monocular depth estimation according to claim 1, characterized in that: Step S400 includes the following steps: S401 adds depth information to each predicted key point according to the monocular depth estimation result to obtain the depth value of the key point; S402 fills in the depth value of the missing key point; S403 performs filtering on the depth values of the key points to remove abnormal values; S404 calculates the position of the key points of the surgical instrument relative to the camera lens based on the camera parameters, and restores the three-dimensional posture of the surgical instrument.

7. The method for detecting three-dimensional posture of surgical instruments based on monocular depth estimation according to claim 6, characterized in that: In step S401, for each key point p k =(x k ,y k ), its depth value Z k Get as: Z k =D(x k ,y k ) Where D is the depth estimation result after boundary expansion.

8. The method for detecting three-dimensional posture of surgical instruments based on monocular depth estimation according to claim 6, characterized in that: In step S402, the depth values of all pixels between two key points are first extracted and arranged in order, and then the points with lost depth are removed and the least squares method is used to fit the remaining data to fill in the lost depth values of the key points; In step S403, based on the coordination of the overall movement of the instrument, a sudden change in the depth of the instrument tip is processed. When the difference between the instrument tip depth value and the previous moment is greater than a set threshold, the depth value is regarded as an abnormal value. The abnormal value is processed as follows: depth err,j =depth err,j-1 +depth c,j -depth c,j-1 Among them, depth err,j is the correction value of the depth of the abnormal point, j is the current time series number, depth err,j-1 is the depth value of the abnormal point at the moment before, depth c,j is the depth value of the center point of the surgical instrument at the current moment, depth c,j-1 It is the depth value of the center point of the surgical instrument at the previous moment.

9. The method for detecting three-dimensional posture of surgical instruments based on monocular depth estimation according to claim 6, characterized in that: In step S404, the position of the surgical instrument key point relative to the camera lens includes: Where, X c and Y c is the coordinate in the camera coordinate system, u and v are the coordinates in the image, c x and c y is the principal point coordinate of the camera, Z c is the depth value obtained by depth estimation, and f is the focal length of the camera.

10. A three-dimensional posture detection system for surgical instruments based on monocular depth estimation, characterized in that: include: The first main control module is used to build a laparoscopic surgery simulation scene and generate a dataset of surgical instrument key points and monocular depth estimation; A second main control module is used to build a monocular absolute depth estimation model and train the monocular absolute depth estimation model using a monocular depth estimation dataset; A third main control module is used to build a two-dimensional key point detection model, and based on the surgical instrument key points, use boundary expansion and random cropping methods to improve the detection stability of the two-dimensional key point detection model and train the two-dimensional key point detection model; The fourth main control module is used to couple the key point detection output by the two-dimensional key point detection model and the depth estimation results output by the monocular absolute depth estimation model, eliminate abnormal depth values of the key points, and restore the three-dimensional posture of the surgical instrument.