Operation tail end instrument positioning method for microscopic laryngoscope simulation training system
Through the multi-sensor fusion method of binocular vision and IMU and extended Kalman filtering, the positioning accuracy and real-time problems of surgical instruments in virtual microlaryngoscope are solved, and high-precision real-time tracking of microsurgical instruments is realized, which is suitable for virtual training systems.
Patent Information
- Application Number
- CN202510731203.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-08-12
AI Technical Summary
The prior art cannot effectively locate surgical instruments in virtual microlaryngoscope scenarios, and traditional binocular vision methods are affected by light conditions and occlusion, and the accuracy and real-timeness are insufficient.
The multi-sensor fusion method of binocular vision and inertial measurement unit (IMU) is adopted, combined with extended Kalman filtering, by obtaining the image and motion data of the operation end instrument, its three-dimensional spatial coordinates and poses are calculated, and the optimal value is tracked in real time in the virtual scene.
It improves the accuracy and real-time calculation of surgical instrument position and reduces costs. It is suitable for real-time tracking of microsurgery instruments and is suitable for accurate positioning of virtual training systems.
Smart Images

Figure CN120472734A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical instruments and posture calculation technology, and in particular to a method for positioning an operating end instrument for a microlaryngoscopy simulation training system. Background Art
[0002] There are currently multiple methods for detecting the position and posture of surgical instruments, including binocular vision, radar, ultrasound, and magnetic fields. Target posture solution based on binocular vision uses a binocular system to obtain target surface features. This has low requirements for target surface features and can obtain more surrounding information, which can improve the target solution accuracy. However, it is easily affected by environmental factors such as lighting conditions and insufficient accuracy. At the same time, using only binocular vision for posture calculation has problems such as low accuracy and low real-time performance.
[0003] Research on surgical simulation systems based on virtual reality technology has led to the design of a virtual training system that uses 3D modeling and real-time image processing to recreate human anatomy and surgical procedures, providing the operator with multi-dimensional sensory feedback, including visual and auditory feedback. However, the positioning and tactile feedback of the end-point manipulator often suffer from insufficient precision and feedback delays, making it difficult to fully meet the posture calculation accuracy requirements for microlaryngoscopy. Summary of the Invention
[0004] The purpose of the present invention is to overcome the deficiencies of the above-mentioned prior art and provide an operating end instrument positioning method for a microlaryngoscopy simulation training system to solve the problem of being unable to position surgical instruments in a virtual microlaryngoscopy scene and to track microsurgical instruments in real time.
[0005] To achieve the above object, the technical solution of the present invention is:
[0006] A method for positioning an operating end instrument for a microlaryngoscopy simulation training system, comprising:
[0007] Acquire a first image, where the first image is an image of an operating terminal device taken by a binocular camera;
[0008] Calculating a three-dimensional spatial coordinate visual observation value of the operating end instrument based on the first image;
[0009] Acquiring motion data of the operating end device measured by an inertial measurement unit;
[0010] fusing the three-dimensional spatial coordinate visual observation value of the operation end instrument and the motion data of the operation end instrument to calculate the optimal value of the operation end instrument posture data;
[0011] Based on the optimal value, the end instrument is tracked in real time in the virtual surgical scene.
[0012] Optionally, calculating the three-dimensional spatial coordinate visual observation value of the operating end instrument based on the first image includes:
[0013] Preprocessing the first image to obtain a second image;
[0014] Predicting the image coordinates of the center of mass of the marker based on the second image to obtain a predicted value of the image coordinates of the center of mass of the marker; the marker is fixed to the operating end instrument;
[0015] Taking the image coordinate prediction value of the center of mass of the marker as the center, cutting out a portion of the second image to obtain a third image;
[0016] Extracting pixel points of the marker from the third image and obtaining image coordinates of the pixel points for pose calculation;
[0017] The three-dimensional space coordinate visual observation value of the operating end instrument is calculated based on the pixel point image coordinates.
[0018] Optionally, extracting pixel points of the marking element from the third image includes:
[0019] According to the HSV average value of the marker in multiple binocular camera images, the color screening threshold in the pixel extraction process is set;
[0020] The HSV value of each pixel point is read from the third image to determine whether it meets the color screening threshold. If it meets the color screening threshold requirement, the pixel point that meets the requirement is determined to belong to the marking part, and the coordinates of the pixel point image are stored at the same time.
[0021] Optionally, the marking member and the inertial measurement unit are both fixed on the operating end instrument; the number of the marking members is set to two.
[0022] Optionally, the average of the center points of the two marking members is consistent with the center point of the inertial measurement unit; the two marking members and the inertial measurement unit are fixed to an unobstructed part of the side of the operating end device by magnetic attraction.
[0023] Optionally, obtaining pixel image coordinates for pose calculation includes:
[0024] Calculate the image coordinates of the center of the minimum enclosing circle that encloses all the pixels of the markers;
[0025] The image coordinates of the centers of the minimum enclosing circles corresponding to the two markers are averaged to obtain the pixel image coordinates used for pose calculation.
[0026] Optionally, the three-dimensional spatial coordinate visual observation value of the operating end instrument is calculated based on the parallax method.
[0027] Optionally, the calculating of the three-dimensional spatial coordinate visual observation value of the operating end instrument according to the parallax method includes:
[0028] The formula for the Z-axis observation coordinate value of the end-of-operation instrument is:
[0029] Where b is the distance between the left and right cameras, u L ,u R represents the u coordinate of the pixel point in the left and right camera images used to calculate the three-dimensional position, and f represents the focal length of the binocular camera;
[0030] The X and Y axis coordinate values of the end-of-line instrument:
[0031] Where, (u L ,v L ) represents the image coordinates of the pixel point used to calculate the three-dimensional position in the left camera image, c x ,c y Indicates the vertical and horizontal offset of the image origin relative to the optical center imaging point, f x ,f y Indicates the horizontal and vertical focal lengths of the binocular camera lens.
[0032] Optionally, an extended Kalman filter is used to fuse the three-dimensional spatial coordinate visual observation value of the operation end instrument and the motion data of the operation end instrument to calculate the optimal value of the operation end instrument posture data
[0033] Compared with the prior art, the present invention has the following beneficial effects:
[0034] (1) There are currently multiple methods for detecting the position and posture of surgical instruments, including binocular vision, radar, ultrasound, and magnetic fields. Compared with radar and ultrasound, binocular vision is cheaper and easier to integrate. At the same time, because modern image processing technology can achieve rapid image analysis, binocular vision can achieve the same good real-time performance as radar and ultrasound technology for position and posture detection, which can meet the real-time tracking requirements of microsurgical instruments.
[0035] (2) The traditional binocular vision posture detection method only uses a single visual method to detect the posture, which is easily affected by changes in lighting conditions and hand occlusion. In addition, the rotation angle error of the operating terminal device calculated by a single visual method is large. The present invention uses a multi-sensor fusion posture calculation method of binocular vision and IMU. It uses IMU components to make up for the disadvantage that the binocular camera is easily affected by lighting conditions, and improves the accuracy of the posture calculation of the operating terminal device.
[0036] (3) The present invention uses a linear regression method to predict the coordinates of the marker image of the current frame, shortening the time to find the center of mass of the marker during the marker pixel screening process, thereby improving the speed of calculating the posture of the end-user device.
[0037] (4) The present invention uses an extended Kalman filter to process posture data, which can provide more stable and reliable posture output under high-speed or unstable motion conditions, and at the same time improve the posture calculation accuracy of the operating end device.
[0038] (5) The present invention uses magnetic attraction to fix the IMU and the marker, which is easy to install. By optimizing the arrangement of sensors, the surgical instrument can be accurately positioned at a micro scale, and the present invention can be conveniently used in other virtual training surgical systems for posture detection of the operating end instrument. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 Schematic diagram of the posture calculation platform and virtual scene for microlaryngoscopy simulation training;
[0040] Figure 2 This is a schematic diagram of detecting the posture of the end-of-operation instrument;
[0041] Figure 3 A flowchart of a method for positioning an operating end instrument for a microlaryngoscopy simulation training system provided in an embodiment of the present application;
[0042] Figure 4 is a flow chart of sub-steps of step 120;
[0043] Figure 5 is a flow chart of sub-steps of step 130;
[0044] Figure 6 is a flow chart of sub-steps of step 150;
[0045] Figure 7 The initial pose image of the end-of-operation instrument model;
[0046] Figure 8 It is an image of the end-of-operation instrument model in motion;
[0047] Figure 9 Images of real-time operation of microlaryngoscopy simulation training; DETAILED DESCRIPTION
[0048] Example:
[0049] The technical solution of the present invention is further described below with reference to the accompanying drawings and embodiments.
[0050] Before implementing the method for positioning the end-of-operation instrument in the microlaryngoscopy simulation training system of the present application, the following preparatory steps are required:
[0051] Steps for building a posture calculation platform and virtual scene for microlaryngoscopy simulation training:
[0052] Build a posture calculation platform and virtual scene for microlaryngoscopy simulation training. The completed environment is as follows: Figure 1 As shown in FIG, it includes a virtual surgical scene 1, a binocular camera 2, a marker 3, an inertial measurement unit 4 (IMU for short) and a microsurgery instrument 5. The process diagram of detecting the posture data of the end-operation instrument is shown in FIG. Figure 2 As shown, the operation end device model in the virtual scene will track the operation end device in reality in real time.
[0053] Fixing assembly step: fix the marker 3, inertial measurement unit 4 (IMU for short) and microsurgical instrument 5 into one body, and use the microsurgical instrument with the marker and IMU fixed as the operating end instrument for microlaryngoscopy simulation training.
[0054] In the specific implementation, two red spherical markers with a strong contrast to the background color are selected, and the diameter of the markers is 10 mm. The markers and the IMU are fixed on the side of the microsurgical instrument by magnetic attraction. The fixed position requires that the average of the center points of the two markers is consistent with the center point of the IMU, so that the object posture data measured by the markers and the IMU are in the same coordinate system, without the need for coordinate system transformation. At the same time, they should be fixed in a position that is not easily blocked to avoid interference from external lighting conditions. In this way, by cleverly designing the operating terminal suitable for microlaryngoscopy simulation training and optimizing the layout of sensors, precise positioning of surgical instruments at a microscale can be achieved.
[0055] After completing the two preparation steps above, refer to Figure 3 As shown, the method for positioning the operating end instrument for the microlaryngoscopy simulation training system provided in the embodiment of the present application mainly includes the following steps:
[0056] 110. Acquire a first image, where the first image is an image of an operating terminal device captured by a binocular camera;
[0057] 120. Calculate a three-dimensional spatial coordinate visual observation value of the operating end instrument based on the first image;
[0058] 130, obtaining motion data of the operating end device measured by an inertial measurement unit;
[0059] 140 , fusing the three-dimensional spatial coordinate visual observation value of the operation end instrument and the motion data of the operation end instrument to calculate the optimal value of the position and posture data of the operation end instrument;
[0060] 150 , based on the optimal value, track the operating end instrument in real time in the virtual surgical scene.
[0061] This shows that this method uses binocular vision to detect the position and posture of surgical instruments, which is lower in cost and easier to integrate than radar, ultrasound, etc. At the same time, because modern image processing technology can achieve rapid image analysis, binocular vision can detect posture with good real-time performance, just like radar and ultrasound technology, and can meet the real-time tracking requirements of microsurgical instruments. Traditional binocular vision posture detection methods only use a single visual method to detect posture, which is easily affected by changes in lighting conditions and hand occlusion. In addition, the rotation angle calculation error of the end-of-operation instrument calculated by a single visual method is large. This method uses a multi-sensor fusion posture calculation method of binocular vision and IMU. It uses IMU components to make up for the disadvantage of binocular cameras being easily affected by lighting conditions, and improves the accuracy of the end-of-operation instrument posture calculation.
[0062] In a specific embodiment, if Figure 4 As shown, the above step 120 includes the following sub-steps:
[0063] 1201, preprocess the first image to obtain a second image;
[0064] In a specific implementation, a binocular camera is used to capture and transmit an image of the end-use instrument. The image of the end-use instrument (i.e., the first image) is pre-processed in real time, including cropping the image size, converting the image resolution, correcting the image, and performing noise reduction. This results in an image that meets the requirements of subsequent operations, i.e., the second image.
[0065] 1202, predicting the image coordinates of the center of mass of the marker based on the second image to obtain a predicted value of the image coordinates of the center of mass of the marker;
[0066] In a specific implementation, the historical image coordinates of the marker in the previous n frames are read, and the image coordinates of the qth marker in the pth image in the i-th group of binocular camera images in the previous n frames are:
[0067] (i=1, 2, ..., n, p=L, indicating the left image, p=R, indicating the right image, q=1, 2); n is an integer;
[0068] The image coordinates of the markers in the current image are predicted using a linear regression method; the linear regression model used is as follows:
[0069] For the parameter calculation formula in the prediction u coordinate formula:
[0070] Therefore, the formula for predicting the centroid image coordinates of the marker at the current moment is as follows:
[0071]
[0072] Thus, the image coordinate prediction value of the center of mass of the marker is obtained.
[0073] Thus, in this sub-step, the linear regression method is used to predict the coordinates of the marker image of the current frame, shortening the time to find the center of mass of the marker during the marker pixel screening process, thereby improving the speed of calculating the posture of the end-operating device.
[0074] 1203, using the image coordinate prediction value of the center of mass of the marker A portion of the image is cut out from the second image with the image being the center, to obtain a third image; the size of the cut out image refers to the average size of the marker in the camera image;
[0075] 1204 , extract pixel points of the marker from the third image, and obtain image coordinates of the pixel points for pose calculation.
[0076] In a specific implementation, extracting pixel points of the marking element from the third image includes:
[0077] According to the HSV average value of the marker in multiple binocular camera images, the color screening threshold in the pixel extraction process is set. Assume that (H min ,S min ,V min ) is the lower color threshold, (H max ,S max ,V max ) is the upper color threshold;
[0078] Read each pixel PX from the binocular camera image i The color value (H i ,S i ,V i ), determine whether the color screening threshold set in the above step is met. If the color screening threshold requirement is met, it can be considered that the pixel point belongs to the marking part, and the coordinates of the pixel point image are stored;
[0079] The formula for filtering pixels is:
[0080] In this way, the pixel points of the marking element can be accurately and efficiently extracted through the above method.
[0081] The obtaining of pixel image coordinates for pose calculation includes:
[0082] Since the marker selected in this method is spherical, to calculate the image coordinates of the centroid of a single marker, it is only necessary to calculate the image coordinates of the center of the minimum enclosing circle that encloses the pixel point of the current marker;
[0083] The formula for calculating the minimum radius circle that encloses all pixels that meet the threshold is:
[0084]
[0085] Where C pq (p=L,R,q=1,2) is the minimum circle that encloses all pixels of the marker q in the photo p, (u pq ,u pq ) is the minimum enclosing circle C pq The image coordinates of the circle center, is the pixel coordinate of the pixel point in the image that meets the threshold requirement, r pq is the minimum enclosing circle C pq radius;
[0086] The image coordinates of the center of the minimum enclosing circle corresponding to the two markers (u pq ,v pq ) and take the mean value to obtain the pixel image coordinates (u p , v p )(p=L,R).
[0087] In this way, the pixel image coordinates used for pose calculation can be accurately obtained through the above method.
[0088] 1205 , calculating the three-dimensional space coordinate visual observation value of the operating end instrument based on the pixel image coordinates.
[0089] In the specific implementation, the three-dimensional space coordinate visual observation value of the end-of-operation instrument is calculated according to the parallax method. The specific calculation formula includes:
[0090] The formula for calculating the Z-axis observation coordinate value of the end-of-operation instrument is:
[0091] Where b is the distance between the left and right cameras, u L ,u R represents the u coordinate of the pixel point in the left and right camera images used to calculate the three-dimensional position, and f represents the focal length of the binocular camera;
[0092] Calculate the X and Y axis observation coordinate values of the end-of-operation instrument:
[0093] Where, (uL ,v L ) represents the image coordinates of the pixel point used to calculate the three-dimensional position in the left camera image, c x ,c y Indicates the vertical and horizontal offset of the image origin relative to the optical center imaging point, f x ,f y Indicates the horizontal and vertical focal lengths of the binocular camera lens.
[0094] In this way, the three-dimensional spatial coordinate visual observation value of the operating end instrument can be calculated quickly and accurately.
[0095] In a specific embodiment, if Figure 5 As shown, the above step 130 includes the following sub-steps:
[0096] 1301, initializing the IMU and measuring the errors of the angular velocity and acceleration obtained by the IMU;
[0097] 1302, using IMU to obtain the angular velocity (ω x ,ω y ,ω z ) and acceleration (a x ,a y ,a z ) and transfer it to your computer.
[0098] In a specific embodiment, in step 140, the extended Kalman filter is used to fuse the three-dimensional spatial coordinate visual observation value of the operation end device (hereinafter referred to as the state observation value of the binocular camera) and the motion data of the operation end device (hereinafter referred to as the IMU data), based on the state x of the operation end device in the previous frame. k-1 , use IMU data to predict the state of the end device in the current frame
[0099] The formula for predicting pose data based on Kalman filtering is:
[0100] The specific formula for predicting pose data is:
[0101] Among them, x k-1 =[p k v k θ k ] T represents the state of the end device in the k-1th frame image, p k Indicates the three-dimensional space coordinates of the end-of-operation instrument, v k Indicates the speed of the end-device operation, θ k represents the posture of the end device, F represents the state transfer matrix, uk =[ω k a k ] T Represents the rotation data of the end-of-operation device measured by the IMU, ω k represents the angular velocity of the end-of-line instrument, a k Indicates the acceleration of the end-of-line instrument;
[0102] Based on the predicted state value of the end-user device predicted by the IMU data and the state observation value of the binocular camera, the optimal state value x of the end-user device in the current frame is calculated. k ;
[0103] The formula for calculating the Kalman gain K is:
[0104] Among them, P k represents the covariance matrix, H k represents the observation matrix, R k represents the observation noise covariance matrix.
[0105] The formula for calculating the optimal value of the end-of-operation instrument posture data is:
[0106] Among them, z k The binocular camera image observation value representing the state of the end-of-line instrument, Represents the observation model.
[0107] Through the above calculations, a multi-sensor fusion pose calculation method using binocular vision and IMU is used, and an extended Kalman filter algorithm is used to fuse and correct the data, improving the real-time performance and accuracy of positioning. This method not only accurately reflects the state of the operator terminal, but also provides more intuitive feedback information for training.
[0108] In a specific embodiment, if Figure 6 The above step 150 includes:
[0109] 1501, transmitting the optimal value of the pose data processed by the extended Kalman filter to the virtual scene;
[0110] 1502, assigning posture data to the end device model in real time to achieve posture tracking.
[0111] So, like Figure 7-9 As shown in the figure, it is possible to track the actual end-of-operation instruments in real time in the virtual surgical scene. Figure 7 is the initial pose image of the end-of-line instrument model, Figure 8 To operate the image of the end instrument model in motion, Figure 9 An image of a live run for microlaryngoscopy simulation training.
[0112] The above embodiments are intended only to illustrate the technical concepts and features of the present invention. Their purpose is to enable those skilled in the art to understand the contents of the present invention and implement them accordingly. They are not intended to limit the scope of protection of the present invention. Any equivalent changes or modifications made based on the essence of the present invention are intended to be covered by the scope of protection of the present invention.
Claims
1. A method for positioning an operating end instrument for a microlaryngoscopy simulation training system, characterized in that: include: Acquire a first image, where the first image is an image of an operating terminal device taken by a binocular camera; Calculating a three-dimensional spatial coordinate visual observation value of the operating end instrument based on the first image; Acquiring motion data of the end-of-operation instrument measured by an inertial measurement unit; fusing the three-dimensional spatial coordinate visual observation value of the operation end instrument and the motion data of the operation end instrument to calculate the optimal value of the operation end instrument posture data; Based on the optimal value, the end instrument is tracked in real time in the virtual surgical scene.
2. The method for positioning an operating end instrument for a microlaryngoscopy simulation training system according to claim 1, wherein: The calculating of the three-dimensional spatial coordinate visual observation value of the operating end instrument based on the first image includes: Preprocessing the first image to obtain a second image; Predicting the image coordinates of the center of mass of the marker based on the second image to obtain a predicted value of the image coordinates of the center of mass of the marker; the marker is fixed to the operating end instrument; Taking the image coordinate prediction value of the center of mass of the marker as the center, cutting out a portion of the second image to obtain a third image; Extracting pixel points of the marker from the third image and obtaining image coordinates of the pixel points for pose calculation; The three-dimensional space coordinate visual observation value of the operating end instrument is calculated based on the pixel point image coordinates.
3. The method for positioning an operating end instrument for a microlaryngoscopy simulation training system according to claim 2, wherein: Extracting pixel points of the marking element from the third image includes: According to the HSV average value of the marker in multiple binocular camera images, the color screening threshold in the pixel extraction process is set; The HSV value of each pixel point is read from the third image to determine whether it meets the color screening threshold. If it meets the color screening threshold requirement, the pixel point that meets the requirement is determined to belong to the marking part, and the coordinates of the pixel point image are stored at the same time.
4. The method for positioning an operating end instrument for a microlaryngoscopy simulation training system according to claim 3, wherein: The marking piece and the inertial measurement unit are both fixed on the operating end instrument; the number of the marking pieces is two.
5. The method for positioning an operating end instrument for a microlaryngoscopy simulation training system according to claim 4, characterized in that: The average of the center points of the two marking members is consistent with the center point of the inertial measurement unit; the two marking members and the inertial measurement unit are fixed to an unobstructed part of the side of the operating end device by magnetic attraction.
6. The method for positioning an operating end instrument for a microlaryngoscopy simulation training system according to claim 4 or 5, characterized in that: The obtaining of pixel image coordinates for pose calculation includes: Calculate the image coordinates of the center of the minimum enclosing circle that encloses all the pixels of the markers; The image coordinates of the centers of the minimum enclosing circles corresponding to the two markers are averaged to obtain the pixel image coordinates used for pose calculation.
7. The method for positioning an operating end instrument for a microlaryngoscopy simulation training system according to claim 5, characterized in that: The three-dimensional spatial coordinate visual observation value of the operating end instrument is calculated based on the parallax method.
8. The method for positioning an operating end instrument for a microlaryngoscopy simulation training system according to claim 7, wherein: The method of calculating the three-dimensional spatial coordinate visual observation value of the operating end instrument according to the parallax method includes: The formula for the Z-axis observation coordinate value of the end-of-operation instrument is: Where b is the distance between the left and right cameras, u L ,u R represents the u coordinate of the pixel point in the left and right camera images used to calculate the three-dimensional position, and f represents the focal length of the binocular camera; The X and Y axis coordinate values of the end-of-line instrument: Where, (u L ,v L ) represents the image coordinates of the pixel point used to calculate the three-dimensional position in the left camera image, c x ,c y Indicates the vertical and horizontal offset of the image origin relative to the optical center imaging point, f x ,f y Indicates the horizontal and vertical focal lengths of the binocular camera lens.
9. The method for positioning an operating end instrument for a microlaryngoscopy simulation training system according to claim 1, wherein: The extended Kalman filter is used to fuse the three-dimensional spatial coordinate visual observation value of the operation end instrument and the motion data of the operation end instrument to calculate the optimal value of the operation end instrument posture data.