Endoscopic surgery target positioning device and method based on multi-modal image fusion

By combining binocular stereoscopic vision and fluorescence imaging technology in laparoscopic surgery, the three-dimensional posture and depth information of fluorescent marking points is solved, and a more efficient and safe surgical process is achieved.

CN120189057APending Publication Date: 2025-06-24BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510298945.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

In existing laparoscopic surgery, it is difficult to obtain the depth information of soft tissues and organs and the three-dimensional positioning of fluorescent marking points, resulting in low positioning accuracy and poor real-time performance, which cannot meet the needs of autonomous surgery.

Method used

The laminoscopic surgical target positioning device using multimodal image fusion is used to combine binocular stereoscopic vision and fluorescence imaging technology to register the information of the fluorescent camera and the binocular camera to obtain the three-dimensional posture and depth information of the fluorescent marking points.

Benefits of technology

It achieves more accurate and real-time visual assistance in laparoscopic surgery, improves the safety and efficiency of the surgery, and meets the high accuracy and high real-time requirements of autonomous surgery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120189057A_ABST
    Figure CN120189057A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal image fusion endoscope operation target positioning device and method, and the method comprises the steps: calibrating internal and external parameters of a binocular camera and a fluorescence camera, and obtaining a pose conversion matrix between the fluorescence camera and the binocular camera; using a U-Net network model to identify and track the position of the fluorescent mark point in the fluorescent image in real time, and obtaining the two-dimensional image coordinate of the fluorescent mark point; processing a left image and a right image of the binocular camera by using a region-based local stereo matching method in binocular stereo imaging to obtain a point cloud image of a target space; and obtaining three-dimensional pose information of the fluorescent mark point by combining the two-dimensional image coordinate of the fluorescent mark point and the point cloud image of the target space. The binocular stereoscopic vision and the fluorescence imaging technology are combined, the three-dimensional pose and depth information of the fluorescence mark point is obtained by registering the information between the fluorescence camera and the binocular camera, more accurate and real-time vision assistance is provided for an operation, and the safety and efficiency of the operation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical devices, and particularly to a laparoscopic surgery target positioning device and method for multi-modal image fusion. Background Art

[0002] In medical surgery, the use of laparoscopic surgical robots can improve the accuracy of surgery and reduce the risks during the surgical process. With the development of intelligent autonomous surgery, using a laparoscope to locate a target and then controlling the robotic arm for surgical operations has become an important development direction in the surgical field. Currently, most of the targets in laparoscopic surgery are soft tissue organs. Most existing medical laparoscopes can only observe two-dimensional images, making it difficult to obtain depth information and unable to be directly used for the observation and positioning of autonomous surgery. Laparoscopic surgical robots, etc., adopt 3D laparoscopes, which can provide depth information through 3D imaging methods and provide three-dimensional positioning information for the target. However, laparoscopic surgical robots face the following core problems: The operating object is a soft tissue organ with weak image texture and unclear features, resulting in difficult feature matching during stereoscopic imaging and making it difficult to give accurate target position information. Moreover, the soft tissue organ has large elasticity. For example, during the suturing process, the front and back changes of the target tissue are large, and it is difficult to identify the surgical environment information and the target area, resulting in great difficulty and long recognition cycle for binocular stereoscopic positioning using a simple 3D laparoscope, and unable to meet the real-time motion control requirements of autonomous surgery.

[0003] In addition, in the existing binocular stereovision technology, the target object images of two cameras are respectively acquired, and the feature points on the two images are matched to calculate the parallax to obtain the three-dimensional depth information of the target point. However, this method relies on high-quality image alignment and accurate matching algorithms. In an environment with low contrast or less texture, the matching accuracy often cannot meet the requirements of high-precision, high real-time frame rate, and high robustness positioning, resulting in large errors in the depth information of the target point and low real-time performance and robustness.

[0004] Currently, the near-infrared fluorescence laparoscopy technology marked with indocyanine green is gradually applied to clinical surgeries. By injecting indocyanine green reagent into the target tissue, under the irradiation of near-infrared light, the target tissue will emit a strong near-infrared light signal. Compared with traditional visible light, the near-infrared light signal has stronger surface penetration, can effectively overcome the problem of surface tissue occlusion, thereby helping the surgeon distinguish the target tissue from the remaining healthy tissues, contributing to reducing the damage to healthy tissues during the surgery and improving the safety of the surgery.

[0005] Taking advantage of the characteristic that only the fluorescence - labeled points are visible in the fluorescence image, the recognition difficulty can be effectively reduced and the real - time performance can be improved. However, the existing fluorescence imaging system and visible - light imaging system are independent of each other. The fluorescence imaging system can only display the positions of the fluorescence - labeled objects on the two - dimensional image plane, without an effective algorithm to extract the two - dimensional position information of the fluorescence - labeled objects, and thus cannot fully utilize the positioning advantage of the fluorescence - labeled objects during the operation. The existing 3D endoscope imaging system can only perform three - dimensional visible - light imaging on the target area, and cannot fuse the depth information of visible light with the fluorescence information, resulting in limited application of 3D endoscopes in high - precision positioning. Summary of the Invention

[0006] Aiming at the problems of the existing technology mainly relying on binocular stereo imaging, being unable to obtain the depth information of the fluorescence - labeled points, having poor 3D imaging ability and low positioning accuracy, etc., the present invention combines binocular stereo vision and fluorescence imaging technology, and provides a laparoscopic surgical target positioning device and method for multi - modal image fusion. By registering the information between the fluorescence camera and the binocular camera, the three - dimensional pose and depth information of the fluorescence - labeled points are obtained, providing more accurate and real - time visual assistance for the operation, and improving the safety and efficiency of the operation.

[0007] To solve the above - mentioned technical problems, the present invention provides the following technical solutions:

[0008] On the one hand, a laparoscopic surgical target positioning device for multi - modal image fusion is provided. The device includes: a binocular camera, a fluorescence camera, a binocular image acquisition card, a fluorescence image acquisition card, a computer, and a display;

[0009] The binocular camera includes a binocular endoscope, which acquires binocular visible - light images of the target space from two different left and right perspectives through double image sensors; the binocular image acquisition card performs digital processing and format conversion on the acquired binocular visible - light images and transmits them to the computer;

[0010] The fluorescence camera includes a fluorescence endoscope, which acquires fluorescence images of the target space through an image sensor; the fluorescence image acquisition card performs digital processing and format conversion on the acquired fluorescence images and transmits them to the computer;

[0011] After receiving the binocular visible - light images and fluorescence images, the computer performs fusion processing, obtains the point - cloud image of the target space and the depth information of the fluorescence - labeled points, and displays them through the display.

[0012] Optionally, the device adopts a binocular endoscope - binocular fluorescence co - axial installation method; the binocular endoscope includes a first visible - light image sensor and a second visible - light image sensor, and the two fluorescence endoscopes respectively include a first fluorescence image sensor and a second fluorescence image sensor;

[0013] Among them, in the binocular endoscope, each monocular endoscope and a fluorescence endoscope are coaxially installed with a light source and share an optical viewing tube, forming a first monocular endoscope module and a second monocular endoscope module. The first monocular endoscope module and the second monocular endoscope module are connected through a binocular endoscope mounting base;

[0014] The normal direction of the first visible light image sensor is coaxial with the first optical viewing tube. The first visible light image sensor and the first fluorescence image sensor are orthogonally arranged, and a first beam splitter is installed at the intersection of the normal lines at the center point. Visible light passes through the first optical viewing tube and penetrates the first beam splitter to reach the first visible light image sensor, and fluorescence reaches the first fluorescence image sensor through the reflection of the first beam splitter;

[0015] The normal direction of the second visible light image sensor is coaxial with the second optical viewing tube. The second visible light image sensor and the second fluorescence image sensor are orthogonally arranged, and a second beam splitter is installed at the intersection of the normal lines at the center point. Visible light passes through the second optical viewing tube and penetrates the second beam splitter to reach the second visible light image sensor, and fluorescence reaches the second fluorescence image sensor through the reflection of the second beam splitter.

[0016] Optionally, the device adopts an independent installation method for binocular endoscope - binocular fluorescence; the binocular endoscope includes a first visible light image sensor and a second visible light image sensor, and the two fluorescence endoscopes respectively include a first fluorescence image sensor and a second fluorescence image sensor;

[0017] Among them, the binocular endoscope includes two monocular endoscopes, which are connected through a binocular endoscope mounting base; the normal direction of the first visible light image sensor is coaxial with the first optical viewing tube, and visible light reaches the first visible light image sensor through the first optical viewing tube; the normal direction of the second visible light image sensor is coaxial with the second optical viewing tube, and visible light reaches the second visible light image sensor through the second optical viewing tube;

[0018] The two fluorescence endoscopes are connected through a binocular fluorescence mounting base; the normal direction of the first fluorescence image sensor is coaxial with the third optical viewing tube, and fluorescence reaches the first fluorescence image sensor through the third optical viewing tube; the normal direction of the second fluorescence image sensor is coaxial with the fourth optical viewing tube, and fluorescence reaches the second fluorescence image sensor through the fourth optical viewing tube; the binocular endoscope mounting base is connected to the binocular fluorescence mounting base.

[0019] Optionally, the device adopts an independent installation method for binocular endoscope - monocular fluorescence; the binocular endoscope includes a first visible light image sensor and a second visible light image sensor, and the fluorescence endoscope includes a fluorescence image sensor;

[0020] Among them, the binocular endoscope includes two monocular endoscopes, which are connected by a binocular endoscope mounting base; the normal direction of the first visible light image sensor is coaxial with the first optical viewing tube, and visible light reaches the first visible light image sensor through the first optical viewing tube; the normal direction of the second visible light image sensor is coaxial with the second optical viewing tube, and visible light reaches the second visible light image sensor through the second optical viewing tube;

[0021] The fluorescence endoscope is mounted on the monocular fluorescence mounting base; the normal direction of the fluorescence image sensor is coaxial with the third optical viewing tube, and fluorescence reaches the fluorescence image sensor through the third optical viewing tube; the binocular endoscope mounting base is connected to the monocular fluorescence mounting base.

[0022] On the other hand, a method for positioning the target of endoscopic surgery based on multimodal image fusion of the above device is provided, and the method includes the following steps:

[0023] S1. Camera calibration: Calibrate the internal and external parameters of the binocular camera and the fluorescence camera respectively to obtain the pose transformation matrix between the fluorescence camera and the binocular camera;

[0024] Among them, the binocular camera includes a binocular endoscope, and the fluorescence camera includes a fluorescence endoscope;

[0025] S2. Fluorescent marker point recognition: Use the U-Net network model to real-time recognize and track the position of the fluorescent marker point in the fluorescence image, and obtain the two-dimensional image coordinates of the fluorescent marker point;

[0026] S3. Binocular stereo imaging: Use the region-based local stereo matching method in binocular stereo imaging to process the left image and the right image of the binocular camera, and obtain the point cloud image of the target space;

[0027] S4. Image fusion: Combine the two-dimensional image coordinates of the fluorescent marker point and the point cloud image of the target space to obtain the three-dimensional pose information of the fluorescent marker point.

[0028] Optionally, the step S1 specifically includes:

[0029] Calibration of the internal and external parameters of the binocular camera: Use the binocular camera to take multiple pictures of the calibration board, and use the Zhang-Zhengyou calibration method to calibrate the internal parameter matrix and the external parameter matrix of the binocular camera;

[0030] Calibration of the internal and external parameters of the fluorescence camera: Use the fluorescence camera to take multiple pictures of the calibration board, and obtain the internal parameter matrix and the external parameter matrix of the fluorescence camera according to the pixel coordinates of the corner points on the calibration board and the actual spatial position coordinates of the corner points;

[0031] Correct the camera distortion;

[0032] The specific calibration process is as follows:

[0033] The point coordinates (x W , y W , z W ) in the world coordinate system are transformed through the external parameter matrix of the camera into the point coordinates (x C , y C , z C ) in the camera coordinate system; the camera is a binocular camera or a fluorescence camera;

[0034] The point coordinates (x C , y C , z C ) in the camera coordinate system are transformed through the internal parameter matrix into the pixel coordinates (u, v) on the image coordinate system;

[0035] The camera distortion is corrected by the distortion parameters to obtain the corrected pixel coordinates (u', v'); the distortion parameters include radial distortion parameters and tangential distortion parameters;

[0036] The internal parameter matrices I S and I N of the binocular camera and the fluorescence camera, the external parameter matrices H S and H N of the binocular camera and the fluorescence camera relative to the same calibration board, and the distortion parameters of the binocular camera and the fluorescence camera are obtained respectively through the above steps; therefrom, the pose transformation matrix H S-N between the fluorescence camera and the binocular camera is obtained:

[0037] H S-N = H N -1 * H S

[0038] The conversion relationship between the point coordinates (x S-N , y CN , z CN ) of the same fluorescence marker point in the fluorescence camera coordinate system and the point coordinates (x CN , y CS , z CS , z CS ) in the binocular camera coordinate system is obtained through the pose transformation matrix H

[0039] Optionally, step S2 specifically includes:

[0040] Data set collection: Collect the image data set of the fluorescence marker point;

[0041] Data set annotation: Use label-me to annotate the collected image data set;

[0042] Label file generation: The annotation information of each image is stored in a JSON file. Based on the original image and the generated JSON file, the target is segmented to obtain the corresponding segmentation label file, and the dataset for training is obtained;

[0043] Dataset division: The dataset is divided into a training set and a validation set according to a predetermined ratio;

[0044] Model training and validation: The U-Net network model is trained using the training set, and the U-Net network model is validated using the validation set.

[0045] Optionally, step S3 specifically includes:

[0046] Epipolar rectification: Geometric transformation is performed on the original left and right images captured by the binocular camera so that the corresponding image points are aligned on the same horizontal line;

[0047] Stereo matching: It includes matching cost calculation and cost aggregation, and the depth information is calculated by obtaining the disparity of the target point in the left and right images;

[0048] The matching cost calculation refers to calculating the matching cost of each pixel point in the left image of the binocular camera within a given disparity range after epipolar rectification to measure the similarity with the corresponding pixel point in the right image; among them, the sum of absolute differences of grayscales is used to calculate the matching cost;

[0049] Cost aggregation refers to using Gaussian filtering for cost aggregation, calculating the weighted average of the matching costs within the specified neighborhood of each pixel point, that is, weighting and averaging the matching cost of each pixel point with the matching costs of its neighborhood pixel points, and the distribution of the weights follows a Gaussian distribution to reduce noise and false matches;

[0050] Depth calculation: After cost aggregation is completed, the disparity value corresponding to the minimum matching cost is selected for each pixel point to obtain the depth information of each pixel point;

[0051] Point cloud acquisition: After calculating the depth information of the pixel points, the point cloud image of the target space is obtained using the internal parameter matrix of the binocular camera.

[0052] Optionally, step S4 specifically includes:

[0053] For the independent installation method of binocular endoscope - monocular fluorescence, use the U-Net deep learning network model to real-time identify the pixel coordinates (u i , v i ) of the fluorescence marker points on the fluorescence image; use the binocular stereo imaging method to obtain the three-dimensional point cloud (x Si , y Si , z Si); Utilize the pose transformation matrix H between the fluorescence camera and the binocular camera obtained during the calibration phase S-N , and transform the original three-dimensional point cloud obtained by binocular stereo imaging into the coordinate system of the fluorescence camera:

[0054]

[0055] Then, transform all three-dimensional coordinates in the coordinate system of the fluorescence camera into the image coordinate system of the fluorescence camera:

[0056]

[0057] Thus, it can be obtained that for each original three-dimensional point cloud coordinate (x Si , y Si , z Si ), a corresponding pseudo-pixel coordinate (u i ’, v i ’) can be found on the fluorescence image; for each pixel coordinate (u i , v i ) of the fluorescence marker point recognized by the U-Net network, use the Euclidean distance to calculate the pseudo-pixel coordinate (u imin ’, v imin ’) with the smallest distance from this pixel coordinate, record the subscript index of this pseudo-pixel coordinate, and then the three-dimensional pose coordinate (x Si , y Si , z Si ) of this fluorescence marker point can be read reversely on the three-dimensional point cloud;

[0058] For the binocular endoscope - binocular fluorescence coaxial installation method, the pixel coordinate of the target fluorescence marker point on the fluorescence image is the same as the pixel coordinate on the depth image of the binocular endoscope. Determine the target position corresponding to the fluorescence marker point on the depth image through the one-to-one correspondence relationship, and further obtain the three-dimensional pose coordinate of the fluorescence marker point;

[0059] For the binocular endoscope - binocular fluorescence independent installation method, during the camera calibration phase, obtain the coordinate transformation relationship H S between the binocular camera and the visible light calibration board, the coordinate transformation relationship H N between the fluorescence camera and the fluorescence marker board, and the transformation relationship H ns between the fluorescence marker board and the visible light marker board, so as to obtain the coordinate transformation relationship H N-S = H S -1 * H ns * H N; In the binocular stereo imaging stage, using the binocular stereo imaging method, three-dimensional reconstruction is performed on the binocular endoscope and the binocular fluorescence endoscope respectively to obtain the three-dimensional point cloud of the visible light image and the three-dimensional point cloud of the fluorescence image; then using the coordinate transformation relationship H N-S , the three-dimensional point cloud of the visible light image is fused with the three-dimensional point cloud of the fluorescence image, so as to obtain three-dimensional point cloud data containing both visible light information and fluorescence information.

[0060] On the other hand, an electronic device is provided, and the electronic device includes:

[0061] A processor;

[0062] A memory, on which computer-readable instructions are stored. When the computer-readable instructions are loaded and executed by the processor, the steps of the endoscopic surgery target localization method for multimodal image fusion as described above are implemented.

[0063] On the other hand, a computer-readable storage medium is provided, and program code is stored in the computer-readable storage medium. The program code can be called by the processor to execute the steps of the endoscopic surgery target localization method for multimodal image fusion as described above.

[0064] The beneficial effects brought by the technical solution provided by the present invention at least include:

[0065] (1) The present invention proposes three different structures of endoscopic surgery target three-dimensional localization devices for multimodal image fusion: binocular endoscope-binocular fluorescence coaxial installation, binocular endoscope-binocular fluorescence independent installation, and binocular endoscope-monocular fluorescence independent installation. The binocular endoscope-binocular fluorescence coaxial installation structure simplifies the fusion method of visible light and fluorescence images by coaxial installation of the binocular endoscope and the binocular fluorescence, sharing an optical viewing tube; the binocular endoscope-binocular fluorescence independent installation can obtain a more independent and flexible split installation when facing different surgical scenarios; the binocular endoscope-monocular fluorescence independent installation structure can ensure the efficient capture of fluorescence signals by the monocular fluorescence while ensuring a certain degree of flexibility, ensuring the real-time nature of intraoperative identification.

[0066] (2) By using the U-Net network, the present invention can detect the two-dimensional coordinate information of the fluorescence marker points on the image in real time and accurately, and adopts binocular stereo localization and three-dimensional point cloud reconstruction technologies, effectively improving the efficiency and accuracy of intraoperative three-dimensional imaging. It can not only obtain the depth information of the target tissue in real time, but also accurately track the three-dimensional position of the fluorescence marker points, helping the surgeon and the robot to more accurately locate the lesion area and improving the surgical efficiency. Description of the Drawings

[0067] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0068] Figure 1 is a schematic structural diagram of a laparoscopic surgery target positioning device for multimodal image fusion provided by an embodiment of the present invention;

[0069] Figure 2 is a schematic diagram of the coaxial installation of binocular laparoscope - binocular fluorescence provided by an embodiment of the present invention;

[0070] Figure 3 is a schematic diagram of the independent installation of binocular laparoscope - binocular fluorescence provided by an embodiment of the present invention;

[0071] Figure 4 is a schematic diagram of the independent installation of binocular laparoscope - monocular fluorescence provided by an embodiment of the present invention;

[0072] Figure 5 is a flowchart of a method for laparoscopic surgery target positioning for multimodal image fusion provided by an embodiment of the present invention;

[0073] Figure 6 is a schematic diagram of the pose relationship between a binocular camera and a fluorescence camera provided by an embodiment of the present invention;

[0074] Figure 7 is a schematic diagram of the U-Net network architecture provided by an embodiment of the present invention;

[0075] Figure 8 is a technical roadmap for using the U-Net network to identify and track fluorescence marker points provided by an embodiment of the present invention;

[0076] Figure 9 is a technical roadmap for binocular stereoscopic imaging provided by an embodiment of the present invention;

[0077] Figure 10 is a schematic diagram of a depth calculation method provided by an embodiment of the present invention. Detailed implementation manners

[0078] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0079] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to give examples, illustrations or explanations. Any embodiment or design described as an "example" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Rather, the use of the word "example" is intended to present concepts in a specific manner.

[0080] The embodiments of the present invention provide a laparoscopic surgery target positioning device for multimodal image fusion. Refer to Figure 1 As shown, the device includes: a binocular camera, a fluorescence camera, a binocular image acquisition card, a fluorescence image acquisition card, a computer, a display, etc.

[0081] Among them, the binocular camera includes a binocular endoscope, which respectively acquires binocular visible light images of the target space from two different left and right perspectives through a dual image sensor (such as a CMOS sensor); the binocular image acquisition card performs digital processing and format conversion on the acquired binocular visible light images and transmits them to the computer;

[0082] The fluorescence camera includes a fluorescence endoscope, which acquires fluorescence images of the target space through an image sensor; the fluorescence image acquisition card performs digital processing and format conversion on the acquired fluorescence images and transmits them to the computer;

[0083] After receiving the binocular visible light images and fluorescence images, the computer performs fusion processing, obtains the point cloud image of the target space and the depth information of the fluorescence marker points, and displays them through the display.

[0084] In the embodiments of the present invention, the laparoscopic surgery target positioning device for multimodal image fusion includes 1 binocular endoscope and at least 1 fluorescence endoscope, and three different installation methods can be adopted.

[0085] The first installation method is as Figure 2 As shown, the device adopts a binocular endoscope-binocular fluorescence coaxial installation method. The binocular endoscope includes a first visible light image sensor (visible light image sensor 1) and a second visible light image sensor (visible light image sensor 2), and the two fluorescence endoscopes respectively include a first fluorescence image sensor (fluorescence image sensor 1) and a second fluorescence image sensor (fluorescence image sensor 2).

[0086] Among them, each monocular endoscope in the binocular endoscope and a fluorescence endoscope are installed coaxially with the light source, sharing an optical viewing tube, forming a first monocular endoscope module (monocular endoscope module 1) and a second monocular endoscope module (monocular endoscope module 2), and the first monocular endoscope module and the second monocular endoscope module are connected through a binocular endoscope mounting base.

[0087] The normal direction of the first visible light image sensor (visible light image sensor 1) is coaxial with the first optical viewing tube (optical viewing tube 1). The first visible light image sensor (visible light image sensor 1) and the first fluorescence image sensor (fluorescence image sensor 1) are orthogonally arranged, and a first beam splitter (beam splitter 1) is installed at the intersection of the normal lines at the center point. Visible light passes through the first optical viewing tube (optical viewing tube 1) and penetrates the first beam splitter (beam splitter 1) to reach the first visible light image sensor (visible light image sensor 1), and fluorescence reaches the first fluorescence image sensor (fluorescence image sensor 1) through the reflection of the first beam splitter (beam splitter 1).

[0088] The normal direction of the second visible light image sensor (visible light image sensor 2) is coaxial with the second optical viewing tube (optical viewing tube 2). The second visible light image sensor (visible light image sensor 2) and the second fluorescence image sensor (fluorescence image sensor 2) are orthogonally arranged, and a second beam splitter (beam splitter 2) is installed at the intersection of the normal lines at the center point. Visible light passes through the second optical viewing tube (optical viewing tube 2) and penetrates the second beam splitter (beam splitter 2) to reach the second visible light image sensor (visible light image sensor 2), and fluorescence reaches the second fluorescence image sensor (fluorescence image sensor 2) through the reflection of the second beam splitter (beam splitter 2).

[0089] The second installation method is as Figure 3 shown. The device adopts a binocular endoscope - binocular fluorescence independent installation method. The binocular endoscope includes a first visible light image sensor (visible light image sensor 1) and a second visible light image sensor (visible light image sensor 2), and the two fluorescence endoscopes respectively include a first fluorescence image sensor (fluorescence image sensor 1) and a second fluorescence image sensor (fluorescence image sensor 2).

[0090] Among them, the binocular endoscope includes two monocular endoscopes, which are connected by a binocular endoscope mounting base. The normal direction of the first visible light image sensor (visible light image sensor 1) is coaxial with the first optical viewing tube (optical viewing tube 1), and visible light reaches the first visible light image sensor (visible light image sensor 1) through the first optical viewing tube (optical viewing tube 1). The normal direction of the second visible light image sensor (visible light image sensor 2) is coaxial with the second optical viewing tube (optical viewing tube 2), and visible light reaches the second visible light image sensor (visible light image sensor 2) through the second optical viewing tube (optical viewing tube 2).

[0091] Two fluorescence endoscopes are connected by a binocular fluorescence mounting base; the normal direction of the first fluorescence image sensor (fluorescence image sensor 1) is coaxial with the third optical viewing tube (optical viewing tube 3), and fluorescence reaches the first fluorescence image sensor (fluorescence image sensor 1) through the third optical viewing tube (optical viewing tube 3); the normal direction of the second fluorescence image sensor (fluorescence image sensor 2) is coaxial with the fourth optical viewing tube (optical viewing tube 4), and fluorescence reaches the second fluorescence image sensor (fluorescence image sensor 2) through the fourth optical viewing tube (optical viewing tube 4); the binocular endoscope mounting base is connected to the binocular fluorescence mounting base.

[0092] The third installation method is as Figure 4 shown, and the device adopts a binocular endoscope - monocular fluorescence independent installation method. The binocular endoscope includes a first visible light image sensor (visible light image sensor 1) and a second visible light image sensor (visible light image sensor 2), and the fluorescence endoscope includes a fluorescence image sensor.

[0093] Among them, the binocular endoscope includes two monocular endoscopes, which are connected by a binocular endoscope mounting base; the normal direction of the first visible light image sensor (visible light image sensor 1) is coaxial with the first optical viewing tube (optical viewing tube 1), and visible light reaches the first visible light image sensor (visible light image sensor 1) through the first optical viewing tube (optical viewing tube 1); the normal direction of the second visible light image sensor (visible light image sensor 2) is coaxial with the second optical viewing tube (optical viewing tube 2), and visible light reaches the second visible light image sensor (visible light image sensor 2) through the second optical viewing tube (optical viewing tube 2).

[0094] The fluorescence endoscope is mounted on a monocular fluorescence mounting base; the normal direction of the fluorescence image sensor is coaxial with the third optical viewing tube (optical viewing tube 3), and fluorescence reaches the fluorescence image sensor through the third optical viewing tube (optical viewing tube 3); the binocular endoscope mounting base is connected to the monocular fluorescence mounting base.

[0095] In the embodiments of the present invention, three multi - modal image fusion endoscopic surgical target three - dimensional positioning devices with different structures are proposed: binocular endoscope - binocular fluorescence coaxial installation, binocular endoscope - binocular fluorescence independent installation, and binocular endoscope - monocular fluorescence independent installation. The binocular endoscope - binocular fluorescence coaxial installation structure simplifies the fusion method of visible light and fluorescence images by coaxially installing the binocular endoscope and binocular fluorescence and sharing an optical viewing tube; the binocular endoscope - binocular fluorescence independent installation can obtain a more independent and flexible split installation in the face of different surgical scenarios; the binocular endoscope - monocular fluorescence independent installation structure can ensure the efficient capture of fluorescence signals by the monocular fluorescence while ensuring a certain degree of flexibility, ensuring the real - time nature of intraoperative identification.

[0096] An embodiment of the present invention also provides a method for positioning the target of a laparoscopic surgery with multimodal image fusion. The method is based on the device for positioning the target of a laparoscopic surgery with multimodal image fusion provided by the present invention, as Figure 5 shown. The processing flow of this method may include the following steps:

[0097] S1. Camera calibration: Calibrate the internal and external parameters of the binocular camera and the fluorescence camera respectively to obtain the pose transformation matrix between the fluorescence camera and the binocular camera. Among them, the binocular camera includes a binocular laparoscope, and the fluorescence camera includes a fluorescence laparoscope.

[0098] In this step, camera calibration refers to calibrating the relationship between the fluorescence camera and the binocular camera, which specifically includes:

[0099] Calibration of the internal and external parameters of the binocular camera: Use the binocular camera to take multiple pictures of the calibration board, and use the Zhang-Zhengyou calibration method to calibrate the internal parameter matrix and the external parameter matrix of the binocular camera.

[0100] Calibration of the internal and external parameters of the fluorescence camera: Use the fluorescence camera to take multiple pictures of the calibration board, and obtain the internal parameter matrix and the external parameter matrix of the fluorescence camera based on the pixel coordinates of the corner points on the calibration board and the actual spatial position coordinates of the corner points.

[0101] Correct camera distortion: Due to the manufacturing errors of the camera itself, the camera has distortion, so it is necessary to correct the distortion of the images taken by the camera.

[0102] Refer to Figure 6 shown. O W 、O S 、O N are the world coordinate system, the binocular camera coordinate system, and the fluorescence camera coordinate system respectively. O UN is the image coordinate system of the fluorescence camera. O UL 、O UR are the image coordinate systems of the left and right images of the binocular camera respectively; (x W , y W , z W ), (x CS , y CS , z CS ), (x CN , y CN , z CN ) are the three-dimensional coordinates of the fluorescence marker point in the world coordinate system, the binocular camera coordinate system, and the fluorescence camera coordinate system respectively; H S is the transformation matrix of the binocular camera coordinate system in the world coordinate system, H N is the transformation matrix of the fluorescence camera coordinate system in the world coordinate system, and H N-S is the pose transformation matrix of the fluorescence camera coordinate system in the binocular camera coordinate system.

[0103] The specific calibration process is as follows:

[0104] The point coordinates (x W , y W , z W ) in the world coordinate system are transformed into the point coordinates (x C , y C , z C ) in the camera coordinate system through the external parameter matrix of the camera; the camera is a binocular camera or a fluorescence camera, and the transformation formula is as follows:

[0105]

[0106] In the formula, is the rotation matrix, is the translation matrix, and the external parameter matrix is composed of the rotation matrix and the translation matrix which is obtained from the pixel coordinates and the actual space coordinates of the corner points on the calibration board.

[0107] The point coordinates (x C , y C , z C ) in the camera coordinate system are transformed into the pixel coordinates (u, v) on the image coordinate system through the internal parameter matrix, and the transformation formula is as follows:

[0108]

[0109] In the formula, f x and f y in the internal parameter matrix are the focal lengths of the camera in the x and y directions respectively, and u x and u y in the internal parameter matrix are the pixel coordinates of the principal point on the camera image in the x and y directions respectively, and z is the scale factor representing the depth scaling.

[0110] Due to the distortion of the camera, there is a deviation between the ideal pixel projection point and the actual pixel projection point on the camera image. Therefore, it is necessary to correct the ideal projection point through the distortion parameters to eliminate the influence of lens distortion on imaging.

[0111] The camera distortion is corrected through the distortion parameters, and the distortion parameters include the radial distortion parameters k1, k2, k3 and the tangential distortion parameters p1, p2, to obtain the corrected pixel coordinates (u’, v’).

[0112] The radial distortion correction formula is as follows:

[0113] u' = u(1 + k1r 2 + k2r 4 + k3r 6 )

[0114] v' = v(1 + k1r 2 + k2r 4 + k3r 6 )

[0115] The tangential distortion correction formula is as follows:

[0116] u' = 2p1uv + p2(r 2 + 2u 2 )

[0117] v' = 2p2uv + p1(r 2 + 2v 2 )

[0118] Where (u', v') are the corrected pixel coordinates, (u, v) are the pixel coordinates before correction, and r 2 = x 2 + y 2 .

[0119] By the above steps, the internal parameter matrices I S and I N of the binocular camera and the fluorescence camera are obtained respectively, and the external parameter matrices H S and H N of the binocular camera and the fluorescence camera with respect to the same calibration board, as well as the distortion parameters of the binocular camera and the fluorescence camera. Thus, the pose transformation matrix H S-N between the fluorescence camera and the binocular camera is obtained:

[0120] H S-N = H N -1 * H S

[0121] Through the pose transformation matrix H S-N , the conversion relationship between the point coordinates (x CN , y CN , z CN ) of the same fluorescence marker point in the fluorescence camera coordinate system and the point coordinates (x CS , y CS , z CS ) in the binocular camera coordinate system can be obtained:

[0122]

[0123] S2. Fluorescence marker point recognition: Use the U-Net deep learning network model to real-time recognize and track the position of the fluorescence marker point in the fluorescence image, and obtain the two-dimensional image coordinates of the fluorescence marker point.

[0124] The U-Net network is a convolutional neural network architecture for image segmentation, suitable for tasks with small samples, imbalanced data, and the need to retain detailed information. As Figure 7 shown in the architecture of the U-Net network, it can be applied to fields such as tumor segmentation, organ segmentation, and cell segmentation, and has become one of the important algorithms in the field of image segmentation. The network structure of U-Net is U-shaped. Multiple convolutional layers inside the model extract rich multi-scale features from the input image. The image is downsampled through max-pooling operations, and this process is repeated until the image is downsampled by 16 times. The decoding operation is the inverse of downsampling and is used to restore the dimensions. Skip connection operations are performed on the encoded and decoded feature maps of the same size, which can effectively fuse shallow features and deep features, thereby using the information of the encoded feature maps during the decoding process to improve the classification ability of the model. These features are crucial for the image segmentation task.

[0125] Figure 8 is the technical roadmap for identifying and tracking fluorescent marker points using the U-Net network, which specifically includes:

[0126] Dataset collection: Collect an image dataset of fluorescent marker points. Among them, collect a predetermined number and format of fluorescent images, such as 50 single-channel grayscale images with a size of 1920×1080, to form an image dataset.

[0127] Dataset annotation: Use label-me to annotate the collected image dataset, and the annotation target is the fluorescent marker points.

[0128] Label file generation: The annotation information of each picture is stored in a json file. According to the original picture and the generated json file, the target is segmented to obtain the corresponding segmentation label file, and the dataset for training is obtained.

[0129] Dataset division: Divide the dataset into a training set (e.g., 45 pictures) and a validation set (e.g., 5 pictures) according to a predetermined ratio (e.g., 9:1).

[0130] Model training and validation: Use the training set to train the U-Net deep learning network model, and use the validation set to validate the U-Net deep learning network model.

[0131] Among them, the training of the U-Net network model refers to carrying out the training work of the U-Net network with the help of a deep learning framework (such as PyTorch). In terms of data settings, the size of the data input into the model is automatically adjusted to 512×512, ensuring that when the U-Net network model performs 2 to 32 times downsampling, the image size is still an integer. During the model training process, the Adam optimizer is used, the learning rate is set to 1e-4, and the learning rate decay method is cos. The weights are saved and evaluated every 5 epochs. Given that the network model belongs to the field of semantic segmentation, the MIoU (mean intersection over union) score is used as the evaluation metric during model evaluation, so as to quantitatively analyze the experimental results. After each training epoch, some metrics are calculated and recorded, including the loss value and accuracy of the training set, and the loss value and accuracy of the validation set. The MIoU value ranges from 0 to 1, and the closer it is to 1, the higher the average matching degree of the model on all categories or samples, and the better the model performance. The loss value of the training result index is stable at about 0.001, and the MIoU value is stable at about 0.92. The prediction result almost completely coincides with the true label, indicating that the performance of the model meets the training requirements.

[0132] After that, the test set is used to finally evaluate the model and check the performance of the model on unseen data. The trained model is applied to the real-time data of the fluorescence image collected on the fluorescence image acquisition card to verify the performance of the model training. As the target moves, the prediction area also moves accordingly, achieving accurate segmentation of the fluorescence area. While segmenting the target, the coordinates of the center point of the fluorescence-labeled area are tracked in real time for subsequent fluorescence-labeling positioning and tracking work.

[0133] S3. Binocular stereo imaging: Use the region-based local stereo matching method in binocular stereo imaging to process the left and right images of the binocular camera and obtain the point cloud image of the target space.

[0134] Reference Figure 9 As shown, this step specifically includes:

[0135] Epipolar rectification: Perform geometric transformation on the original left and right images taken by the binocular camera to align the corresponding image points on the same horizontal line, improving the efficiency of subsequent feature matching of the two images.

[0136] Stereo matching: It includes matching cost calculation and cost aggregation. By obtaining the disparity d of the target point P in the left and right images P to calculate the depth information Z P . Stereo matching includes a series of steps to calculate the disparity value of each pixel point in the image, and finally obtains the disparity map.

[0137] Match cost calculation refers to calculating the match cost for each pixel point in the left image of a stereo camera within a given disparity range after epipolar rectification, which is used to measure the similarity with the corresponding pixel point in the right image.

[0138] Among them, the sum of absolute differences of grayscale is used to calculate the match cost. For each pixel point (u, v) in the left image, a window of size W×H is selected, which is located around this pixel point (u, v). The corresponding window on the right image is obtained by adding the disparity d to the window on the left image. The grayscale value difference of the corresponding pixel points is calculated using the formula:

[0139]

[0140] In the formula, I L (x + i, y + j) is the grayscale value of the point (x + i, y + j) in the left image, and I R (x + i - d, y + j) is the grayscale value of the point (x + i, y + j) in the right image. d is the disparity, and W and H are the width and height of the window respectively.

[0141] For problems such as noise, reflection, lack of texture information, and unclear object surface features in the image, there are often multiple cases where the match costs are the same, that is, there are multiple matching points in the right image corresponding to the left image, resulting in matching ambiguity and affecting the calculation of the depth of stereo imaging. Cost aggregation aims at the problem that the disparity values are the same or similar in a local area, and calculates the weighted average value or smooth processing of the match cost within the specified neighborhood of each pixel point to reduce the influence of noise and false matches.

[0142] In the embodiments of the present invention, Gaussian filtering is used for cost aggregation, and the weighted average value of the match cost is calculated within the specified neighborhood of each pixel point, that is, the match cost of each pixel point is weighted and averaged with the match costs of its neighboring pixel points. The distribution of the weights follows a Gaussian distribution to reduce noise and false matches.

[0143] The formula of Gaussian filtering is as follows:

[0144]

[0145] In the formula, C′(x, y, d) is the result of cost aggregation for each pixel point; G(x' - x, y′ - y) is the Gaussian weight function; SAD(x′, y′, d) is the grayscale value difference calculated based on the sum of absolute differences of grayscale; σ is the standard deviation of the Gaussian distribution.

[0146] Through cost aggregation, the accuracy of depth calculation can be significantly improved in stereo image matching. Especially in areas with low contrast and lack of texture, it can effectively eliminate the false match problems caused by image defects.

[0147] Depth calculation: After cost aggregation is completed, the disparity value corresponding to the minimum matching cost is selected for each pixel point to obtain the depth information of each pixel point. As Figure 10 shown in the schematic diagram of the depth calculation method based on the triangulation principle, the left and right imaging planes of the binocular camera are parallel to each other. Using the principle of similar triangles, the relational expression can be obtained:

[0148]

[0149] In the formula, |P l P r | is the pixel distance between the pixel points P l and P r on the left and right images of a point P in the target space; Z P is the depth of the point P in the target space to the image plane; b is the length of the baseline between the left and right image planes; f is the focal length of the binocular camera.

[0150] P l and P r The pixel distance |P l P r | is calculated by the following formula:

[0151] |P l P r | = b - (x1 - w / 2) - (w / 2 - x2) = b - (x1 - x2) = b - d P (1.2)

[0152] In the formula, x1 is the pixel coordinate in the x direction of the pixel point P l in the image coordinate system of the left image; x2 is the pixel coordinate in the x direction of the pixel point P r in the image coordinate system of the right image; w is the width of the image; d P is the disparity of the point P.

[0153] Substituting formula (1.2) into formula (1.1) gives:

[0154]

[0155] For a given binocular imaging system, the focal length f and the length b of the baseline are both fixed. Therefore, by determining the disparity d P of a point P in the space on the left and right image planes, the depth Z P of the point P can be calculated.

[0156] Point cloud acquisition: After calculating the depth information of the pixel points, the point cloud image of the target space is obtained using the internal parameter matrix of the binocular camera. For a pixel point image coordinate (u, v) and depth information Z P, the three-dimensional coordinates (x C , y C , z C ) of this point in the camera coordinate system are calculated using the following formula:

[0157]

[0158] z c = Z p

[0159] In the above formula, u x , u y , f x , f y are all internal parameters of the binocular camera, and the specific values of each parameter have been obtained in the camera calibration step. For each pixel point on each binocular camera image, its three-dimensional coordinates (x C , y C , z C ) are calculated, and the three-dimensional point cloud of the image can be obtained.

[0160] S4. Image fusion: Combine the two-dimensional image coordinates of the fluorescence marker points and the point cloud image of the target space to obtain the three-dimensional pose information of the fluorescence marker points.

[0161] For the independent installation method of binocular endoscope - monocular fluorescence, use the U-Net deep learning network model to real-time identify the pixel coordinates (u i , v i ) of the fluorescence marker points on the fluorescence image; use the binocular stereo imaging method to obtain the three-dimensional point cloud (x Si , y Si , z Si ) of the target position; use the pose transformation matrix H S-N between the fluorescence camera and the binocular camera obtained in the calibration stage to transform the original three-dimensional point cloud obtained by binocular stereo imaging into the fluorescence camera coordinate system:

[0162]

[0163] Then transform all the three-dimensional coordinates in the fluorescence camera coordinate system into the image coordinate system of the fluorescence camera:

[0164]

[0165] It can be concluded that for each original three-dimensional point cloud coordinate (x Si , y Si , z Si ), a corresponding pseudo-pixel coordinate (u i ’, v i’); For the pixel coordinates (u i , v i ) of each fluorescent marker point identified by the U-Net network, the pseudo-pixel coordinates (u imin ’, v imin ’) with the minimum distance from this pixel coordinate are calculated using the Euclidean distance. Record the subscript index of this pseudo-pixel coordinate, and then the three-dimensional pose coordinates (x Si , y Si , z Si ) of this fluorescent marker point can be read reversely from the three-dimensional point cloud.

[0166] For the binocular endoscope - binocular fluorescence coaxial installation method, the pixel coordinates of the target fluorescent marker point on the fluorescence image are the same as those on the depth image of the binocular endoscope. The target position corresponding to the fluorescent marker point on the depth image is determined through a one-to-one correspondence relationship, and further the three-dimensional pose coordinates of the fluorescent marker point can be obtained.

[0167] For the binocular endoscope - binocular fluorescence independent installation method, in the camera calibration stage, obtain the coordinate transformation relationship H S between the binocular camera and the visible light calibration board, the coordinate transformation relationship H N between the fluorescence camera and the fluorescence marker board, and the transformation relationship H ns between the fluorescence marker board and the visible light marker board, so as to obtain the coordinate transformation relationship H N-S = H S -1 * H ns * H N ; In the binocular stereo imaging stage, use the binocular stereo imaging method described above to perform three-dimensional reconstruction on the binocular endoscope and the binocular fluorescence endoscope respectively to obtain the three-dimensional point cloud of the visible light image and the three-dimensional point cloud of the fluorescence image; then use the coordinate transformation relationship H N-S to fuse the three-dimensional point cloud of the visible light image and the three-dimensional point cloud of the fluorescence image, so as to obtain the three-dimensional point cloud data containing both visible light information and fluorescence information.

[0168] In the embodiments of the present invention, by introducing a fluorescence camera and combining an image recognition algorithm to accurately detect the image information of the fluorescent marker point, it is ensured that the endoscope system can detect the pose information of the fluorescent marker point in real time. At the same time, the binocular stereo positioning and three-dimensional point cloud reconstruction technologies are used to perform binocular stereo positioning on the left and right images of the binocular visible light camera, and the three-dimensional point cloud of the target area is obtained using the three-dimensional point cloud reconstruction technology, significantly improving the environmental perception ability of the endoscope system.

[0169] Compared with the prior art, by using the U-Net network, the present invention can detect the two-dimensional coordinate information of the fluorescence marker points on the image in real time and accurately, and adopt binocular stereo positioning and three-dimensional point cloud reconstruction technologies, effectively improving the efficiency and accuracy of intraoperative three-dimensional imaging. It can not only obtain the depth information of the target tissue in real time, but also accurately track the three-dimensional position of the fluorescence marker points, helping surgeons and robots to more accurately locate the lesion area and improve the surgical efficiency.

[0170] In an exemplary embodiment, the present invention further provides an electronic device, which includes:

[0171] A processor;

[0172] A memory, on which computer-readable instructions are stored. When the computer-readable instructions are loaded and executed by the processor, the steps of the endoscopic surgery target positioning method for multimodal image fusion as described above are implemented.

[0173] In an exemplary embodiment, the present invention further provides a computer-readable storage medium, in which at least one instruction is stored. The at least one instruction is loaded and executed by a processor to implement the steps of the endoscopic surgery target positioning method for multimodal image fusion as described above. For example, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0174] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or terminal device. Without more limitations, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or terminal device including the element.

[0175] When referring to "an embodiment", "embodiment", "exemplary embodiment", "some embodiments" and the like in the specification, it indicates that the described embodiment may include specific features, structures or characteristics, but not necessarily every embodiment includes the specific feature, structure or characteristic. In addition, when combining an embodiment to describe a specific feature, structure or characteristic, implementing such feature, structure or characteristic in combination with other embodiments (whether explicitly described or not) should be within the knowledge scope of those skilled in the relevant art.

[0176] It should be understood that the term "and / or" in this text is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. Additionally, the character " / " in this text generally represents an "or" relationship between the preceding and following associated objects, but it may also represent an "and / or" relationship, which can be specifically understood by referring to the context before and after.

[0177] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.

[0178] It should be understood that in various embodiments of the present invention, the magnitude of the serial numbers of the above processes does not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0179] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.

[0180] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0181] In addition, in each embodiment of the present invention, the functional units can be integrated in one processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0182] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0183] The present invention covers any alternatives, modifications, equivalent methods, and solutions made within the spirit and scope of the present invention. To enable the public to have a thorough understanding of the present invention, specific details are described in detail in the following preferred embodiments of the present invention. However, those skilled in the art can also fully understand the present invention without the description of these details. Additionally, to avoid unnecessary confusion to the essence of the present invention, well-known methods, processes, procedures, components, and circuits, etc. are not described in detail.

[0184] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A multi-modal image fusion laparoscopic surgery target positioning device, characterized in that: The device comprises: a binocular camera, a fluorescence camera, a binocular image acquisition card, a fluorescence image acquisition card, a computer, and a display; The binocular camera includes a binocular cavity mirror, and uses a dual image sensor to collect binocular visible light images of the target space from two different left and right viewing angles; the binocular image acquisition card performs digital processing and format conversion on the collected binocular visible light images, and transmits them to the computer; The fluorescence camera includes a fluorescence cavity mirror, and collects the fluorescence image of the target space through an image sensor; the fluorescence image acquisition card performs digital processing and format conversion on the collected fluorescence image, and transmits it to the computer; After receiving the binocular visible light image and the fluorescent image, the computer performs fusion processing to obtain the point cloud image of the target space and the depth information of the fluorescent marking points, and displays them through the display.

2. The multimodal image fusion laparoscopic surgery target positioning device according to claim 1, characterized in that: The device adopts a binocular cavity mirror-binocular fluorescence coaxial installation mode; the binocular cavity mirror includes a first visible light image sensor and a second visible light image sensor, and the two fluorescence cavity mirrors respectively include a first fluorescence image sensor and a second fluorescence image sensor; Each monocular laparoscope and a fluorescent laparoscope in the binocular laparoscope are coaxially installed with a light source and share an optical sight tube to form a first monocular laparoscope module and a second monocular laparoscope module, and the first monocular laparoscope module and the second monocular laparoscope module are connected through a binocular laparoscope mounting base; The normal direction of the first visible light image sensor is coaxial with the first optical sight tube, the first visible light image sensor and the first fluorescent image sensor are arranged orthogonally, and a first spectroscope is installed at the intersection of the normal lines of the center point; visible light passes through the first optical sight tube and penetrates the first spectroscope to reach the first visible light image sensor, and fluorescent light is reflected by the first spectroscope to reach the first fluorescent image sensor; The normal direction of the second visible light image sensor is coaxial with the second optical sight tube, the second visible light image sensor and the second fluorescence image sensor are arranged orthogonally, and a second beam splitter is installed at the intersection of the normals of the center point; visible light passes through the second optical sight tube and penetrates the second beam splitter to reach the second visible light image sensor, and fluorescence is reflected by the second beam splitter to reach the second fluorescence image sensor.

3. The multimodal image fusion laparoscopic surgery target positioning device according to claim 1, characterized in that: The device adopts a binocular cavity mirror-binocular fluorescence independent installation mode; the binocular cavity mirror includes a first visible light image sensor and a second visible light image sensor, and the two fluorescence cavity mirrors include a first fluorescence image sensor and a second fluorescence image sensor respectively; The binocular laparoscope includes two monocular laparoscopes connected by a binocular laparoscope mounting base; the normal direction of the first visible light image sensor is coaxial with the first optical sight tube, and the visible light reaches the first visible light image sensor through the first optical sight tube; the normal direction of the second visible light image sensor is coaxial with the second optical sight tube, and the visible light reaches the second visible light image sensor through the second optical sight tube; The two fluorescence cavity mirrors are connected through a binocular fluorescence mounting base; the normal direction of the first fluorescence image sensor is coaxial with the third optical sight tube, and the fluorescence reaches the first fluorescence image sensor through the third optical sight tube; the normal direction of the second fluorescence image sensor is coaxial with the fourth optical sight tube, and the fluorescence reaches the second fluorescence image sensor through the fourth optical sight tube; the binocular cavity mirror mounting base is connected to the binocular fluorescence mounting base.

4. The multimodal image fusion laparoscopic surgery target positioning device according to claim 1, characterized in that: The device adopts a binocular cavity mirror-monocular fluorescence independent installation mode; the binocular cavity mirror includes a first visible light image sensor and a second visible light image sensor, and the fluorescence cavity mirror includes a fluorescence image sensor; The binocular laparoscope includes two monocular laparoscopes connected by a binocular laparoscope mounting base; the normal direction of the first visible light image sensor is coaxial with the first optical sight tube, and the visible light reaches the first visible light image sensor through the first optical sight tube; the normal direction of the second visible light image sensor is coaxial with the second optical sight tube, and the visible light reaches the second visible light image sensor through the second optical sight tube; The fluorescent cavity mirror is installed on the monocular fluorescent mounting base; the normal direction of the fluorescent image sensor is coaxial with the third optical sight tube, and the fluorescence reaches the fluorescent image sensor through the third optical sight tube; the binocular cavity mirror mounting base is connected to the monocular fluorescent mounting base.

5. A multimodal image fusion laparoscopic surgery target positioning method, the method is based on the device according to any one of claims 1 to 4, characterized in that: The method comprises the following steps: S1, camera calibration: calibrate the internal and external parameters of the binocular camera and the fluorescence camera respectively, and obtain the pose conversion matrix between the fluorescence camera and the binocular camera; Wherein, the binocular camera comprises a binocular cavity mirror, and the fluorescence camera comprises a fluorescence cavity mirror; S2, fluorescent marker point recognition: use the U-Net network model to recognize and track the position of the fluorescent marker point in the fluorescent image in real time, and obtain the two-dimensional image coordinates of the fluorescent marker point; S3, binocular stereo imaging: use the region-based local stereo matching method in binocular stereo imaging to process the left image and the right image of the binocular camera to obtain the point cloud image of the target space; S4, image fusion: Combine the two-dimensional image coordinates of the fluorescent marker point and the point cloud image of the target space to obtain the three-dimensional pose information of the fluorescent marker point.

6. The multimodal image fusion laparoscopic surgery target positioning method according to claim 5, characterized in that: The step S1 specifically includes: Calibration of the internal and external parameters of the binocular camera: Use the binocular camera to take multiple pictures of the calibration plate, and use the Zhang Zhengyou calibration method to calibrate the internal and external parameter matrices of the binocular camera; Internal and external parameter calibration of fluorescence camera: Use the fluorescence camera to take multiple pictures of the calibration plate, and obtain the internal and external parameter matrix of the fluorescence camera based on the pixel coordinates of the corner points on the calibration plate and the actual spatial position coordinates of the corner points; Correct camera distortion; The specific calibration process is as follows: The point coordinates in the world coordinate system (x W ,y W ,z W ), and then transformed to the point coordinates (x C ,y C ,z C ); the camera is a binocular camera or a fluorescent camera; The point coordinates in the camera coordinate system (x C ,y C ,z C ), transformed to the pixel coordinates (u, v) in the image coordinate system through the intrinsic parameter matrix; Correcting the camera distortion by using distortion parameters to obtain corrected pixel coordinates (u', v'); the distortion parameters include radial distortion parameters and tangential distortion parameters; Through the above steps, the internal parameter matrix I of the binocular camera and the fluorescence camera is obtained respectively. S and I N , the external parameter matrix H of the binocular camera and the fluorescence camera relative to the same calibration plate S and H N , and the distortion parameters of the binocular camera and the fluorescence camera; thus, the pose transformation matrix H between the fluorescence camera and the binocular camera is obtained S-N : H S-N =H N -1 *H S Through the pose transformation matrix H S-N Get the point coordinates (x CN ,y CN ,z CN ) and the point coordinates (x CS ,y CS ,z CS ) between them.

7. The multimodal image fusion laparoscopic surgery target positioning method according to claim 5, characterized in that: The step S2 specifically includes: Dataset collection: Collect image datasets of fluorescent markers; Dataset annotation: Use label-me to annotate the collected image dataset; Label file generation: The annotation information of each image is stored in a json file. According to the original image and the generated json file, the target is segmented to obtain the corresponding segmentation label file and the data set for training. Dataset division: Divide the dataset into training set and validation set according to a predetermined ratio; Model training and verification: Use the training set to train the U-Net network model, and use the verification set to verify the U-Net network model.

8. The multimodal image fusion laparoscopic surgery target positioning method according to claim 5, characterized in that: The step S3 specifically includes: Epipolar correction: geometrically transform the original left and right images taken by the binocular camera so that the corresponding image points are aligned on the same horizontal line; Stereo matching: including matching cost calculation and cost aggregation, and calculating depth information by obtaining the disparity of the target point in the left image and the right image; Matching cost calculation refers to calculating the matching cost of each pixel in the left image of the binocular camera within a given disparity range after completing epipolar correction, which is used to measure the similarity between the corresponding pixel in the right image; the matching cost is calculated using the sum of the grayscale absolute value difference; Cost aggregation refers to the use of Gaussian filtering for cost aggregation, which calculates the weighted average of the matching cost within the specified neighborhood of each pixel, that is, the matching cost of each pixel and the matching cost of its neighboring pixels are weighted averaged according to the weight. The distribution of weights follows Gaussian distribution to reduce noise and mismatches. Depth calculation: After completing the cost aggregation, select the disparity value corresponding to the minimum matching cost for each pixel to obtain the depth information of each pixel; Point cloud acquisition: After calculating the depth information of the pixel points, the point cloud image of the target space is obtained using the intrinsic parameter matrix of the binocular camera.

9. The multimodal image fusion laparoscopic surgery target positioning method according to claim 5, characterized in that: The step S4 specifically includes: For the binocular laparoscope-monocular fluorescence independent installation mode, the U-Net deep learning network model is used to recognize the pixel coordinates of the fluorescent marker points on the fluorescence image in real time (u i ,v i );Use binocular stereo imaging method to obtain the three-dimensional point cloud of the target position (x Si ,y Si ,z Si );Use the pose transformation matrix H between the fluorescence camera and the binocular camera obtained in the calibration phase S-N , convert the original 3D point cloud obtained by binocular stereo imaging into the fluorescent camera coordinate system: Then transform all three-dimensional coordinates in the fluorescence camera coordinate system to the image coordinate system of the fluorescence camera: It can be concluded that for each original 3D point cloud coordinate (x Si ,y Si ,z Si ), we can find the corresponding pseudo pixel coordinates (u i ',v i '); For each pixel coordinate of the fluorescent marker identified by the U-Net network (u i ,v i ), use the Euclidean distance to calculate the pseudo pixel coordinates with the shortest distance from the pixel coordinates (u imin ',v imin '), record the subscript index of the pseudo pixel coordinates, and then read the 3D pose coordinates (x Si ,y Si ,z Si ); For the binocular laparoscope-binocular fluorescence coaxial installation method, the pixel coordinates of the target fluorescent marker point on the fluorescent image are consistent with the pixel coordinates on the depth image of the binocular laparoscope. The target position corresponding to the fluorescent marker point on the depth image is determined through a one-to-one correspondence relationship, and the three-dimensional pose coordinates of the fluorescent marker point can be further obtained; For the binocular cavity mirror-binocular fluorescence independent installation mode, in the camera calibration stage, the coordinate transformation relationship H between the binocular camera and the visible light calibration board is obtained. S , the coordinate transformation relationship between the fluorescence camera and the fluorescence marker plate H N , and the conversion relationship between the fluorescent label plate and the visible light label plate H ns , thus obtaining the coordinate transformation relationship H between the binocular camera and the fluorescence camera N-S =H S -1 *H ns *H N ; In the binocular stereo imaging stage, the binocular cavity mirror and the binocular fluorescence cavity mirror are reconstructed in three dimensions respectively using the binocular stereo imaging method to obtain a three-dimensional point cloud of the visible light image and a three-dimensional point cloud of the fluorescence image; and then the coordinate transformation relationship H is used N-S , the three-dimensional point cloud of the visible light image is fused with the three-dimensional point cloud of the fluorescence image, so as to obtain three-dimensional point cloud data containing both visible light information and fluorescence information.

10. An electronic device, characterized in that: The electronic device comprises: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are loaded and executed by the processor, the method according to any one of claims 5 to 9 is implemented.

Citation Information

Cited By

  • Rigid-flexible coupling force-position hybrid sensing method, device, equipment, medium and product

    CN121089964A

  • Fluorescence and visible light image fusion imaging method for intraoperative environment

    CN122200268A