Monocular endoscope stereo imaging system and three-dimensional measurement method

By using an optical prism to generate two types of image information with horizontal parallax in a monocular endoscope stereo imaging system, the problems of small field of view of traditional monocular endoscopes and large hardware cost of binocular endoscopes are solved, and three-dimensional measurement and image quality improvement are achieved.

CN116091521BActive Publication Date: 2026-03-31TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-29
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Traditional monocular endoscopes have a small field of view and cannot provide three-dimensional perception information of depth and direction. Binocular endoscope systems have large hardware size and high cost, and existing single-camera stereo imaging technology has the problem of large hardware size.

Method used

A monocular endoscope stereo imaging system is adopted, which uses an optical prism to generate two types of image information with horizontal parallax through the reflection of light. Based on these image information, three-dimensional measurement information is calculated. The system is equipped with an optical prism to acquire two types of image information by using its reflection ability, which reduces hardware size and lowers costs.

Benefits of technology

In minimally invasive surgery, it provides doctors with three-dimensional perception information, reduces hardware size, lowers costs, improves image quality, and enables three-dimensional measurement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116091521B_ABST
    Figure CN116091521B_ABST
Patent Text Reader

Abstract

The application provides a monocular endoscope stereo imaging system and a three-dimensional measurement method, and relates to the technical field of endoscope imaging. The monocular endoscope stereo imaging system comprises a monocular endoscope body and an optical prism body, the monocular endoscope body is internally provided with an imaging unit and a lens unit, first incident light and second incident light correspondingly form first image information and second image information on the imaging unit after being reflected by the optical prism body, the first image information and the second image information have horizontal parallax, and the first image information and the second image information after stereo correction have no vertical parallax. The monocular endoscope stereo imaging system provided by the application can realize stereo imaging based on a single camera structure, overcome the defects of the prior art, reduce the hardware size, reduce the development cost, and improve the image quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of endoscopic imaging technology, and in particular to a monocular endoscopic stereoscopic imaging system and a three-dimensional measurement method. Background Technology

[0002] Compared to traditional surgery, minimally invasive surgery offers advantages such as less trauma and shorter patient recovery time. However, in minimally invasive surgery, traditional monocular endoscopes have a limited field of view and cannot provide surgeons with three-dimensional perception information such as depth and orientation. From a hardware design perspective, existing technologies that can provide stereoscopic image information to endoscopes mainly include binocular endoscope systems and single-camera stereoscopic imaging systems.

[0003] Binocular endoscope systems typically include two side-by-side entrance lenses, two optical transmission systems, two image sensors, and two control units. They may also require two prisms and two adapter optics, resulting in a generally large hardware size. Furthermore, ensuring the synchronization of images acquired by the two cameras necessitates the development of costly two-camera synchronization technology. Single-camera stereo imaging technology, on the other hand, utilizes specific optical structures and a single camera to acquire stereo image information. Existing single-camera stereo imaging technologies mainly include prism-based methods, diffraction grating-based methods, plane mirror-based methods, and methods that utilize the red and blue channels of the original image to achieve full-resolution stereo imaging. Specifically in the field of endoscopy, considering hardware size limitations, the most commonly used method is the prism-based method. For example, using the refractive and reflective capabilities of two prisms placed between the lens and the imaging unit to achieve stereo imaging in a monocular endoscope. However, utilizing the refractive capabilities of prisms also leads to a large hardware size for the imaging system. Summary of the Invention

[0004] This invention aims to at least solve one of the technical problems existing in the background art. To this end, this invention provides a monocular endoscope stereo imaging system and a three-dimensional measurement method, which can realize stereo imaging based on a single-camera structure, reduce hardware size, lower development costs, and improve image quality. The specific technical solution is as follows:

[0005] A first aspect of the present invention provides a monocular endoscope stereo imaging system, comprising:

[0006] A monocular endoscope body, wherein an imaging unit and a lens unit are disposed within the monocular endoscope body;

[0007] An optical prism body includes a first light-transmitting surface, a second light-transmitting surface, and a reflective inclined surface. The first light-transmitting surface is adapted to receive a first incident light ray and a second incident light ray of the imaging target point, and the second light-transmitting surface faces the lens unit.

[0008] The lens unit is disposed between the imaging unit and the second light-transmitting surface;

[0009] After the first incident light ray passes through the first light-transmitting surface, the reflective inclined surface and the second light-transmitting surface in sequence, it passes through the lens unit and reaches the imaging unit, forming first image information on the imaging unit, and forming a first virtual camera on the extension line of the first incident light ray before entering the optical prism body;

[0010] After the second incident light ray passes through the first light-transmitting surface, the second light-transmitting surface and the reflective inclined surface in sequence, it passes through the second light-transmitting surface again, passes through the lens unit and reaches the imaging unit, forming second image information on the imaging unit, and forming a second virtual camera on the extension line of the second incident light ray before entering the optical prism body;

[0011] The first image information and the second image information have horizontal parallax, and the first image information and the second image information after stereo correction have no vertical parallax.

[0012] According to a first aspect embodiment of the present invention, a monocular endoscope stereo imaging system is provided. An optical prism body is set up, and the reflectivity of the optical prism body to reflect light is used to acquire two types of image information with horizontal parallax relationship. Based on these two types of image information, three-dimensional measurement information is further calculated. Specifically, a first incident ray emitted from the imaging target point enters the interior of the optical prism body after passing through the first light-transmitting surface. The first incident ray undergoes a first refraction at the first light-transmitting surface. Inside the optical prism body, the first incident ray reaches the reflective inclined surface along the path after the first refraction and is reflected. Then, it reaches the second light-transmitting surface along the reflected path. After a second refraction at the second light-transmitting surface, it exits the optical prism body and finally reaches the imaging unit after passing through the lens unit. The imaging unit senses and generates first image information based on the first incident ray and forms a first virtual camera on the extension line of the first incident ray before entering the optical prism body. The second incident ray emitted from the imaging target point... The incident light ray enters the optical prism body after passing through the first transparent surface. The second incident light ray undergoes a first refraction at the first transparent surface. Inside the optical prism body, the second incident light ray follows the path after the first refraction to the second transparent surface and undergoes total internal reflection. Then, the second incident light ray follows the path of total internal reflection to the reflective inclined surface and undergoes reflection. It then follows the reflected path to the second transparent surface, undergoes a second refraction at the second transparent surface, and exits the optical prism body. Finally, it passes through the lens unit to reach the imaging unit. The imaging unit senses and generates second image information based on the second incident light ray, and forms a second virtual camera on the extension line of the second incident light ray before it enters the optical prism body. In this way, two types of image information corresponding to a single imaging target point are simultaneously generated on the same imaging unit. The first and second image information are coplanar and have a certain horizontal parallax. Based on these two types of image information, three-dimensional measurement information can be further calculated. Furthermore, the monocular endoscopic stereo imaging system can capture all points of the imaging target scene within the current field of view. Thus, on the same imaging unit, first image information and second image information corresponding to all points of the imaging target scene within the current field of view are generated. Based on these two types of image information, three-dimensional measurement information can be further calculated. Applying the monocular endoscopic stereo imaging system provided by this invention, a single optical prism body converges the light from the imaging target scene onto a single imaging unit, generating first image information and second image information. Based on the first and second image information, three-dimensional measurement information can be further calculated, thereby providing doctors with three-dimensional perception information in minimally invasive surgery. The monocular endoscopic stereo imaging system eliminates the need for developing costly two-camera synchronization technology, reducing hardware size, lowering costs, and improving image quality.

[0013] This explanation uses the transmission path of light emitted from a single imaging target point as an example. The first and second incident rays represent different types of light rays that successfully reach the imaging unit, with each type containing multiple rays. The transmission paths of the first and second incident rays within the optical prism are determined by the optical properties of the prism. The total internal reflection of the second incident ray upon its first arrival at the second transparent surface is determined by the relative positions of the monocular endoscope and the optical prism in the monocular endoscope stereoscopic imaging system.

[0014] According to one embodiment of the present invention, the monocular endoscope body further includes an image processor, the image processor being connected to the imaging unit via a first electrical signal transmission line;

[0015] The optical center of the lens unit, the photosensitive center point of the imaging unit, and the image processor are connected to form a first connecting line, and the first electrical signal transmission line is located at the part of the first connecting line where the photosensitive center point of the imaging unit and the image processor are connected.

[0016] According to one embodiment of the present invention, the monocular endoscope body further includes an image processor, the image processor being connected to the imaging unit via a second electrical signal transmission line;

[0017] A second connecting line is formed between the optical center of the lens unit and the photosensitive center point of the imaging unit, and the second electrical signal transmission line is located on a straight line perpendicular to the second connecting line.

[0018] According to one embodiment of the present invention, the first light-transmitting surface and the second light-transmitting surface are perpendicular to each other, the angle between the reflective inclined surface and the first light-transmitting surface is 45 degrees, and the angle between the reflective inclined surface and the second light-transmitting surface is 45 degrees.

[0019] According to one embodiment of the present invention, a reflective film structure is provided on the reflective inclined surface.

[0020] A second aspect of the present invention provides a three-dimensional measurement method based on a monocular endoscope stereo imaging system as described in any of the above embodiments, comprising:

[0021] The monocular endoscope stereo imaging system is used to acquire image sequences, and the image sequences are preprocessed to generate stereo image sequence pairs.

[0022] Based on the stereo image sequence, the surgical instruments are three-dimensionally located, and the three-dimensional coordinates of multiple surgical instruments in the first virtual camera coordinate system are obtained in real time.

[0023] Based on the stereo image pairs, a 3D scene reconstruction is performed, and the 3D point cloud data of all pixels in the common field of view of the first virtual camera and the second virtual camera are acquired in real time.

[0024] According to an embodiment of the present invention, the step of acquiring an image sequence using the monocular endoscope stereo imaging system, preprocessing the image sequence, and generating a stereo image sequence pair includes:

[0025] The image sequence is segmented to generate a first image information sequence and a second image information sequence;

[0026] Mirror and flip the first image information sequence;

[0027] The first image information sequence after being mirrored and flipped is cropped so that the first image information sequence and the second image information sequence have the same size;

[0028] The monocular endoscope stereo imaging system is used to acquire checkerboard images. Based on the checkerboard images, the first virtual camera and the second virtual camera are calibrated. The calibration parameters are then used to perform distortion correction and stereo correction on the first image information sequence and the second image information sequence.

[0029] According to an embodiment of the present invention, the step of performing three-dimensional positioning of surgical instruments based on the stereo image sequence pair and acquiring the three-dimensional coordinates of multiple surgical instruments in the first virtual camera coordinate system in real time includes:

[0030] The stereoscopic image sequence pairs are respectively input into the surgical instrument front end region localization network and the surgical instrument front end region detection network, wherein the surgical instrument front end region localization network generates localization region information of each surgical instrument front end in the first image information sequence and the second image information sequence, and the surgical instrument front end region detection network generates detection boxes of each surgical instrument front end region and the category of each surgical instrument in the first image information sequence and the second image information sequence.

[0031] The positioning area information of the front end of each surgical instrument is corrected by using the detection box of the front end area of ​​each surgical instrument to obtain more accurate positioning area information of the front end of each surgical instrument.

[0032] Based on the precise positioning area information of each surgical instrument's front end, the center coordinates of the two-dimensional positioning area of ​​each surgical instrument in the first image information sequence and the second image information sequence are obtained;

[0033] By combining the surgical instrument category and coordinate difference matching method, two-dimensional coordinate point matching pairs of the front ends of each surgical instrument in the first image information sequence and the second image information sequence are obtained;

[0034] Using the two-dimensional coordinate point matching pairs of the front ends of each surgical instrument, and the calibration parameters of the first virtual camera and the second virtual camera, the real-time three-dimensional coordinates of each surgical instrument in the coordinate system of the first virtual camera are calculated by the parallax method.

[0035] According to one embodiment of the present invention, the surgical instrument tip region positioning network includes:

[0036] The surgical instrument front-end region localization network is designed as a pseudo-twin network structure so that the surgical instrument front-end region localization network has the ability to process the first image information sequence and the second image information sequence simultaneously;

[0037] The pseudo-twin network structure consists of two network branches with the same structure but without sharing weights. Each network branch includes an encoder-decoder branch and a full-resolution feature map generator branch to obtain more accurate positioning area information for each surgical instrument.

[0038] According to one embodiment of the present invention, the coordinate difference matching method includes:

[0039] The sum of the differences in the x-coordinate and y-coordinate of the center coordinates of the two-dimensional positioning areas of each surgical instrument in the first image information sequence and the second image information sequence is calculated to obtain an accurate matching pair of two-dimensional coordinate points of the front end of each surgical instrument in the first image information sequence and the second image information sequence.

[0040] According to an embodiment of the present invention, the step of performing scene 3D reconstruction based on the stereo image pair and acquiring 3D point cloud data of all pixels in the common field of view of the first virtual camera and the second virtual camera in real time includes:

[0041] The disparity map of the stereo image pair is obtained using a binocular disparity estimation algorithm;

[0042] Based on the disparity map and the calibration parameters of the first and second virtual cameras, three-dimensional coordinate point cloud data is obtained, and the first image information is used to perform texture coverage on the corresponding three-dimensional coordinate point cloud data.

[0043] The three-dimensional measurement method provided by the second aspect of the present invention is implemented based on the monocular endoscope stereo imaging system of the first aspect of the present invention. Specifically, the monocular endoscope stereo imaging system is provided with an optical prism body, and the reflectivity of the optical prism body to reflect light is used to acquire two types of image information with horizontal parallax relationship, and three-dimensional measurement information is further calculated based on the two types of image information. The process of acquiring two types of image information with horizontal parallax relationship is as follows: the first incident light emitted from the imaging target point enters the interior of the optical prism body after passing through the first light-transmitting surface of the optical prism body. The first incident light undergoes a first refraction at the first light-transmitting surface. Inside the optical prism body, the first incident light reaches the reflective inclined surface and is reflected along the path after the first refraction, and then reaches the second light-transmitting surface along the reflected path. After a second refraction at the second light-transmitting surface, it exits the optical prism body, and then passes through the lens unit and finally reaches the imaging unit. The imaging unit senses and generates the first image information based on the first incident light, and forms a first virtual camera on the extension line of the first incident light before entering the optical prism body. A second incident ray emitted from the target point enters the optical prism body after passing through the first transparent surface. At the first transparent surface, the ray undergoes a first refraction. Inside the prism, the ray travels along the path of the first refraction to the second transparent surface and undergoes total internal reflection. It then travels along the path of total internal reflection to the reflective inclined surface and is reflected again. Following this reflection, the ray travels back to the second transparent surface, undergoes a second refraction, and exits the prism body. Finally, it passes through the lens unit and reaches the imaging unit. The imaging unit senses and generates second image information based on the second incident ray, forming a second virtual camera along the extension of the second incident ray before it enters the prism body. Thus, two types of image information corresponding to a single imaging target point are simultaneously generated on the same imaging unit. The first and second image information are coplanar and exhibit a certain horizontal parallax. Based on these two types of image information, three-dimensional measurement information can be further calculated. Furthermore, the monocular endoscope stereo imaging system can capture all points of the imaging target scene within the current field of view. In this way, first image information and second image information corresponding to all points of the imaging target scene within the current field of view are generated on the same imaging unit.Furthermore, after acquiring images of the imaging target scene at different times within the current field of view, the image sequence is divided into a first image information sequence and a second image information sequence. Since the light corresponding to the first image information sequence and the second image information sequence undergoes different numbers of reflections inside the optical prism body—that is, the light corresponding to the first image information sequence undergoes one reflection inside the optical prism body, while the light corresponding to the second image information sequence undergoes two reflections inside the optical prism body—there is a mirror flip relationship between the first image information sequence and the second image information sequence. Therefore, it is necessary to mirror flip the first image information sequence to ensure the consistency of the orientation of the first image information sequence and the second image information sequence. Then, a checkerboard image is acquired using a monocular endoscope stereo imaging system, and the acquired checkerboard image is used to calibrate the first virtual camera and the second virtual camera. Then, the calibration parameters are used to perform distortion correction and stereo correction on the first image information sequence and the second image information sequence to obtain a stereo image sequence pair. Based on the stereo image sequence pair, the three-dimensional measurement method in this invention can realize the three-dimensional positioning of surgical instruments and the three-dimensional reconstruction of the endoscopic scene. Attached Figure Description

[0044] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0045] Figure 1 This is a schematic diagram of the structure of the monocular endoscope stereo imaging system provided in the embodiment of the present invention. Figure 1 ;

[0046] Figure 2 This is a schematic diagram of the structure of the monocular endoscope stereo imaging system provided in the embodiment of the present invention. Figure 2 ;

[0047] Figure 3 This is a flowchart illustrating the three-dimensional measurement method provided in an embodiment of the present invention;

[0048] Figure 4 This is a schematic diagram of the light transmission path at the imaging boundary of the monocular endoscopic stereoscopic imaging system provided in this embodiment of the invention. Figure 1 ;

[0049] Figure 5 This is a schematic diagram of the light transmission path at the imaging boundary of monocular endoscopic stereoscopic imaging provided in this embodiment of the invention. Figure 2 ;

[0050] Figure 6This is a flowchart illustrating the three-dimensional positioning algorithm for surgical instruments provided in an embodiment of the present invention.

[0051] Figure label:

[0052] 10. Imaging unit;

[0053] 20. Optical prism body; 201. First light-transmitting surface; 202. Second light-transmitting surface; 203. Reflective slope;

[0054] 30. Imaging target point; 301. First incident ray; 302. Second incident ray;

[0055] 60. Lens unit;

[0056] 70. First electrical signal transmission line;

[0057] 80. Second electrical signal transmission line;

[0058] 91. First virtual camera; 92. Second virtual camera; 910. First image information; 920. Second image information. Detailed Implementation

[0059] The embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are for illustrative purposes only and should not be construed as limiting the scope of the invention.

[0060] In the description of the embodiments of the present invention, it should be noted that the terms "center," "longitudinal," "lateral," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing the embodiments of the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the embodiments of the present invention. In addition, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0061] In the description of the embodiments of the present invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "connected" and "linked" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium. Those skilled in the art can understand the specific meaning of the above terms in the embodiments of the present invention based on the specific circumstances.

[0062] In embodiments of the present invention, unless otherwise explicitly specified and limited, "above" or "below" the second feature can mean that the first feature is in direct contact with the second feature, or that the first feature is in indirect contact with the second feature through an intermediate medium. Furthermore, "above," "on top of," and "over" the second feature can mean that the first feature is directly above or diagonally above the second feature, or simply that the first feature is at a higher horizontal level than the second feature. "Below," "below," and "under" the second feature can mean that the first feature is directly below or diagonally below the second feature, or simply that the first feature is at a lower horizontal level than the second feature.

[0063] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0064] like Figure 1As shown, a first aspect of the present invention provides a monocular endoscope stereoscopic imaging system, including a monocular endoscope body and an optical prism body 20. An imaging unit 10 and a lens unit 60 are disposed within the monocular endoscope body. The optical prism body 20 includes a first light-transmitting surface 201, a second light-transmitting surface 202, and a reflective inclined surface 203. The first light-transmitting surface 201 is adapted to receive a first incident light ray 301 and a second incident light ray 302 from an imaging target point 30. The second light-transmitting surface 202 faces the lens unit 60. The lens unit 60 is disposed between the imaging unit 10 and the second light-transmitting surface 202. The first incident light ray 301 passes sequentially through the first light-transmitting surface 201, the reflective inclined surface 203, and the second light-transmitting surface 202, and then through the lens unit 60 to reach the imaging unit 10. A first image information 910 is formed on the imaging unit 10, and a first virtual camera 91 is formed on the extension line of the first incident ray 301 before entering the optical prism body 20; the second incident ray 302 passes through the first light-transmitting surface 201, the second light-transmitting surface 202 and the reflective inclined surface 203 in sequence, then passes through the second light-transmitting surface 202 again, and then passes through the lens unit 60 to reach the imaging unit 10, where a second image information 920 is formed on the imaging unit 10, and a second virtual camera 92 is formed on the extension line of the second incident ray 302 before entering the optical prism body 20; wherein, the first image information 910 and the second image information 920 have horizontal parallax, and the first image information 910 and the second image information 920 after stereo correction have no vertical parallax.

[0065] Imaging unit 10 refers to a component capable of converting optical signals into electronic signals and performing analog-to-digital signal conversion. For example, it may include charge-coupled devices (CCDs) and complementary metal-oxide-semiconductor (CMOS) devices. Horizontal parallax, also called left-right parallax, lateral parallax, or X-parallax, refers to the difference in the abscissa values ​​of corresponding image points in a stereo image pair, similar to the parallax caused by the horizontal distance between the two eyes when observing an object. For ease of understanding, in this invention, the horizontal parallax relationship between the first image information 910 and the second image information 920 can be defined as follows: the first image information 910 and the second image information 920 are arranged side-by-side, with the first image information 910 located to the left of the second image information 920; there is a certain horizontal parallax between the first image information 910 and the second image information 920; and after stereo correction, there is no vertical parallax between the first image information 910 and the second image information 920.

[0066] According to the first aspect of the present invention, the monocular endoscope stereo imaging system provides an optical prism body 20 and uses the light reflection capability of the optical prism body 20 to acquire two types of image information with horizontal parallax relationship, and further calculates three-dimensional measurement information based on the two types of image information. Specifically, the first incident light ray 301 emitted from the imaging target point 30 enters the interior of the optical prism body 20 after passing through the first light-transmitting surface 201 of the optical prism body 20. The first incident light ray 301 undergoes a first refraction at the first light-transmitting surface 201. Inside the optical prism body 20, the first incident light ray 301 reaches the reflective inclined surface 203 along the path after the first refraction and is reflected. Then, it reaches the second light-transmitting surface 202 along the reflected path, undergoes a second refraction at the second light-transmitting surface 202, and exits the optical prism body 20. It then passes through the lens unit 60 and finally reaches the imaging unit 10. The imaging unit 10 senses and generates first image information 910 based on the first incident light ray 301, and forms a first virtual camera 91 on the extension line of the first incident light ray 301 before entering the optical prism body 20. The second incident light emitted from the imaging target point 30... After passing through the first light-transmitting surface 201 of the optical prism body 20, the second incident ray 302 enters the interior of the optical prism body 20. The second incident ray 302 undergoes a first refraction at the first light-transmitting surface 201. Inside the optical prism body 20, the second incident ray 302 reaches the second light-transmitting surface 202 along the path after the first refraction and undergoes total internal reflection. Then, the second incident ray 302 reaches the reflective inclined surface 203 along the path of total internal reflection and undergoes reflection. Then, it reaches the second light-transmitting surface 202 again along the path of reflection. After undergoing a second refraction at the second light-transmitting surface 202, it exits the optical prism body 20 and finally reaches the imaging unit 10 through the lens unit 60. The imaging unit 10 senses and generates second image information 920 based on the second incident ray 302, and forms a second virtual camera 92 on the extension line of the second incident ray 302 before entering the optical prism body 20. In this way, two types of image information corresponding to a single imaging target point are simultaneously generated on the same imaging unit 10. The first image information 910 and the second image information 920 are coplanar and exhibit a certain horizontal parallax. Based on these two types of image information, three-dimensional measurement information can be further calculated. Furthermore, the monocular endoscope stereo imaging system can capture all points of the imaging target scene within the current field of view. Thus, on the same imaging unit 10, first image information 910 and second image information 920 corresponding to all points of the imaging target scene within the current field of view are generated. Based on these two types of image information, three-dimensional measurement information can be further calculated.The monocular endoscopic stereo imaging system provided by this invention uses a single optical prism body 20 to converge the light of the imaging target scene onto a single imaging unit 10, generating first image information 910 and second image information 920. Based on the first image information 910 and the second image information 920, three-dimensional measurement information can be further calculated, thereby providing doctors with three-dimensional perception information in minimally invasive surgery. The monocular endoscopic stereo imaging system does not require the development of costly two-camera synchronization technology, reducing hardware size, lowering costs, and improving image quality.

[0067] This explanation uses the transmission path of light emitted from a single imaging target point 30 as an example. The first incident ray 301 and the second incident ray 302 represent different types of light rays that can successfully reach the imaging unit 10, and each type of light ray contains multiple rays. The light transmission paths of the first incident ray 301 and the second incident ray 302 within the optical prism body 20 are determined based on the optical characteristics of the optical prism. The total internal reflection that occurs when the second incident ray 302 first reaches the second light-transmitting surface 202 is determined by the relative positions of the monocular endoscope body and the optical prism body 20 in the monocular endoscope stereoscopic imaging system. This relative position is calculated.

[0068] like Figure 1 As shown, in an embodiment of the present invention, the monocular endoscope body further includes an image processor, which is connected to the imaging unit 10 via a first electrical signal transmission line 70. A first connecting line is formed between the optical center of the lens unit 60, the photosensitive center point of the imaging unit 10, and the image processor. The first electrical signal transmission line 70 is located at the portion of the first connecting line where the photosensitive center point of the imaging unit 10 and the image processor are connected. The first electrical signal transmission line 70 can be understood as an endoscope signal transmission line. In conjunction with the above, to facilitate the acquisition of the incident light from the imaging target point 30, in one embodiment of the present invention, the endoscope signal transmission line is located behind the imaging unit 10.

[0069] like Figure 2As shown, in an embodiment of the present invention, the monocular endoscope body further includes an image processor, which is connected to the imaging unit 10 via a second electrical signal transmission line 80. A second connecting line is formed between the optical center of the lens unit 60 and the photosensitive center point of the imaging unit 10, and the second electrical signal transmission line 80 is located on a straight line perpendicular to the second connecting line. In this embodiment, the second electrical signal transmission line 80 is an endoscope signal transmission line. By adjusting the way the endoscope signal transmission line is led out from the back of the imaging unit 10 to the side, the imaging target point can be located on the side in the direction pointed to by the endoscope signal transmission line. This allows the viewing angle of the monocular endoscope stereoscopic imaging system to be 0 degrees, that is, the imaging target is located in front of the lens of the monocular endoscope body, which is more in line with the conventional method of endoscope image acquisition and facilitates operation.

[0070] like Figure 1 As shown, in an embodiment of the present invention, the first light-transmitting surface 201 and the second light-transmitting surface 202 are perpendicular, the angle between the reflective inclined surface 203 and the first light-transmitting surface 201 is 45 degrees, and the angle between the reflective inclined surface 203 and the second light-transmitting surface 202 is 45 degrees. 45° is a typical angle for the design of the optical prism body 20. When the prism angle is 45°, it can ensure that the first image information 910 and the second image information 920 are parallel.

[0071] like Figure 4 As shown, the imaging boundary rays of two virtual cameras in a monocular endoscope stereo imaging system are illustrated. Due to... Figure 4 There is a lot of light in the medium, so in order to ensure image clarity without affecting the understanding of the optical path, Figure 4 The two first incident rays 301 and the upper second incident ray 302 do not show refraction when entering the optical prism body 20, while the lower second incident ray 302 does show refraction when entering the optical prism body 20. Figure 4 The optical paths corresponding to the two first incident rays 301 are the transmission paths of the imaging boundary rays of the first virtual camera 91, and the optical paths corresponding to the two second incident rays 302 are the transmission paths of the imaging boundary rays of the second virtual camera 92. The imaging boundary rays of the two virtual cameras are very important for calculating the width ratio of the first image information and the second image information. Figure 5 The light shown is Figure 4 The light in the middle corresponds to, but Figure 5 Only the rays useful for calculating the width ratio of the first and second image information are retained. For example... Figure 5 As shown, l c The shortest distance between the optical center of the lens unit 60 and the intersection line of the reflective slope 203 and the second light-transmitting surface 202 of the optical prism body 20, l pLet be the side length of the second light-transmitting surface 202 of the optical prism body 20, and ∠CBL be the field of view angle of the endoscope, denoted by θ below. θ is bisected by the vertical line KB, therefore we have Since ray EB is simultaneously the imaging boundary ray of the first virtual camera 91 and the second virtual camera 92, this ray becomes the ray that divides the field of view θ of the real endoscope camera into the field of view angles of the first virtual camera 91 and the second virtual camera 92. Therefore, ∠CBE and ∠EBL are the field of view angles of the first virtual camera 91 and the second virtual camera 92, respectively, and have... and Let α represent ∠EBK in the following text. Based on the above definition and the characteristics of the imaging system, the field of view of the two virtual cameras can be calculated, and the width ratio of the first image information and the second image information corresponding to the two virtual cameras can be further calculated as shown in formula (1):

[0072]

[0073] Because it is necessary to ensure Therefore, the endoscope body must be positioned slightly to the left of the lower end of the optical prism body 20, and must meet the following requirements:

[0074] In an embodiment of the present invention, a reflective film structure is provided on the reflective slope 203. Utilizing the reflective effect of the reflective film structure, light arriving at this reflective slope can be fully reflected, ensuring the reflection effect. The reflective film structure is made of an opaque material, such as reinforced aluminum or silver-plated material.

[0075] like Figure 3 and Figure 6 As shown, a second aspect of the present invention provides a three-dimensional measurement method based on a monocular endoscopic stereo imaging system according to any of the above embodiments, comprising:

[0076] S100. Use a monocular endoscope stereo imaging system to acquire image sequences, preprocess the image sequences, and generate stereo image sequence pairs.

[0077] S200. Perform three-dimensional positioning of surgical instruments based on stereo image sequence pairs, and obtain the three-dimensional coordinates of multiple surgical instruments in the coordinate system of the first virtual camera 91 in real time.

[0078] S300: Perform 3D reconstruction of the scene based on stereo image pairs, and acquire 3D point cloud data of all pixels in the common field of view of the first virtual camera 91 and the second virtual camera 92 in real time.

[0079] The three-dimensional measurement method provided by this invention can realize the three-dimensional positioning algorithm and experiment of surgical instruments, as well as the three-dimensional reconstruction experiment of the scene. Through the above steps, the three-dimensional positioning experiment of surgical instruments and the three-dimensional reconstruction experiment of the endoscopic scene can be carried out to obtain the three-dimensional coordinates of each surgical instrument in the 91 coordinate system of the first virtual camera in the monocular endoscopic stereo imaging system and the reconstructed three-dimensional surface of the endoscopic scene.

[0080] In an embodiment of the present invention, the steps of acquiring image sequences using a monocular endoscope stereo imaging system, preprocessing the image sequences, and generating stereo image sequence pairs include:

[0081] Segment the image sequence to generate a first image information 910 sequence and a second image information 920 sequence;

[0082] Mirror and flip the first image information 910 sequence;

[0083] The first image information 910 sequence after being mirrored and flipped is cropped so that the first image information 910 sequence and the second image information 920 sequence have the same size;

[0084] A checkerboard image is acquired using a monocular endoscope stereo imaging system. Based on the checkerboard image, the first virtual camera 91 and the second virtual camera 92 are calibrated. The parameters obtained from the calibration are used to perform distortion correction and stereo correction on the first image information 910 sequence and the second image information 920 sequence.

[0085] After constructing the monocular endoscope stereo imaging system, the relative position and angle between the imaging target and the monocular endoscope stereo imaging system need to be adjusted, and then images are acquired using the monocular endoscope stereo imaging system.

[0086] Image segmentation is performed on the acquired original image sequence to generate a first image information sequence 910 and a second image information sequence 920. Since the monocular endoscopic stereoscopic imaging system of this invention acquires stereoscopic image pairs by dividing the plane of an imaging unit 10 into two sub-planes, the acquired original image simultaneously contains both the first image information 910 and the second image information 920. Therefore, the original image sequence needs to be segmented to obtain the first image information sequence 910 and the second image information sequence 920. Furthermore, since the light rays corresponding to the first virtual camera 91 in the monocular endoscopic stereoscopic imaging system of this invention undergo one reflection inside the optical prism body 20, while the light rays corresponding to the second virtual camera 92 undergo two reflections inside the optical prism body 20, a mirror image relationship exists between the first image information sequence 910 and the second image information sequence 920. Therefore, the first image information sequence 910 needs to be mirror-flipped to ensure the directional consistency of the first image information sequence 910 and the second image information sequence 920.

[0087] After acquiring two types of image information with consistent orientation, in order to ensure the smooth progress of subsequent three-dimensional measurement tasks based on endoscopic images, it is necessary to calibrate the two virtual cameras to obtain their respective intrinsic parameters and extrinsic parameters. This invention uses Zhang's calibration method to calibrate the two virtual cameras.

[0088] The calibration process for the two virtual cameras includes: capturing 25-30 pairs of checkerboard stereo images at different distances and angles using a moving checkerboard plane or a monocular endoscopic stereo imaging system; identifying feature points in the images; and estimating the five intrinsic parameters {f} of each virtual camera using a closed-form solution. x ,f y ,c x ,c y ,γ} and camera extrinsic parameters, optimize all parameters of each camera including lens distortion parameters using the maximum likelihood estimation algorithm, and use the extrinsic parameters {R} of the first virtual camera 91 L ,t L} and the extrinsic parameters of the second virtual camera 92 {R R ,t R Calculate the extrinsic parameters {R,t} between the two virtual cameras, and optimize {R,t} using the reprojection method, where f x and f y The focal lengths are the horizontal and vertical focal lengths, respectively. x ,c y ) represents the intersection of the camera's optical axis and the image sensor, γ is the unbiased constraint parameter, and R L and t L These are the rotation and translation matrices of the first virtual camera 91, R. R and t R R and t are the rotation and translation matrices of the second virtual camera 92, respectively, and the rotation and translation matrices between the two virtual cameras are respectively.

[0089] After obtaining the calibration parameters of the two virtual cameras, the present invention uses the image correction function in OpenCV to perform image distortion correction and stereo correction to remove image distortion and ensure the coplanarity and row alignment of the first image information 910 and the second image information 920.

[0090] In an embodiment of the present invention, the step of performing three-dimensional positioning of surgical instruments based on stereo image sequences and acquiring the three-dimensional coordinates of multiple surgical instruments in the coordinate system of the first virtual camera 91 in real time includes:

[0091] The stereo image sequence pairs are respectively input into the surgical instrument front end region localization network and the surgical instrument front end region detection network. The surgical instrument front end region localization network generates localization region information of each surgical instrument front end in the first image information 910 sequence and the second image information 920 sequence. The surgical instrument front end region detection network generates detection boxes of each surgical instrument front end region in the first image information 910 sequence and the second image information 920 sequence, as well as the category of each surgical instrument.

[0092] The positional information of the front-end positioning area of ​​each surgical instrument is corrected by using the detection frame of the front-end area of ​​each surgical instrument to obtain more accurate positioning information of the front-end area of ​​each surgical instrument.

[0093] Based on the precise positioning area information of each surgical instrument's front end, the center coordinates of the two-dimensional positioning area of ​​each surgical instrument in the first image information 910 sequence and the second image information 920 sequence are obtained.

[0094] The two-dimensional coordinate point matching pairs of the front ends of each surgical instrument in the first image information 910 sequence and the second image information 920 sequence are obtained by combining surgical instrument category and coordinate difference matching method;

[0095] Using the two-dimensional coordinate point matching pairs of the front ends of each surgical instrument, and the calibration parameters of the first virtual camera 91 and the second virtual camera 92, the real-time three-dimensional coordinates of each surgical instrument in the coordinate system of the first virtual camera 91 are calculated by the parallax method.

[0096] Figure 6 The flowchart of the three-dimensional positioning algorithm for surgical instruments designed in this invention is shown below. The specific implementation process of the three-dimensional positioning algorithm for surgical instruments is as follows:

[0097] Step (1):

[0098] Corrected stereo image pairs {(I 1L ,I 1R ),(I 2L ,I 2R ),(I 3L ,I 3R ),...,(I NL ,I NR The data are respectively input into the surgical instrument front end region localization network and the surgical instrument front end region detection network. The surgical instrument front end region localization network is used to obtain the two-dimensional localization region of each surgical instrument front end in the first image information 910 sequence and the second image information 920 sequence. The surgical instrument front end region detection network is used to obtain the two-dimensional detection box of each surgical instrument front end region in the first image information 910 sequence and the second image information 920 sequence, as well as the category of each surgical instrument.

[0099] In embodiments of the present invention, the surgical instrument tip region positioning network includes:

[0100] The surgical instrument front-end region localization network is designed as a pseudo-twin network structure so that the surgical instrument front-end region localization network has the ability to process the first image information sequence and the second image information sequence simultaneously.

[0101] The pseudo-twin network structure consists of two network branches with the same structure but without sharing weights. Each network branch includes an encoder-decoder branch and a full-resolution feature map generator branch to obtain more accurate positioning information for each surgical instrument.

[0102] Specifically, to enable the surgical instrument tip region localization network to simultaneously process the first image information 910 sequence and the second image information 920 sequence, this invention designs it as a pseudo-Twin network structure. This structure consists of two identical network branches that do not share weight parameters, namely the first branch and the second branch. Each network branch includes an encoder-decoder branch and a full-resolution feature map generator branch to fully utilize the full-resolution and multi-scale information of the image, thereby obtaining accurate two-dimensional localization results of the surgical instrument tip region. The surgical instrument tip region localization network is trained using the mean square loss function shown in formula (2) in a regression manner:

[0103]

[0104] Where n is the number of surgical instruments in a single image, and y is the ground truth value of the tip region of the surgical instrument. This is the predicted value for the tip region of the surgical instrument.

[0105] Specifically, the surgical instrument front-end region detection network mainly consists of a multi-head attention module and a multi-head self-attention module. The network introduces a tracking query mechanism on top of the target query mechanism to fully utilize the temporal information between images at different times, thereby obtaining a more accurate surgical instrument front-end region detection box and surgical instrument category. The loss function of the surgical instrument front-end region detection network is shown in formula (3):

[0106]

[0107] Where y and These represent the true value and the predicted value, respectively. N is used to characterize the class matching and bounding box matching between predicted and ground truth values. object and N track L represents the number of detected targets and the number of tracked targets in each image, respectively. object Let L be the target detection loss function. objectThe prediction loss of surgical instrument category and the prediction loss of the detection box of the front end region of surgical instrument were evaluated at the same time. The specific calculation method is shown in formula (4):

[0108]

[0109] in This indicates that the predicted value matches the true value, while This indicates that the predicted value and the true value do not match. This represents the probability that the predicted surgical instrument category matches the true category. This represents the probability that the predicted surgical instrument category is inconsistent with the true category, L. box Let the loss function be the target detection bounding box. This represents a truth check box. This represents the predicted detection box. L box The specific calculation method is shown in formula (5):

[0110]

[0111] The first term evaluates the L1 distance loss between the predicted and ground truth bounding boxes, and the second term evaluates the cross-union ratio loss between them. λ l1 and λ iou These are the weight parameters for the L1 distance loss and the cross-union ratio loss, respectively.

[0112] In this invention, the surgical instrument tip region localization network and the surgical instrument tip region detection network are trained in a multi-task manner, and the total loss function during training is loss = λ1 × loss localization +λ2×loss detection λ1 and λ2 are the weight parameters of the localization loss function term and the detection loss function term, respectively. In this algorithm, the detection bounding boxes output by the detection network are used to correct the results output by the localization network, thereby improving the accuracy of the two-dimensional localization results of the surgical instrument tip region.

[0113] Step (2):

[0114] After outputting the front-end positioning area of ​​each surgical instrument in the first image information 910 sequence and the second image information 920 sequence in step (1), the present invention obtains the center coordinates of each two-dimensional positioning area by using the centroid calculation formula shown in formula (6).

[0115]

[0116]

[0117] Step (3):

[0118] In embodiments of the present invention, the coordinate difference matching method includes:

[0119] Calculate the sum of the x-coordinate difference and y-coordinate difference of the center coordinates of the two-dimensional positioning areas of each surgical instrument in the first image information 910 sequence and the second image information 920 sequence, so as to obtain an accurate two-dimensional coordinate point matching pair of the front end of each surgical instrument in the first image information 910 sequence and the second image information 920 sequence.

[0120] Furthermore, considering that a scene captured by an endoscope may contain multiple surgical instruments, in order to obtain the three-dimensional coordinates of each surgical instrument, it is necessary to know the matching status of the center coordinates of the front end region of each surgical instrument in the first image information 910 sequence and the second image information 920 sequence, as shown in formula (7). This invention proposes a matching method that combines surgical instrument category information and the coordinate difference information of surgical instruments in the first image information 910 sequence and the second image information 920 sequence to obtain the two-dimensional coordinate point matching pairs of each surgical instrument in the first image information 910 sequence and the second image information 920 sequence:

[0121] match = λ class ×class+λ coord_diff ×coord_diff (7) where, λ class and λ coord_diff These are the weights of the surgical instrument category information and the coordinate difference information of the surgical instrument in the first image information 910 sequence and the second image information 920 sequence, respectively. `class` represents the surgical instrument category provided by the surgical instrument front-end region detection network, and `coord_diff` is the coordinate difference between the center coordinates of the surgical instrument front-end region in the first image information 910 sequence and the second image information 920 sequence. The specific calculation method of `coord_diff` is shown in formula (8):

[0122] coord_diff=x coord_diff +y coord_diff

[0123] x coord_diff =||x l -x r ||,y coord_diff =||y l -y r || (8) where x coord_diff and y coord_diff These are the differences in x-coordinate and y-coordinate between the center positions of the surgical instrument tip region in the first image information 910 sequence and the second image information 920 sequence, respectively. l and y lThese represent the x and y components of the center coordinates of the surgical instrument tip region in the first image information 910 sequence, respectively. r and y r These represent the x and y components of the center coordinates of the surgical instrument tip region in the second image information 920 sequence, respectively.

[0124] Step (4):

[0125] In step (3), the center coordinates of the front end regions of each surgical instrument are obtained by matching the coordinate points in the first image information 910 sequence and the second image information 920 sequence {{(x 1L ,y 1L ),(x 1R ,y 1R )},{(x 2L ,y 2L ),(x 2R ,y 2R )},...,{(x NL ,y NL ),(x NR ,y NR Afterwards, this invention utilizes coordinate point matching pairs to calculate the disparity of the center position of the front end region of each surgical instrument in the first image information 910 sequence and the second image information 920 sequence. The calculation formula is d. N =x NR -x NL , where d N For parallax, x NR and x NL The x-coordinates of the center positions of the surgical instrument tip regions in the second and first image information are respectively given. Then, the three-dimensional coordinates of the center positions of the surgical instrument tip regions in the first virtual camera 91 coordinate system are calculated according to formula (9):

[0126]

[0127] Where N represents the number of surgical instruments in the image, b is the baseline distance between the two virtual cameras obtained from the calibration, and f is the focal length of the first virtual camera in the x-direction obtained from the calibration.

[0128] The surgical instrument three-dimensional positioning algorithm in this invention, based on the spatial geometric information between the first image information 910 sequence and the second image information 920 sequence, also introduces temporal information between images at different times through the tracking query module in the detection network, thereby improving the performance of the algorithm.

[0129] In an embodiment of the present invention, the step of performing 3D scene reconstruction based on stereo image pairs and acquiring 3D point cloud data of all pixels in the common field of view of the first virtual camera 91 and the second virtual camera 92 in real time includes:

[0130] Disparity maps of stereo image pairs are obtained using a binocular disparity estimation algorithm;

[0131] Based on the disparity map and the calibration parameters of the first virtual camera 91 and the second virtual camera 92, three-dimensional coordinate point cloud data is obtained, and the corresponding three-dimensional coordinate point cloud data is textured using the first image information 910.

[0132] Specifically, considering the difficulty in obtaining ground truth depth data in endoscopic scenarios, this invention employs a self-supervised binocular disparity estimation algorithm that does not utilize external annotation information to obtain the disparity map of endoscopic stereo image pairs. The self-supervised binocular disparity estimation algorithm achieves self-supervised training as follows: it learns to reconstruct one image from one of the first image information 910 and the second image information 920. During the reconstruction process, a disparity map between the two image information is generated. The algorithm's self-supervised constraints are then achieved by calculating the error between the disparity map and the target reconstructed image information. The specific training process of the self-supervised binocular disparity estimation algorithm is as follows:

[0133] The network structure of the self-supervised binocular disparity estimation algorithm is built based on the pseudo-twin network structure. The corrected first image information 910 and second image information 920 are used as the input data of the network. This network consists of two branches, the first and the second, with the same structure but without sharing weight parameters. The main body of the network is an encoder-decoder structure, which is used to extract multi-scale features of the image and obtain a disparity map with the same resolution as the input image. The loss function of the two network branches is calculated in the same way. Therefore, the first branch is used as an example to illustrate the calculation method of the loss function: First, the first image information 910 and the first network branch are used to generate the first disparity map. Then, the loss between the first disparity map and the second image information 920 is calculated to ensure the accuracy of the predicted disparity map. The loss function consists of a structural similarity measure and a pixel consistency measure. The calculation method is shown in formula (10):

[0134] loss=α×l SSIM +(1-α)×l L1

[0135]

[0136]

[0137] Where α is used to adjust the weights of the two loss function terms, N represents the number of stereo image pairs, and y is the second image information 920. To predict the disparity map.

[0138] Subsequently, based on formula (11), the three-dimensional coordinates of each pixel in the image in the coordinate system of the first virtual camera 91 are calculated according to the disparity map and the calibration parameters of the first virtual camera 91 and the second virtual camera 92:

[0139]

[0140] Where f is the focal length of the virtual camera, b is the baseline distance between the two virtual cameras, (x L ,y L ) represents the coordinates of each pixel in the first image information 910, d represents the predicted disparity value of each pixel, and (X,Y,Z) represents the three-dimensional coordinates of each pixel in the coordinate system of the first virtual camera 91.

[0141] After obtaining the three-dimensional coordinates of each pixel, the three-dimensional coordinate point cloud data of the endoscope scene is obtained. Then, the first image information 910 is used to perform texture coverage on the three-dimensional coordinate point cloud data. Texture coverage can add color information to each point cloud data, thereby obtaining the three-dimensional reconstructed surface of the endoscope scene.

[0142] The three-dimensional measurement method provided by the second aspect of the present invention is implemented based on the monocular endoscopic stereo imaging system of the first aspect of the present invention. Specifically, the monocular endoscopic stereo imaging system includes an optical prism body 20, and utilizes the light reflection capability of the optical prism body 20 to acquire two types of image information with horizontal parallax relationship, and further calculates three-dimensional measurement information based on the two types of image information. The process of acquiring two types of image information with horizontal parallax is as follows: A first incident ray 301 emitted from the imaging target point 30 enters the optical prism body 20 after passing through the first light-transmitting surface 201. The first incident ray 301 undergoes a first refraction at the first light-transmitting surface 201. Inside the optical prism body 20, the first incident ray 301 reaches the reflecting slope 203 along the path after the first refraction and is reflected. Then, it reaches the second light-transmitting surface 202 along the reflected path. After a second refraction at the second light-transmitting surface 202, it exits the optical prism body 20 and finally reaches the imaging unit 10 through the lens unit 60. The imaging unit 10 senses and generates first image information 910 based on the first incident ray 301, and forms a first virtual camera 91 on the extension line of the first incident ray 301 before entering the optical prism body 20, thus imaging the target. The second incident ray 302 emitted from point 30 enters the optical prism body 20 after passing through the first light-transmitting surface 201. The second incident ray 302 undergoes a first refraction at the first light-transmitting surface 201. Inside the optical prism body 20, the second incident ray 302 reaches the second light-transmitting surface 202 along the path after the first refraction and undergoes total internal reflection. Then, the second incident ray 302 reaches the reflective inclined surface 203 along the path of total internal reflection and undergoes reflection. Then, it reaches the second light-transmitting surface 202 along the path of reflection. After a second refraction at the second light-transmitting surface 202, it exits the optical prism body 20 and finally reaches the imaging unit 10 through the lens unit 60. The imaging unit 10 senses and generates a second image information 920 based on the second incident ray 302, and forms a second virtual camera 92 on the extension line of the second incident ray 302 before entering the optical prism body 20. In this way, two types of image information corresponding to a single imaging target point are simultaneously generated on the same imaging unit 10. The first image information 910 and the second image information 920 are coplanar and have a certain horizontal parallax. Based on these two types of image information, three-dimensional measurement information can be further calculated. Furthermore, the monocular endoscope stereo imaging system can capture all points of the imaging target scene within the current field of view. Thus, the first image information 910 and the second image information 920 corresponding to all points of the imaging target scene within the current field of view are generated on the same imaging unit 10.Furthermore, after acquiring image sequences of the imaging target scene at different times within the current field of view, the image sequences are divided into a first image information sequence 910 and a second image information sequence 920. Since the light rays corresponding to the first image information sequence 910 and the second image information sequence 920 undergo different numbers of reflections inside the optical prism body 20—that is, the light rays corresponding to the first image information sequence 910 undergo one reflection inside the optical prism body 20, while the light rays corresponding to the second image information sequence 920 undergo two reflections inside the optical prism body 20—there is a mirror image flip between the first image information sequence 910 and the second image information sequence 920. Therefore, the first image information 910 sequence needs to be mirrored to ensure the consistency of the orientation of the first image information 910 sequence and the second image information 920 sequence. Then, a checkerboard image is acquired using a monocular endoscope stereo imaging system, and the acquired checkerboard image is used to calibrate the first virtual camera 91 and the second virtual camera 92. Then, the calibration parameters are used to perform distortion correction and stereo correction on the first image information 910 sequence and the second image information 920 sequence to obtain a stereo image sequence pair. Based on the stereo image sequence pair, the three-dimensional positioning of surgical instruments and the three-dimensional reconstruction of the endoscopic scene can be realized by the three-dimensional measurement method in this invention.

[0143] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A monocular endoscopic stereo imaging system, comprising: The single endoscope body is internally provided with an imaging unit and a lens unit. The single optical prism body includes a first light-transmitting surface, a second light-transmitting surface and a reflective inclined surface, the first light-transmitting surface is adapted to receive first incident light rays and second incident light rays of an imaging target point, and the second light-transmitting surface faces the lens unit. The lens unit is arranged between the imaging unit and the second light-transmitting surface. The first incident light rays enter the optical prism body after passing through the first light-transmitting surface, are refracted at the first light-transmitting surface for the first time, reach the reflective inclined surface inside the optical prism body and are reflected, then reach the second light-transmitting surface again, pass out of the optical prism body after being refracted at the second light-transmitting surface for the second time, pass through the lens unit and reach the imaging unit, form first image information on the imaging unit, and form a first virtual camera on the extension line of the first incident light rays before entering the optical prism body. The second incident light rays enter the optical prism body after passing through the first light-transmitting surface, are refracted at the first light-transmitting surface for the first time, reach the second light-transmitting surface inside the optical prism body and are totally reflected, then reach the reflective inclined surface again and are reflected, then reach the second light-transmitting surface again, pass out of the optical prism body after being refracted at the second light-transmitting surface for the second time, pass through the lens unit and reach the imaging unit, form second image information on the imaging unit, and form a second virtual camera on the extension line of the second incident light rays before entering the optical prism body. The single endoscope body further includes an image processor, which is connected to the imaging unit through a first electrical signal transmission line.

2. The monocular endoscopic stereo imaging system of claim 1, wherein, An optical center of the lens unit, a photosensitive center point of the imaging unit and the image processor are connected to form a first connection line, and the first electrical signal transmission line is located on the part of the first connection line where the photosensitive center point of the imaging unit and the image processor are connected. The single endoscope body further includes an image processor, which is connected to the imaging unit through a second electrical signal transmission line.

3. The monocular endoscopic stereo imaging system of claim 1, wherein, An optical center of the lens unit and a photosensitive center point of the imaging unit are connected to form a second connection line, and the second electrical signal transmission line is located on a straight line perpendicular to the second connection line. The first light-transmitting surface and the second light-transmitting surface are perpendicular to each other, the angle between the reflective inclined surface and the first light-transmitting surface is 45 degrees, and the angle between the reflective inclined surface and the second light-transmitting surface is 45 degrees.

4. The monocular endoscopic stereo imaging system of claim 1, wherein, The reflective inclined surface is provided with a reflective film structure.

5. The monocular endoscopic stereo imaging system of claim 1, wherein, ​ 6. A three-dimensional measurement method based on the monocular endoscopic stereoscopic imaging system according to any one of claims 1 to 5, characterized by, The method comprises the following steps: acquiring an image sequence by using the monocular endoscope stereo imaging system, preprocessing the image sequence, and generating a stereo image sequence pair; based on the stereo image sequence pair, performing three-dimensional positioning of surgical instruments, and acquiring real-time three-dimensional coordinates of multiple surgical instruments in the first virtual camera coordinate system; based on the stereo image pair, performing three-dimensional reconstruction of the scene, and acquiring real-time three-dimensional point cloud data of all pixels in the common field of view of the first virtual camera and the second virtual camera.

7. The three-dimensional measurement method according to claim 6, characterized in that, The step of acquiring an image sequence by using the monocular endoscope stereo imaging system, preprocessing the image sequence, and generating a stereo image sequence pair comprises the following steps: segmenting the image sequence to generate a first image information sequence and a second image information sequence; mirroring and flipping the first image information sequence; cropping the first image information sequence after mirroring and flipping to make the first image information sequence and the second image information sequence have the same size; acquiring a checkerboard image by using the monocular endoscope stereo imaging system, performing calibration of the first virtual camera and the second virtual camera based on the checkerboard image, and performing distortion correction and stereo correction on the first image information sequence and the second image information sequence by using the parameters obtained by calibration.

8. The three-dimensional measurement method according to claim 7, characterized by, The step of acquiring real-time three-dimensional coordinates of multiple surgical instruments in the first virtual camera coordinate system based on the stereo image sequence pair comprises the following steps: inputting the stereo image sequence pair into the surgical instrument front-end region positioning network and the surgical instrument front-end region detection network respectively, wherein the surgical instrument front-end region positioning network generates positioning region information of each surgical instrument front end in the first image information sequence and the second image information sequence, and the surgical instrument front-end region detection network generates each surgical instrument front-end region detection frame and the category of each surgical instrument in the first image information sequence and the second image information sequence; using the each surgical instrument front-end region detection frame to correct the position of the positioning region information of each surgical instrument front end, and obtaining more accurate each surgical instrument front-end positioning region information; based on the accurate each surgical instrument front-end positioning region information, acquiring the center coordinates of the two-dimensional positioning region of each surgical instrument in the first image information sequence and the second image information sequence; jointly using the surgical instrument category and the coordinate difference value matching method to acquire the two-dimensional coordinate point matching pair of each surgical instrument front end in the first image information sequence and the second image information sequence; using the two-dimensional coordinate point matching pair of each surgical instrument front end and the calibration parameters of the first virtual camera and the second virtual camera, and calculating the real-time three-dimensional coordinates of each surgical instrument in the first virtual camera coordinate system by using the parallax method.

9. The three-dimensional measurement method according to claim 8, characterized in that, The surgical instrument front-end region positioning network comprises the following steps: designing the surgical instrument front-end region positioning network as a pseudo-twin network structure, so that the surgical instrument front-end region positioning network has the ability to process the first image information sequence and the second image information sequence at the same time. The pseudo-twin network structure is composed of two network branches with the same structure but without sharing weights, each of the network branches comprises an encoding-decoding branch and a full-resolution feature map generator branch, so as to obtain more accurate positioning area information of each surgical instrument.

10. The three-dimensional measurement method according to claim 8, characterized by, The coordinate difference matching method comprises: The sum of the x-coordinate difference and the y-coordinate difference of the center coordinates of the two-dimensional positioning area of each surgical instrument in the first image information sequence and the second image information sequence is calculated, so as to obtain a two-dimensional coordinate point matching pair of the front end of each surgical instrument in the first image information sequence and the second image information sequence.

11. The three-dimensional measurement method according to claim 6, characterized by, The step of performing scene three-dimensional reconstruction based on the stereoscopic image pair and obtaining three-dimensional point cloud data of all pixels in the common field of view of the first virtual camera and the second virtual camera in real time comprises: A binocular disparity estimation algorithm is used to obtain a disparity map of the stereoscopic image pair; Based on the disparity map and the calibration parameters of the first virtual camera and the second virtual camera, three-dimensional coordinate point cloud data is obtained, and corresponding three-dimensional coordinate point cloud data is covered with texture by using the first image information.

Citation Information

Patent Citations

  • Systems and methods for tracking a position of a robotically-manipulated surgical instrument

    CN112672709A

  • Stereoscopic hard endoscope

    JP1995236610A