Imaging system and laparoscope for imaging an object
By combining a dual-sensor imaging system and a control unit, the problem of obtaining detailed images and true depth information in laparoscopy is solved, the sensor installation and operation are simplified, it adapts to narrow environments, and improves the flexibility and operability of the imaging system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- JOSHI INNOVATIONS GMBH
- Filing Date
- 2021-11-11
- Publication Date
- 2026-07-31
AI Technical Summary
Existing imaging systems struggle to acquire detailed images of objects in laparoscopy, especially to provide accurate depth information. Furthermore, the installation and operation of sensors are complex and difficult to adapt to confined environments.
A dual-sensor imaging system is adopted, which acquires image data through the first and second sensors along optical paths with different focal offsets. The optical path is guided to the same viewpoint by the optical channel to achieve automatic image registration, and depth information is generated by the control unit.
It enables rapid acquisition of images with true depth information in laparoscopy, simplifies sensor installation and operation, adapts to narrow environments, and improves the flexibility and operability of the imaging system.
Smart Images

Figure CN116830009B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to imaging systems, laparoscopes, and methods for imaging objects. Background Technology
[0002] An imaging system is used to acquire images of an object to be imaged. This system can be housed within a laparoscope or endoscope used as a video endoscope (often referred to as a video endoscope or video probe). In such a laparoscope, an image sensor is attached to the objective lens and configured to acquire images of the object being examined; this object can be an internal organ or other hard-to-access object such as the interior of a machine. Therefore, a laparoscope is a long, narrow device with a very small internal space. Consequently, it is difficult to include an imaging system capable of acquiring detailed images of the object to be imaged within the laparoscope.
[0003] EP 1691667 A1 discloses a stereoscopic laparoscopic device comprising a laparoscope, a computer adapted to convert and store image information of the patient's affected area from the laparoscope, and a monitor for outputting the image information. The laparoscope includes: a support unit having a manipulator and a pair of parallel left and right support rods; a flexible tube unit having a pair of left and right flexible tubes adapted to be spaced apart from each other within a predetermined angle range; and a binocular camera assembly having a pair of left and right cameras mounted at the front end of the flexible tube unit so that they can capture images of the affected area in the abdominal cavity under the operation of the manipulator. Using this configuration, image information of the patient's affected area can be processed into stereoscopic photographs, thereby enabling precise diagnosis and laparoscopic surgery.
[0004] However, stereoscopic technology is a technique that creates the illusion of depth in photographs solely through the stereoscopic effect of binocular vision. Nevertheless, for detailed analysis of the object to be imaged, it is desirable to provide realistic depth information. Summary of the Invention
[0005] Therefore, the purpose of embodiments of the present invention is to provide an imaging system, a laparoscope, and a method for imaging an object, which can acquire an image that can be used to determine the true depth information of the object.
[0006] This objective is achieved by an imaging system comprising the features of claim 1, a laparoscopy comprising the features of claim 12, and a method for imaging an object comprising the features of claim 13.
[0007] According to one aspect of the invention, an imaging system is provided, comprising: an optical channel configured to transmit light; a first sensor configured to generate first image data by imaging an object along a first optical path; and a second sensor configured to generate second image data by imaging an object along a second optical path, wherein the first sensor and the second sensor are focally offset, and wherein the first optical path and the second optical path are at least partially guided through the optical channel.
[0008] The imaging system can be an optical imaging system and can be used to create image data of an object using a first sensor and a second sensor. That is, the first image data and the second image data can image the same part of the object from the same viewpoint. This is achieved by providing a dual-sensor imaging system that can rapidly generate the first and second image data continuously or simultaneously. The imaging data preferably includes an image sequence, such as a video stream. Further, since the first and second optical paths are guided through the same optical channel, the directions (i.e., viewpoints) in which the first and second sensors image the object are the same. As a result, the first and second image data are automatically registered to each other so that they completely overlap without requiring any additional alignment or registration process. Due to a focus shift between the first and second sensors, the first image data has a different focus (i.e., the point where the object is clearly depicted by the image data) compared to the second image data. That is, the first sensor can focus on a point at a first distance d1 from the far end of the optical channel, and the second sensor can focus on another point at a second distance d2 from the far end of the optical channel. The difference between the first distance d1 and the second distance d2 can be the focus shift of the two sensors. The focus shift can be determined by the hardware of the imaging system. More specifically, focus shift can be provided by the arrangement of sensors within the imaging system and / or by the optical system within the imaging system. Therefore, the focal length and / or optical path can be shortened or lengthened by the specific arrangement of the sensors or by the optical system (explained in more detail below).
[0009] An imaging system can be housed within a laparoscopy, which can provide essential information for the diagnosis and treatment of internal organs in humans or animals. Furthermore, laparoscopy can be used for the repair of large machinery, such as checking whether hard-to-access gears must be replaced without disassembling the entire machine. The sensors of the imaging system can be housed within the camera of the laparoscopy, allowing for a dual-sensor camera, where each sensor has its own focal point. In other words, the sensors can be housed individually within the camera, leaving no sensors in the shaft. Therefore, the shaft can be made, for example, elongated, enabling it to operate without problems in narrow cavities. Thus, during operation, the sensors can be located outside the cavity being inspected. Furthermore, this allows for advantageously altering the weight distribution of the endoscope, resulting in a relatively low weight of the shaft compared to the camera. The operation of the endoscope can be further simplified as the user or robot holds it at or near the camera. That is, each sensor can have its own focal point, causing the sensors to shift in focus relative to each other. For example, each sensor may have its own focal length. That is, the focal lengths of the sensors may differ. Furthermore, the laparoscope can have an axis protruding from the camera, which can be brought near the object to be imaged. That is, for example, the axis can be at least partially inserted into the human body or a narrow cavity, and the axis can include an optical channel. The axis can be connected to the camera at its proximal end. The camera can be removed from the axis. In other words, the sensor can be detached from the axis. Therefore, the sensor can be easily replaced, modified, repaired, or maintained without disassembling the axis itself. This means that the axis can be used continuously, while only the camera can be replaced. The distal end of the axis can face the object to be imaged. The axis and the optical channel can be constructed to be at least partially flexible. Furthermore, the axis can be controllably movable (e.g., controllably bent to redirect the distal end of the axis). As a result, the laparoscope can be adapted to any environment in which it is used and can reach areas located behind other objects. An optical channel (e.g., in the form of a waveguide) can be provided within the axis, which is configured to guide a first optical path and a second optical path from the proximal end of the axis to the distal end of the axis, regardless of the bending of the axis. Furthermore, optical devices (e.g., lenses or lens arrays) can be respectively disposed at the far end and / or near end of the axis to properly guide the first and second optical paths into and out of the waveguide.
[0010] The offset focus of the first and second sensors (e.g., different focal lengths of each sensor) can be achieved by providing at least one additional lens in at least one optical path of each sensor. However, only one additional lens can be provided in at least one optical path. Furthermore, more lenses or optical systems can be provided to appropriately guide light from the object to the sensors. An imaging system using an embodiment of the invention provides first and second image data that are completely overlapping each other and can be directly further processed without any additional registration process. For example, completely overlapping images with different focal points (e.g., focal points) can be used to determine the depth information of an object (described in further detail below).
[0011] The optical channel can be an elongated body. Further, the optical channel can be hollow or made of a transparent material. The optical channel can have a proximal end facing the first and second sensors and a distal end facing the object to be imaged. The distal end can protrude from the imaging system and can be the exit end of the first and second optical paths. Thus, the first and second optical paths are aligned with each other at the distal end, such that the first and second sensors have the same viewing angle. The optical channel can be configured such that the optical path exiting the optical channel at the distal end extends in a divergent manner, preferably having an extension angle of 30° to 90°, and more preferably 50° to 70° (i.e., the optical path can become conical once it exits the optical channel to capture a larger scene). For example, the optical channel can be made of optical fiber. Furthermore, the optical channel can be at least partially flexible, so that it can guide light even if the optical channel is formed as a curve. Preferably, the optical channel is flexible according to the axis of the laparoscope in which the optical channel is disposed. Therefore, objects located behind obstacles can be easily imaged by means of an axis arranged around an obstacle. An imaging system may have only one optical channel. Therefore, the imaging system can be compact, requiring less space to operate. On the other hand, due to inherent system requirements, a stereo imaging system may have two or more optical channels, and therefore may require a larger operating space.
[0012] The first and second sensors can be photoelectric sensors (also known as image sensors or sensor chips), preferably complementary metal-oxide-semiconductor (CMOS) sensors (also known as complementary symmetric metal-oxide-semiconductor (COS-MOS) sensors). Alternatively, the first and second sensors can be charge-coupled device (CCD) sensors. The first and second sensors can be sensors of the same type. Alternatively, the first sensor can be a different sensor than the second sensor. By providing different types of sensors, different image qualities and / or different image information can be obtained. For example, low-resolution image data can be used to provide coarse information about an object, while high-resolution images can be used for further processing and to provide detailed information. Preferably, the minimum resolution of at least one sensor is 1080 × 1920 pixels. Preferably, each sensor has the same resolution. Hereinafter, they are referred to as the first sensor and the second sensor; however, the imaging system can have more than two sensors, each with its own optical path. That is, if three sensors are provided, three optical paths are also provided, and so on. For example, the imaging system can have three, four, five, or more sensors. In this case, each optical path of each sensor is at least partially guided through the same optical channel. Using multiple sensors is particularly useful when the object to be imaged has a large spatial extent, or when very detailed information about the object is required. Furthermore, each sensor can have its own shutter configured to control the amount of light applied to the sensor. As a result, by adjusting the shutter speed, the sensors can adapt to different lighting conditions. Alternatively, the sensors can acquire a continuous video stream of image data.
[0013] When using a CMOS sensor, image data can be a voltage signal output by the sensor. Preferably, each pixel of the sensor can have a specific voltage signal. That is, the CMOS sensor can output a digital signal. When using a CCD sensor, image data can be a charge signal. Voltage signals are less susceptible to degradation by electromagnetic fields, therefore, a CMOS sensor is preferred as the sensor in the imaging system.
[0014] Image data can be the output of a sensor. Furthermore, image data can include brightness information for one or more color channels. A color channel can represent a specific spectrum. For example, image data can include information for a green channel, RGB color channels, and / or NIR (near-infrared) color channels. RGB color channels are considered to include a green channel, a red channel, and a blue channel. Moreover, each sensor can provide image data including information from different color channels. That is, a first sensor can provide first image data including green channel information, while a second sensor can provide second image data including NIR channel information. Further combinations of different color channels from each sensor are also possible, such as NIR-NIR, RGB-RGB, green-green, NIR-RGB, green-NIR, or green-RGB. Each color channel can be defined by a different wavelength band. For further processing of the image data, the imaging system can be connected to or is capable of being connected to a processing unit (e.g., a computer), or can have a control unit. The control unit can be configured to further process the image data. For example, the control unit can determine information contained in the image data, such as a histogram displaying the brightness distribution. Furthermore, image data can include images or photographs depicting objects. To cope with harsh lighting conditions, the imaging system can have a light source configured to illuminate the object to be imaged. More specifically, the light source can be coupled to an additional waveguide configured to guide the light from the light source, either within or parallel to the optical path, to the object to be imaged.
[0015] An optical path is the path of light from an object to a sensor. That is, light can be reflected by the object to be imaged, introduced into the optical channel, guided through the optical channel, output by the optical channel, and captured by the first and second sensors. The optical path can be defined by the light received by the respective sensors. In other words, the length of the optical path from the sensor to the object to be imaged can be measured. Furthermore, lenses or lens arrays can be provided within each optical path, configured to appropriately guide light from the object to each sensor. Additionally, the imaging system can include at least one prism configured to separate the first optical path and / or the second optical path. Preferably, the prism is configured as a beam-splitting prism. The prism can be configured to filter specific wavelengths, i.e., transmit only light of specific wavelengths. Therefore, the wavelengths transmitted to the sensors can be predetermined. For example, only wavelengths corresponding to the green channel can be transmitted by the prism. Furthermore, the prism can increase the length of one optical path relative to the other. Thus, by using a prism with multiple functions, the imaging system can be implemented with a minimal number of components. Furthermore, the imaging system can include at least one aperture in the optical path or optical channel, configured to control the aperture of the imaging system. In other words, the larger the aperture, the shallower the depth of field, and vice versa. Generally speaking, the larger the aperture, the better the quality of the image data and the more information it contains. Therefore, it is preferable to adjust the aperture to its maximum opening. As a result, the depth of field range may be relatively narrow. Furthermore, the imaging system can have one aperture for each optical path.
[0016] Offset focal lengths for the first and second sensors can be provided by offering different focal lengths for each sensor. The focal length of the first and second sensors can be the distance along the respective optical path from the principal axis of an optical lens disposed within the imaging system to the focal point. The optical lens can be configured to project an image onto the sensor. Each sensor can have a separate lens. That is, in the presence of two sensors, the imaging system can have two optical lenses. Furthermore, focal length can be a measure of the intensity of convergent or divergent light in the imaging system. A positive focal length indicates that the system converges light, while a negative focal length indicates that the system diverges light. A system with a shorter focal length bends light more significantly, causing it to focus at a shorter distance or diverge more quickly. For example, the focal length ratio of the first and second sensors can be in the range of 0.1 to 1, preferably in the range of 0.3 to 0.7. Within this range, optimal imaging of the object can be achieved. In other words, with the aforementioned ratio, the distance between the focal points of the first and second sensors is within an optimal range. More specifically, if the first image data and the second image data are combined or compared with each other, the aforementioned ratio ensures that information about the object is not lost due to complete misfocus. In other words, the ratio ensures that the distance between the focused parts of the object is not too large.
[0017] According to the present invention, at least two image data sets can be received or generated that are fully registered with each other without any additional registration process. That is, image data of an object generated by at least two sensors can have the same size and position, and can be imaged from the same viewpoint. Therefore, some parts of the object can be focused on the first image data, while other parts can be focused on the second image data. Furthermore, some parts of the object can be partially focused on the first image data and partially focused on the second image data. Thus, by combining or comparing the two image data sets, a complete image or complete data set can be generated that includes more information about the object (e.g., more detail, depth information, etc.) compared to a single image of the object or a stereoscopic image of the object. Since the two image data sets are registered to be fully overlapping, the combination or comparison of the image data is straightforward. That is, it is not necessary to register the image data sets with each other before further processing. In other words, the first and second image data sets can include exactly the same parts of the object (i.e., the scene). As a result, the present invention provides an efficient way to generate fully overlapping image data that can therefore be easily processed. Preferably, the focus offset of the sensor is determined in advance by the imaging system (i.e., before imaging the object). In other words, the focus offset can be set by the hardware of the imaging system.
[0018] Preferably, the first and second optical paths have different lengths. Therefore, the focus of the first sensor is located at a different position compared to the focus of the second sensor. Preferably, the distance between at least one sensor and the proximal end of the optical channel, and / or between the sensors, can be adjusted. As a result, the focal length of each sensor may be different, i.e., the distance from the respective sensor to the point where the object is clearly depicted by the respective image data may be different. Providing optical paths of different lengths is a simple and robust way to achieve different focal lengths (i.e., focus offset between the first and second sensors) for the first and second sensors. For example, the sensors can have different distances from the proximal end of the optical channel. That is, the first sensor can be arranged within the imaging system to be located at a position further away from the proximal end of the optical channel compared to the second sensor. As a result, the same sensor and the same image settings (e.g., aperture, shutter speed, etc.) can be used, while the different focus (i.e., focus offset between the first and second sensors) is ensured by the different positions of the sensors within the imaging system. As a result, the same components can be used to acquire the first and second image data. Therefore, the imaging system can be simplified, and manufacturing costs can be reduced.
[0019] Preferably, the first and second sensors are also configured to image the object simultaneously. That is, first image data and second image data can be generated or acquired simultaneously. As a result, the first and second image data are perfectly registered to each other because the object cannot move or change its shape between the time the first image data is acquired and the time the second image data is acquired. In other words, the first and second image data can differ from each other only in their respective focal points (i.e., focus / blur). Therefore, the imaging system can include only one shutter for the two sensors positioned within the imaging system. This further simplifies the system while ensuring complete overlap of the first and second image data.
[0020] Preferably, the system further includes a focusing system arranged in the first optical path and / or the second optical path and configured to change the focus of the first sensor and / or the second sensor. Therefore, the focus (e.g., focal point or focal length) of at least one sensor can be adjusted. In other words, the distance between the focus of the first sensor and / or the focus of the second sensor can be changed. As a result, the imaging system can be applied to different objects with different spatial extents. Furthermore, the imaging system can be applied to different applications. For example, if the imaging system is used during laparoscopy (i.e., in conjunction with a laparoscopy), the first sensor is focused at a distance of 6 cm measured from the distal end of the axis, while the second sensor can be focused at a distance of 9 cm measured from the distal end of the axis. Therefore, the focus offset of the sensors is 3 cm. Depending on the specific application, the focus offset can be positive or negative. Therefore, the imaging system can be used in a variety of applications, and multiple objects can be imaged using the imaging system. For example, the focusing system can be configured to adjust the focus of at least one sensor such that the ratio defined above between the focus of the first sensor and the focus of the second sensor can be obtained.
[0021] Preferably, the system further includes a focusing device configured to control the focusing system such that the focus (e.g., focal length) of the first and / or second sensors can be adjusted. The focusing device may be a focusing ring configured to be gripped by the user's hand. Therefore, the focusing device can protrude outside the imaging system (and outside the laparoscope) for the user to grip. The position (i.e., rotation) of the focusing ring can depend on the distance between the object to be imaged and the sensors. The greater the distance to the object, the more the focusing ring must rotate to depict at least a portion of the object clearly on one of the sensors, and vice versa. The size of the focusing device can be determined such that only a few of the user's fingers (e.g., two fingers) can contact it, while the other fingers of the user's hand can grip the imaging system (i.e., the laparoscope). As a result, the imaging system is advantageously operable with one hand. That is, the user does not need to use both hands to operate the imaging system to adjust the focus and grip the imaging system. As a result, the imaging system exhibits better operability.
[0022] Furthermore, the focusing device can be operated automatically by the control unit. More specifically, the focusing device can be controlled by the control unit to adjust the focus of each sensor according to the specific application before acquiring image data. Alternatively or supplementarily, the focusing device can be controlled by the control unit to adjust the focal length of each sensor in predetermined steps (i.e., increments). At each step, image data can be acquired by the first and second sensors. The step size can be in the range of 0.1 mm to 5 mm, preferably in the range of 0.5 mm to 2 mm. The focusing device can be located at the camera of the endoscope. Therefore, the focusing device does not need to be located at the axis. This helps to keep the weight of the axis at a low level. Thus, the operability of the entire endoscope can be improved. In addition, during the operation of the endoscope, the focusing device is placed outside the cavity being examined (e.g., the human body). Therefore, the user or robot can easily access the focusing device.
[0023] Preferably, the first image data and the second image data represent the same scene, such as an exact portion of an object. Depending on the sensor's focus (e.g., focal length) and / or the type of sensor used, the field of view of the first image data may be smaller than that of the second image data, and vice versa. To provide the same field of view of the object within the first and second image data, the image data covering the larger field of view can be adjusted to accurately include the same field of view of the object (e.g., by cutting a portion of the image data, such as the outer edge). As a result, the two image data can be compared or combined in a simple manner. For example, the edges of the image data can be used as reference points in the two image data. In other words, the object portions represented by the first and second image data completely overlap. Therefore, further processing of the image data can be further improved and simplified.
[0024] Preferably, the imaging system further includes a control unit configured to generate depth information of the object based on the first image data and the second image data. In other words, the imaging system may include a control unit for further processing the image data. The control unit may be a computer-like device and may be configured to receive input data, process the input data, and output processed data. Specifically, image data including the 2D coordinates of the object (e.g., multiple image data) may be the input data, and the object's third coordinates (i.e., depth information or 3D shape information) may be the output data. Generating depth information is an example of using multiple image data. That is, the control unit may output a depth map of the object (e.g., 3D image data). Specifically, the depth map may be a scatter plot of points, each point having three coordinates describing the spatial location or coordinates of each point on the object. That is, the control unit may be configured to determine the depth map based on the first image data and the second image data.
[0025] Furthermore, the control unit can be configured to divide the image data into segments (referred to as patches) of one or more (e.g., >9) pixels. The control unit can also be configured to compare patches of the first image data with patches of the second image data. Patches can be rectangular or square. Specifically, the control unit can be configured to determine how sharp an object is depicted in each patch (i.e., determine the sharpness of each patch). Patches can have various sizes. That is, the size of the patches set in the various image data can vary depending on the object depicted in the image data and / or depending on the specific application (e.g., the type of surgery). Specifically, in areas with a high degree of texture, the patch size can be smaller (e.g., 5×5 pixels), while in areas of image data with a more uniform pixel density (i.e., little texture), the pixel size can be larger (e.g., 50×50 or up to 1000 pixels). For example, a high degree of texture might be present in areas of image data where a large number of edges are depicted. Therefore, the imaging system can operate in a highly efficient manner because in areas with a lot of texture, the patch size is smaller (i.e., high resolution) to obtain high-precision information about that area, while in areas with less texture, the patch size is larger to accelerate the processing performed by the control unit.
[0026] Preferably, the position of at least one or each first patch in the first image data corresponds to the position of at least one or each second patch in the second image data. Preferably, at least one first patch preferably has the same size as at least one second patch, preferably 20×20 pixels. That is, the position of the first patch within the first image data is the same as the position of the second patch within the second image data. More specifically, if the first and second image data are registered with each other, the first and second patches overlap. Therefore, for each patch, the depth value (i.e., the z-coordinate) can be determined by the control unit (i.e., for each patch, the x, y, and z coordinates in 3D space are determined). A patch can have a size of one or more pixels. However, the required computational resources depend on the number of patches contained in the image data. Therefore, each patch can preferably have a size of 20×20 pixels. This size ensures both high accuracy of the depth map and system efficiency.
[0027] Furthermore, the control unit can be configured to determine the sharpness of each patch of the first image data and each patch of the second image data. It is important to note that multiple patches can have the same sharpness. The sharpness of each patch is proportional to the entropy of the corresponding patch. In other words, sharpness is equivalent to entropy. Patches of the first image data can be considered as first patches, and patches of the second image data can be considered as second patches. The first and second patches can completely overlap each other (i.e., depicting the same scene / part of an object). In other words, the first and second patches (also referred to as a pair of patches) can be located at the same position within the first and second image data. As mentioned above, the focal lengths of the first and second sensors can be preset. Therefore, the focal lengths of the first and second sensors are known. The control unit can use the following formula to determine the depth (i.e., z-coordinate) of the portion of the object depicted by the first and second patches.
[0028]
[0029] in,
[0030] d is the unknown distance (i.e., depth or z-coordinate) of the object portion depicted in the first and second patches;
[0031] d1 is the focal length of the first sensor;
[0032] I1 is the sharpness of the first segment;
[0033] d2 is the focal length of the second sensor;
[0034] I2 is the resolution of the second patch.
[0035] The focal lengths of the first and second sensors, as well as the patch size, must be carefully selected to obtain useful depth information using the formula described above. Preferably, the focal length depends on the specific application being performed (e.g., the specific surgical procedure or laparoscopy). For example, the focal length of the first sensor could be 6 cm measured from the distal end of the axis, and the focal length of the second sensor could be 9 cm measured from the distal end of the axis. The focal length of the first sensor could be an empirical value determined based on the type of surgery (e.g., for laparoscopy, the focal length could be 6 cm measured from the tip of the laparoscope (i.e., the distal end of the axis)).
[0036] Image sharpness can be represented by the information density (e.g., information entropy) of a specific patch of image data. The control unit can be configured to apply the above formula to each pair of patches in the first and second image data (i.e., each first patch and its corresponding second patch). As a result, the control unit can determine the depth value for each pair of patches in the image data. With the depth information, the control unit can be configured to create a depth map using the depth information and the x and y coordinates of the corresponding patch pair. In other words, the control unit can be configured to create a 3D model based on the x, y, and z coordinates of each pair of patches.
[0037] Alternatively or supplementarily, the control unit can be configured to create a 3D model using a focus stacking method. Another method that can be used to create a depth map is so-called depth from focus / defocus ranging. Furthermore, the control unit can be configured to perform one or more of the above methods to create depth information (i.e., 3D image data) based on at least two 2D image data sets (i.e., first image data and second image data). Specifically, the control unit can be configured to combine at least two methods for creating a depth map. In particular, the above formulas can be applied to patches of image data that have relatively high texture compared to other patches of image data. Furthermore, the shape-from-lightning method can be applied to other areas of image data that have relatively low texture compared to other areas of image data. As a result, the efficiency of depth map generation can be significantly improved by applying different methods to different parts of the image data. Therefore, at least one embodiment of the present invention can convert the relative difference in focus between image data obtained from two sensors into 3D shape information. In other words, the 3D coordinates of all patches within the image data can be calculated / determined by comparing the focus / blur of the corresponding patch (i.e., a pixel or an array of pixels) generated by each sensor. Optionally, the control unit can be configured to further process the generated depth map by filtering it. More specifically, to remove erroneous depth information, the control unit can be configured to compare the depth information of adjacent depth maps. If a depth is too high or too low compared to its neighboring depths (i.e., exceeding a pre-termination threshold), the control unit can be configured to delete that depth information, as it is likely erroneous. As a result, the depth map output by the control unit can have higher accuracy.
[0038] Preferably, the control unit is further configured to generate depth information by comparing the information entropy of at least one first patch of first image data and at least one second patch of second image data. Image sharpness is equivalent to information entropy. Information entropy is an example of information included within image data (i.e., a measure of the sharpness of patches of image data).
[0039] In other words, the first image data and the second image data can be compared with each other based on their information entropy. Therefore, the control unit can be configured to determine the information entropy of the first image data and the second image data. Specifically, the higher the information entropy, the clearer the corresponding image data. Therefore, the lower the information entropy, the less clear (i.e., blurry) the corresponding image data. Information entropy can be determined by the following formula:
[0040]
[0041] Where k is the number of levels for a channel or frequency band (e.g., the observed gray or green channel), and p k This is the probability associated with the gray level k. Furthermore, RGB channels and / or NIR channels can be used as frequency bands. The sharpness I1 and I2 used to determine the depth of a specific part of the object in the above formula can be replaced by information entropy, respectively. That is, using the above formula, information entropy H1 is determined for the first patch, and information entropy H2 is determined for the second patch (and information entropy is determined for each other patch pair). Then, in the formula used to determine the depth of a specific part of the object depicted by the first and second patches, H1 and H2 are used instead of I1 and I2. In other words, information entropy is equivalent to sharpness.
[0042] The following section outlines three optional features for improving depth estimation accuracy:
[0043] Preferably, the control unit can be configured to monitor the adjustment of the focus of at least one sensor. For example, the control unit can detect the operation of the focusing device. The focus can be adjusted by a human user and / or by a robot operating the imaging system. That is, if the focus of at least one sensor is adjusted, the control unit can monitor the amount of adjustment. In other words, the control unit can detect the offset (i.e., distance) of the focus point of each sensor. Based on the image data, the control unit can detect which portion of the image data is clearly depicted (e.g., using edge detection). Furthermore, the control unit can, for example, know the distance between the clearly depicted portion of the object being examined and the tip of the endoscope. By combining the knowledge of the position of the focus point within the image data and the clearly depicted portion of the image data, the control unit can improve depth estimation. In other words, the control unit can obtain information about the currently clearly depicted portion of the image data and the position of the focus point of each sensor. As a result, the control unit can determine the depth of the clearly depicted portion of the image data. Monitoring can be performed in real time. Therefore, continuous improvement of depth estimation is possible. Monitoring of focus adjustment can be achieved by providing optical encoders, electrical measurements, and / or other mechanical sensors. Focus adjustment can be effectively measured by using one of the aforementioned devices. Furthermore, measurement accuracy can be improved by combining these measuring devices.
[0044] Preferably, the control unit can be configured to monitor the movement of the imaging system. That is, the spatial position of the imaging system can be monitored by the control unit. For example, if the imaging system is placed inside an endoscope, the control unit can detect the spatial position of the endoscope. Preferably, this position is detected in real time. This information can then be used to improve the accuracy of depth estimation. The spatial position can be detected by an optical tracker, an electromagnetic tracker, and / or by monitoring an actuator. The actuator can be used by a robot to operate the imaging system. That is, if the robot moves the imaging system during operation, the control unit can detect this movement and can therefore determine the current spatial position of the imaging system. Furthermore, the imaging system may include an accelerometer and / or a rotation sensor. Thus, the movement of the imaging system can be monitored accurately. Based on the known spatial position of the imaging system and information about clearly depicted portions of the image data (see the overview above), the accuracy of depth estimation can be further improved.
[0045] Preferably, the control unit can be configured to infer depth information using the dimensions of a known object within the image data. That is, the control unit can be configured to determine the distance to a known object (e.g., from the tip of an endoscope) by knowing the object's dimensions. This object can be a marker present within the cavity being examined. Furthermore, the marker can be placed on additional instruments used during the examination (e.g., during surgery). Additionally, the object can be the instrument itself or other objects present within the cavity being examined. Therefore, optical calibration can be performed to improve the accuracy of depth estimation.
[0046] To reduce erroneous depth estimation results, it may be useful to improve depth estimation accuracy. Because depth estimation may be based on a partial comparison of first and second image data, erroneous results can sometimes occur (see the overview above). For example, the three features previously defined can be used to check whether the depth estimation is correct. The latter feature may be advantageous because there are no additional structural devices necessary to perform optical calibration. In other words, optical calibration can be performed by the control unit without requiring any additional sensors, etc. However, two or all three of the features described above can be implemented to improve depth estimation accuracy.
[0047] Furthermore, the information obtained as described above can also be used to compensate for distortions in imaging systems (e.g., endoscopes that include imaging systems). At least one of the features described above can be implemented using computer vision techniques. Computer vision can be an interdisciplinary scientific field that involves how computers gain high-level understanding from digital images or videos. From an engineering perspective, it may seek to understand and automate tasks that the human visual system is capable of performing. Computer vision tasks may include methods for acquiring, processing, analyzing, and understanding digital images, as well as extracting high-dimensional data from the real world to generate digital or symbolic information (e.g., in the form of decisions). In this context, "understanding" may mean converting visual images (image data) into situational descriptions that are meaningful to thought processes and can elicit appropriate actions. Such image understanding can be viewed as separating symbolic information from image data using models built with the aid of geometry, physics, statistics, and learning theories.
[0048] Preferably, the optical channel is rotatable, such that the field of view of the first and second sensors is variable. Specifically, only the distal end of the optical channel can be configured to rotate. That is, the optical channel can rotate only partially. The sensors can be fixed in rotation relative to the optical channel. Alternatively, an optical device (e.g., a lens or lens array) positioned at the distal end of the optical channel can be configured to change the field of view by rotation. Thus, the proximal end of the optical channel can be fixedly held within the imaging system, while the distal end can rotate to face the object to be imaged. The rotatable portion of the optical channel can protrude from the imaging system. The optical channel can have an initial position in which it extends straight without any curvature. Furthermore, the optical channel can rotate about its axis in the initial position. Specifically, the optical channel can be configured such that the optical path can be tilted about 20° to 40°, preferably about 30°, relative to the axis of the optical channel in the initial position. Furthermore, once the optical path leaves the optical channel, it can have a conical shape. The cone can be tilted by rotating at least a portion of the optical channel about a central axis. Specifically, by rotation, the optical path can be offset by 30° relative to the central axis of the optical channel. Furthermore, at least the distal end of the optical channel can be movable relative to the sensor. As a result, the imaging system can be used for a variety of applications, even in confined spaces or cavities. That is, only a portion of the optical channel can be oriented towards the object to be imaged in order to acquire first and second image data. Therefore, even hard-to-access objects can be imaged by the imaging system. Moreover, by being able to rotate the field of view around the axis of the optical channel and by being able to tilt the field of view relative to the axis of the optical channel, a full field of view of the cavity being inspected can be provided. For example, to obtain an overall impression of the cavity, the endoscope, including the imaging system, does not need to move within the cavity.
[0049] Preferably, the first and second sensors are arranged such that they are tilted relative to each other. That is, at least two sensors are tilted relative to each other. For example, the sensors are tilted relative to each other when the photosensitive surfaces of the sensors are tilted relative to each other. This tilting provides a more compact imaging system compared to a configuration where the sensors are coplanar. Preferably, the sensors are housed in a camera, which is separate from an axial element configured to insert into the cavity being inspected. Therefore, the camera can have a compact size due to the tilted arrangement of the sensors. That is, the extension of the camera along the axial direction can be limited to ensure the compactness of the camera. In particular, the sensors can be tilted substantially 90° relative to each other. "Substantially 90°" can refer to an arrangement where the sensors form an angle of 90° ± 10° between them. Therefore, the imaging system can be manufactured efficiently with acceptable tolerance levels.
[0050] Preferably, the first and second sensors are configured to generate first and second image data using RGB, green, and / or NIR spectra. Therefore, depending on the specific application, a spectrum containing a wealth of information under given conditions can be used. Specifically, the green spectrum contains a wealth of information under relatively poor lighting conditions. On the other hand, the NIR spectrum provides a wealth of information under almost dark lighting conditions. As a result, the image data can be easily identified or its shape and information determined under different lighting conditions. Furthermore, the first and second sensors can use different spectra / channels. For example, the first sensor can use the green spectrum, while the second sensor can use the NIR spectrum. As a result, depth information can be determined even under varying lighting conditions.
[0051] According to another aspect of the invention, a laparoscopy including the aforementioned imaging system is provided. In other words, the laparoscopy can be an endoscope including a camera with the aforementioned imaging system. The camera housing the imaging system can be connected to the laparoscopy via a camera adapter. A focusing ring can be arranged on the camera adapter of the laparoscopy to bridge the camera and the telescope of the laparoscopy. The telescope can be the axis of the laparoscopy. An optical channel can be arranged within the axis. Furthermore, the axis can be flexible to allow the optical channel to rotate at the distal end of the laparoscopy (e.g., the telescope) or to protrude from the distal end of the laparoscopy. Additionally, the laparoscopy can include a light source. The light source can be adjustable to illuminate the area oriented towards the distal end of the optical channel. Therefore, first image data and second image data can be appropriately generated even under adverse lighting conditions. The light source can preferably be an LED light source. As a result, a laparoscopy can be provided that can acquire image data with at least two focal shifts using a single camera. Therefore, the laparoscopy can have a highly integrated design and can be used in a wide range of applications. The laparoscopy can be configured for operation by a robot. In other words, a laparoscopy can have an interface configured to receive operational commands from a robot and transmit information (e.g., image data) to the robot. Therefore, a laparoscopy can be used in at least partially automated surgical environments.
[0052] According to another aspect of the present invention, a method for imaging an object is provided, wherein the method includes: generating first image data of the object by imaging the object along a first optical path using a first sensor; and generating second image data of the object by imaging the object along a second optical path using a second sensor, wherein the first sensor and the second sensor are focally offset, and wherein the first optical path and the second optical path are at least partially guided through the same optical channel.
[0053] Preferably, the method further includes the step of dividing the first image data and the second image data into patches, wherein the patches of the first image data and the patches of the second image data correspond to each other. In other words, the patch of the first image data covers the exact same area as the corresponding patch of the second image data. The corresponding patches of the first image data and the second image data can be regarded as a pair of patches.
[0054] Preferably, the method further includes the step of comparing the first image data and the second image data to generate depth information.
[0055] Preferably, the images are compared with each other by comparing at least one pair of patches.
[0056] Preferably, the first image data and the second image data are generated simultaneously.
[0057] The advantages and features described in connection with the device also apply to the method, and vice versa. Unless explicitly stated otherwise, individual embodiments or their individual aspects and features may be combined or interchanged with each other, provided that such combination or exchange is meaningful and in accordance with the spirit of the invention, without limiting or expanding the scope of the invention. Where applicable, advantages described with respect to one aspect of the invention are also advantages of other aspects of the invention. Attached Figure Description
[0058] Useful embodiments of the invention will now be described with reference to the accompanying drawings. In the drawings, the same reference numerals denote similar elements or features.
[0059] Figure 1 This is a schematic diagram of an imaging system according to an embodiment of the present invention.
[0060] Figure 2 This is a schematic diagram of an imaging system according to an embodiment of the present invention.
[0061] Figure 3 This is a schematic diagram of an imaging system according to an embodiment of the present invention.
[0062] Figure 4 This is a perspective view of a laparoscopy according to an embodiment of the present invention. Detailed Implementation
[0063] Figure 1 An imaging system 1 according to an embodiment of the present invention is schematically illustrated. Furthermore, in Figure 1The image depicts a coordinate system with x-axis, y-axis, and z-axis. Imaging system 1 is configured to image object 3. Imaging system 1 includes a first sensor 10 and a second sensor 20 configured to generate first image data and second image data. Furthermore, the first sensor 10 and the second sensor 20 are focally offset. In this embodiment, sensors 10 and 20 are both CMOS sensors. The image data includes spatial information (i.e., 2D information) in the xy-plane. Therefore, imaging system 1 includes a first optical path 11 extending between the first sensor 10 and object 3 and a second optical path 21 extending between the second sensor 20 and object 3. Furthermore, imaging system 1 includes an optical channel 2 having a distal end 5 facing object 3 and a proximal end 6 facing the first sensor 10 and the second sensor 20. Optical channel 2 is an elongated transparent body extending along its central axis C. In this embodiment, the central axis C extends along the z-axis. The first and second optical paths 11 and 21 extend from the first and second sensors 10 and 20, respectively, to a beam-splitting prism 8 disposed within imaging system 1. At the beam-splitting prism 8, the first optical path 11 and the second optical path 21 are combined / separated by the beam-splitting prism 8 and further extended to the proximal end 6 of the optical channel 2. That is, the beam-splitting prism 8 is disposed in the optical paths 11 and 21 between the sensors 10 and 20 and the optical channel 2. Then, the first and second optical paths are guided by optical lenses (not shown in the figure). In another embodiment, each optical path has its own optical lens. Subsequently, the first and second optical paths 11 and 21 are both guided through the optical channel 2. The two optical paths are guided from the distal end 5 of the optical channel toward the object 3 in the same manner. As a result, the first sensor 10 and the second sensor 20 can image the same part of the object having the same viewpoint (i.e., the same field of view). In addition, the imaging system 1 includes a focusing system 4 configured to adjust the focus (e.g., focal length) of the first sensor 10 and the second sensor 20. In this embodiment, the focusing system is a lens device and is disposed in the optical channel 2. That is, the focusing system 4 is positioned such that it can adjust the focus of the two sensors 10 and 20 simultaneously. In other words, the optical channel can consist of two parts, and the focusing system 4 can be disposed between these two parts. In another embodiment, not shown in the figure, a focusing system is provided for each sensor, allowing the focus of each sensor to be adjusted individually. In this embodiment, the focus of the two sensors 10 and 20 is set before acquiring the first image data and the second image data.
[0064] Furthermore, the first sensor 10 and the second sensor 20 are arranged within the imaging system 1 such that their focal points are shifted. That is, the first sensor 10 has a different focal point (i.e., a different focal length) compared to the second sensor 20. In other words, the focal point of the first sensor 10 is located at a different position than that of the second sensor 20. In this embodiment, the focus shift is achieved by providing an additional lens 8 in the second optical path 21. That is, in addition to the optical system (not shown) used to guide the first and second optical paths, an additional lens 8 is also provided. Therefore, the same part of the object 3 cannot be imaged clearly on both the first sensor 10 and the second sensor 20. Figure 1 In the illustrated embodiment, the first portion 31 of object 3 is imaged clearly by the first sensor 10, while the first portion 31 imaged by the second sensor 20 is blurred (i.e., cannot be shown clearly).
[0065] Figure 2 An imaging system 1 according to an embodiment of the present invention is illustrated schematically. Figure 2 The imaging system 1 shown is Figure 1 The imaging system shown is the same. (Similar to...) Figure 1 The difference lies in the depiction of a second portion 32 of object 3 located at different positions along the z-axis. Furthermore, the second portion 32 is imaged clearly by the second sensor 20, while the second portion 32 is imaged blurred by the first sensor 10. It should be noted that the first portion 31 and the second portion 32 depicted in the figure are parts of object 3, and these parts are omitted in the figure for simplicity. That is, one part of the object 3 to be imaged is clear on the first sensor 10, while another part of the object 3 to be imaged is clear on the second sensor 20. As a result, a dual-sensor optical imaging system is provided, wherein the image axes (first optical path 11 and second optical path 21) of the two sensors are at least partially identical, while the focal points of the sensors are offset. Therefore, fully registered image data can be obtained. However, according to another embodiment not shown in the figure, the imaging system includes more than two sensors, each constructed and arranged in a manner similar to the two sensors in the above embodiment.
[0066] Furthermore, the imaging system 1 includes a control unit configured to divide the first image data and the second image data into multiple patches. In this embodiment, each patch has a size of 20×20 pixels. The control unit also determines the image information (i.e., information entropy) of each patch. That is, the first image data is divided into the same patches as the second image data. As a result, a pair of patches can be determined. This pair of patches can consist of a first patch of the first image data and a corresponding second patch of the second image data. Then, the depth (i.e., z-coordinate) of the corresponding patch pair is determined using the information entropy of each patch and the known focal lengths of the first sensor 10 and the second sensor 20. This is done for each pair of patches.
[0067] In some embodiments, the focusing system 4 is operated such that the focus of sensors 10, 20 is moved in predetermined increments. In some embodiments, the focusing system 4 is automatically controlled and operated by a control unit. In another embodiment, the focusing system 4 may be operated by a user via a focusing device (e.g., a focusing ring) 104 (see [link to relevant documentation]). Figure 3 Manual operation. In each increment, image data is generated by the first sensor 10 and the second sensor 20. The image data can be generated simultaneously or sequentially by the first sensor 10 and the second sensor 20. In this embodiment, the control unit uses the following formula to determine the depth of each pair of patches:
[0068]
[0069] in,
[0070] d is the unknown distance (i.e., depth or z-coordinate) of the object portion depicted in the first and second patches;
[0071] d1 is the focal length of the first sensor;
[0072] I1 is the sharpness of the first segment;
[0073] d2 is the focal length of the second sensor;
[0074] I2 is the resolution of the second patch.
[0075] It is important to note that for the algorithm to function, the focal lengths of the first sensor 10 and the second sensor 20, as well as the patch size, need to be carefully selected. Therefore, these factors should be carefully chosen based on the specific structure of the object to be imaged. The focal lengths of each sensor 10, 20 are known. Therefore, the depth information for each patch can be determined using the formula described above.
[0076] exist Figure 3 The image depicts a portion of an imaging system 1 according to an embodiment of the present invention. Specifically, in Figure 3Two pairs of schematic diagrams are depicted. In each diagram, the optical channel 2 is schematically depicted. Furthermore, the first optical path 11 and the second optical path 21 are schematically depicted (i.e., the first optical path 11 and the second optical path 21 overlap). The optical paths extend from the distal end 6 of the optical channel, which has a cone shape (i.e., extends in a divergent manner). In this embodiment, the cone has an acute angle α of 60°. Therefore, in this embodiment, the optical channel 2 includes a lens at its distal end, which is configured to diverge the first optical path 11 and the second optical path 21 to a divergence angle of 60°.
[0077] In the first row of figures, optical channel 2 is depicted in its initial position. In the second row of figures, the optical channel is at least partially rotated to guide optical paths 11 and 21 in different directions. Specifically, according to this embodiment, optical channel 2 can be rotated at least partially about the z-axis. Therefore, the first optical path 11 and the second optical path 21 can rotate about the z-axis. In this embodiment, the optical paths are tilted such that the central axis D of the optical path forms a 30° angle with the z-axis (see...). Figure 3 (See the second row of figures). The left side of the second row shows optical paths 11 and 21 redirected to the lower side of optical path 2. On the other hand, the right side of the second row shows optical paths 11 and 21 redirected to the upper side of optical path 2. That is, optical paths 11 and 21 can rotate completely around the z-axis.
[0078] exist Figure 4 The image depicts a laparoscope 100 according to an embodiment of the present invention. The laparoscope 100 includes an imaging system 1 located within a housing 102. The laparoscope includes an axis 101 having a distal end 103. An optical channel 2 is disposed within the axis 101 of the laparoscope 100. Furthermore, the laparoscope 100 includes a light source (not shown) configured to emit light from the distal end 103 of the laparoscope along a first optical path 11 and a second optical path 12 toward an object 3.
[0079] The foregoing discussion is intended to illustrate the system only and should not be construed as limiting the appended claims to any particular embodiment or group of embodiments. Therefore, while the system has been described in particular detail with reference to exemplary embodiments, it should be understood that many modifications and alternative embodiments can be devised by those skilled in the art without departing from the broader and contemplated spirit and scope of the system as set forth in the claims. Consequently, the specification and drawings are considered illustrative and are not intended to limit the scope of the appended claims.
[0080] List of reference numerals
[0081] 1 Imaging System
[0082] 2 optical channels
[0083] 3 objects
[0084] 4-Focusing System
[0085] 5-channel far end
[0086] 6-channel near end
[0087] 7-beam prism
[0088] 8 lenses
[0089] 10 First Sensor
[0090] 11 First Optical Path
[0091] 20 Second Sensor
[0092] 21 Second optical path
[0093] 31 The first part of the object
[0094] The second part of the 32 objects
[0095] 100 Laparoscopes
[0096] 101 axis
[0097] 102 housing
[0098] 103 Distal end of laparoscopy
[0099] 104 focusing device
[0100] α-angle of the light path
[0101] C-channel axis
[0102] The axis of the D-path
[0103] xx axis
[0104] yy axis
[0105] zz axis.
Claims
1. An imaging system (1), comprising: Optical channel (2), which is configured to transmit light; A first sensor (10) is configured to generate first image data by imaging an object (3) along a first optical path (11); as well as The second sensor (20) is configured to generate second image data by imaging the object (3) along the second optical path (21). Among them, the focus of the first sensor (10) and the second sensor (20) are offset, and Wherein, the first optical path (11) and the second optical path (21) are at least partially guided through the optical channel (2), and The first sensor (10) and the second sensor (20) are arranged such that they are tilted relative to each other. The imaging system (1) further includes a control unit configured to generate depth information of the object based on the first image data and the second image data. The control unit is further configured to generate the depth information by comparing the information entropy of at least one first patch of the first image data and at least one second patch of the second image data. Wherein, the position of the at least one first patch in the first image data corresponds to the position of the at least one second patch in the second image data.
2. The imaging system (1) according to claim 1, wherein, The first optical path (11) and the second optical path (21) have different lengths.
3. The imaging system (1) according to claim 1, wherein, The first sensor (10) and the second sensor (20) are also configured to simultaneously image the object (3).
4. The imaging system (1) according to any one of claims 1 to 3, wherein, The system (1) further includes a focusing system (4) arranged in the first optical path (11) and / or the second optical path (21) and configured to change the focus of the first sensor (10) and / or the second sensor (20).
5. The imaging system (1) according to claim 4, wherein, The system (1) further includes a focusing device (104) configured to control the focusing system (4) such that the focus of the first sensor (10) and / or the second sensor (20) can be adjusted.
6. The imaging system (1) according to any one of claims 1 to 3, wherein, The first image data and the second image data represent the same area of the object (3).
7. The imaging system (1) according to claim 1, wherein, The at least one first piece has the same size as the at least one second piece.
8. The imaging system (1) according to claim 7, wherein, The at least one first piece has a size of 20×20 pixels.
9. The imaging system (1) according to any one of claims 1 to 3, wherein, The optical channel (2) is rotatable, so that the field of view of the first sensor (10) and the second sensor (20) can be changed.
10. A laparoscopy (100) comprising an imaging system (1) according to any one of claims 1 to 9.
11. A method for imaging an object, comprising: First image data of the object is generated by imaging the object along the first optical path (11) using the first sensor (10); as well as A second image of the object is generated by imaging the object along the second optical path (21) using a second sensor (20). Among them, the focus of the first sensor (10) and the second sensor (20) are offset, and Wherein, the first optical path (11) and the second optical path (21) are at least partially guided through the same optical channel (2), and The first sensor (10) and the second sensor (20) are arranged such that they are tilted relative to each other. Specifically, depth information of the object is generated based on the first image data and the second image data. The depth information is generated by comparing the information entropy of at least one first patch of the first image data and at least one second patch of the second image data. Wherein, the position of the at least one first patch in the first image data corresponds to the position of the at least one second patch in the second image data.
12. The method according to claim 11, wherein, Simultaneously, the first image data and the second image data are generated.