Methods for robust surface and depth estimation

Image processing algorithms using registration and reconstruction techniques address the challenge of inconsistent size representation in medical imaging systems by generating accurate three-dimensional representations, enhancing procedural efficiency.

JP7827256B2Active Publication Date: 2026-03-10BOSTON SCIENTIFIC SCIMED INC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-09-08
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing medical imaging systems, particularly monocular optical imaging systems, struggle to accurately represent the size of objects within the human body due to fixed pixel sizes and varying object distances, leading to inconsistent image scaling and challenging depth perception during procedures like lithotripsy.

Method used

Implement image processing algorithms that utilize image registration and reconstruction techniques, leveraging illumination data and camera movement to generate accurate three-dimensional representations of the imaged scene, minimizing computational requirements.

Benefits of technology

Enhances clarity and accuracy of the field of view during medical procedures by providing precise depth and size estimation of anatomical features, improving procedural efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007827256000058
    Figure 0007827256000058
  • Figure 0007827256000059
    Figure 0007827256000059
  • Figure 0007827256000060
    Figure 0007827256000060
Patent Text Reader

Abstract

Systems and methods related to estimating distance of a body structure from a medical device are disclosed. An exemplary method includes illuminating the body structure with a light source on the medical device, capturing a first input image of the body structure with a digital camera disposed on the medical device, representing the first image with a first plurality of pixels including one or more pixels displaying intensity maxima, and determining a first group of pixels from the one or more pixels displaying the intensity maxima that correspond to a plurality of surface points of the body structure and further includes a first image intensity. The method further includes calculating a relative distance from the digital camera to a first of the plurality of surface points.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of and priority to U.S. Provisional Patent Application No. 63 / 242,547, filed September 10, 2021, the disclosure of which is incorporated herein by reference.

[0002] The present disclosure relates to image processing techniques, and more particularly to reconstructing multiple images captured during a medical procedure whereby the process of aligning and reconstructing the imaged scene takes advantage of unique features of the scene to accurately display the captured images while minimizing computational requirements. [Background technology]

[0003] A variety of medical device technologies are available to medical professionals for viewing and imaging the internal organs and systems of the human body. For example, medical endoscopes equipped with digital cameras may be used by physicians in many medical fields to view the interior of parts of the human body for examination, diagnosis, and during treatment. For example, physicians may utilize a digital camera coupled to an endoscope to view the treatment of kidney stones during a lithotripsy procedure. Summary of the Invention [Problem to be solved by the invention]

[0004] During a stone-breaking procedure, a physician may observe a raw video stream captured by a digital camera positioned adjacent to the laser fiber being used to break up the kidney stone. To ensure that a medical procedure is performed efficiently, the physician (or other operator) may recognize the need to visualize the kidney stone in an appropriate field of view. For example, the image captured by the digital camera positioned adjacent to the kidney stone must accurately reflect the size of the kidney stone. Knowing the physical size of the kidney stone (and / or any remaining stone fragments) may directly affect procedural decision-making and overall procedural efficiency. In some optical imaging systems (e.g., monocular optical imaging systems), the image sensor pixel size may be fixed, and therefore the physical size of a displayed object depends on the object's distance from the collection optics. In such cases, two objects of the same size may appear different in the same image, such that an object further from the optics may appear smaller than a second object. Therefore, when analyzing video images in a medical procedure, it may be useful to accumulate data from multiple image frames that can include changes to the image "scene" in addition to changes in the camera viewpoint. This accumulated data can be used to reconstruct a three-dimensional representation of the imaged area (e.g., the size and volume of a kidney stone or other anatomical feature). It would therefore be desirable to develop image processing algorithms that register video frames and reconstruct the imaged environment, thereby improving the clarity and accuracy of the field of view observed by a physician during a medical procedure. An image processing algorithm is disclosed that utilizes image registration and reconstruction techniques to enhance multi-exposure images (while minimizing computer processing requirements). [Means for solving the problem]

[0005] The present disclosure provides design, materials, manufacturing methods, and use alternatives for medical devices. An exemplary method for estimating a distance of a body structure from a medical device includes illuminating the body structure with a light source positioned on a distal end region of the medical device, capturing a first input image of the body structure with a digital camera positioned on the distal end region of the medical device, representing the first image with a first plurality of pixels including one or more pixels representing local intensity maxima, and determining a first group of pixels from the one or more pixels representing the local intensity maxima that correspond to a plurality of surface points of the body structure and further including a first image intensity. The method further includes: JPEG0007827256000001.jpg63 and Assume that JPEG0007827256000002.jpg65 are parallel at the first surface point, A is a constant, and the following relation for r: calculating a relative distance r from the digital camera to a first surface point of the plurality of surface points by solving JPEG0007827256000003.jpg10150, wherein: I is the image intensity, L is the illumination intensity, A is the surface albedo coefficient, JPEG0007827256000004.jpg63 is the vector from the first surface point to the camera, JPEG0007827256000005.jpg65 is a vector perpendicular to the first surface point, and r is the distance from the digital camera to the first surface point.

[0006] Alternatively or additionally to any of the above embodiments, further comprising calculating a relative distance from the digital camera to each of the surface points of the plurality of surface points.

[0007] Alternatively or additionally to any of the above embodiments, further comprising calculating an accurate surface albedo coefficient using the relative distance from the digital camera to each of the surface points and the pixel intensity average over pixels with similar hues.

[0008] Alternatively or additionally to any of the above embodiments, further comprising calculating accurate relative distances from the digital camera to each of the surface points using accurate surface albedo coefficients.

[0009] Alternatively or additionally to any of the above embodiments, the precise distance and precise surface albedo coefficient are constrained by a map of three-dimensional position uncertainty vectors derived from one or more registered frames.

[0010] Alternatively or additionally to any of the above embodiments, further comprising calculating a second surface point position from a weighted average of the accurate distance values, the accurate surface albedo coefficient, and a previous estimate of the second surface point position.

[0011] Alternatively or additionally to any of the above embodiments, the weighted average of the precise distance values, the precise surface albedo coefficient, and the previous estimate of the second surface point position is inversely proportional to the magnitude of the uncertainty vector.

[0012] Alternatively or additionally to any of the above embodiments, the following relationship may be used: JPEG0007827256000006.jpg1230 JPEG0007827256000007.jpg1430 Uncertainty vector according to JPEG0007827256000008.jpg631: JPEG0007827256000009.jpg65, a weighted average, and a position, w p = weighted average of previous distance values, w m = weighted average of new measured node coordinates, JPEG0007827256000010.jpg63 = Previous model node coordinates, JPEG0007827256000011.jpg64=Updated model node coordinates, U m = new model node uncertainty measure, JPEG0007827256000012.jpg65 = Previous model node uncertainty vector, and JPEG0007827256000013.jpg64=New measurement node coordinates.

[0013] Alternatively or additionally to any of the above embodiments, further comprising generating a surface texture map associated with the surface position map, the surface texture map satisfying the following relationship: Generated using JPEG0007827256000014.jpg1363.

[0014] Another method of estimating a distance of a body structure from a medical device includes illuminating the body structure with a light source positioned on a distal end region of the medical device; using a digital camera positioned on the distal end of the region of the medical device to capture a first image of the body structure at a first time, wherein the image capture device is positioned at a first position when it captures the first image at the first time; representing the first image with a first plurality of pixels including one or more pixels representing local intensity maxima; and determining a first group of pixels from the one or more pixels representing the local intensity maxima corresponding to a plurality of surface points of the body structure and further including a first image intensity. JPEG0007827256000015.jpg63 and Assuming that JPEG0007827256000016.jpg65 are parallel at the first surface point, A is a constant, and the following relation for r: calculating a relative distance r from the digital camera to a first surface point of the plurality of surface points by solving JPEG0007827256000017.jpg10150, wherein: I is the image intensity, L is the illumination intensity, A is the surface albedo coefficient, JPEG0007827256000018.jpg63 is the vector from the first surface point to the camera, JPEG0007827256000019.jpg65 is a vector perpendicular to the first surface point, and r is the distance from the digital camera to the first surface point.

[0015] Alternatively or additionally to any of the above embodiments, further comprising calculating a relative distance from the digital camera to each of the surface points of the plurality of surface points.

[0016] Alternatively or additionally to any of the above embodiments, further comprising calculating an accurate surface albedo coefficient using the relative distance from the digital camera to each of the surface points and the pixel intensity average over pixels with similar hues.

[0017] Alternatively or additionally to any of the above embodiments, further comprising calculating accurate relative distances from the digital camera to each of the surface points using accurate surface albedo coefficients.

[0018] Alternatively or additionally to any of the above embodiments, the precise distance and precise surface albedo coefficient are constrained by a map of three-dimensional position uncertainty vectors derived from one or more registered frames.

[0019] Alternatively or additionally to any of the above embodiments, further comprising calculating a second surface point position from a weighted average of the accurate distance values, the accurate surface albedo coefficient, and a previous estimate of the second surface point position.

[0020] Alternatively or additionally to any of the above embodiments, the weighted average of the precise distance values, the precise surface albedo coefficient, and the previous estimate of the second surface point position is inversely proportional to the magnitude of the uncertainty vector.

[0021] Alternatively or additionally to any of the above embodiments, the following relationship may be used: JPEG0007827256000020.jpg1230 JPEG0007827256000021.jpg1430 Uncertainty vector according to JPEG0007827256000022.jpg631 JPEG0007827256000023.jpg65, a weighted average, and a position, w p = weighted average of previous distance values, w m = weighted average of new measured node coordinates, JPEG0007827256000024.jpg63 = Previous model node coordinates, JPEG0007827256000025.jpg64=Updated model node coordinates, U m = new model node uncertainty measure, JPEG0007827256000026.jpg65 = previous model node uncertainty vector, and JPEG0007827256000027.jpg64=New measurement node coordinates.

[0022] Alternatively or additionally to any of the above embodiments, further comprising generating a surface texture map associated with the surface position map, the surface texture map satisfying the following relationship: Generated using JPEG0007827256000028.jpg1363.

[0023] An exemplary system for estimating a distance of a body structure from a medical device includes a processor and a non-transitory computer-readable storage medium including code configured to execute a method for estimating a distance of a body structure from a medical device. The method for estimating a distance of a body structure from a medical device includes illuminating the body structure with a light source positioned on a distal end region of the medical device, capturing a first input image of the body structure with a digital camera positioned on the distal end region of the medical device, representing the first image with a first plurality of pixels including one or more pixels representing local intensity maxima, and determining a first group of pixels from the one or more pixels representing the local intensity maxima corresponding to a plurality of surface points of the body structure and further including a first image intensity. The method further includes: JPEG0007827256000029.jpg63 and Assume that JPEG0007827256000030.jpg65 are parallel at the first surface point, A is a constant, and the following relation for r: JPEG0007827256000031.jpg10150, where: I is the image intensity, L is the illumination intensity, A is the surface albedo coefficient, JPEG0007827256000032.jpg63 is the vector from the first surface point to the camera, JPEG0007827256000033.jpg65 is a vector perpendicular to the first surface point, and r is the distance from the digital camera to the first surface point.

[0024] Alternatively or additionally to any of the above embodiments, further comprising calculating a relative distance from the digital camera to each of the surface points of the plurality of surface points.

[0025] The above summary of some embodiments is not intended to describe each disclosed embodiment or every implementation of the present disclosure. The following drawings and detailed description more particularly exemplify these embodiments.

[0026] The present disclosure can be more fully understood from the following detailed description considered in conjunction with the accompanying drawings. [Brief explanation of the drawings]

[0027] [Figure 1] FIG. 1 is a schematic diagram of an exemplary endoscopic system. [Figure 2] FIG. 1 illustrates a sequence of images collected by a digital camera over a period of time. [Figure 3] FIG. 1 illustrates an exemplary optical imaging system for capturing images in an exemplary medical procedure. [Figure 4A] FIG. 4 illustrates a first image captured by the exemplary optical imaging system of FIG. 3 at a first time point. [Figure 4B] FIG. 4 shows a second image captured by the exemplary optical imaging system of FIG. 3 at a second time point. [Figure 5] 1 is a schematic diagram of an exemplary optical imaging system with an adjacent exemplary light source illuminating a surface of an object. [Figure 6] FIG. 1 is a block diagram of an image processing algorithm for estimating the distance of an object surface from a light source.

[0028] The present disclosure is susceptible to various modifications and alternative forms, details of which are shown by way of example in the drawings and described in detail below. It should be understood, however, that the intention is not to limit the present disclosure to the particular embodiments described. On the contrary, the intention is to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0029] For the following defined terms, these definitions shall be applied, unless a different definition is given in the claims or elsewhere in this specification.

[0030] All numerical values, whether expressly stated herein or not, are assumed to be modified by the term "about." The term "about" generally refers to a range of numerical values ​​that one of ordinary skill in the art would consider equivalent to the recited value (e.g., having the same function or result). In many instances, the term "about" can include numbers that are rounded to the nearest significant figure.

[0031] The recitation of numerical ranges by endpoints includes all numbers within that range (eg, 1 to 5 includes 1, 1.5, 2, 2.75, 3, 3.80, 4, and 5).

[0032] As used in this specification and claims, the singular forms "a," "an," and "the" include plural referents unless the content clearly dictates otherwise. As used in this specification and claims, the term "or" is generally used in its sense including "and / or" unless the content clearly dictates otherwise.

[0033] It should be noted that references herein to "embodiments," "some embodiments," "other embodiments," etc., indicate that the described embodiments may include one or more particular features, structures, and / or characteristics. However, such references do not necessarily imply that all embodiments include the particular feature, structure, and / or characteristic. In addition, when a particular feature, structure, and / or characteristic is described in connection with one embodiment, it should be understood that such feature, structure, and / or characteristic can also be used in connection with other embodiments, whether explicitly described or not, unless clearly stated to the contrary.

[0034] The following detailed description should be read with reference to the drawings, in which like elements in different drawings are numbered the same. The drawings, which are not necessarily to scale, depict illustrative embodiments and are not intended to limit the scope of the present disclosure.

[0035] Described herein are image processing methods performed on images collected through a medical device (e.g., an endoscope) during a medical procedure. Furthermore, the image processing methods described herein can include image registration and reconstruction algorithms. Various embodiments are disclosed for generating accurate image registration and reconstruction methods that accurately reconstruct three-dimensional images of an imaging area while minimizing computer processing requirements. Specifically, various embodiments are directed to utilizing illumination data to provide information regarding image scene depth and surface orientation. For example, the methods disclosed herein can use algorithms to extract vessel axis locations and utilize angle matching techniques to optimize registration between two or more images. Furthermore, because the medical device (e.g., an endoscope) collecting the images shifts position while collecting images (over the course of a medical procedure), the degrees of freedom (DOF) inherent in the endoscope's motion can be utilized to improve the optimization of the registration algorithm. For example, the image processing algorithms disclosed herein can utilize data representing camera movement over time, thereby utilizing data representing the camera's positional changes to reconstruct a three-dimensional representation of the imaged scene.

[0036] During medical procedures (e.g., ureteroscopy procedures), accurately representing depth perception in digital images is important to procedural efficiency. For example, obtaining an accurate representation of an object within the imaged field of view (e.g., the size of a kidney stone in a displayed image) is crucial for procedural decision-making. Furthermore, size estimation through digital imaging is directly related to depth estimation. For example, images obtained from digital sensors are inherently two-dimensional. To obtain accurate volumetric estimates and / or accurate scene reconstructions, collected images may need to be evaluated from multiple viewpoints. Furthermore, after collecting multiple images from various viewpoints (including camera positioning), the multiple image frames can be registered and aligned with each other to generate a three-dimensional representation of the anatomical scene. It can be appreciated that the process of registering multiple image frames with each other can be complicated by the inherent movement of the operator (e.g., physician) operating the image collection device (e.g., a digital camera positioned within the patient's body) as well as the movement of the patient's anatomy. As described above, understanding the camera movement from frame to frame can provide accurate depth estimates for each pixel used to represent the three-dimensional scene.

[0037] In any imaging system, to accurately interpret images, it may be important for an operator (e.g., a physician) to know the actual physical size of the displayed object. In the case of a stereoscopic optical imaging system containing multiple imaging sensors at fixed relative positions, this is achieved by calibrating the system's optical parameters (e.g., focal length, distance between sensors, distortion), registering image features from the sensors, and using that information to calculate the distance from the imaging system for each pixel, which in turn allows for the calculation of pixel size (often displayed using a scale bar). However, this may not be possible in "monocular" optical imaging systems that image three-dimensional scenes with significant depth. In these systems, the pixel size of the image sensor may be fixed, but the physical size of the displayed object will depend on the object's distance from the collection optics (e.g., the distance from the distal end of an endoscope to the object). For example, in some optical imaging systems, two objects of the same size may appear different in an image, such that an object farther from the collection optics may appear smaller than an object closer to the collection optics. Therefore, when analyzing video images, it can be useful to accumulate data from multiple image frames, which may contain changes in the imaged scene along with changes in the camera viewpoint.

[0038] In some imaging systems, the size of the field of view is estimated by comparing an object of unknown size to an object of known size. For example, during a lithotripsy procedure, the size of the field of view can be estimated by comparing the size of a laser fiber to the size of a kidney stone. However, developing the ability to make comparative estimates can be quite time-consuming for a physician. Furthermore, particularly in the case of endoscopic imaging systems, this method is inherently limited because endoscopes are relatively small and, as a result, may include cameras that (1) use a single objective lens and (2) rely on fixed focal length optics. These limitations result in imaging configurations that are prone to variable magnification of objects in the scene, which can result in each pixel detected by the camera's sensor representing a different physical size on the object.

[0039] As mentioned above, when analyzing video images, it may be useful to accumulate data from multiple image frames (which may include changes to the captured scene) and / or changes in camera viewpoint. For example, a change in camera position between two frames may allow a relative depth measurement of a scene object if pixels corresponding to the scene object's features are identified in both frames. While mapping corresponding pixels in the two images is very useful, it is often difficult and computationally complex to perform for any number of image features.

[0040] However, while collecting images with relatively small medical devices (such as endoscopes) can present challenges, endoscopic imaging can also offer unique advantages that can be utilized for efficient multiple image registration. For example, because an endoscopic scene (e.g., image collection of a kidney stone within a kidney) is typically illuminated by a single light source that has a known and fixed relationship to the camera, illumination data can provide an additional source of information regarding image depth and plane orientation.

[0041] A description of a system for combining multiple exposure images to reconstruct and register multiple images is provided below. FIG. 1 illustrates an exemplary endoscopic system that can be used in conjunction with other aspects of the present disclosure. In some embodiments, the endoscopic system can include an endoscope 10. The endoscope 10 can be specialized for a particular endoscopic procedure, such as ureteroscopy or lithotripsy, or can be a general-purpose device suitable for a wide range of procedures. In some embodiments, the endoscope 10 can include a handle 12 and an elongated shaft 14 extending distally from the handle 12, where the handle 12 includes a port configured to accept a laser fiber 16 extending within the elongated shaft 14. As shown in FIG. 1 , the laser fiber 16 can be threaded into a working channel of the elongated shaft 14 through a connector 20 (e.g., a Y-connector) or through another port positioned along the distal region of the handle 12. It can be appreciated that the laser fiber 16 can deliver laser energy to a target site within the body. For example, during a lithotripsy procedure, the laser fiber 16 can deliver laser energy to fragment a kidney stone.

[0042] Additionally, the endoscopic system shown in FIG. 1 may include a camera and / or lens positioned at the distal end of the elongate shaft 14. The elongate shaft and / or camera / lens may be deflectable and / or articulating in one or more directions to view the patient's anatomy. In some embodiments, the endoscope 10 may be a ureteroscope. However, other medical devices, such as different endoscopes or related systems, may be used in addition to or instead of a ureteroscope. Furthermore, in some embodiments, the endoscope 10 may be configured to deliver fluid from a fluid management system through the elongate shaft 14 to a treatment site. The elongate shaft 14 may include one or more working lumens for fluid flow and / or for receiving other medical devices therethrough. In some embodiments, the endoscope 10 may be connected to the fluid management system through one or more supply lines.

[0043] In some embodiments, the handle 12 of the endoscope 10 can include multiple elements configured to facilitate an endoscopic procedure. In some embodiments, a cable 18 can extend from the handle 12 and be configured to attach to an electronic device (e.g., a computer system, console, microcontroller, etc.) to provide power, analyze endoscopic data, control endoscopic procedures, or perform other functions. In some embodiments, the electronic device to which the cable 18 is connected can have the capability to recognize and exchange data with other endoscopic accessories.

[0044] In some embodiments, image signals can be transmitted from a camera at the distal end of the endoscope through cable 18 and displayed on a monitor. For example, as described above, the endoscopic system shown in FIG. 1 can include at least one camera that provides a visual feed to a user on a display screen of a computer workstation. Although not explicitly shown, it can be appreciated that elongate shaft 14 can include one or more working lumens through which a data transmission cable (e.g., fiber optic cable, optical cable, connector, wire, etc.) can extend. The data transmission cable can be connected to the camera described above. Further, the data transmission cable can be coupled to cable 18. Further, cable 18 can be coupled to a computer processing system and a display screen. Images collected by the camera can be transmitted through a data transmission cable positioned within elongate shaft 14, whereby image data passes through cable 18 to the computer processing workstation.

[0045] In some embodiments, the workstation may include, among other features, a touch panel computer, an interface box for accepting a wired connection (e.g., cable 18), a cart, and a power source. In some embodiments, the interface box may include a wired or wireless communication connection to a controller of the fluid management system. The touch panel computer may include at least a display screen and an image processor, and in some embodiments, may include and / or define a user interface. In some embodiments, the workstation may be a multi-use component (e.g., used for more than one procedure), while the endoscope 10 may be a single-use device, although this is not required. In some embodiments, the workstation may be omitted, and the endoscope 10 may be electronically coupled directly to the controller of the fluid management system.

[0046] 2 illustrates a plurality of images 100 captured sequentially by a camera over a period of time. It can be appreciated that the images 100 may represent a sequence of images captured during a medical procedure. For example, the images 100 may represent a sequence of images captured during a lithotripsy procedure in which a physician utilizes a laser fiber to treat kidney stones.

[0047] It may further be appreciated that the images 100 may be collected by an image processing system that may include, for example, a computer workstation, laptop, tablet, or other computer platform including a display that allows a physician to visualize the procedure in real time. During real-time collection of the images 100, the image processing system may be designed to process and / or enhance a given image based on fusion of one or more images taken subsequently to the given image. The enhanced image may then be viewed by the physician during the procedure.

[0048] As mentioned above, it can be appreciated that the images 100 shown in FIG. 2 can include images captured by an endoscopic device (i.e., an endoscope) during a medical procedure (e.g., during a lithotripsy procedure). Furthermore, the images 100 shown in FIG. 2 can represent a sequence 100 of images captured over time. For example, image 112 can represent an image captured by time T1, while image 114 can represent an image captured by time T2, such that image 114 captured by time T2 occurs after image 112 captured by time T1. Furthermore, image 116 can represent an image captured by time T3, such that image 116 captured by time T3 occurs after image 114 captured by time T2. This sequence can progress through images 118, 120, and 122 taken at time points T4, T5, and T6, respectively, where time point T4 occurs after time point T5, which occurs after time point T4, which occurs after time point T6.

[0049] It can further be appreciated that image 100 can be captured by a camera of an endoscopic device positioned during a live event. For example, image 100 can be captured by a digital camera positioned within a body vessel during a medical procedure. Thus, it can further be appreciated that while the camera's field of view remains constant during the procedure, the images generated during the procedure can change due to the dynamic nature of the procedure being captured. For example, image 112 can represent an image taken just before a laser fiber emits laser energy to fragment a kidney stone. Furthermore, image 114 can represent an image taken just after the laser fiber emits laser energy to fragment the kidney stone. It can further be appreciated that various particles from the kidney stone can move quickly within the camera's field of view after the laser imparts energy to the kidney stone. Additionally, it can be appreciated that the position of the camera can change over the time (while collecting image 100) that the camera collects image 100. As described herein, the changing camera position can provide data that can contribute to the generation of an accurate three-dimensional reconstructed image scene.

[0050] It can be appreciated that a digital image (such as any one of the images 100 shown in FIG. 1) can be represented as a collection of pixels (or individual pixels) arranged in a two-dimensional grid represented using squares. Furthermore, each individual pixel comprising an image can be defined as the smallest element of information within the image. Each pixel is a small sample of the original image, and typically, the more samples, the more accurately the original image is represented.

[0051] Figure 3 illustrates an example of an endoscope 110 positioned within a kidney 129. While Figure 3 and the associated discussion may be directed to images taken within the kidney, the techniques, algorithms, and / or methods disclosed herein may be applied to images acquired and processed in any body structure (e.g., body lumens, cavities, organs, etc.).

[0052] The exemplary endoscope 110 shown in FIG. 3 may be similar in form and function to the endoscope 10 described above with respect to FIG. 1. For example, FIG. 3 illustrates that the distal end region of the elongated shaft 160 of the endoscope 110 may include a digital camera 124. As described above, the digital camera 124 may be utilized to capture images of objects positioned within the exemplary kidney 129. In particular, FIG. 3 illustrates a kidney stone 128 positioned downstream (within the kidney 129) of the distal end region of the elongated shaft 160 of the endoscope 110. Thus, the camera 124 positioned on the distal end region of the shaft 160 may be utilized to capture images of the kidney stone 128 as a physician performs a medical procedure (such as a lithotripsy procedure to break up the kidney stone 128). In addition, FIG. 3 illustrates one or more calyces (cup-shaped extensions) dispersed within the kidney 129.

[0053] Additionally, it can be appreciated that as a physician operates the endoscope 110 during a medical procedure, the digital camera 124, the kidney 129, and / or the kidney stone 128 may shift position as the digital camera 124 captures images over time. Thus, images captured by the camera 124 over time may vary slightly relative to one another.

[0054] FIG. 4A illustrates a first image 130 captured by the digital camera 124 of the endoscope 110 along line 4-4 in FIG. 3. It can be seen that the image 130 illustrated in FIG. 4A illustrates a cross-sectional image of the cavity of a kidney 129 captured along line 4-4 in FIG. 3. Accordingly, FIG. 4A illustrates a kidney stone 128 positioned within the internal cavity of the kidney 129 at a first time point. Furthermore, FIG. 4B illustrates a second image 132 captured after the first image 130. In other words, FIG. 4B illustrates the second image 132 captured at a second time point that occurs after the first time point (the first time point corresponding to the time at which the image 130 was captured). It can be seen that the position of the digital camera 124 may have changed during the time lapse between the first and second time points. Accordingly, it can be seen that the change in position of the digital camera 124 is reflected in the difference between the first image 130 captured at the first time point and the second image 132 captured at a later time point.

[0055] The detailed view of FIG. 4A further illustrates that the kidney 129 can include a first blood vessel 126 that includes a central longitudinal axis 136. The blood vessel 126 may be adjacent to a kidney stone 128 and may be visible on the surface of the interior cavity of the kidney 129. It can be appreciated that the central longitudinal axis 136 can represent an approximate central location of a cross-section of the blood vessel 126 (taken at any point along the length of the blood vessel 126). For example, as shown in FIG. 4A, a dashed line 136 is depicted tracing the central longitudinal axis of the blood vessel 126. FIG. 4A also illustrates another exemplary blood vessel 127 that branches off from the blood vessel 136, the blood vessel 127 including the central longitudinal axis 137.

[0056] It can further be appreciated that generating an accurate real-time representation of the location and size of the kidney stone 128 within the cavity of the kidney 129 requires the inclusion of a “hybrid” image using data from both the first image 130 and the second image 132. In particular, the first image 130 can be registered with the second image 132 to reconstruct a hybrid image that accurately represents the location and size of the kidney stone 128 (or other structure) within the kidney 129. An exemplary method for generating a hybrid image that accurately represents the location and size of the kidney stone 128 within the kidney 129 is provided below. Additionally, as described herein, this hybrid image generation can represent a stage in generating an accurate three-dimensional reconstruction of the imaging scene represented in FIGS. 4A and 4B . For example, the hybrid image can be utilized in conjunction with position change data of the medical device 110 to generate a three-dimensional representation of the image scene.

[0057] Estimating tissue depth from object intensity As described herein, accurate estimation of the depth and surface shape of imaged tissue is important when reconstructing the imaging environment (thereby improving the clarity and accuracy of the view observed by a physician during a medical procedure). For example, determining the change in position of the medical device 110 from a body structure (e.g., change in depth from the imaged body structure) as the medical device 110 shifts position while acquiring successive images can be used in one or more algorithms to generate an accurate three-dimensional representation of the image scene. This section describes a modified Lambertian model that can be used to estimate tissue depth from object intensity.

[0058] The modified Lambertian model described herein assumes that a light source illuminating an object (e.g., a body structure) is substantially collocated with a digital camera (e.g., the light source and digital camera are positioned adjacent to the distal end of an endoscope). It can be appreciated that conventional algorithms generally include a Lambertian reflection model, which assumes that the light source is displaced at infinity from the illuminated object structure and can therefore illuminate near and far elements of the scene with substantially consistent parallel light rays. However, in the examples disclosed herein, surface reflection can be correlated with the dot product of the camera ray and the surface normal, in addition to a quartic function of the surface distance from the camera. Furthermore, as described below, assumptions regarding illumination intensity, surface albedo, and minimum inter-reflection can enable direct calculation of depth and orientation at singular points. Additionally, smoothness constraints and geometric heuristics (e.g., the assumption of a globally concave object structure) can enable the generation of an internally consistent model of the surface shape of an object (e.g., a body structure).

[0059] 5 shows a schematic diagram 200 depicting an example surface shape of an object 218 (e.g., the interior surface of a body structure) illuminated by a medical device 210. It can be appreciated that the medical device 210 can include an endoscope (such as the endoscope 10 described herein). Additionally, FIG. 5 shows that a distal end region 220 of the endoscope can include a digital camera 214 positioned next to a light source 216.

[0060] 5 further illustrates a single point 212 (out of multiple surface points) along the interior surface of a body lumen 218, where distance r represents the distance from the single point 212 to the camera 214 and light source 216. JPEG0007827256000034.jpg63 and line of sight It can be seen that the angle between the single point 212 and JPEG0007827256000035.jpg63 is minimized. It can further be seen that although a single point 212 will have a local maximum image intensity, an image obtained from a surface point can have one or more local maximum image intensity corresponding to one or more surface points.

[0061] It can be seen that the above assumptions lead to a simplified reflection model given by the following relationship:

[0062] (1) JPEG0007827256000036.jpg10150

[0063] where I is the image intensity, L is the illumination intensity, A is the surface albedo coefficient, JPEG0007827256000037.jpg63 is a line of sight vector (e.g., a vector from a surface point 212 on the body structure 218 to the camera 214), JPEG0007827256000038.jpg63 is the surface normal (eg, the vector perpendicular to the surface point 212 on the body structure 218 ), and r is the distance from the camera 214 to the surface point 212 .

[0064] Additionally, assuming A is a constant, using an initial estimate of L (described in more detail below), and assuming a surface normal parallel to the line of sight vector, an estimate of distance r (e.g., the distance from camera 214 to single point 212) can be obtained from a given intensity. It can be further recognized that the estimate of r may effectively be a maximum constraint, since relaxing assumptions about surface orientation only reduces the distance estimate. Additionally, by requiring consistency between the imaged distance, intensity, and surface orientation in equation (1) above, while introducing a smoothness assumption, a depth estimate can be derived from the first imaged point. Furthermore, the calculated depth estimate can be used to derive an estimated 3D surface orientation for the imaged object. It can also be recognized that a concave scene can be assumed, in which case the depth estimate may decrease with angle from the endoscope axis.

[0065] 6 shows an exemplary processing method 300, which includes calculating filtered intensities starting from a scalar image 302, estimating illumination parameters 304, determining depth at any central maxima 306, applying surface consistency and geometric norms to obtain depth estimates around and between maxima 308, and finally refining the illumination parameters 310. The method can derive an estimate of the distance of the imaged object (potentially each of the imaged pixels) from the camera (e.g., camera 210) that collects the image.

[0066] The surface topology, illumination parameters, and albedo estimates are interdependent models of patient imaging, each helping to inform the others. These models can be optimized together in an iterative process such as that described above, refining one model at a time and then using it to improve the estimates of the next model. This optimization cycle can be repeated multiple times for a given image set, or the model can be retained as a starting point for calculations on subsequent captured images.

[0067] Robust surface estimation The techniques described above for inferring the shape of a scene from digital images can generate a series of depth estimates from the camera 210 to the plane of anatomy over time, each of which is often subject to uncertainty or varying levels of confidence. It can be appreciated that collecting data from multiple techniques and / or integrating data from multiple images of the same scene over time can improve the overall imaged scene. For example, Kalman filtering can provide a system that optimizes depth estimates when combined with multiple measurements with varying uncertainties. However, a data representation that allows for efficient integration of multiple measurement sets with varying degrees of overlap and varying orientations may be required. This type of data representation may also be required to maintain adequate information about past measurements while supporting rapid computation of a three-dimensional surface mesh.

[0068] An example of a data representation that allows efficient integration of multiple measurement sets with varying degrees of overlap and orientation may include a triplet or quadruple tessellated surface mesh that accumulates a complex texture map to incorporate image data from all measurements, where each mesh node includes: 1) 3D position (in a scalable, potentially unitless coordinate system), 2) 3D normal vector, 3) 2D texture coordinates or other surface color information, and 4) A three-dimensional uncertainty vector whose direction points in the direction of maximum uncertainty and whose magnitude correlates with the measurement uncertainty.

[0069] Additionally, the associated texture map may contain, in addition to color information, an additional value for each pixel that counts the number of color observations factored into the pixel. Furthermore, the algorithm for combining the new measurements may include a single gain parameter that controls the stability of the system in the presence of varying measurements.

[0070] An exemplary update process might include a color bitmap as input, in addition to a 3D surface mesh, a line-of-sight origin (reflecting the camera position), and an uncertainty scalar. The uncertainty scalar can also be pre-divided by a reliability factor if the depth estimation provides independent measures of error and reliability. A vector from the origin through each node of the input mesh can be intersected with the model mesh to determine whether the previously modeled surface is within the range of a new measurement, determined by adding the uncertainty of the new measurement to the uncertainty value of the previous mesh node closest to the line-of-sight vector. If not, a mesh can be added, with each node getting an uncertainty vector equal to the normalized line-of-sight vector multiplied by its uncertainty value. If the existing mesh is within the uncertainty range of the new mesh, each node of the new mesh has a chance of intersecting with a model mesh surface and is either associated with the nearest node if within a threshold distance or added as a subdivision of the model surface. The node's position would then be set to a weighted average according to the following formula: JPEG0007827256000039.jpg1230 JPEG0007827256000040.jpg1430 JPEG0007827256000041.jpg631

[0071] where: JPEG0007827256000042.jpg63 and JPEG0007827256000043.jpg64 shows the node coordinates of the previous model and the updated model, respectively. m is the uncertainty of the new measurement, JPEG0007827256000044.jpg65 is the model node uncertainty vector, JPEG0007827256000045.jpg64 is the new measurement node coordinate.

[0072] The uncertainty is then updated according to: JPEG0007827256000046.jpg1363

[0073] The corresponding texture pixels are similarly integrated according to the weighted average described above, and each color component c p is updated using: JPEG0007827256000047.jpg737

[0074] where c m is the corresponding color component in the bitmap associated with the new measurement.

[0075] It should be understood that the present disclosure is, in many respects, merely illustrative. Changes may be made in details, particularly in matters of shape, size, and arrangement of steps, without exceeding the scope of the present disclosure. This may include, to the extent appropriate, using any of the features of one exemplary embodiment in other embodiments. The scope of the present disclosure is, of course, defined by the language in which the appended claims are expressed. [Explanation of symbols]

[0076] 200 Schematic diagram showing the surface shape of an object 212 Single points along the inner surface of a body cavity 214 Digital Camera 216 Light source 220 Distal end region of endoscope

Claims

1. 1. An endoscopy system for estimating a distance of a body structure from a medical device, comprising: a single light source disposed on a distal end region of the medical device and configured to illuminate the body structure; a digital camera disposed on the distal end region of the medical device and configured to capture a first input image of the body structure; the endoscope system is configured to represent the first input image with a first plurality of pixels, the first plurality of pixels including one or more pixels representing local intensity maxima; the endoscope system is configured to define a first group of pixels from the one or more pixels representing local intensity maxima, the first group of pixels corresponding to a plurality of surface points of the body structure, and further wherein the first group of pixels includes a first image intensity; The endoscope system includes: and are parallel at a first one of the surface points, A is a constant, and the following relationship for r is satisfied: is configured to calculate a relative distance r from the digital camera to the first surface point by solving I is the image intensity, L is the illumination intensity provided by the single light source, A is the surface albedo coefficient, is a vector from the first surface point to the digital camera, is a vector perpendicular to the first surface point, This is an endoscopy system.

2. The endoscopic system of claim 1 , further configured to calculate the relative distance from the digital camera to each of the surface points of the plurality of surface points.

3. The endoscopic system of claim 1 , wherein the medical device includes an endoscope.

4. The endoscopic system of claim 3 , wherein the endoscope includes a handle and an elongated shaft extending distally from the handle.

5. The endoscopic system of claim 4 , wherein the handle includes a port configured to receive a laser fiber extending within the elongate shaft.

6. The endoscopic system of claim 5 , wherein the laser fiber is configured to deliver laser energy to a target site of the body structure.

7. The endoscopic system of claim 1 , wherein the body structure includes a kidney stone.

8. The endoscopic system of claim 1, wherein the endoscopic system is further configured to capture multiple input images of the body structure over a period of time using the digital camera, and generate a series of depth estimates from the digital camera to a surface of the body structure based on the multiple input images.

9. The endoscopic system of claim 8 , wherein the endoscopic system is further configured to derive an estimated three-dimensional surface orientation of the body structure based on the series of depth estimates.

10. 1. An endoscopy system for estimating a distance of a body structure from a medical device, comprising: a single light source disposed on a distal end region of the medical device and configured to illuminate the body structure; a digital camera disposed on the distal end region of the medical device, the digital camera configured to capture a first image of the body structure at a first time point and positioned at a first location when capturing the first image at the first time point; the endoscope system is configured to represent the first image using a first plurality of pixels, the first plurality of pixels including one or more pixels representing local intensity maxima; the endoscope system is configured to define a first group of pixels from the one or more pixels representing local intensity maxima, the first group of pixels corresponding to a plurality of surface points of the body structure, and further wherein the first group of pixels includes a first image intensity; The endoscope system includes: and are parallel at a first one of the surface points, A is a constant, and the following relationship for r is satisfied: is configured to calculate a relative distance r from the digital camera to the first surface point by solving I is the image intensity, L is the illumination intensity provided by the single light source, A is the surface albedo coefficient, is a vector from the first surface point to the digital camera, is a vector perpendicular to the first surface point, This is an endoscopy system.

11. The endoscopic system of claim 10 , wherein the endoscopic system is further configured to calculate the relative distance from the digital camera to each of the surface points of the plurality of surface points.

12. the endoscopic system uses the digital camera to capture a plurality of input images of the body structure over a period of time, each of the plurality of input images being captured at a second time point within the period of time that is different from the first time point; The endoscopic system of claim 10 , wherein the endoscopic system is further configured to generate a series of depth estimates from the digital camera to a plane of the body structure based on the plurality of input images.

13. The endoscopic system of claim 12 , wherein the endoscopic system is further configured to derive an estimated three-dimensional surface orientation of the body structure based on the series of depth estimates.

14. The endoscopic system of claim 10 , wherein the medical device includes an endoscope.

15. The endoscopic system of claim 10 , wherein the body structure includes a kidney stone.

Citation Information

Patent Citations

  • Medical apparatus

    JP2008264539A

  • Non-realistic rendering of augmented reality

    JP2010532035A

  • Video imaging system with image processing optimized for small-diameter endoscopes

    US5475420A