Method for multi-image reconstruction and registration

The image processing algorithm addresses the challenge of accurately representing three-dimensional scenes in medical imaging by registering and reconstructing multiple images, improving procedural efficiency through enhanced depth perception and size estimation.

JP2026020175APending Publication Date: 2026-02-06BOSTON SCIENTIFIC SCIMED INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025180407
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-09-10
Filing Date
2025-10-27
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

Medical imaging systems face challenges in accurately representing the size and volume of objects within a three-dimensional scene due to varying exposure conditions and camera positions, leading to inefficiencies in procedural decision-making during medical procedures like lithotripsy.

Method used

An image processing algorithm that registers and reconstructs multiple images by generating feature distance maps and calculating camera position changes to create a three-dimensional surface approximation, utilizing illumination data and chamfer matching techniques to optimize registration.

Benefits of technology

Enhances the clarity and accuracy of the viewed scene during medical procedures by providing accurate depth perception and size estimation of anatomical features, minimizing computational requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026020175000001_ABST
    Figure 2026020175000001_ABST
Patent Text Reader

Abstract

To provide systems and methods related to combining multiple images.SOLUTION: An exemplary method of combining multiple images of a body structure includes capturing a first input image with a digital camera positioned at a first location at a first time, representing the first image with a first plurality of pixels, capturing a second input image with the digital camera positioned at a second location at a second time, representing the second image with a second plurality of pixels, generating a first feature distance map of the first input image, generating a second feature distance map of the second input image, calculating a change in position of the digital camera between the first time and the second time, and generating a three dimensional surface approximation of the body structure utilizing the first feature distance map, the second feature distance map, and the change in position of the digital camera.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of and priority to U.S. Provisional Patent Application No. 63 / 242,540, filed September 10, 2021, the disclosure of which is incorporated herein by reference.

[0002] The present disclosure relates to image processing techniques, and more particularly to registering and reconstructing multiple images captured during a medical procedure, whereby the process of registering and reconstructing an imaged scene takes advantage of unique features of the scene to accurately display the captured images, while minimizing computational requirements. [Background technology]

[0003] A variety of medical device technologies are available to medical professionals for use in viewing and imaging the internal organs and systems of the human body. For example, medical endoscopes equipped with digital cameras may be used by physicians in many medical fields to view parts of the human body internally for examination, diagnosis, and during treatment. For example, physicians may utilize a digital camera coupled to an endoscope to view the treatment of kidney stones during a lithotripsy procedure. Summary of the Invention [Problem to be solved by the invention]

[0004] However, images captured by a camera during some portions of a medical procedure may be subject to various complex exposure sequences and different exposure conditions. For example, during a stone-breaking procedure, a physician may view a raw video stream captured by a digital camera positioned adjacent to a laser fiber being used to break up a kidney stone. To ensure that the medical procedure is performed efficiently, the physician (or other operator) may recognize the need to visualize the kidney stone in an appropriate field of view. For example, images captured by a digital camera positioned adjacent to the kidney stone need to accurately reflect the size of the kidney stone. Knowing the physical size of the kidney stone (and / or remaining stone fragments) may directly affect procedural decision-making and overall procedural efficiency. In some optical imaging systems (e.g., monocular optical imaging systems), the image sensor pixel size may be fixed, and therefore the physical size of a displayed object depends on the object's distance from the collection optics. In such cases, two objects of the same size may appear differently in the same image, such that an object further from the optics may appear smaller than a second object. Thus, when analyzing video images in a medical procedure, it may be useful to accumulate data from multiple or multiple image frames, which may contain changes in the image "scene" in addition to changes in camera viewpoint. This accumulated data can be used to reconstruct a three-dimensional representation of the imaged area (e.g., the size and volume of a kidney stone or other anatomical feature). It would therefore be desirable to develop an image processing algorithm that registers video frames and reconstructs the imaged environment, thereby improving the clarity and accuracy of the view observed by a physician during a medical procedure. An image processing algorithm is disclosed that utilizes image registration and reconstruction techniques to enhance multi-exposure images (while minimizing computer processing requirements). [Means for solving the problem]

[0005] The present disclosure provides design, materials, manufacturing methods, and use alternatives for medical devices. An exemplary method for combining multiple images of a body structure includes capturing a first input image with a digital camera positioned at a first location at a first time point, representing the first image with a first plurality of pixels, capturing a second input image with a digital camera positioned at a second location at a second time point, representing the second image with a second plurality of pixels, generating a first feature distance map of the first input image, generating a second feature distance map of the second input image, calculating a position change of the digital camera between the first and second time points, and generating a three-dimensional surface approximation of the body structure using the first feature distance map, the second feature distance map, and the position change of the digital camera.

[0006] Alternatively or additionally to any of the above embodiments, the first image corresponds to a body structure, and generating the first feature distance map includes selecting one or more pixels from the first plurality of pixels, the one or more pixels from the first plurality of pixels selected based on their proximity to features in the first image.

[0007] Alternatively or additionally to any of the above embodiments, one or more pixels from the first plurality of pixels are selected based on their proximity to a central longitudinal axis of the anatomical structure.

[0008] Alternatively or additionally to any of the above embodiments, the second image corresponds to a body structure, and generating the second feature distance map includes selecting one or more pixels from the second plurality of pixels, the one or more pixels from the second plurality of pixels being selected based on their proximity to features in the second image.

[0009] Alternatively or additionally to any of the above embodiments, one or more pixels from the second plurality of pixels are selected based on their proximity to a central longitudinal axis of the anatomical structure.

[0010] Alternatively or additionally to any of the above embodiments, generating the first feature distance map includes calculating a straight-line distance from a portion of the body structure to one or more pixels of the first image.

[0011] Alternatively or additionally to any of the above embodiments, generating the second feature distance map includes calculating a straight-line distance from a portion of the body structure to one or more pixels of the second image.

[0012] Alternatively or additionally to any of the above embodiments, generating the first feature distance map includes assigning a numerical value to one or more pixels of the first plurality of pixels.

[0013] Alternatively or additionally to any of the above embodiments, generating the second feature distance map includes assigning a numerical value to one or more pixels of the second plurality of pixels.

[0014] Alternatively or additionally to any of the above embodiments, the first plurality of pixels are arranged in a first coordinate grid and the second plurality of pixels are arranged in a second coordinate grid, the coordinate locations of the first plurality of pixels being at the same respective locations as the coordinate locations of the second plurality of pixels.

[0015] Alternatively or additionally to any of the above embodiments, the method further includes generating a hybrid feature distance map by aligning the first feature distance map with the second feature distance map using one or more degrees of freedom corresponding to the digital camera motion configuration parameters.

[0016] Alternatively or additionally to any of the above embodiments, the digital camera motion parameters include one or more of a positional change and a rotational change of the digital camera along the endoscope axis.

[0017] Alternatively or additionally to any of the above embodiments, further comprising assessing the reliability of the hybrid distance map by comparing distance values ​​calculated in the hybrid distance map with a distance threshold.

[0018] Another exemplary method for combining multiple images of a body structure includes using an image capture device to acquire a first image at a first time point and acquire a second image at a second time point, where the image capture device is positioned at a first location when it captures the first image at the first time point and the image capture device is positioned at a second location when it captures the second image at the second time point, the second time point occurring after the first time point. The exemplary method further includes representing the first image using a first plurality of pixels, representing the second image using a second plurality of pixels, generating a first feature distance map of the first input image, generating a second feature distance map of the second input image, calculating a positional change of the digital camera between the first time point and the second time point, and generating a three-dimensional surface approximation of the body structure using the first feature distance map, the second feature distance map, and the positional change of the digital camera.

[0019] Alternatively or additionally to any of the above embodiments, the first image corresponds to a body structure, and generating the first feature distance map includes selecting one or more pixels from the first plurality of pixels, the one or more pixels from the first plurality of pixels selected based on their proximity to a feature of the first image that correlates to the body structure.

[0020] Alternatively or additionally to any of the above embodiments, one or more pixels from the first plurality of pixels are selected based on their proximity to a central longitudinal axis of the anatomical structure.

[0021] Alternatively or additionally to any of the above embodiments, the second image corresponds to a body structure, and generating the second feature distance map includes selecting one or more pixels from the second plurality of pixels, the one or more pixels from the second plurality of pixels being selected based on their proximity to features of the body structure in the second image.

[0022] Alternatively or additionally to any of the above embodiments, one or more pixels from the second plurality of pixels are selected based on their proximity to a central longitudinal axis of the anatomical structure.

[0023] Alternatively or additionally to any of the above embodiments, the method further includes generating a hybrid feature distance map by aligning the first feature distance map with the second feature distance map using one or more degrees of freedom corresponding to the digital camera motion parameters and the endoscope state configuration parameters.

[0024] Another exemplary system for generating a fused image from multiple images includes a processor and a non-transitory computer-readable storage medium containing code configured to execute a method for fusing images. The method also includes capturing a first input image with a digital camera positioned at a first location at a first time point, representing the first image with a first plurality of pixels, capturing a second input image with a digital camera positioned at a second location at a second time point, representing the second image with a second plurality of pixels, generating a first feature distance map of the first input image, generating a second feature distance map of the second input image, calculating a positional change of the digital camera between the first and second time points, and generating a three-dimensional surface approximation of the body structure using the first feature distance map, the second feature distance map, and the positional change of the digital camera.

[0025] The above summary of some embodiments is not intended to describe each disclosed embodiment or every implementation of the present disclosure. The following drawings and detailed description more particularly exemplify these embodiments.

[0026] The present disclosure can be more fully understood from the following detailed description considered in conjunction with the accompanying drawings. [Brief explanation of the drawings]

[0027] [Figure 1] FIG. 1 is a schematic diagram of an exemplary endoscopic system. [Figure 2] FIG. 1 illustrates a sequence of images collected by a digital camera over a period of time. [Figure 3] FIG. 1 illustrates an exemplary optical imaging system for capturing images in an exemplary medical procedure. [Figure 4A] FIG. 4 illustrates a first image captured by the exemplary optical imaging system of FIG. 3 at a first time point. [Figure 4B] FIG. 4 shows a second image captured by the exemplary optical imaging system of FIG. 3 at a second time point. [Figure 5] FIG. 1 is a block diagram of an image processing algorithm that utilizes two image feature image maps to generate an optimized feature map. [Figure 6] FIG. 1 is a block diagram of an image processing algorithm for multi-image registration.

[0028] While the present disclosure is susceptible to various modifications and alternative forms, details thereof have been shown by way of example in the drawings and will be described in detail below. It should be understood, however, that the intention is not to limit the present disclosure to the particular embodiments described. On the contrary, the intention is to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0029] For the terms defined below, these definitions shall be applied unless a different definition is given in the claims or elsewhere in this specification.

[0030] All numerical values, whether explicitly stated or not, are assumed herein to be modified by the term "about." The term "about" generally refers to a range of numbers that one of ordinary skill in the art would consider equivalent to the recited value (e.g., having the same function or result). In many instances, the term "about" may include numbers that are rounded to the nearest significant figure.

[0031] The recitation of numerical ranges by endpoints includes all numbers within that range (eg, 1 to 5 includes 1, 1.5, 2, 2.75, 3, 3.80, 4, and 5).

[0032] As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the content clearly dictates otherwise. As used in this specification and the appended claims, the term "or" is generally used in its sense, including "and / or," unless the content clearly dictates otherwise.

[0033] It should be noted that references herein to "embodiments," "some embodiments," "other embodiments," etc., indicate that the described embodiment may include one or more particular features, structures, and / or characteristics. However, such a description does not necessarily imply that all embodiments include the particular feature, structure, and / or characteristic. In addition, when a particular feature, structure, and / or characteristic is described in connection with one embodiment, it should be understood that such feature, structure, and / or characteristic may also be used in connection with other embodiments, whether or not explicitly described, unless clearly stated to the contrary.

[0034] The following detailed description should be read with reference to the drawings, in which like elements in different drawings are numbered the same. The drawings, which are not necessarily to scale, depict illustrative embodiments and are not intended to limit the scope of the present disclosure.

[0035] Described herein are image processing methods performed on images collected through a medical device (e.g., an endoscope) during a medical procedure. Furthermore, the image processing methods described herein may include image registration and reconstruction algorithms. Various embodiments are disclosed for generating improved image registration and reconstruction methods that accurately reconstruct three-dimensional images of an imaging area while minimizing computer processing requirements. Specifically, various embodiments relate to utilizing illumination data to provide information regarding image scene depth and surface orientation. For example, the methods disclosed herein can use algorithms to extract vessel center axis locations and utilize chamfer matching techniques to optimize registration between two or more images. Furthermore, because the medical device (e.g., an endoscope) collecting images shifts position as it collects images (over the course of a medical procedure), the degrees of freedom (DOF) inherent in objects moving with the endoscope's field of view can be utilized to improve the optimization of the registration algorithm. For example, the image processing algorithms disclosed herein can utilize data representing camera movement over time, whereby data representing the camera's positional change can be utilized to reconstruct a three-dimensional representation of the imaging scene.

[0036] During medical procedures (e.g., ureteroscopy procedures), accurate representation of depth perception in digital images is important to procedural efficiency. For example, having an accurate representation of an object within the imaged field of view (e.g., the size of a kidney stone in a displayed image) is crucial for procedural decision-making. Furthermore, size estimation through digital imaging is directly related to depth estimation. For example, images acquired from digital sensors are essentially only two-dimensional. To obtain accurate volume estimation and / or accurate scene reconstruction, collected images may need to be evaluated from multiple viewpoints. Furthermore, after collecting multiple images from various viewpoints (including camera position changes), the multiple image frames may be registered with each other to generate a three-dimensional representation of the anatomical scene. It can be appreciated that the process of registering multiple image frames with each other may be exaggerated by the movement of the patient's anatomy as well as the inherent movement of the operator (e.g., physician) operating the image acquisition device (e.g., a digital camera positioned within the patient). As discussed above, understanding the camera's movement from frame to frame can provide accurate depth estimation for each pixel utilized to represent the three-dimensional scene.

[0037] With any imaging system, it is considered important for an operator (e.g., a physician) to know the actual physical size of the displayed object in order to accurately interpret the image. For optical imaging systems that image a two-dimensional scene at a fixed point in space, this is typically achieved by calibrating the system's optical parameters (e.g., focal length and distortion) and using that information to calculate pixel size (which can often be displayed using a scale bar). However, this may not be possible with "monocular" optical imaging systems that image three-dimensional scenes with significant depth. In these systems, the image sensor pixel size may be fixed, but the physical size of the displayed object will depend on the object's distance from the collection optics (e.g., the object's distance from the distal end of an endoscope). For example, in some optical imaging systems, two objects of the same size may appear different in the image, such that an object farther from the collection optics may appear smaller than an object closer to the collection optics. Therefore, when analyzing video images, it may be beneficial to collect data from multiple image frames that can include changes in the imaged scene as well as changes in the camera viewpoint.

[0038] In some imaging systems, the size of the field of view is estimated by comparing an object of unknown size to an object of known size. For example, during a lithotripsy procedure, the size of the field of view can be estimated by comparing the size of a laser fiber to that of a kidney stone. However, due to the size limitations inherent in conventional camera systems utilized in endoscopic procedures, it can take a significant amount of time for a physician to develop the ability to make comparative estimates. These limitations can result in imaging configurations with variable magnification of objects across a scene, whereby each pixel detected by the camera's sensor can represent a different physical size on the object.

[0039] As discussed above, when analyzing video images, it can be useful to accumulate data from multiple image frames (which may include changes to the captured scene) and / or changes in camera viewpoint. For example, a change in camera position between two frames can allow relative depth measurements of scene objects to be made if pixels corresponding to features of those objects are identified in both frames. While mapping corresponding pixels in two images can be very useful, it is often difficult and computationally complex to do for a significant number of image features.

[0040] However, while collecting images using relatively small medical devices (such as endoscopes) can present challenges, endoscopic imaging may also offer unique advantages that can be exploited for efficient multi-image registration. For example, because an endoscopic scene (e.g., collecting images of a kidney stone within a kidney) is typically illuminated by a single light source that has a known, fixed relationship to the camera, illumination data can provide an additional source of information about image depth and surface orientation. Furthermore, alternative techniques may be utilized that incorporate the local environment (such as the surface vasculature of the body cavity in which the image collection device is positioned).

[0041] A description of a system for combining multi-exposure images, registering the multiple images, and reconstructing the multiple images is provided below. FIG. 1 illustrates an exemplary endoscopic system that can be used in connection with other aspects of the present disclosure. In some embodiments, the endoscopic system can include an endoscope 10. The endoscope 10 can be specialized for a particular endoscopic procedure, such as ureteroscopy or lithotripsy, or can be a general-purpose device suitable for a wide range of procedures. In some embodiments, the endoscope 10 can include a handle 12 and an elongated shaft 14 extending distally from the handle 12, the handle 12 including a port configured to accept a laser fiber 16 extending within the elongated shaft 14. As shown in FIG. 1 , the laser fiber 16 can be threaded into a working channel of the elongated shaft 14 through a connector 20 (e.g., a Y-connector) or through another port positioned along the distal section of the handle 12. It can be appreciated that the laser fiber 16 can deliver laser energy to a target site within the body. For example, during a lithotripsy procedure, the laser fiber 16 can deliver laser energy to fragment a kidney stone.

[0042] Additionally, the endoscopic system shown in FIG. 1 may include a camera and / or lens positioned at the distal end of the elongate shaft 14. The elongate shaft and / or camera / lens may be deflectable and / or articulating in one or more directions to view the patient's anatomy. In some embodiments, the endoscope 10 may be a ureteroscope. However, other medical devices, such as different endoscopes or related systems, may be used in addition to or instead of a ureteroscope. Furthermore, in some embodiments, the endoscope 10 may be configured to deliver fluid from a fluid management system through the elongate shaft 14 to a treatment site. The elongate shaft 14 may include one or more working lumens for fluid flow and / or for receiving other medical devices therethrough. In some embodiments, the endoscope 10 may be connected to the fluid management system through one or more supply lines.

[0043] In some embodiments, the handle 12 of the endoscope 10 can include multiple elements configured to facilitate an endoscopic procedure. In some embodiments, a cable 18 can extend from the handle 12 and be configured to attach to an electronic device (not depicted) (e.g., a computer system, console, microcontroller, etc.) for providing power, analyzing endoscopic data, controlling endoscopic intervention, or performing other functions. In some embodiments, the electronic device to which the cable 18 is connected can have the capability to recognize and exchange data with other endoscopic accessories.

[0044] In some embodiments, image signals can be transmitted through cable 18 from a camera at the distal end of the endoscope for display on a monitor. For example, as discussed above, the endoscopic system shown in FIG. 1 can include at least one camera that provides a visual feed to a user on a display screen of a computer workstation. Although not explicitly shown, it can be appreciated that elongate shaft 14 can include one or more working lumens through which a data transmission cable (e.g., fiber optic cable, optical cable, connector, wire, etc.) can extend. The data transmission cable can be connected to the camera described above. Further, the data transmission cable can be coupled to cable 18. Still further, cable 18 can be coupled to a computer processing system and a display screen. Images collected by the camera can be transmitted through a data transmission cable positioned within elongate shaft 14, whereby image data passes through cable 18 to the computer processing workstation.

[0045] In some embodiments, the workstation may include, among other features, a touch panel computer, an interface box for accepting a wired connection (e.g., cable 18), a cart, and a power source. In some embodiments, the interface box may include a wired or wireless communication connection to a controller of the fluid management system. The touch panel computer may include at least a display screen and an image processor, and in some embodiments, may include and / or define a user interface. In some embodiments, the workstation is a multi-use component (e.g., used for more than one procedure), while the endoscope 10 may be a single-use device, although this is not required. In some embodiments, the workstation may be omitted, and the endoscope 10 may be electronically coupled directly to the controller of the fluid management system.

[0046] 2 illustrates a plurality of images 100 sequentially captured by a camera over a period of time. It can be appreciated that the images 100 can represent a sequence of images captured during a medical procedure. For example, the images 100 can represent a sequence of images captured during a lithotripsy procedure in which a physician utilizes a laser fiber to treat kidney stones. In some cases, the images captured by the digital camera can be captured in a green channel. Capturing images in the green channel can be beneficial because the green channel can contain the best spatial resolution within a typical colored camera filter (e.g., a Bayer filter).

[0047] It may further be appreciated that the image 100 may be collected by an image processing system, which may include, for example, a computer workstation, laptop, tablet, or other computer platform including a display that allows a physician to visualize the procedure in real time. During real-time collection of the image 100, the image processing system may be designed to process and / or enhance a given image based on fusion of one or more images taken subsequently to the given image. The enhanced image may then be viewed by the physician during the procedure.

[0048] As discussed above, it can be appreciated that the images 100 shown in FIG. 2 can include images captured by an endoscopic device (i.e., an endoscope) during a medical procedure (e.g., during a lithotripsy procedure). Furthermore, the images 100 shown in FIG. 2 can represent a sequence 100 of images captured over time. For example, image 112 can represent an image captured by time T1, while image 114 can represent an image captured by time T2, such that image 114 captured by time T2 occurs after image 112 captured by time T1. Furthermore, image 116 can represent an image captured by time T3, such that image 116 captured by time T3 occurs after image 114 captured by time T2. This order can also continue for images 118, 120, and 122 taken at time points T4, T5, and T6, respectively, such that time point T4 occurs after time point T5, which occurs after time point T4, which occurs after time point T6.

[0049] It can be further appreciated that image 100 can be captured using a camera of an endoscopic device positioned during a live event. For example, image 100 can be captured using a digital camera positioned within a body canal during a medical procedure. Thus, it can be further appreciated that while the camera's field of view remains constant during the procedure, the images generated during the procedure may change due to the dynamic nature of the procedure being captured. For example, image 112 can represent an image captured just before a laser fiber emits laser energy to fragment a kidney stone. Furthermore, image 114 can represent an image captured just after the laser fiber emits laser energy to fragment the kidney stone. It can also be appreciated that various particles from the kidney stone can move quickly within the camera's field of view after the laser applies energy to the kidney stone. It can also be appreciated that the position of the camera can change over the period of time during which the camera collects image 100 (as image 100 is collected). As described herein, the changing camera position can provide data that can contribute to the generation of an accurate three-dimensional reconstructed image scene.

[0050] It can be appreciated that a digital image (such as any one of the images 100 shown in FIG. 1) can be represented as a collection of pixels (or individual picture elements) arranged in a two-dimensional grid represented using squares. Furthermore, each individual pixel that makes up an image can be defined as the smallest information element within the image. Each pixel is a small sample of the original image, and generally, the more samples there are, the more accurately the original image is represented.

[0051] Figure 3 illustrates an example of an endoscope 110 positioned within a kidney 129. While Figure 3 and the associated discussion may be directed to images taken within the kidney, the techniques, algorithms, and / or methodologies disclosed herein may be applied to images acquired and processed in any body structure (e.g., a lumen, cavity, organ, etc.).

[0052] The exemplary endoscope 110 shown in FIG. 3 may be similar in form and function to the endoscope 10 described above with respect to FIG. 1. For example, FIG. 3 illustrates that the distal end region of the elongated shaft 160 of the endoscope 110 may include a digital camera 124. As discussed above, the digital camera 124 may be utilized to capture images of objects positioned within the exemplary kidney 129. In particular, FIG. 3 illustrates a kidney stone 128 positioned downstream (within the kidney 129) of the distal end region of the elongated shaft 160 of the endoscope 110. Thus, the camera 124 positioned at the distal end region of the shaft 160 may be utilized to capture images of the kidney stone 128 as a physician performs a medical procedure (such as a lithotripsy procedure to break up the kidney stone 128). Additionally, FIG. 3 illustrates one or more renal calyces (cup-shaped extensions) dispersed within the kidney 129.

[0053] Additionally, it can be appreciated that as the physician manipulates the endoscope 110 while performing a medical procedure, the digital camera 124, kidney 129, and / or kidney stone 128 may shift position as the digital camera 124 captures images over time. Thus, images captured by the camera 124 over time may vary slightly relative to one another.

[0054] FIG. 4A illustrates a first image 130 captured by the digital camera 124 of the endoscope 110 along line 4-4 in FIG. 3. It can be seen that the image 130 shown in FIG. 4A illustrates a cross-sectional image of the cavity of a kidney 129 captured along line 4-4 in FIG. 3. Accordingly, FIG. 4A illustrates a kidney stone 128 positioned within the internal cavity of the kidney 129 at a first time point. Additionally, FIG. 4B illustrates a second image 132 captured after the first image 130. In other words, FIG. 4B illustrates the second image 132 captured at a second time point that occurs after the first time point (the first time point corresponding to the time at which the image 130 was captured). It can be seen that the position of the digital camera 124 may change during the time lapse between the first and second time points. Accordingly, it can be seen that the change in position of the digital camera 124 is reflected in the difference between the first image 130 captured at the first time point and the second image 132 captured at a later time point.

[0055] The detailed view of FIG. 4A further illustrates that the kidney 129 can include a first blood vessel 126 that includes a central longitudinal axis 136. The blood vessel 126 may be adjacent to a kidney stone 128 and may be visible on the surface of the interior cavity of the kidney 129. It can be appreciated that the central longitudinal axis 136 can represent an approximate center location of a cross-section of the blood vessel 126 (taken at any point along the length of the blood vessel 126). For example, as shown in FIG. 4A , a dashed line 136 is depicted tracing the central longitudinal axis of the blood vessel 126. FIG. 4A also illustrates another exemplary blood vessel 127 that branches off from the blood vessel 136, the blood vessel 127 including the central longitudinal axis 137.

[0056] It can be appreciated that generating an accurate real-time representation of the location and size of the kidney stone 128 within the cavity of the kidney 129 may require the inclusion of a "hybrid" image using data from both the first image 130 and the second image 132. In particular, the first image 130 can be registered with the second image 132 to reconstruct a hybrid image that accurately represents the location and size of the kidney stone 128 (or other structure) within the kidney 129. An exemplary methodology for generating a hybrid image that accurately represents the location and size of the kidney stone 128 within the kidney 129 is provided below. Additionally, as described herein, this hybrid image generation can represent a stage in generating an accurate three-dimensional reconstruction of the imaging scene represented in FIGS. 4A and 4B . For example, the hybrid image can be utilized in conjunction with position change data of the medical device 110 to generate a three-dimensional representation of the image scene.

[0057] High-Performance Feature Maps for Registration 5 illustrates an exemplary algorithm for aligning image 130 with image 132 to generate a composite image of a scene shown over the first and second time periods described above with respect to FIGS. 4A and 4B. For simplicity, the following description assumes that image 130 and image 132 were captured at a first and second time point (whereby the second time point follows the first time point). However, it can be appreciated that the following algorithm can utilize images captured at various time points. For example, the following algorithm can be used to align image 130 with a third image captured after the second image 132.

[0058] Generally, the registration algorithms described herein extract vessel central axis locations (e.g., vessel central axis locations 136 described above) and use chamfer matching to calculate the transformation between a first image (e.g., image 130) and a second image (e.g., image 132). Furthermore, it can be appreciated that branching vasculature is a prominent feature within the endoscopic landscape, and thus the registration algorithms described herein can focus on identifying and utilizing features unique to vasculature, such as curved segments of a particular size range with light-dark-light transitions. As discussed above, these features can be best visualized in the green channel of a color image. However, it can be further appreciated that vessel edges are less well-defined and stable when considering changes in viewpoint or lighting conditions relative to the estimation of the central longitudinal axis. Therefore, a “feature detection” algorithm that locates clusters of vessel central axis locations and simultaneously builds a map of linear (“Manhattan”) distances to those features can minimize both the number of computer operations and pixel data accesses required to sufficiently register the images. Furthermore, pairs of image frames can be efficiently aligned using chamfer matching techniques by evaluating the distance map in a first image frame (e.g., image 130) at the location of the central axis cluster in a subsequent image frame (e.g., image 132). This process can enable rapid evaluation of feature alignments that can be efficiently iterated over a variety of frame alignment candidates. Bilinear interpolation of distances can be utilized when the cluster points and distance map are not perfectly aligned.

[0059] FIG. 5 shows a graphical representation of the "cluster alignment" processing step described above. It can be appreciated that a portion of a digital image can be represented as a collection of pixels (or individual picture elements) arranged in a two-dimensional grid, represented using squares. FIG. 5 shows a first pixel grid 138 that can correspond to a collection of pixels clustered around a central axis 136 of a blood vessel 126 shown in image 130. For example, the cluster of pixels shown in grid 138 can correspond to a portion of image 130 centered on central axis 136 of blood vessel 126. Similarly, FIG. 5 shows a second pixel grid 140 that can correspond to a collection of pixels centered on central axis 136 of blood vessel 126 shown in image 132. For example, the cluster of pixels shown in grid 140 can correspond to a portion of image 132 centered on central axis 136 of lumen 134. It can be appreciated that blood vessel 126 can be the same blood vessel in both first image 130 and second image 132. However, because the position of the camera 124 may change, the set of pixels (e.g., grid 138) representing the image of the blood vessels 126 in the first image 130 may differ from the set of pixels (e.g., grid 140) representing the image of the blood vessels 126 in the second image 132.

[0060] Furthermore, while the grids 138 / 140 represent selected portions of the overall images captured by the medical device (e.g., the grids 138 / 140 represent selected portions of the overall images 130 / 132 shown in FIGS. 4A and 4B, respectively), the algorithms described herein may be simultaneously applied to all of the pixels utilized to define each of the images 130 / 132. It may be appreciated that each individual pixel comprising an image may be defined as the smallest element of information within the image. However, while each pixel is a small sample of the original image, the more samples (e.g., the entire set of pixels defining the image) the more accurately the original image may be represented.

[0061] It can further be appreciated that, for simplicity, the grid for each of sub-images 130 and 132 is sized 8x8. In other words, the two-dimensional grid of images 130 / 132 includes eight columns of pixels extending vertically and eight rows of pixels extending horizontally. It can be appreciated that the image sizes depicted in FIG. 5 are exemplary. The size (total number of pixels) of a digital image can vary. For example, the size of the pixel grid representing the entire image can be approximately several hundred by several hundred pixels (e.g., 250x250, 400x400).

[0062] It can be appreciated that individual pixel locations can be identified by their coordinates (X,Y) on a two-dimensional image grid. Additionally, comparison of neighboring pixels within a given image may provide desirable information regarding which portions of a given image an algorithm should utilize when performing the registration process. For example, FIG. 5 illustrates that each grid 138 / 140 can include one or more "feature" pixels. These feature pixels may represent pixels within each grid 138 / 140 that are closest to the central axis 136 of each image 130 / 132, respectively. It can be appreciated that the pixels that make up a feature pixel are substantially darker than the pixels adjacent to the feature pixel. For example, feature pixel 142 within grid 138 and feature pixel 144 within grid 140 are black. All other pixels adjacent to feature pixel 142 / 144 are depicted as various shades of gray (or white).

[0063] FIG. 5 further illustrates that each pixel in each respective pixel grid 138 / 140 can be assigned a numerical value corresponding to its relative distance to a feature pixel. In some cases, this methodology of assigning a numerical value corresponding to the distance from a given pixel to a feature pixel can include constructing a feature map using a linear "Manhattan" distance. For example, FIG. 5 illustrates that feature pixels 142 / 144 in grids 138 / 140 can each be assigned a numerical value of "0," gray variations can be assigned a value of "1," and white pixels can be assigned a value of "2," whereby larger numerical values ​​correspond to greater distances from a given feature pixel. Furthermore, FIG. 5 illustrates that a numerical representation of pixel grid 138 is displayed in grid 146, while a numerical representation of pixel grid 138 is displayed in grid 148.

[0064] As described herein, because images 130 and 132 were captured at different times, feature pixels in image 130 may be located at different coordinates than feature pixels in image 132. Therefore, to generate a hybrid image utilizing feature pixel data from both images 130 and 132, a registration process can be used to generate a hybrid numeric grid 150 having feature pixels generated by the sum of coordinate locations of numeric grids 138 and 148. The feature pixel locations in hybrid numeric grid 150 will include overlapping locations of feature pixels 142 / 144 in each grid 138 / 10, respectively (e.g., coordinates within each grid 138 / 140 that share the feature pixel). For example, FIG. 5 shows that coordinate location 4,6 (row 4, column 6) in hybrid grid 150 includes feature pixel 152 (identified by having a numeric value of "0"), which is the sum of the value "0" (at coordinate location 4,6 in grid 138) and the value "0" (at coordinate location 4,6 in grid 140). The remaining feature pixels of the hybrid grid 150 can be identified by repeating this process across all coordinates of the grid 150. Furthermore, as described herein, while Figure 5 illustrates a summation process performed on a selected group of pixel locations for each image 130 / 132, it can be appreciated that this alignment process can be performed across all pixels in the images 130 / 132.

[0065] Additionally, it can be further appreciated that the feature pixels of the hybrid grid 150 can be optimized across multiple frames. For example, the first iteration of generating the hybrid numerical grid can provide an initial estimate of how "misaligned" the first image 130 is from image 132. By continuing to iterate the algorithm across multiple registration hypotheses in conjunction with an optimization process (e.g., simplex method), the individual parameters of the registration (e.g., translation, scale, and rotation for rigid registration) are adjusted to identify the combination with the best results.

[0066] Camera-based DOF-limited registration It can be appreciated that, within the "high performance feature map alignment" process described herein, the computationally intensive phase can be an iterative optimization "loop" that scales exponentially with the number of degrees of freedom (DOF) applied to the alignment process. For example, with reference to images 130 and 132 described herein, an object in the image (e.g., a kidney stone being crushed) has six degrees of freedom, including three dimensions (X, Y, Z) along which relative motion can occur, plus an additional three degrees of freedom corresponding to each axis of rotation along each dimension. However, the degrees of freedom inherent in a moving object can be utilized to improve the computational efficiency of the optimization loop. An exemplary process flow methodology 200 for improving the computational efficiency of the optimization loop is described with reference to FIG. 6.

[0067] 6 shows that an exemplary first stage methodology for improving the computational efficiency of the optimization loop can include step 202, in which a digital camera (e.g., positioned at the distal end of the endoscope) captures a first image. An exemplary second stage can include step 204, in which a "feature" map is calculated, as discussed above in the previous section, "High Performance Feature Maps for Registration." The output of feature map calculation 204 can include a grid containing numerical representations of feature elements (e.g., pixel clusters) or similar image features positioned closest to the central axes of surface vessels, as discussed above.

[0068] An exemplary next step in the methodology can include an initial estimate 208 of the depths of various objects in the first image. This step can provide a preliminary approximation of the three-dimensional surface of the first image, such that the calculation of the initial depth estimate can incorporate the six degrees of freedom properties described herein. The preliminary approximation can include utilizing intensity data to calculate a coarse approximation of the three-dimensional surface of the first image.

[0069] An exemplary next step in the methodology may include collecting 210 subsequent image frames from a digital camera. Similar to that described above with respect to the first image, an exemplary next step may include computing 214 a "feature" map of the second image, as discussed above in "High-Performance Feature Maps for Registration." The output of computing 214 the feature map may include a grid containing numerical representations of feature elements (e.g., pixel clusters) located closest to the central axis of the vessel lumen in the second image.

[0070] An exemplary next step 216 in the methodology can include chamfer-matching pixel clusters of the first image feature map with pixel clusters of the second image feature map. This step can include the chamfer-matching process described above in the "High-Performance Feature Map for Registration" section. Furthermore, this step can include aligning the first image with the second image using four degrees of freedom, including the most likely motion of the endoscopic camera 124, such as endoscope advancement / retraction, rotation, and bending (changes in bending are motion parameters), and / or the current bending angle (the bending angle is an endoscope state estimate that is adjusted from frame to frame). It can be appreciated that this step can provide an initial approximation of the three-dimensional surface across the frames.

[0071] An exemplary next step in the methodology may include step 218, which evaluates the confidence of the initial registration between the first and second images calculated in step 216. In some examples, this evaluation may be performed using a threshold on the value of the optimized cost function. For example, this threshold may include the total Chamfer distance determined in the "High Performance Feature Map for Registration" section. As shown in process 200, if the minimum threshold is not met, a new "scene" may be initiated, whereby the new scene reinitializes the data structure while maintaining the surface shape and feature cluster correspondences. However, instead of starting a new scene, it is contemplated that the process may simply reject the current frame and proceed to the next frame until a predetermined maximum number of frames have been dropped, triggering a scene reset.

[0072] After assessing the reliability of the initial registration (and if the assessment satisfies a predetermined threshold), an exemplary next step 220 of the methodology can include repeating the registration process for each feature pixel cluster using only pixels positioned in the immediate vicinity of each cluster. The degrees of freedom explored in the registration can be set according to one of three strategies. A first strategy can include a "fixed strategy," which may allow for the use of many degrees of freedom (e.g., six degrees of freedom) with limited cluster size and robust initial estimates. Another strategy can include a "context" strategy, which adapts the degrees of freedom to the expected cluster distortion as a result of the registration obtained from the initial registration step 216 (including the current endoscope flexion angle estimate). For example, if the initial registration step 216 results in a dominant change in flexion angle, a two-degree-of-freedom registration can simply utilize image translation in the X and Y directions. Further, another strategy may include an "adaptive" strategy, whereby alignments with additional degrees of freedom are used to iterate over alignments with fewer degrees of freedom based on a registration quality metric (e.g., an optimization cost function). A well-parameterized alignment, when started from accurate initial estimates, may converge much faster than an alignment using more parameters. The resulting alignment (using any of the strategies described above) may be reducible as a set of independent affine cluster alignments using an interpolation strategy.

[0073] An exemplary next step in the methodology can include step 222 of evaluating the confidence of each cluster alignment against a threshold, whereby the number of clusters passing the threshold can itself be compared to a threshold. It can be appreciated that a given number of high-quality cluster matches is presumed to indicate a reliable alignment within the scene. As shown in FIG. 6, if a minimum threshold is not met, if the number of omitted frames exceeds a maximum value, or alternatively, if the number of omitted frames is equal to or less than the maximum threshold, processing can abandon the current frame and proceed to the next frame to begin computing feature maps (as described above with respect to step 214).

[0074] If the thresholds for the individual cluster alignments are met in the evaluation step 222 described above, an exemplary next step may include step 224 of combining the cluster alignments to determine the most likely camera pose changes that occur between frames and the resulting new endoscope bending angle estimate.

[0075] An exemplary next step in the methodology may include calculating a depth estimate for each cluster center 224. The depth estimate for each cluster center may be converted to a three-dimensional surface position.

[0076] An exemplary final step in the methodology may include estimating depth maps between and across clusters and incorporating them into the scene plane description.

[0077] After the cluster depth maps have been estimated, the missing information needed to accurately represent the three-dimensional surface of the image can be approximated and filled in between cluster centers using image intensity data in step 228. One possible methodology is to parameterize this 2D interpolation and extrapolation using the sum of intensity gradients along paths separating clusters, which assumes that depth changes occur primarily in areas of varying image intensity. Finally, when a new image is acquired, the process can be repeated, beginning with step 214, to compute a feature map for the new image.

[0078] It should be understood that the present disclosure is in many respects merely illustrative. Changes may be made in details, particularly in matters of shape, size, and arrangement of steps, without exceeding the scope of the present disclosure. This may include, to the extent appropriate, using any of the features of one exemplary embodiment in other embodiments. The scope of the present disclosure, of course, is defined by the language in which the appended claims are expressed. [Explanation of symbols]

[0079] 138 First Pixel Grid 140 Second Pixel Grid 142, 144, 152 feature pixels 146, 148 Numerical grid 150 Hybrid Numerical Grid

Claims

1. 1. A method for combining multiple images of a body structure, comprising: capturing a first input image with a digital camera positioned at a first location at a first time; representing the first image using a first plurality of pixels; capturing a second input image with the digital camera positioned at a second location at a second time; representing the second image using a second plurality of pixels; generating a first feature distance map of the first input image; generating a second feature distance map of the second input image; calculating a change in position of the digital camera between the first point in time and the second point in time; generating a three-dimensional surface approximation of the body structure using the first feature distance map, the second feature distance map, and the positional changes of the digital camera; A method comprising:

2. the first image corresponds to the body structure; generating the first feature distance map includes selecting one or more pixels from the first plurality of pixels; The method of claim 1 , wherein the one or more pixels from the first plurality of pixels are selected based on their proximity to features of the first image.

3. 3. The method of claim 1 or 2, wherein the one or more pixels from the first plurality of pixels are selected based on their proximity to a central longitudinal axis of the body structure.

4. the second image corresponds to the body structure; generating the second feature distance map includes selecting one or more pixels from the second plurality of pixels; The method of any one of claims 1 to 3, wherein the one or more pixels from the second plurality of pixels are selected based on their proximity to features of the second image.

5. The method of any one of claims 1 to 4, wherein the one or more pixels from the second plurality of pixels are selected based on their proximity to a central longitudinal axis of the body structure.

6. 6. The method of claim 1, wherein generating the first feature distance map comprises calculating a straight-line distance from a portion of the body structure to one or more pixels of the first image.

7. 7. The method of claim 1, wherein generating the second feature distance map comprises calculating a straight-line distance from a portion of the body structure to the one or more pixels of the second image.

8. The method of any one of claims 1 to 7, wherein generating the first feature distance map comprises assigning a numerical value to the one or more pixels of the first plurality of pixels.

9. The method of any one of claims 1 to 8, wherein generating the second feature distance map comprises assigning a numerical value to the one or more pixels of the second plurality of pixels.

10. the first plurality of pixels are arranged in a first coordinate grid; the second plurality of pixels are arranged in a second coordinate grid; A method according to any preceding claim, wherein the coordinate locations of the first plurality of pixels are at the same respective locations as the coordinate locations of the second plurality of pixels.

11. The method of any one of claims 1 to 10, further comprising generating a hybrid feature distance map by aligning the first feature distance map with the second feature distance map using one or more degrees of freedom corresponding to digital camera motion configuration parameters.

12. The method of claim 11 , wherein the digital camera motion configuration parameters include one or more of a positional change and a rotational change of the digital camera along an endoscope axis.

13. The method of claim 11 , further comprising assessing the reliability of the hybrid feature distance map by comparing distance values ​​calculated in the hybrid feature distance map to a distance threshold.

14. 1. A method for combining multiple images of a body structure, comprising: acquiring a first image at a first time point and acquiring a second image at a second time point using an image capture device, wherein the image capture device is positioned at a first location when capturing the first image at the first time point and the image capture device is positioned at a second location when capturing the second image at the second time point, the second time point occurring after the first time point; representing the first image using a first plurality of pixels; representing the second image using a second plurality of pixels; generating a first feature distance map of the first input image; generating a second feature distance map of the second input image; calculating a change in position of the digital camera between the first point in time and the second point in time; generating a three-dimensional surface approximation of the body structure using the first feature distance map, the second feature distance map, and the positional changes of the digital camera; A method comprising:

15. the first image corresponds to the body structure; generating the first feature distance map includes selecting one or more pixels from the first plurality of pixels; The method of claim 14 , wherein the one or more pixels from the first plurality of pixels are selected based on their proximity to features of the first image that correlate to body structures.