System for generating a fused image from a plurality of images acquired from an endoscope and method for combining a plurality of images acquired from an endoscope
An image processing algorithm for medical imaging combines weighted metrics to enhance clarity and accuracy by fusing images with varying exposures, addressing issues of overexposure, underexposure, and motion blur in medical procedures.
Patent Information
- Application Number
- JP2023553251
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-03-03
- Filing Date
- 2022-03-02
- Publication Date
- 2025-07-24
- Estimated Expiration
- 2042-03-02
AI Technical Summary
Medical imaging during procedures like lithotripsy is hindered by complex exposure sequences and varying exposure conditions, leading to unclear images due to overexposed and underexposed regions, halo effects, and motion blur from fast-moving particles.
An image processing algorithm that combines multiple images using weighted metrics for contrast, saturation, exposure, and motion to generate a fused image, enhancing clarity and accuracy by preserving desirable image features while minimizing undesirable effects.
The algorithm effectively improves image clarity and accuracy by fusing multi-exposure images, maintaining details in overexposed, underexposed, and blurred regions with minimal degradation.
Smart Images

Figure 0007713022000002 
Figure 0007713022000003 
Figure 0007713022000004
Abstract
Description
Technical Field
[0001] [Cross - Reference to Related Applications] This application claims the benefit and priority of U.S. Provisional Patent Application No. 63 / 155,976, filed on March 3, 2021, the disclosure of which is incorporated herein by reference.
[0002] The present disclosure relates to image processing technology, and more specifically, to fusing a plurality of images captured during a medical procedure, whereby fusing a plurality of images includes combining selected portions of various weighted images having various contrast levels, saturation levels, exposure levels, and motion blur.
Background Art
[0003] Various medical device technologies are available to medical professionals for observing and imaging the internal organs and body systems of the human body. For example, in many medical fields, a physician may use a medical endoscope equipped with a digital camera to observe a part of the human body inside the body for examination, diagnostic purposes, and during treatment. For example, a physician may use a digital camera coupled to an endoscope to observe the treatment of kidney stones during a lithotripsy procedure.
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, during some part of a medical procedure, the images captured by a camera may encounter various complex exposure sequences and various exposure conditions. For example, during a lithotripsy procedure, a physician may observe a raw video stream captured by a digital camera positioned adjacent to a laser fiber used to crush a kidney stone. During the procedure, the physician's view of the kidney stone may become unclear due to the flashing of the laser and / or fast-moving kidney stone particles. Specifically, the raw images captured by the camera may include overexposed and / or underexposed regions. Further, portions of the image that include overexposed and / or underexposed regions may lose details of highlight and shadow regions and may also exhibit other undesirable effects such as a halo effect. Therefore, it is considered desirable to develop an image processing algorithm that enhances the images collected by the camera, thereby improving the clarity and accuracy of the field of view observed by the physician during a medical procedure. An image processing algorithm that utilizes image fusion to enhance multi-exposure images is disclosed.
Means for Solving the Problem
[0005] This disclosure provides designs, materials, manufacturing methods, and alternative uses for medical devices. An exemplary method of combining a plurality of images includes obtaining a first input image formed from a first plurality of pixels, wherein a first pixel of the first plurality of pixels includes a first characteristic having a first value. The method also includes obtaining a second input image formed from a second plurality of pixels, wherein a second pixel of the second plurality of pixels includes a second characteristic having a second value. The method also includes subtracting the first value from the second value to generate a motion metric and generating a weighted metric map of the second input image using the motion metric.
[0006] Instead of or in addition to any of the above-described embodiments, the method further comprises converting the first input image into a first grayscale image and converting the second input image into a second grayscale image.
[0007] Instead of or in addition to any of the above embodiments, the first characteristic is a first grayscale intensity value, and the second characteristic is a second grayscale intensity value.
[0008] Instead of or in addition to any of the above embodiments, the first input image is formed at a first time point, and the second input image is formed at a second time point that occurs after the first time point.
[0009] Instead of or in addition to any of the above embodiments, the first image and the second image have different exposures.
[0010] Instead of or in addition to any of the above embodiments, the first image and the second image are captured by a digital camera, and the digital camera is positioned at the same location when it captures the first image and the second image.
[0011] Instead of or in addition to any of the above embodiments, the first plurality of pixels are arranged in a first coordinate grid, the first pixel is positioned at a first coordinate location of the first coordinate grid, the second plurality of pixels are arranged in a second coordinate grid, the second pixel is positioned at a second coordinate location of the second coordinate grid, and the first coordinate location is at the same respective location as the second coordinate location.
[0012] Instead of or in addition to any of the above embodiments, the step of generating a motion metric further comprises the step of weighting the motion metric using a power function.
[0013] Instead of or in addition to any of the above embodiments, the method further comprises the step of generating a contrast metric, a saturation metric, and an exposure metric.
[0014] Instead of or in addition to any of the above embodiments, further comprising the step of multiplying a contrast metric, a saturation metric, an exposure metric, and a motion metric with each other to generate a weighted metric map.
[0015] Instead of or in addition to any of the above embodiments, further comprising the step of generating a fused image from the first image and the second image using the weighted metric map.
[0016] Instead of or in addition to any of the above embodiments, the step of generating a fused image from the first image and the second image using the weighted metric map further includes the step of normalizing the weighted metric map.
[0017] Another method of combining a plurality of images includes the steps of acquiring a first image at a first time point and a second image at a second time point using an endoscopic image capture device, the image capture device being positioned at the same location when it captures the first image at the first time point and the second image at the second time point, and the second time point occurring after the first time point. The method also includes the steps of converting the first input image into a first grayscale image, converting the second input image into a second grayscale image, and generating a motion metric based on the characteristics of the pixels of both the first grayscale image and the second grayscale image, the pixels of the first grayscale image having the same coordinate location within their respective images as the pixels of the second grayscale image. The method also includes the step of generating a weighted metric map using the motion metric.
[0018] Instead of or in addition to any of the above embodiments, the characteristic of the pixel of the first image is the first grayscale intensity value, and the characteristic of the pixel of the second image is the second grayscale intensity value.
[0019] Instead of or in addition to any of the above-described embodiments, the step of generating a motion metric further comprises the step of weighting the motion metric using a cost function.
[0020] Instead of or in addition to any of the above-described embodiments, the method further comprises the step of generating a contrast metric, a saturation metric, and an exposure metric based on the characteristics of the pixels of the second image.
[0021] Instead of or in addition to any of the above-described embodiments, the method further comprises the step of multiplying the contrast metric, the saturation metric, the exposure metric, and the motion metric with each other to generate a weighted metric map.
[0022] Instead of or in addition to any of the above-described embodiments, the method further comprises the step of normalizing the weighted metric map across the first image and the second image.
[0023] Instead of or in addition to any of the above-described embodiments, the method further comprises the step of generating a fused image from the first image and the second image using the normalized weighted metric map.
[0024] An exemplary system for generating a fused image from a plurality of images acquired from an endoscope includes a processor operably connected to the endoscope and a non-transitory computer-readable storage medium comprising code configured to execute a method for fusing images. The method comprises the step of acquiring a first input image formed from a first plurality of pixels from the endoscope, wherein a first pixel of the plurality of pixels comprises a first characteristic having a first value. The method also comprises the step of acquiring a second input image formed from a second plurality of pixels from the endoscope, wherein a second pixel of the second plurality of pixels comprises a second characteristic having a second value. The method also comprises the step of subtracting the first value from the second value to generate a motion metric for the second pixel. The method also comprises the step of generating a weighted metric map using the motion metric.
[0025] The above summary of some embodiments is not intended to describe each disclosed embodiment or any implementation of the present disclosure. The drawings and the following "Detailed Description of the Invention" more specifically illustrate these embodiments.
[0026] The present disclosure can be more fully understood by considering the following detailed description in conjunction with the accompanying drawings.
Brief Description of the Drawings
[0027]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
[0028] The present disclosure has room for various modifications and alternative forms. Some of them are shown as an example in the drawings and will be described in detail. However, it should be understood that the present disclosure is not intended to be limited to the specific embodiments described. Instead, the present disclosure is intended to cover all modifications, equivalents, and alternatives that fall within its spirit and scope.
Detailed Description of the Invention
[0029] For the terms defined below, unless otherwise defined in the claims and other parts of this specification, the following definitions shall apply.
[0030] In this specification, all numerical values are assumed to be modified by the term "about", whether explicitly indicated or not. The term "about" generally means a numerical range that is considered equivalent (e.g., having the same function or result) to the value enumerated by those skilled in the art. In many cases, the term "about" can include numbers rounded to the nearest significant digit.
[0031] The recitation of a numerical range by endpoints includes all numbers within that range (e.g., from 1 to 5 includes 1, 1.5, 2, 2.75, 3, 3.80, 4, and 5).
[0032] As used in this specification and the claims, the singular forms "a", "an", and "the" include both singular and plural referents unless the context clearly dictates otherwise. As used in this specification and the claims, the term "or" is generally used in the sense of "and / or" unless the context clearly dictates otherwise.
[0033] Note that references in this specification to "embodiments", "some embodiments", "other embodiments", etc. indicate that the embodiments being described can include one or more specific features, structures, or characteristics. However, such recitations do not necessarily mean that all embodiments include the specific features, structures, and / or characteristics. In addition, when a specific feature, structure, and / or characteristic is described in relation to one embodiment, it should be understood that such feature, structure, and / or characteristic can also be used in relation to other embodiments, whether explicitly described or not, unless the contrary is explicitly stated.
[0034] The following detailed description should be read with reference to the drawings in which like elements are given the same number in the various drawings. These drawings are not necessarily to scale and are illustrative of exemplary embodiments and are not intended to limit the scope of the present disclosure.
[0035] This specification describes an image processing method performed on an image collected by a medical device (e.g., an endoscope) during a medical procedure. Further, the image processing method described in this specification can include an image fusion method. Disclosed are various embodiments of an improved image fusion method that maintains desirable portions of a given image (e.g., such as edges, textures, saturated colors) while minimizing undesirable portions of the image (e.g., such as distorted images or artifacts from moving particles, laser blinking, halo effects). In particular, the various embodiments relate to a step of selecting desirable portions of a multi-exposure image and a step of generating a weighted metric map for the purpose of improving the overall resolution of a given image. For example, a fused image can be generated that depicts overexposed regions, underexposed regions, and / or blurred regions created by moving particles with only a slight degradation.
[0036] The following describes a system for generating a fused image as a result of combining multi-exposure images. FIG. 1 shows an exemplary endoscope system that can be used in combination with other aspects of the present disclosure. In some embodiments, the endoscope system can include an endoscope 10. The endoscope 10 can be dedicated to specific endoscopy procedures such as ureteroscopy, lithotripsy, or can be a general-purpose device for a wide variety of procedures. In some embodiments, the endoscope 10 can include a handle 12 and an elongated shaft 14 extending distally therefrom, and the handle 12 includes a port configured to receive a laser fiber 16 extending through the elongated shaft 14. As illustrated in FIG. 1, the laser fiber 16 can be passed through the working channel of the elongated shaft 14 through a connector 20 (e.g., a Y-shaped connector) or other port positioned along the distal region of the handle 12. It can be recognized that the laser fiber 16 can deliver laser energy to a target site in the body. For example, during a lithotripsy procedure, the laser fiber 16 can deliver laser energy to crush kidney stones.
[0037] Furthermore, the endoscope system shown in FIG. 1 can include a camera and / or lens positioned at the distal end of the elongated shaft 14. The elongated shaft and / or the camera / lens can have deflection and / or articulation in one or more directions for observing the patient's anatomical structure. In some embodiments, the endoscope 10 can be a ureteroscope. However, in addition to or instead of a ureteroscope, other medical devices such as different endoscopes or related systems can be used. Additionally, in some embodiments, the endoscope 10 can be configured to deliver fluid through the elongated shaft 14 from a fluid management system to the treatment site. The elongated shaft 14 can include one or more working lumens for receiving fluid flow and / or other medical devices therethrough. In some embodiments, the endoscope 10 can be connected to a fluid management system through one or more supply lines.
[0038] In some embodiments, the handle 12 of the endoscope 10 can include a plurality of elements configured to facilitate an endoscopic procedure. In some embodiments, the cable 18 can extend from the handle 12 and be configured for attachment to an electronic device (not depicted), such as a computer system, a console, a microcontroller, to supply power, to analyze endoscopic data, to control an endoscopic intervention, or to perform other functions. In some embodiments, the electronic device to which the cable 18 is connected can have the function of recognizing data and exchanging it with other endoscopic accessories.
[0039] In some embodiments, an image signal can be transmitted through the cable 18 from a camera at the distal end of the endoscope and displayed on a monitor. For example, as described above, the endoscopic system shown in FIG. 1 can include at least one camera for presenting a visual feed to a user on a display screen of a computer workstation. Although not explicitly shown, it can be appreciated that the elongate shaft 14 can include one or more working lumens through which a data transmission cable (such as an optical fiber cable, an optical cable, a connector, a wire, etc.) can extend. The data transmission cable can be connected to the camera described above. Further, the data transmission cable can be coupled to the cable 18. Further, the cable 18 can be coupled to a computer processing system and a display screen. The image collected by the camera can be transmitted through the data transmission cable positioned within the elongate shaft 14, whereby the image data subsequently passes through the cable 18 and reaches the computer processing workstation.
[0040] In some embodiments, the workstation can include, among other features, a touch panel computer, an interface box for receiving a wired connection (e.g., cable 18), a cart, and a power supply. In some embodiments, the interface box can be configured to have a wired or wireless communication connection with a controller of the fluid management system. The touch panel computer can include at least a display screen and an image processor, and in some embodiments, can include and / or define a user interface. In some embodiments, the workstation can be a multi-purpose component (e.g., used in more than one procedure), whereas the endoscope 10 can be a single-use device, although this is not required. In some embodiments, the workstation can be excluded, and the endoscope 10 can be electronically directly coupled to the controller of the fluid management system.
[0041] FIG. 2 illustrates a plurality of images 100 sequentially captured by a camera over a time interval. It can be appreciated that the images 100 can represent an image sequence captured during a medical procedure. For example, the images 100 can represent an image sequence captured during a lithotripsy procedure in which a physician uses a laser fiber to treat a kidney stone. The images 100 can be collected, for example, by an image processing system that can include a computer workstation, a laptop, a tablet, or other computing platform that includes a display that enables a physician to visualize the procedure in real time. It can further be appreciated that during real-time collection of the images 100, the image processing system can be designed to process and / or enhance a given image based on the fusion of one or more subsequently acquired images. The enhanced images can subsequently be visualized by a physician during the procedure.
[0042] As described above, it can be recognized that the image 100 illustrated in FIG. 2 can include an image captured using an endoscopic device (i.e., an endoscope) during a medical procedure (e.g., during a lithotripsy procedure). Further, it can be recognized that the image 100 illustrated in FIG. 2 can represent an image sequence 100 captured over time. For example, the image 112 can represent an image captured at time point T1, whereas the image 114 can represent an image captured at time point T2, and thus, the image 114 captured at time point T2 occurs after the image 112 captured at time point T1. Further, the image 116 can represent an image captured at time point T3, and thus, the image 116 captured at time point T3 occurs after the image 114 captured at time point T2. This sequence can proceed with respect to the images 118, 120, and 122 respectively acquired at time points T4, T5, and T6, where in this case, time point T4 occurs after time point T5, time point T5 occurs after time point T4, and time point T6 occurs after time point T5.
[0043] It can further be recognized that the image 100 can be captured during a live event by a camera of an endoscope device having a fixed position. For example, the image 100 can be captured during a medical procedure by a digital camera having a fixed position. Thus, it can further be recognized that although the field of view of the camera remains constant during the procedure, the images generated during the procedure may vary due to the dynamic nature of the procedure being captured by the images. As a simple example, the image 112 can represent an image acquired just before the laser fiber emits laser energy to pulverize a kidney stone. Further, the image 114 can represent an image acquired immediately after the laser fiber emits laser energy to pulverize a kidney stone. Since the laser emits high-intensity flashing light, it can be recognized that the image 112 captured just before the laser emits light may be very different in terms of saturation, contrast, exposure, etc. compared to the image 114. In particular, the image 114 may include undesirable characteristics compared to the image 112 due to the sudden emission of the laser light. In addition, it can further be recognized that after the laser has imparted energy to the kidney, various particles from the kidney may quickly move through the field of view of the camera. These fast-moving particles may manifest as local regions having undesirable image features (such as overexposure, underexposure, etc.) through a series of images over time.
[0044] It can be recognized that a digital image (such as any one of the plurality of images 100 shown in FIG. 1) can be represented as a set of pixels (or individual picture elements) arranged in a two-dimensional grid and represented using squares. Further, each individual pixel constituting the image can be defined as the smallest information item within the image. Each pixel is a small sample of the original image, and generally, the more samples there are, the more accurate the representation of the original object is given.
[0045] For example, FIG. 3 shows the exemplary digital image 112 and digital image 114 described in FIG. 2, represented as a set of pixels arranged in a two-dimensional grid. For purposes of simplification, each grid of image 112 and image 114 is sized 18x12. In other words, the two-dimensional grid for image 112 / 114 includes 18 pixel columns extending vertically and 12 pixel rows extending horizontally. It can be recognized that the size of the images represented in FIG. 3 is exemplary. The size (total number of pixels) for digital images can vary. For example, common sizes for digital images can include images having 1080 pixel columns × 720 pixel rows (e.g., a frame size of 1080×720).
[0046] It can be recognized that individual pixel locations can be identified by coordinates (X,Y) on a two-dimensional image grid. Further, comparison of adjacent pixels within a given image can provide desirable information regarding which portions of a given image an algorithm may attempt to preserve when performing image processing (e.g., image fusion). For example, FIG. 3 illustrates an image feature 170 generally centered at the pixel location (16,10) in the lower right corner of the image. It can be recognized that the pixels constituting feature 170 are substantially darker compared to the pixels surrounding feature 170. Thus, these pixels representing feature 170 can have a high contrast value compared to the pixels surrounding feature 170. This high contrast can be an example of valuable information in the image. Other valuable information can include high saturation, edges, and texture.
[0047] Furthermore, it can be further recognized that the information represented by a pixel at a given coordinate can be compared across multiple images. Comparison of pixels at the same coordinate across multiple images can provide desirable information regarding portions of a given image that an image processing algorithm may attempt to discard when performing image processing (e.g., image fusion).
[0048] For example, image 114 in FIG. 3 illustrates that the feature portion 170 can represent moving particles captured over time by two images. Specifically, FIG. 3 illustrates the location where the feature portion 170 (shown at the lower right corner of image 112) has moved to the upper left corner of image 114. Note that images 112 / 114 were captured pixel by pixel from the same image capture location and map using the same image capture device (e.g., a digital camera). Subtraction of the two images at the pixel level can generate a new image in which all pixel values except at locations (16, 10) and (5, 4) are close to zero, and this image can provide information that can be utilized by the algorithms of the image processing system to reduce the undesirable effects of the moving particles. For example, the image processing system can generate a fused image of image 112 and image 114, in which case the fused image retains the desirable information of each image while discarding the undesirable information.
[0049] The basic mechanism for generating a fused image can be described as multiple pixels with different exposures within the input images being weighted according to various "metrics" such as contrast, saturation, exposure, and motion. Using these metric weightings, it is possible to determine to what extent a given pixel within the image to be fused with one or more other images will contribute to the final fused image.
[0050] Methods for enabling the fusion of multiple differently exposed images (1...N) of an event (e.g., a medical procedure) into a single fused image by an image processing algorithm are disclosed in FIGS. 4 - 5. For purposes of simplicity, the image processing algorithm illustrated in FIGS. 4 - 5 will be described using as an example the live image 114 and its "immediately preceding" image 112 (both shown in FIGS. 2 - 3). However, it can be recognized that the image processing algorithm can utilize any number of images captured during an event to generate one or more fused images.
[0051] The fusion process can be performed by an image processing system and output to a display device. In this case, the final fused image must be recognized as maintaining the desirable features of the image metrics (contrast, saturation, exposure, motion) from the input images from 1 to N.
[0052] Figures 4 - 5 illustrate exemplary steps for generating an exemplary weight map for a fused image. An exemplary first step can include step 128 of selecting an initial input image. For the purposes of the explanation herein, the exemplary input image 114 will be described as the initial input image that will be fused with the immediately preceding image 112 (note that with reference to the above explanation, image 114 was captured after image 112).
[0053] After selecting the input image 114, the individual contrast metric, saturation metric, exposure metric, and motion metric will be calculated at each of the individual pixel coordinates that make up the image 114. In this exemplary embodiment, there are 216 individual pixel coordinates that make up the image 114 (referring back to FIG. 3, the exemplary image 114 includes 216 individual pixels in an 18x12 grid). As will be explained in more detail below, each of the contrast metric, saturation metric, exposure metric, and motion metric will be multiplied with each other at each pixel location (i.e., pixel coordinate) with respect to the image 114, thereby generating a "pre - weight map" with respect to the image 114 (for reference purposes, this multiplication step is illustrated by the text box 146 in FIG. 4). However, before generating the pre - weight map with respect to the image, it is necessary to calculate the individual contrast metric, saturation metric, exposure metric, and motion metric for each pixel location. The following explanation describes the calculation of the contrast metric, saturation metric, exposure metric, and motion metric for each pixel location.
[0054] Calculation of the Contrast Metric
[0055] The calculation of the contrast metric is represented by text box 132 in FIG. 4. To calculate the contrast metric for each pixel location, a Laplacian filter can be applied to the image to generate a grayscale version of the image. After the grayscale version of the image is generated, the absolute value of the filter response can be obtained. Note that the contrast metric can assign high weights to important elements such as edges and textures in the image. In this specification, the contrast metric (calculated for each pixel) may be represented as (W c ).
[0056] Calculation of the saturation metric
[0057] The calculation of the saturation metric is represented by text box 134 in FIG. 4. To calculate the saturation metric for each pixel location, the standard deviation is obtained among the R, G, and B channels (at each given pixel). A lower standard deviation means that the R value, G value, and B value are close to each other (e.g., the pixel is somewhat gray). It can be recognized that gray pixels may not contain as valuable information as other colors such as red (e.g., due to the human anatomical structure containing reddish colors). However, pixels tend to have a reddish, greenish, or bluish tint when they have a higher standard deviation, and these tints contain more valuable information. Note that saturated colors are desirable and make the image look sharper. In this specification, the saturation metric (calculated for each pixel) may be represented as (W s ).
[0058] Calculation of the exposure metric
[0059] The calculation of the exposure metric is represented in text box 136 of FIG. 4. The raw light intensity for a given pixel channel can present an indication of how well the pixel is exposed. Each individual pixel can include three channels separated with respect to red, green, and blue, and it should be noted that all of these channels can be displayed at a certain “intensity”. In other words, the total intensity of the pixel location (i.e., coordinates) can be the sum of the red channel (at a certain intensity), the green channel (at a certain intensity), and the blue channel (at a certain intensity). These RGB light intensities at the pixel location can fall within the range of 0 to 1, where a value close to zero indicates under - exposure and a value close to 1 indicates over - exposure. It is desirable to maintain an intensity that is not at the ends of the spectrum (not close to 0 or 1). Thus, to calculate the exposure metric for a given pixel, each channel is weighted based on how close the channel is to the median of a unimodal distribution with proper symmetry. For example, each channel can be weighted based on how close the channel is to a given value on a Gaussian curve or any other type of distribution. After determining the weighting values for each RGB color channel at an individual pixel location, these three weighting values are multiplied together to obtain the total exposure metric for the pixel. In this specification, the exposure metric (calculated for each pixel) may be represented as (W e ).
[0060] Calculation of the motion metric
[0061] The motion metric calculation is represented by text box 138 in FIG. 4. A more detailed discussion of the motion metric calculation is provided in the block diagram shown in FIG. 6. The motion metric is designed to suppress distraction caused by flying particles, laser flicker, or dynamic fast changing scenes that occur over multiple frames. The motion metric for a given pixel location is calculated using an input image (e.g., 114) and the image immediately preceding it (e.g., image 112). An exemplary first step in calculating the motion metric for each pixel location in image 114 is step 160 of converting all colored pixels in image 114 to grayscale values. For example, it can be appreciated that in some cases, raw data for each pixel location in image 114 can be represented as integer data, in which case each pixel location can be converted to a grayscale value between 0 and 255. However, in other examples, the raw integer values can be converted to other data types. For example, the pixel grayscale integer values can be converted to floating point types, in which case each pixel can be assigned a grayscale value between 0 and 1, followed by dividing by 255. In these examples, completely black pixels may be assigned a value of 0, and completely white pixels may be assigned a value of 1. Additionally, while the above examples have described representing the raw image data as integer data or floating point data, it can be appreciated that the raw image data may be represented as any data type.
[0062] After converting the exemplary frame 114 to grayscale values, an exemplary second step in calculating motion metrics for each pixel location in the image 114 is step 162 of converting all colored pixels in the frame 112 to grayscale values using the same methodology described above.
[0063] After converting each pixel location of frame 114 and frame 112 to grayscale, an exemplary third step 164 in the stage of calculating a motion metric for each pixel location of image 114 is the step of subtracting the grayscale value of each pixel location of image 112 from the grayscale value of the corresponding pixel location of image 114. For example, if the grayscale value of the pixel coordinates (16, 10) of image 114 (shown in Figure 3) is equal to 0.90 and the grayscale value of the corresponding grayscale pixel coordinates (16, 10) of image 112 (shown in Figure 3) is equal to 0.10, then the motion subtraction value for pixel (16, 10) of image 114 is equal to 0.80 (0.90 minus 0.10). This post-subtraction motion value (e.g., 0.80) will be stored as a representative motion metric value for pixel (16, 10) within image 114. As explained above, the motion metric values will be used to generate a fused image based on the algorithm explained below. Further, it can be recognized that the subtraction value may be negative, which can indicate that a distorted image or artifact is present within image 112. Negative values will be set to zero to facilitate weight coefficient calculation.
[0064] For a fast moving object, since the change in pixel grayscale value becomes dramatic (e.g., when a fast moving object moving through multiple images is captured, the grayscale color of the individual pixels defining the object will change dramatically from image to image), it can be recognized that the post-subtraction motion value for a given pixel coordinate may be large. Conversely, for a slow moving object, since the change in pixel grayscale value becomes more gradual (e.g., when a slow moving object moving through multiple images is captured, the grayscale color of the individual pixels defining the object will only change slowly from image to image), the motion value for a given pixel coordinate may be small.
[0065] After the post-subtraction motion value is calculated for each pixel location in the most recent image (e.g., image 114), an exemplary fourth step 166 in the stage of calculating the motion metric for each pixel location in image 114 is the step of weighting each post-subtraction motion value for each pixel location based on how close the value is to zero. One weight calculation using a power function is shown in Equation 1 below. (1)W m =(1 - [post-subtraction motion value])^100
[0066] Other weight calculations are possible and are not limited to the power function described above. Any function that makes the weight close to 1 for a zero motion value and rapidly transitions to 0 as the motion value increases is possible. For example, it is considered possible to use the following weight calculation shown in Equation 2. (2)W m = JPEG0007713022000001.jpg12150
[0067] These values are the weighted motion metric values for each pixel in image 114 and may be represented as (W m ).
[0068] For pixels representing stationary objects, the post-subtraction motion value is close to zero, and thus it can be recognized that W m is close to 1. Conversely, for moving objects, the post-subtraction motion value is close to 1, and thus W m is close to 0. Returning to the example of the fast-moving dark circle 170 moving through images 114 / 112, pixel (16, 10) rapidly changes from a darker color (e.g., a grayscale value of 0.90) in image 112 to a lighter color (e.g., a grayscale value of 0.10) in image 114. The resulting post-subtraction motion value is equal to 0.80 (close to 1), and the weighted motion metric for pixel location (16, 10) in image 114 is (1 - 0.80)^100, i.e., very close to 0.
[0069] W for the immediately previous image (e.g., image 112) mNote that it will be set to 1. In this way, in the area of the stationary object, W m is close to 1 for both Image 112 and Image 114, and thus has the same weight with respect to the motion metric. The final weight map of the stationary object will depend on the contrast metric, the exposure metric, and the saturation metric. In the area of the moving object or the laser on-off, W m remains 1 unchanged, but W m for Image 114 approaches 0, and thus the final weight of these areas in Image 114 becomes much smaller than those in Image 112. A small weight value will result in discarding the corresponding pixel of the frame from the fused frame, and a relatively large weight value will bring the corresponding pixel of the frame into the fused image. In this way, the distorted image or artifacts will be removed from the final fused image.
[0070] FIG. 4 further illustrates that after the contrast metric (W c ), the saturation metric (W s ), the exposure metric (W e ), and the motion metric (W m ) are calculated for each pixel of Image 114, a pre-weight map for Image 114 can be generated by multiplying all four metric values (for a given pixel) with each other at each individual pixel location in Image 114 (this calculation is shown in Equation 3 below) 146. As shown in FIG. 4, Equation 3 is executed at each individual pixel location (e.g., (1,1), (1,2), (1,3) …, and so on) for all 216 exemplary pixels in Image 114 to generate the pre-weight map. (3)W=(W c ) * (W s ) * (W e ) * (W m )
[0071] Figure 4 illustrates that after a prior weight map for an exemplary image (e.g., Image 114) is generated, an exemplary next step 150 can include normalizing the weight maps across multiple images. The purpose of this step can be to make the sum of the weighted values for each pixel location (within any two images being fused) equal to 1. Expressions for normalizing the pixel location (X,Y) for exemplary "Image 1" and exemplary "Image 2" are shown in Expressions 4 and 5 below. (4)W 画像1 (X,Y)=W 画像1 (X,Y) / (W 画像1 (X,Y)+W 画像2 (X,Y)) (5)W 画像2 (X,Y)=W 画像2 (X,Y) / (W 画像1 (X,Y)+W 画像2 (X,Y))
[0072] As an example, assume two images, each having a prior weight map with a pixel location (14,4), i.e., Image 1 and Image 2. In this case, the prior weight value for pixel (14,4) of Image 1 is = 0.05, and the prior weight value for pixel (14,4) of Image 2 is = 0.15. The normalized value for pixel location (14,4) for Image 1 is shown in Expression 6 below, and the normalized value for pixel location (14,4) for Image 2 is shown in Expression 7 below. (6)W 画像1 (14,4)=0.05 / (0.05 + 0.15)=0.25 (7)W 画像2 (14,4)=0.15 / (0.05 + 0.15)=0.75 Note that as explained above, the sum of the normalized weighted values of Image 1 and Image 2 at pixel location (14,4) is equal to 1.
[0073] FIG. 5 illustrates an exemplary next step in generating a fused image from two or more exemplary images (e.g., image 112 / 114) (shown in FIGS. 4 - 5, and in the algorithms described above, at this point, each exemplary image 112 / 114 can have the original input image and the normalized weight map). The exemplary step can include step 152 of generating a smoothed pyramid from the normalized weight map (for each image). The smoothed pyramid can be calculated as follows. G1 = W G2 = downsample(G1, filter) G3 = downsample(G2, filter) The above calculation continues until the size of Gx is smaller than the size of the smoothing filter which is a low - pass filter (e.g., Gaussian filter).
[0074] FIG. 5 shows that the exemplary step may further include step 154 of generating an edge map from the input image (for each image). The edge map contains texture information at various scales of the input image. For example, let the input image be denoted as "I". The edge map at level x can be denoted as Lx and is calculated in the following steps. I1 = downsample(I, filter) L1 = I - upsample(I1) I2 = downsample(I1, filter) L2 = I1 - upsample(I2) I3 = downsample(I2, filter) L3 = I2 - upsample(I3) This calculation continues until the size of Ix is smaller than the size of the edge filter which is a high - pass filter (e.g., Laplacian filter).
[0075] Figure 5 illustrates another exemplary next step in generating a fused image from two or more exemplary images (e.g., image 112 / 114). For example, for each input image, the edge map (described above) can be multiplied by the smoothing map (described above). The maps obtained as a result of multiplying the edge map by the smoothing map for the images can be added together to form a fusion pyramid. In Figure 5, this step is represented by text box 156.
[0076] Figure 5 shows the last example in the stage of generating a fused image, as represented by text box 158 described therein. This stage includes the stage of performing subsequent calculations on the fusion pyramid. IN = LN RN = High-resolution processing(IN, filter) IN-1 = LN-1 + RN RN-1 = High-resolution processing(IN-1, filter) IN-2 = LN-2 + RN-1 RN-2 = High-resolution processing(IN-2, filter) IN-3 = LN-3 + RN-2 This calculation continues until the calculated I1 reaches the L1 level which is the final image.
[0077] It must be understood that the disclosure of the present invention is merely exemplary in many respects. Changes can be made without exceeding the scope of the disclosure of the present invention, particularly with regard to details, especially the shape, size, and arrangement of the steps. Such changes can include the use in any other embodiments of the features of an exemplary embodiment used, as long as appropriate. The scope of the disclosure of the present invention is, of course, defined by the language expressing the claims.
Explanation of Reference Numerals
[0078] Step of selecting 128 frames N. Step of selecting individual pixel (X, Y) from frame 130 N 132 Contrast metric (W cThe step of calculating 134 saturation metric (W s The step of calculating The step of generating a pre-weight map for frame N
Claims
1. A system for generating a fused image from a plurality of images acquired from an endoscope, a processor operatively connected to the endoscope, a non-transitory computer-readable storage medium comprising code configured to execute a method for fusing images, comprising, the method comprising: obtaining a first input image formed from a first plurality of pixels from the endoscope, each of the first pixels of the first plurality of pixels including a first characteristic having a first value, the obtaining of the first input image; obtaining a second input image formed from a second plurality of pixels from the endoscope, each of the second pixels of the second plurality of pixels including a second characteristic having a second value, the second input image being formed at a second time after a first time, the obtaining of the second input image; subtracting the first value from the second value to generate a motion metric for the second pixel; calculating a post-subtraction motion value for each pixel location and setting negative values to zero, and then generating a weighted metric map using the motion metric by weighting each post-subtraction motion value for each pixel location based on how close this value is to zero; forming a fused image from the first input image and the second input image using the weighted metric map; A system characterized by comprising the above.
2. The method further comprises: converting the first input image into a first grayscale image; converting the second input image into a second grayscale image; The system according to claim 1.
3. The first characteristic is a first grayscale intensity value, The second characteristic is a second grayscale intensity value. The system according to claim 2.
4. The first plurality of pixels are arranged in a first coordinate grid, and the first pixel is positioned at a first coordinate location of the first coordinate grid, the second plurality of pixels are arranged in a second coordinate grid, and the second pixel is positioned at a second coordinate location of the second coordinate grid, The first coordinate location is at the same respective location as the second coordinate location. The system according to any one of claims 1 to 3.
5. The system according to any one of claims 1 to 3, wherein the step of generating the motion metric further comprises a step of weighting the motion metric using a power function.
6. The system according to any one of claims 1 to 3, wherein the method further comprises a step of generating a contrast metric, a saturation metric, and an exposure metric.
7. The system according to claim 6, wherein the method further comprises a step of multiplying the contrast metric, the saturation metric, the exposure metric, and the motion metric with each other to generate the weighted metric map.
8. The system according to any one of claims 1 to 3, wherein the step of generating a fused image from the first input image and the second input image using the weighted metric map further includes a step of normalizing the weighted metric map.
9. A method of combining a plurality of images acquired by an image capture device of an endoscope, wherein an image processing system obtains an image signal including a first input image at a first time point and a second input image at a second time point transmitted from the image capture device to the image processing system, the image capture device being positioned at the same location when capturing the first input image at the first time point and the second input image at the second time point, the second time point occurring after the first time point, the obtaining step; the image processing system converting the first input image into a first grayscale image; the image processing system converting the second input image into a second grayscale image; the image processing system generating a motion metric based on the grayscale intensity values of the pixels of both the first grayscale image and the second grayscale image, the pixels of the first grayscale image having the same coordinate location in their respective images as the pixels of the second grayscale image, the generating step including subtracting the grayscale value of each pixel location of the first input image from the grayscale value of each pixel location of the second input image corresponding to each pixel location of the first input image to obtain a post-subtraction motion value. After the image processing system calculates the post-subtraction motion value at each pixel location and sets negative values to zero, the image processing system weights each post-subtraction motion value at each pixel location based on how close this value is to zero, thereby generating a weighted metric map using the motion metric; The image processing system forms a fused image from the first input image and the second input image using the weighted metric map; A method characterized by comprising the above. **Claim 10** The method according to claim 9, wherein weighting the post-subtraction motion value includes the image processing system using a power function. **Claim 11** The method according to claim 9, further comprising the image processing system generating a contrast metric, a saturation metric, and an exposure metric. **Claim 12** The method according to claim 11, further comprising the image processing system multiplying the contrast metric, the saturation metric, the exposure metric, and the motion metric with each other to generate the weighted metric map. **Claim 13** The method according to any one of claims 9 to 12, wherein the step of forming a fused image from the first input image and the second input image using the weighted metric map further includes the image processing system normalizing the weighted metric map.
Citation Information
Patent Citations
Apparatus and method for generating high dynamic range image from which ghost blur is removed using multi-exposure fusion base
JP2013031174A
Imaging apparatus, control method of the same and program
JP2013106149A
Optical level adaptive filter and method
JP2019512178A
Method and apparatus for generating high dynamic range image
US20190318460A1
Medical system, information processing device and information processing method
WO2020045014A1