Method of image enhancement by interference subtraction in endoscopic surgery
By generating a weighted metric map through an image fusion algorithm, the problems of uneven image exposure and motion interference in medical surgery are solved, improving image clarity and accuracy and achieving higher observation quality.
Patent Information
- Application Number
- CN202280018899.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-03-03
- Filing Date
- 2022-03-02
- Publication Date
- 2026-03-03
- Estimated Expiration
- 2042-03-02
AI Technical Summary
During medical procedures, uneven exposure and motion interference can cause captured images to become blurry and lose detail, affecting the clarity and accuracy of the doctor's observation.
Image fusion algorithms generate weighted metric maps that combine contrast, saturation, exposure, and motion metrics from multiple images to produce a fused image. This reduces unwanted artifacts and halo effects, and improves the overall resolution of the image.
It enhances the clarity and accuracy of images during medical procedures, preserving the desired image portions while minimizing the effects of uneven exposure and motion interference.
Smart Images

Figure CN116964620B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims the benefit and priority of U.S. Provisional Patent Application Serial No. 63 / 155,976, filed March 3, 2021, the disclosure of which is incorporated herein by reference. Technical Field
[0003] This disclosure relates to image processing techniques, and more specifically to the fusion of multiple images captured during medical procedures, wherein fusing multiple images includes merging selected portions of various weighted images with different contrast levels, saturation levels, exposure levels and motion interference. Background Technology
[0004] Various medical device technologies are available for medical professionals to view and image the internal organs and systems of the human body. For example, medical endoscopes equipped with digital cameras can be used by doctors in many medical fields to view internal parts of the body for examination, diagnosis, and during treatment. For instance, a doctor can use a digital camera coupled to an endoscope to monitor the treatment of kidney stones during lithotripsy.
[0005] However, during certain parts of a medical procedure, images captured by a camera may undergo various complex exposure sequences and different exposure conditions. For example, during lithotripsy, a physician may view a live video stream captured by a digital camera positioned near the laser fiber used to break up kidney stones. During the procedure, the physician's view of the kidney stones may become blurred due to laser flicker and / or rapidly moving kidney stone particles. Specifically, the live image captured by the camera may include overexposed and / or underexposed areas. Furthermore, image portions including overexposed and / or underexposed areas may lose detail in highlight and shadow areas and may also exhibit other undesirable effects, such as halo effects. Therefore, it may be desirable to develop image processing algorithms that enhance images collected by the camera, thereby improving the sharpness and accuracy of the field of view observed by the physician during medical procedures. An image processing algorithm utilizing image fusion to enhance multi-exposure images is disclosed. Summary of the Invention
[0006] This disclosure provides designs, materials, manufacturing methods, and alternatives for medical devices. An example method for combining multiple images includes acquiring a first input image formed by a first plurality of pixels, wherein the first pixel of the plurality of pixels includes a first characteristic having a first value. The method also includes acquiring a second input image formed by a second plurality of pixels, wherein a second pixel of the second plurality of pixels includes a second characteristic having a second value. The method further includes subtracting the first value from the second value to generate a motion metric, and using the motion metric to generate a weighted metric map of the second input image.
[0007] Alternatively or in addition to any of the above embodiments, it may also include converting the first input image into a first grayscale image and converting the second input image into a second grayscale image.
[0008] Alternatively or in addition to any of the above embodiments, the first characteristic is a first grayscale intensity value, and the second characteristic is a second grayscale intensity value.
[0009] Alternatively or in addition to any of the above embodiments, the first input image is formed at a first time point, and the second input image is formed at a second time point that occurs after the first time point.
[0010] Alternatively or as an addition to any of the above embodiments, wherein the first image and the second image have different exposures.
[0011] Alternatively or in addition to any of the above embodiments, the first image and the second image are captured by a digital camera, and the digital camera is positioned at the same location when it captures the first image and the second image.
[0012] In an alternative to or in addition to any of the above embodiments, a first plurality of pixels are arranged in a first coordinate grid, and the first pixel is located at a first coordinate position in the first coordinate grid, and a second plurality of pixels are arranged in a second coordinate grid, and the second pixel is located at a second coordinate position in the second coordinate grid, and the first coordinate position and the second coordinate position are at the same corresponding position.
[0013] Alternatively or in addition to any of the above embodiments, generating the movement metric may also include using a power-weighted movement metric.
[0014] In addition to or as an alternative to any of the above embodiments, the generation of contrast, saturation, and exposure metrics is also included.
[0015] Alternatively or in addition to any of the above embodiments, it may also include multiplying the contrast metric, saturation metric, exposure metric, and motion metric together to generate a weighted metric graph.
[0016] Alternatively or in addition to any of the above embodiments, it may also include creating a fused image from the first image and the second image using a weighted metric map.
[0017] In addition to or as an alternative to any of the above embodiments, the use of a weighted metric map to create a fused image from the first image and the second image may further include a normalized weighted metric map.
[0018] Another method for combining multiple images includes acquiring a first image at a first time point and a second image at a second time point using an endoscopic image capture device, wherein the image capture device is positioned at the same location when capturing the first image at the first time point and the second image at the second time point, and wherein the second time point occurs after the first time point. The method also includes converting the first input image to a first grayscale image, converting the second input image to a second grayscale image, and generating a motion metric based on the characteristics of pixels in both the first and second grayscale images, wherein pixels in the first grayscale image and pixels in the second grayscale image have the same coordinate positions in their respective images. The method further includes using the motion metric to generate a weighted metric map.
[0019] Alternatively or in addition to any of the above embodiments, the characteristics of the pixels of the first image are first grayscale intensity values, and the characteristics of the pixels of the second image are second grayscale intensity values.
[0020] Alternatively or in addition to any of the above embodiments, generating the movement metric may also include using a power function to weight the movement metric.
[0021] In addition to or as an alternative to any of the above embodiments, it also includes generating contrast, saturation, and exposure metrics based on the characteristics of the pixels in the second image.
[0022] Alternatively or in addition to any of the above embodiments, it may also include multiplying the contrast metric, saturation metric, exposure metric, and motion metric together to generate a weighted metric graph.
[0023] In addition to or as an alternative to any of the above embodiments, a normalized weighted metric map across the first and second images may also be included.
[0024] Alternatively or in addition to any of the above embodiments, it may also include creating a fused image from the first image and the second image using a normalized weighted metric map.
[0025] An example system for generating a fused image from multiple images acquired by an endoscope includes a processor operatively connected to an endoscope and a non-transitory computer-readable storage medium including code configured to perform a method for fusing the image. The method includes acquiring a first input image from the endoscope, the first input image being formed from a first plurality of pixels, wherein the first pixel among the plurality of pixels includes a first characteristic having a first value. The method also includes acquiring a second input image from the endoscope, the second input image being formed from a second plurality of pixels, wherein a second pixel among the second plurality of pixels includes a second characteristic having a second value. The method further includes subtracting the first value from the second value to generate a motion metric for the second pixel. The method also includes using the motion metric to generate a weighted metric map.
[0026] The above overview of some embodiments is not intended to describe every disclosed embodiment or every implementation of this disclosure. The following drawings and detailed description illustrate these embodiments in more detail. Attached Figure Description
[0027] This disclosure can be understood more fully in consideration of the following specific embodiments in conjunction with the accompanying drawings, wherein:
[0028] Figure 1 This is a schematic diagram of an example endoscope system;
[0029] Figure 2 It shows a sequence of images collected by a digital camera over a period of time;
[0030] Figure 3 It shows Figure 2 The first and second images of the image set shown;
[0031] Figures 4-5 This is a block diagram of an image processing algorithm that uses weighted image graphs to generate a fused image from multiple original images;
[0032] Figure 6 This is a block diagram of an image processing algorithm used to generate a weighted moving metric map of an image.
[0033] While this disclosure is open to various modifications and alternatives, its details have been shown by way of example in the accompanying drawings and will be described in detail. However, it should be understood that this disclosure is not intended to be limited to the specific embodiments described. Rather, it is intended to cover all modifications, equivalents, and alternatives that fall within the spirit and scope of this disclosure. Detailed Implementation
[0034] For the terms defined below, these definitions shall apply unless otherwise specified in the claims or elsewhere in this specification.
[0035] All numerical values herein are assumed to be modified by the term "approximately," whether explicitly stated or not. The term "approximately" generally refers to a range of numbers that a person skilled in the art would consider equivalent to the stated value (e.g., having the same function or result). In many cases, the term "approximately" may include numbers rounded to the nearest significant digit.
[0036] Describing a numerical range by endpoints includes all numbers within that range (for example, 1 to 5 includes 1, 1.5, 2, 2.75, 3, 3.80, 4, and 5).
[0037] As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” include plural references unless otherwise clearly stated. As used in this specification and the appended claims, the term “or” is generally used in its sense to include “and / or” unless otherwise clearly stated.
[0038] Note that references to "embodiments," "some embodiments," "other embodiments," etc., in the specification indicate that the described embodiments may include one or more specific features, structures, and / or characteristics. However, this description does not necessarily mean that all embodiments include the specific features, structures, and / or characteristics. Furthermore, when a specific feature, structure, and / or characteristic is described in conjunction with an embodiment, it should be understood that such feature, structure, and / or characteristic may also be used with other embodiments, whether explicitly described or not, unless clearly stated to the contrary.
[0039] The following detailed description should be read with reference to the accompanying drawings, in which similar elements in different drawings are numbered the same. The drawings (which are not necessarily drawn to scale) depict illustrative embodiments and are not intended to limit the scope of this disclosure.
[0040] This document describes an image processing method performed on images acquired via medical devices (e.g., endoscopes) during medical procedures. Furthermore, the image processing method described herein may include an image fusion method. Various embodiments of an improved image fusion method are disclosed, which preserve desired portions of a given image (e.g., edges, textures, saturated colors, etc.) while minimizing unwanted portions of the image (e.g., artifacts from moving particles, laser flash, halo effects, etc.). Specifically, various embodiments target desired portions of selected multi-exposure images and generate a weighted metric map to improve the overall resolution of a given image. For example, a fused image may be generated, thereby representing overexposed and / or underexposed and / or blurred areas caused by moving particles with minimal degradation.
[0041] The following describes a system for combining multiple exposure images to generate a resulting fused image. Figure 1 An example endoscopic system that can be used in conjunction with other aspects of this disclosure is shown. In some embodiments, the endoscopic system may include an endoscope 10. The endoscope 10 may be dedicated to a specific endoscopic procedure, such as, for example, ureteroscopy, lithotripsy, etc., or may be a general-purpose device suitable for a variety of procedures. In some embodiments, the endoscope 10 may include a handle 12 and an elongated shaft 14 extending distally therefrom, wherein the handle 12 includes a port configured to receive a laser fiber 16 extending within the elongated shaft 14. Figure 1As shown, the laser fiber 16 can be fed into the working channel of the elongated shaft 14 via a connector 20 (e.g., a Y-connector) or another port located along the distal region of the handle 12. It is understood that the laser fiber 16 can deliver laser energy to a target site within the body. For example, during lithotripsy, the laser fiber 16 can deliver laser energy to break up kidney stones.
[0042] in addition, Figure 1 The illustrated endoscopic system may include a camera and / or lens positioned at the distal end of an elongated shaft 14. The elongated shaft and / or camera / lens may have deflection and / or articulation capabilities in one or more directions for viewing patient anatomy. In some embodiments, endoscope 10 may be a ureteroscope. However, other medical devices, such as different endoscopes or related systems, may be used in addition to or in place of a ureteroscope. Furthermore, in some embodiments, endoscope 10 may be configured to deliver fluid from a fluid management system to a treatment site via the elongated shaft 14. The elongated shaft 14 may include one or more working lumens for receiving fluid flow and / or other medical devices passing through it. In some embodiments, endoscope 10 may be connected to the fluid management system via one or more supply lines.
[0043] In some embodiments, the handle 12 of the endoscope 10 may include multiple elements configured to facilitate endoscopic surgery. In some embodiments, a cable 18 may extend from the handle 12 and be configured to attach to an electronic device (not shown), such as, for example, a computer system, console, microcontroller, etc., for providing power, analyzing endoscopic data, controlling endoscopic interventions, or performing other functions. In some embodiments, the electronic device to which the cable 18 is connected may have the capability to identify data and exchange data with other endoscopic accessories.
[0044] In some embodiments, image signals can be transmitted from a camera at the distal end of the endoscope via cable 18 for display on a monitor. For example, as described above, Figure 1 The endoscopic system shown may include at least one camera to provide visual feeds to a user on a display screen of a computer workstation. It will be understood that, although not explicitly shown, the elongated shaft 14 may include one or more working cavities within which a data transmission cable (e.g., fiber optic cable, optical fiber, connector, wire, etc.) may extend. The data transmission cable may be connected to the aforementioned camera. Furthermore, the data transmission cable may be coupled to cable 18. Further still, cable 18 may be coupled to a computer processing system and a display screen. Images collected by the camera can be transmitted via the data transmission cable positioned within the elongated shaft 14, whereby the image data is then transmitted to the computer processing workstation via cable 18.
[0045] In some embodiments, among other features, the workstation may include a touch panel computer, an interface box for receiving wired connections (e.g., cable 18), a trolley, and a power supply. In some embodiments, the interface box may be configured for wired or wireless communication with the controller of the fluid management system. The touch panel computer may include at least a display screen and an image processor, and in some embodiments, may include and / or define a user interface. In some embodiments, the workstation may be a multi-purpose component (e.g., for more than one type of surgery), while the endoscope 10 may be a single-purpose device, although this is not necessary. In some embodiments, the workstation may be omitted, and the endoscope 10 may be directly and electronically coupled to the controller of the fluid management system.
[0046] Figure 2 Multiple images 100 captured sequentially by a camera over a period of time are shown. It is understood that images 100 may represent a sequence of images captured during a medical procedure. For example, images 100 may represent a sequence of images captured during a lithotripsy procedure performed by a physician using laser fibers to treat kidney stones. It is further understood that images 100 may be collected by an image processing system, which may include, for example, a computer workstation, laptop computer, tablet computer, or other computing platform, including a display through which the physician can visualize the procedure in real time. During the real-time collection of images 100, the image processing system may be designed to process and / or enhance a given image based on the fusion of one or more images taken after the given image. The enhanced image can then be visualized by the physician during the procedure.
[0047] As discussed above, this is understandable. Figure 2 The image 100 shown may include images captured using an endoscopic device (i.e., an endoscope) during a medical procedure (e.g., during lithotripsy). Furthermore, it can be understood that... Figure 2 The image 100 shown can represent a sequence of images 100 captured over time. For example, image 112 can represent an image captured at time point T1, and image 114 can represent an image captured at time point T2, thus image 114 captured at time point T2 occurs after image 112 captured at time point T1. Similarly, image 116 can represent an image captured at time point T3, thus image 116 captured at time point T3 occurs after image 114 captured at time point T2. This sequence can be for images 118, 120, and 122 captured at time points T4, T5, and T6 respectively, where time point T4 occurs after time point T5, time point T5 occurs after time point T4, and time point T6 occurs after time point T5.
[0048] It is also understood that during the event, image 100 can be captured by a camera of an endoscope with a fixed position. For example, image 100 can be captured by a digital camera with a fixed position during a medical procedure. Therefore, it is further understood that although the camera's field of view remains constant during the procedure, the image generated during the procedure may change due to the dynamic nature of the procedure captured by the image. As a simple example, image 112 may represent an image taken just before the laser fiber emits laser energy to crush the kidney stone. Furthermore, image 114 may represent an image taken just after the laser fiber emits laser energy to crush the kidney stone. Because the laser emits a bright flash, it is understood that image 112, captured just before the laser emits light, may differ significantly from image 114 in terms of saturation, contrast, exposure, etc. In particular, due to the sudden release of the laser, image 114 may include undesirable characteristics compared to image 112. Additionally, it is further understood that after the laser applies energy to the kidney, various particles from the kidney can move rapidly across the camera's field of view. These fast-moving particles may appear as localized areas with unwanted image features (e.g., overexposure, underexposure, etc.) over time through a series of images.
[0049] It is understandable that digital images (such as...) Figure 1 Any one of the multiple images 100 shown can be represented as a set of pixels (or individual picture elements) arranged in a two-dimensional grid, using squares. Furthermore, each individual pixel constituting an image can be defined as the smallest piece of information in the image. Each pixel is a small sample of the original image, where more samples generally provide a more accurate representation of the original image.
[0050] For example, Figure 3 It shows Figure 2 Example digital images 112 and 114 are represented as a collection of pixels arranged in a two-dimensional grid. For simplicity, the grid of each of images 112 and 114 is fixed at 18x12. In other words, the two-dimensional grid of images 112 / 114 consists of 18 vertically extending columns of pixels and 12 horizontally extending rows of pixels. It can be understood that... Figure 3 The image size shown is exemplary. The size of a digital image (the total number of pixels) can vary. For example, a common size for a digital image may include an image with 1080 columns by 720 rows (e.g., a frame size of 1080x720).
[0051] It is understandable that the location of a single pixel can be identified via its coordinates (X, Y) on a two-dimensional image grid. Furthermore, comparisons of neighboring pixels within a given image can generate desired information about which parts of the given image an algorithm can seek to preserve when performing image enhancement (e.g., image fusion). For example, Figure 3 Image feature 170 is shown, generally centered at pixel position (16, 10) at the lower right corner of the image. It can be understood that the pixels constituting feature 170 are substantially darker compared to the pixels surrounding it. Therefore, these pixels representing feature 170 can have high contrast values compared to the pixels surrounding it. This high contrast can be an example of valuable information in the image. Other valuable information might include high saturation, edges, and texture.
[0052] Furthermore, it can be understood that the information represented by a pixel at a given coordinate can be compared across multiple images. Comparisons of pixels with the same coordinates across multiple images can produce desired information about portions of a given image, which image processing algorithms can seek to discard when performing image enhancement (e.g., image fusion).
[0053] For example, Figure 3 Image 114 shows that feature 170 can represent moving particles captured by the two images over time. Specifically, Figure 3 Feature 170 (shown in the lower right corner of image 112) has been moved to the upper left corner of image 114. Note that images 112 / 114 were captured from the same image capture location using the same image capture device (e.g., a digital camera) and are mapped pixel-by-pixel to each other. Pixel-based subtraction of the two images can generate a new image where all pixel values except for positions (16, 10) and (5, 4) are close to zero. This can provide information that can be utilized by algorithms in image processing systems to reduce the undesirable effects of moving particles. For example, an image processing system could create a fused image of images 112 and 114, whereby the fused image retains the desired information of each image while discarding undesirable information.
[0054] The basic mechanism for generating a merged image can be described as pixels in input images with different exposures, weighted according to different "metrics" such as contrast, saturation, exposure, and motion. These metric weights can be used to determine how much a given pixel in an image being merged with one or more images will contribute to the final merged image.
[0055] Figures 4-5 A method is disclosed in which multiple differently exposed images (1…N) of an event (e.g., medical surgery) can be fused into a single fused image using an image processing algorithm. For simplicity, Figures 4-5The image processing algorithm shown will use the original image 114 and its "previous" image 112 (both in... Figures 2-3 (As shown in the image) is described as an example. However, it is understood that image processing algorithms can utilize any number of images captured during an event to generate one or more fused images.
[0056] It should be understood that the fusion process can be performed by an image processing system and output to a display device, thereby ultimately fusing the image to retain the desired characteristics of the image metrics (contrast, saturation, exposure, motion) from the input image 1 to N.
[0057] Figures 4-5 Example steps for generating an example weight map of the fused image are shown. The example first step may include selecting an initial input image 128. For the purposes of this discussion, example input image 114 will be described as the initial input image to be fused with its previous image 112 (referring to the discussion above, note that image 114 was captured after image 112).
[0058] After selecting the input image 114, individual contrast, saturation, exposure, and motion metrics are calculated at each of the individual pixel coordinates constituting image 114. In an exemplary embodiment, image 114 is composed of 216 individual pixel coordinates (refer to reference). Figure 3 Example image 114 comprises 216 individual pixels in an 18x12 grid. As will be described in more detail below, each of the contrast metric, saturation metric, exposure metric, and motion metric is multiplied together at each pixel location (i.e., pixel coordinate) of image 114, thereby generating a “preliminary weight map” of image 114 (for reference, this multiplication step is performed by...). Figure 4 (See text box 146 in the text box). However, before generating the initial weight map of the image, each individual contrast metric, saturation metric, exposure metric, and motion metric needs to be calculated for each pixel location. The following discussion describes the calculation of the contrast metric, saturation metric, exposure metric, and motion metric for each pixel location.
[0059] Calculation of contrast measurement
[0060] The calculation of contrast measurement is by Figure 4 Text box 132 in the image represents the contrast metric. To calculate the contrast metric at each pixel location, a Laplacian filter can be applied to the image to create a grayscale version of the image. After creating the grayscale version of the image, the absolute value of the filter response can be obtained. Note that the contrast metric can assign high weights to important elements (such as image edges and textures). The contrast metric (calculated per pixel) can be represented in this paper as (W c ).
[0061] Calculation of saturation measure
[0062] The calculation of saturation measure is by Figure 4 The text box 134 in the text box represents the saturation measure at each pixel location. To calculate the saturation measure at each pixel location, the standard deviation within the R, G, and B channels is taken (at each given pixel). A lower standard deviation means that the R, G, and B values are close together (e.g., the pixel is more gray). It can be understood that gray pixels may not contain as much valuable information as other colors (such as red) (because, for example, human anatomy includes more reddish hues). However, if a pixel has a higher standard deviation, the pixel tends to be more reddish, greenish, or bluish, and it contains more valuable information. Note that saturated colors are desired and make the image look vibrant. The saturation measure (calculated per pixel) can be represented in this paper as (W s ).
[0063] Calculation of exposure metric
[0064] The calculation of exposure metric is by Figure 4 Text box 136 in the text box represents the raw light intensity of a given pixel channel, which provides an indication of how much the pixel is exposed. Note that each individual pixel can include three separate channels: red, green, and blue (RGB), all of which can be displayed with a certain “intensity.” In other words, the overall intensity of a pixel location (i.e., coordinates) can be the sum of the red channel (with a certain intensity), the green channel (with a certain intensity), and the blue channel (with a certain intensity). These RGB light intensities at a pixel location can fall within the range of 0-1, with values close to zero being underexposed and values close to 1 being overexposed. It is desirable to maintain intensities not at the ends of the spectrum (not close to 0 or close to 1). Therefore, to calculate the exposure metric for a given pixel, each channel is weighted based on how close it is to the median of a reasonably symmetrical unimodal distribution. For example, each channel can be weighted based on how close it is to a given value on a Gaussian curve or any other type of distribution. After determining the weight of each RGB color channel at an individual pixel location, the three weights are multiplied together to obtain the overall exposure metric for that pixel. The exposure metric (calculated per pixel) can be represented in this paper as (W e ).
[0065] Calculation of movement metric
[0066] The calculation of the movement metric is by Figure 4 The text box 138 represents this. Figure 6The block diagram illustrates a more detailed discussion of the calculation of motion metrics. Motion metrics are designed to suppress interference caused by flying particles, laser flashes, or dynamic, rapidly changing scenes occurring across multiple frames. The motion metric for a given pixel location is calculated using the input image (e.g., 114) and its preceding image (e.g., image 112). An example first step in calculating the motion metric for each pixel location in image 114 is to convert all color pixels in image 114 to grayscale values. For example, it can be understood that in some instances, the raw data for each pixel location in image 114 can be represented as integer data, thus each pixel location can be converted to a grayscale value between 0 and 255. However, in other examples, the raw integer values can be converted to other data types. For example, the grayscale integer values of a pixel can be converted to a floating-point type, thus each pixel can be assigned a grayscale value from 0 to 1, and then divided by 255. In these examples, a completely black pixel can be assigned the value 0, and a completely white pixel can be assigned the value 1. Furthermore, while the examples above discussed representing raw image data as integer or floating-point types, it is understood that raw image data can be represented by any data type.
[0067] After converting example frame 114 to grayscale, the example second step for calculating the motion metric for each pixel position in image 114 is to convert all color pixels in frame 112 to grayscale using the same method as described above.
[0068] After converting each pixel location in frames 114 and 112 to grayscale, the example third step 164 for calculating the motion metric for each pixel location in image 114 is to subtract the grayscale value of each pixel location in image 112 from the grayscale value of the corresponding pixel location in image 114. For example, if image 114 ( Figure 3 The grayscale value of pixel coordinates (16, 10) in image 112 is equal to 0.90, and the image 112 (shown) Figure 3 If the grayscale pixel coordinates (16, 10) corresponding to the image 114 are equal to 0.10, then the shift subtraction value for pixel (16, 10) in image 114 is equal to .80 (0.90 minus 0.10). This subtracted shift value (e.g., .80) will be stored as a representative shift metric for pixel (16, 10) in image 114. As discussed above, the shift metric will be used to generate the fused image based on the algorithm discussed below. Furthermore, it is understood that the subtraction value can be negative, which can indicate artifacts in image 112. For ease of weighting coefficient calculation, negative values will be set to zero.
[0069] It's understandable that for fast-moving objects, the subtraction shift for a given pixel coordinate might be large because the change in pixel grayscale value will be significant (e.g., when a fast-moving object is captured moving through multiple images, the grayscale color of the individual pixels defining the object will change significantly from image to image). Conversely, for slow-moving objects, the shift for a given pixel coordinate might be small because the change in pixel grayscale value will be more gradual (e.g., when a slow-moving object is captured moving through multiple images, the grayscale color of the individual pixels defining the object will change slowly from image to image).
[0070] After the subtracted shift values have been calculated for each pixel location in the most recent image (e.g., image 114), example step 4 166 for calculating the shift metric for each pixel in image 114 is to weight each subtracted shift value for each pixel location (based on its proximity to zero). A first-order weighting calculation using a power function is illustrated in Equation 1 below:
[0071] (1) Wm=(1-[subtraction shift value])^100
[0072] Other weight calculations are also possible, and not limited to the power functions described above. Any function that makes the weights close to 1 when the shift value is zero and quickly jumps to 0 as the shift value increases is possible. For example, one can use the following weight calculation set in Equation 2 below:
[0073] (2)
[0074] These values are weighted shift metrics for each pixel in image 114, and can be represented as (W m ).
[0075] It's understandable that for pixels representing stationary objects, the subtraction shift value will be closer to zero, therefore W m The difference will be closer to 1. Conversely, for moving objects, the subtraction value will be closer to 1, therefore W m This will be close to 0. Returning to our example of the fast-moving dark circle 170 activity through images 114 / 112, pixel (16, 10) rapidly changes from a darker color (e.g., a grayscale value of 0.90) in image 112 to a lighter color (e.g., a grayscale value of 0.10) in image 114. The resulting subtractive shift value is equal to 0.80 (closer to 1), and the weighted shift metric for pixel position (16, 10) in image 114 will be: (1 - 0.80)^100, or very close to 0.
[0076] It should be noted that the W of the previous image (e.g., image 112) mIt will be set to 1. In this way, for the region of a stationary object, W in images 112 and 114... m All are close to 1, therefore they have the same weight with respect to the motion metric. The final weight map for a stationary object will depend on the contrast, exposure, and saturation metrics. For a moving object or a region of laser flash, the W of image 112... m It will still be 1, and the W of image 114 m The weights will be close to 0, so the final weights of these regions in image 114 will be much smaller than the final weights in image 112. Smaller weight values discard pixels from one frame into the fused frame, while larger weight values bring pixels from another frame into the fused image. In this way, artifacts will be removed from the final fused image.
[0077] Figure 4 This further illustrates that the contrast metric (W) has been calculated for each pixel of image 114. c ), Saturation measure (W) s Exposure measurement (W) e ) and movement metrics (W m Then, a preliminary weight map of 146 can be generated for image 114 by multiplying all four metrics (for a given pixel) together at each individual pixel location of image 114 (this calculation is illustrated in Equation 3 below). Figure 4 As indicated, Equation 3 will be executed at each individual pixel location (e.g., (1,1), (1,2), (1,3)...etc.) of all 216 example pixels in image 114 to generate a preliminary weight map.
[0078] (3) W = (W c )* (W s )* (W e )* (W m )
[0079] Figure 4 The example shows that after a preliminary weight map has been generated for the example image (e.g., image 114), the next step may include normalizing the weight map across multiple images. The purpose of this step may be to make the sum of the weights for each pixel location (in any two images being fused) equal to 1. The equations for normalizing the pixel locations (X, Y) of example "Image 1" and example "Image 2" (whose preliminary weight maps have already been generated using the method described above) are shown in Equations 4 and 5 below:
[0080] (4) W Image1 (X,Y) = W Image1 (X,Y) / (W Image1 (X,Y) + W Image2(X,Y)
[0081] (5) W Image2 (X,Y) = W Image2 (X,Y) / (W Image1 (X,Y) + W Image2 (X,Y)
[0082] As an example, suppose there are two images (Image 1 and Image 2), each with an initial weight map having a pixel position (14, 4), whereby the initial weight of pixel (14, 4) in Image 1 is 0.05, and the initial weight of pixel (14, 4) in Image 2 is 0.15. The normalized value of pixel position (14, 4) in Image 1 is shown in Equation 6, and the normalized value of pixel position (14, 4) in Image 2 is shown in Equation 7 below:
[0083] (6) W Image1 (14,4) = 0.05 / (0.05 + 0.15) = 0.25
[0084] (7) W Image2 (14,4) = 0.15 / (0.05 + 0.15) = 0.75
[0085] As discussed above, note that the sum of the normalized weights of Image 1 and Image 2 at pixel position (14, 4) equals 1.
[0086] Figure 5 This illustrates an example next step when generating a fused image from two or more example images (e.g., images 112 / 114). (Note that in...) Figures 4-5 At this point in the algorithm shown and described above, each example image 112 / 114 can have the original input image and a normalized weighted map. An example step may include generating a 152-smooth pyramid from the normalized weighted map (for each image). The smoothed pyramid can be calculated as follows:
[0087] G1 = W
[0088] G2 = downsample(G1, filter)
[0089] G3 = downsample(G2, filter)
[0090] The above calculations continue until the size of Gx is smaller than the size of the smoothing filter, which is a low-pass filter (e.g., a Gaussian filter).
[0091] Figure 5The example steps shown can further include generating 154 edge maps from the input images (for each image). The edge maps contain texture information from the input images at different scales. For example, suppose the input image is represented as "I". The edge map at layer x can be represented as Lx and calculated as follows:
[0092] I1 = downsample(I, filter)
[0093] L1 = I – upsample(I1)
[0094] I2 = downsample(I1, filter)
[0095] L2 = I1 – upsample(I2)
[0096] I3 = downsample(I2, filter)
[0097] L3 = I2 – upsample(I3)
[0098] The calculation continues until the size of Ix is smaller than the size of the edge filter, which is a high-pass filter (e.g., a Laplace filter).
[0099] Figure 5 This illustrates another example of the next step when generating a fused image from two or more example images (e.g., images 112 / 114). For example, for each input image, the edge map (as described above) can be multiplied by the smoothing map (as described above). The resulting maps generated by multiplying the edge maps by the smoothing maps of the images can be summed together to form a fused pyramid. This step is... Figure 5 The text box 156 in the text box represents this.
[0100] Figure 5 The final example shown is when generating the fused image, such as Figure 5 The text box 158 in the image represents this step. This step involves performing the following calculations on the fusion pyramid:
[0101] IN = LN
[0102] RN = upsample(IN, filter)
[0103] IN-1 = LN-1 + RN
[0104] RN-1 = upsample(IN-1, filter)
[0105] IN-2 = LN-2 + RN-1
[0106] RN-2 = upsample(IN-2, filter)
[0107] IN-3 = LN-3 + RN-2
[0108] The calculations continue until we reach the L1 layer, where the calculated I1 is the final image.
[0109] It should be understood that this disclosure is illustrative in many respects. Changes may be made in detail, particularly in terms of shape, size, and arrangement of steps, without departing from the scope of this disclosure. To the extent appropriate, this may include the use of any of the features of an exemplary embodiment used in other embodiments. The scope of this disclosure is, of course, defined in the language set forth in the appended claims.
Claims
1. A system for generating a fused image based on multiple images acquired from an endoscope, the system comprising: A processor operatively connected to the endoscope; and A non-transitory computer-readable storage medium, including code configured to perform a method of image fusion, the method comprising: A first input image is acquired from the endoscope, the first input image being formed by a first plurality of pixels arranged in a first coordinate grid, wherein each of the first plurality of pixels is located at a coordinate position in the first coordinate grid, and the first pixel of the plurality of pixels includes a first characteristic having a first value; A second input image is acquired from the endoscope. The second input image is formed by a second plurality of pixels arranged in a second coordinate grid. Each of the second plurality of pixels is located at a coordinate position in the second coordinate grid. The coordinate positions of the first plurality of pixels in the first coordinate grid and the coordinate positions of the second plurality of pixels in the second coordinate grid are located at the same corresponding positions. The second pixel of the second plurality of pixels includes a second characteristic having a second value. Each pixel in the first plurality of pixels of the first input image is converted into a grayscale intensity value to create a first grayscale image; Each pixel in the second plurality of pixels of the second input image is converted into a grayscale intensity value to create a second grayscale image; A movement metric is generated by subtracting the grayscale intensity value of each pixel coordinate position of the first plurality of pixels from the grayscale intensity value of each pixel coordinate position of the second plurality of pixels to generate a subtractive movement value for each pixel coordinate, and setting the negative subtraction value to zero; and After calculating the subtraction shift value for each pixel location and setting negative values to zero, a weighted metric map is generated using the shift metric to remove artifacts present in one of the first and second input images from the fused image, based on how close the subtraction shift value for each pixel location is to zero. The weighted metric map is used to create a fused image based on the first input image and the second input image.
2. The system according to claim 1, wherein the weighted motion metric W of the first input image m The weighted motion metric W of the first input image is set to 1 for areas containing moving objects or laser flashes. m The weighted shift metric W of the second input image remains 1. m The weights will be close to 0, so the final weights of these regions in the second input image will be much smaller than the final weights in the first input image. Smaller weight values will discard the pixels of the frame into the fused frame, while larger weight values will bring the pixels of the frame into the fused image. In this way, artifacts will be removed from the final fused image.
3. A method for combining multiple images, the method comprising: Acquire a first input image formed by a first plurality of pixels arranged in a first coordinate grid, wherein each of the first plurality of pixels is located at a coordinate position in the first coordinate grid, and the first pixel of the plurality of pixels includes a first characteristic having a first value; Acquire a second input image formed by a second plurality of pixels arranged in a second coordinate grid, wherein each pixel of the second plurality of pixels is located at a coordinate position in the second coordinate grid, the coordinate positions of the first plurality of pixels in the first coordinate grid and the coordinate positions of the second plurality of pixels in the second coordinate grid are located at the same corresponding positions, and the second pixel of the second plurality of pixels includes a second characteristic having a second value; Each pixel in the first plurality of pixels of the first input image is converted into a grayscale intensity value to create a first grayscale image; Each pixel in the second plurality of pixels of the second input image is converted into a grayscale intensity value to create a second grayscale image; The gray intensity value of each pixel coordinate position of the first plurality of pixels is subtracted from the gray intensity value of each pixel coordinate position of the second plurality of pixels to generate a subtraction shift value for each pixel coordinate position, and the negative subtraction value is set to zero; as well as After calculating the subtractive shift value for each pixel location and setting negative values to zero, a weighted metric map is generated using the shift metric to remove artifacts present in one of the first and second input images from the fused image, based on how close the subtractive shift value for each pixel location is to zero. The weighted metric map is used to create a fused image based on the first input image and the second input image.
4. The method according to claim 3, wherein the weighted motion metric W of the first input image m The weighted motion metric W of the first input image is set to 1 for areas containing moving objects or laser flashes. m The weighted shift metric W of the second input image remains 1. m The weights will be close to 0, so the final weights of these regions in the second input image will be much smaller than the final weights in the first input image. Smaller weight values will discard the pixels of the frame into the fused frame, while larger weight values will bring the pixels of the frame into the fused image. In this way, artifacts will be removed from the final fused image.
5. The method according to claim 1 or 3, wherein, The first input image is formed at a first time point, and the second input image is formed at a second time point that occurs after the first time point.
6. The method according to claim 1 or 3, wherein, The first input image and the second input image are captured by a digital camera, wherein the digital camera is positioned at the same location when it captures the first input image and the second input image.
7. The method according to claim 1 or 3, wherein, Generating a movement metric also includes using a power function to weight the movement metric.
8. The method according to claim 1 or 3 further includes generating a contrast metric, a saturation metric, and an exposure metric.
9. The method of claim 8, further comprising multiplying the contrast metric, the saturation metric, the exposure metric, and the motion metric together to generate the weighted metric map.
10. The method according to claim 3, wherein, Creating a fused image from the first input image and the second input image using the weighted metric map also includes normalizing the weighted metric map.
11. A method for combining multiple images, the method comprising: An image capture device is used to acquire a first input image at a first time point and a second input image at a second time point, wherein the image capture device is positioned at the same location when it captures the first input image at the first time point and the second input image at the second time point, and wherein the second time point occurs after the first time point, wherein the first input image is formed by a first plurality of pixels arranged in a first coordinate grid, each of the first plurality of pixels being located at a coordinate position in the first coordinate grid, and wherein the second input image is formed by a second plurality of pixels arranged in a second coordinate grid, each of the second plurality of pixels being located at a coordinate position in the second coordinate grid; Convert the first input image into a first grayscale image; Convert the second input image into a second grayscale image; A motion metric is generated based on the gray intensity values of pixels in both the first grayscale image and the second grayscale image, wherein the pixels of the first grayscale image and the pixels of the second grayscale image have the same coordinate positions in their respective images. The generation of the motion metric includes subtracting the gray value of each pixel position of the first plurality of pixels from the gray value of the corresponding pixel position of the second plurality of pixels for each pixel position of the second input image to obtain a subtracted motion value and setting the negative subtracted motion value to zero. as well as After calculating the subtraction shift value for each pixel location and setting negative values to zero, a weighted metric map is generated using the shift metric based on how close the subtraction shift value for each pixel location is to zero. This shift metric is then used to remove artifacts present in one of the first and second input images from the fused image. The weighted metric map is used to create a fused image based on the first input image and the second input image.
12. The method according to claim 11, wherein, Generating a movement metric also includes using a power function to weight the movement metric.
13. The method of claim 11, wherein the weighted motion metric W of the first input image m The weighted motion metric W of the first input image is set to 1 for areas containing moving objects or laser flashes. m The weighted shift metric W of the second input image remains 1. m The weights will be close to 0, so the final weights of these regions in the second input image will be much smaller than the final weights in the first input image. Smaller weight values will discard the pixels of the frame into the fused frame, while larger weight values will bring the pixels of the frame into the fused image. In this way, artifacts will be removed from the final fused image.
Citation Information
Patent Citations
Image processing method and apparatus
US20130070965A1