Endoscope image optimization method and device, storage medium and electronic equipment

By performing edge region restoration processing on the luminance component map of the endoscopic image and combining it with chrominance component map fusion, the problem that existing noise reduction algorithms cannot take into account diagnostic features is solved, achieving a balance between the clarity and realism of the endoscopic image and improving the clinical diagnostic effect.

CN122115268APending Publication Date: 2026-05-29ZHUHAI TAIKE MEDICAL TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHUHAI TAIKE MEDICAL TECH CO LTD
Filing Date
2026-04-28
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing noise reduction algorithms struggle to retain key diagnostic features of endoscopic images while suppressing noise, leading to loss of detail and impacting clinical diagnostic results.

Method used

The luminance component map of the endoscopic image is obtained for noise reduction, the distorted pixels in the edge region are identified and restored, the reference pixels in the original luminance component map are combined for restoration, and the image is fused with the chrominance component map to generate an optimized real-time image.

Benefits of technology

While preserving the noise reduction effect, it restores critical diagnostic details that have been over-smoothed, ensuring the clarity and realism of endoscopic images and improving the accuracy of clinical diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122115268A_ABST
    Figure CN122115268A_ABST
Patent Text Reader

Abstract

The application provides an endoscope image optimization method and device, a storage medium and an electronic device, and relates to the field of image processing. The electronic device obtains a luminance component image after noise reduction processing, and identifies pixels in an edge region that are distorted due to noise reduction. Then, the distorted pixels are restored using reference pixels at corresponding positions in the original luminance component image, thereby preserving the noise reduction effect while restoring key diagnostic details that have been excessively smoothed. Finally, the optimized luminance component image is fused with an unmodified chrominance component image to form a real-time image, thereby effectively resolving the contradiction between noise reduction and detail preservation, and enabling the real-time image of the endoscope to remain clear while retaining its authenticity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing, and more specifically, to an endoscopic image optimization method, apparatus, storage medium, and electronic device. Background Technology

[0002] Endoscopy, as a commonly used visual diagnostic and treatment tool in clinical practice, has been widely applied to the examination and treatment of organs and cavities such as the digestive tract, respiratory tract, and urinary system. Endoscopic video can continuously present information on the morphology, color, and fine structure of tissue surfaces, serving as an important basis for lesion detection, boundary interpretation, biopsy localization, intraoperative navigation, and efficacy evaluation.

[0003] To meet the aforementioned clinical decision-making needs, endoscopic videos require high clarity, detail fidelity, and color consistency. Specifically, physicians need to clearly identify anatomical structures such as mucosal texture, fine blood vessel morphology, glandular openings, and minute depressions or bulges. These structures often exist with high-frequency edges, local gradient variations, and low-contrast details. Simultaneously, accurate tissue color reproduction is crucial, as color abnormalities (e.g., congestion, pallor, cyanosis) are themselves important pathological indicators. Therefore, the image quality of endoscopic videos directly affects clinical diagnostic outcomes.

[0004] However, raw endoscopic images are susceptible to noise and image quality degradation during acquisition and transmission due to various factors. For example, sensor thermal noise in the imaging chain, high-gain amplification noise in low light, environmental electromagnetic interference, blockiness and ringing artifacts introduced by image compression, and motion blur caused by endoscope movement, breathing, or instrument contact during endoscopic operation can all lead to increased graininess, masking of weak textures, blunted edges, decreased contrast, and localized ghosting in raw endoscopic images. These problems manifest as flickering, ghosting, and visual instability during real-time display.

[0005] To address the aforementioned shortcomings of raw endoscopic images, current technologies commonly employ noise reduction algorithms to process these images. However, in practice, it has been found that while current noise reduction algorithms suppress noise, they often fail to preserve the key image features relied upon for clinical diagnosis, leading to the loss of details, particularly edge information crucial for diagnosis, thus impacting the effectiveness of clinical diagnosis. Summary of the Invention

[0006] In order to overcome at least one deficiency in the prior art, this application provides an endoscope image optimization method, apparatus, storage medium and electronic device, which can effectively alleviate the contradiction between noise reduction and detail preservation, so that the real-time image of the endoscope remains clear without losing its authenticity.

[0007] In a first aspect, this application provides an endoscopic image optimization method, the method comprising: The luminance component map and chrominance component map of the image to be processed are obtained, wherein the image to be processed is the original image acquired by the endoscope, and the luminance component map is the original luminance component map of the original image after noise reduction processing. Distorted pixels are determined from the edge regions of the luminance component map, wherein the distorted pixels represent pixels whose information is distorted due to the noise reduction process; The distorted pixel is restored using the reference pixel in the original luminance component image to obtain an optimized luminance component image, wherein the reference pixel is the pixel in the original luminance component image that corresponds to the position of the distorted pixel; The optimized luminance component map and the chrominance component map are fused together to form a real-time image for display.

[0008] Secondly, this application provides an endoscopic image optimization device, the device comprising: The image separation module is used to obtain the luminance component map and chrominance component map of the image to be processed, wherein the image to be processed is the original image acquired by the endoscope, and the luminance component map is the original luminance component map of the original image after noise reduction processing. A pixel recognition module is used to determine distorted pixels from the edge region of the luminance component map, wherein the distorted pixels represent pixels whose information is distorted due to the noise reduction process; An image optimization module is used to restore the distorted pixel using a reference pixel in the original luminance component image to obtain an optimized luminance component image, wherein the reference pixel is the pixel in the original luminance component image that corresponds to the position of the distorted pixel; An image fusion module is used to fuse the optimized luminance component image and the chrominance component image into a real-time image for display.

[0009] Thirdly, this application provides a storage medium storing a computer program that, when executed by a processor, implements the endoscopic image optimization method.

[0010] Fourthly, this application provides an electronic device, which includes a processor and a memory, wherein the memory stores a computer program, and the computer program, when executed by the processor, implements the endoscopic image optimization method.

[0011] Compared with the prior art, this application has the following beneficial effects: The endoscopic image optimization method, apparatus, storage medium, and electronic device provided in this application involve the electronic device acquiring a denoised luminance component image and identifying pixels in the edge region that have been distorted due to denoising. Then, using reference pixels at corresponding positions in the original luminance component image, these distorted pixels are restored, thereby preserving the denoising effect while restoring overly smoothed diagnostic details. Finally, the optimized luminance component image is fused with the unmodified chrominance component image to form a real-time image, effectively alleviating the contradiction between denoising and detail preservation, ensuring that the real-time image of the endoscope remains clear without losing its authenticity. Attached Figure Description

[0012] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 This is one of the endoscopic image optimization methods provided in the embodiments of this application; Figure 2 This is a second method for optimizing endoscopic images provided in the embodiments of this application; Figure 3 This is the third method for optimizing endoscopic images provided in the embodiments of this application; Figure 4 This is a schematic diagram of the structure of the endoscopic image optimization device provided in the embodiments of this application; Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0014] To make the objectives, technical solutions, and advantages of the embodiments of this application (hereinafter referred to as "the embodiments") clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0015] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0016] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0017] In the description of this application, it should be noted that the terms "first," "second," "third," etc., are used only for distinguishing descriptions and should not be construed as indicating or implying relative importance. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0018] Based on the above statement, as introduced in the background section, current technologies commonly employ noise reduction algorithms to process raw endoscopic images. However, in practice, it has been found that while current noise reduction algorithms suppress noise, they often fail to capture the key image features relied upon for clinical diagnosis, leading to a loss of detail.

[0019] For example, spatial domain denoising methods (such as Gaussian smoothing, bilateral filtering, and nonlocal means) tend to over-smooth images, causing blurring of mucosal edges, rupture of small blood vessels, and closure of gland openings. While temporal domain denoising methods (such as multi-frame averaging and recursive filtering) can improve the signal-to-noise ratio in static areas, they induce ghosting and blurring in moving areas, thus weakening the spatial continuity and geometric realism of tissue structures. More importantly, existing denoising algorithms typically do not prioritize diagnostically significant edge information in endoscopic images as a core optimization target. This results in a decrease in the overall noise level of the original endoscopic image after denoising, but a reduction in the edge sharpness, texture contrast, and structural discernibility of lesion areas.

[0020] It should be noted that the defects in the solutions in the prior art are the result of practice and careful research. Therefore, the discovery process of the above problems and the solutions proposed by the embodiments of this application in the following text should be regarded as contributions to this application in the process of invention and creation, and should not be understood as technical content known to those skilled in the art.

[0021] Based on the discovery of the above-mentioned technical problems, this embodiment provides an endoscopic image optimization method. For example... Figure 1 As shown, the method includes: S1, obtain the luminance component map and chrominance component map of the image to be processed.

[0022] The image to be processed is the original image acquired by the endoscope, and the brightness component image is the original brightness component image of the original image after noise reduction processing.

[0023] S2, determine the distorted pixels from the edge region of the luminance component map.

[0024] Among them, distorted pixels represent pixels whose information is distorted due to noise reduction processing; S3. Use the reference pixels in the original luminance component map to restore the distorted pixels and obtain the optimized luminance component map.

[0025] The reference pixel is the pixel in the original luminance component image that corresponds to the position of the distorted pixel.

[0026] S4 merges the optimized luminance component map and chrominance component map into a real-time image for display.

[0027] Thus, by acquiring the denoised luminance component map and identifying pixels in the edge region that have been distorted due to denoising, these distorted pixels are then restored using reference pixels at corresponding positions in the original luminance component map. This process preserves the denoising effect while restoring critical diagnostic details that have been over-smoothed. Finally, the optimized luminance component map is fused with the unmodified chrominance component map to form a real-time image. This effectively alleviates the contradiction between denoising and detail preservation, ensuring that the real-time image of the endoscope remains clear without losing its authenticity.

[0028] It should be noted that, for the endoscopic image optimization method provided in this embodiment, the electronic device implementing the method can be, but is not limited to, an endoscope host (hereinafter referred to as the host) or an integrated endoscope display screen. Specifically, the host can be a desktop computer, laptop, tablet computer, etc., that conforms to medical device standards.

[0029] Taking the endoscope host as an example, after the endoscope camera inputs the raw image signal into the host via cable, the host first performs analog-to-digital conversion and color space analysis on the signal to separate the raw luminance component image and chrominance component image. Then, it calls a preset algorithm to perform noise reduction processing on the raw luminance component image to generate a noise-reduced luminance component image. It also identifies pixels that are distorted due to over-smoothing. Then, it extracts the reference pixels at the corresponding positions from the raw luminance component image and restores the distorted pixels point by point. Finally, it merges the restored and optimized luminance component image with the unprocessed chrominance component image to form a real-time image that meets clinical display standards and outputs it to the monitor for display.

[0030] To make the solution provided in this embodiment clearer, the host is used as the electronic device implementing the method below, and in conjunction with... Figure 1Each step of the method is described in detail. However, it should be understood that the operations in the flowchart may not be implemented in sequence, and steps without logical contextual relationships may be reversed in order or performed simultaneously. Furthermore, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowchart, or remove one or more operations from the flowchart. See also... Figure 1 The method includes: S1, obtain the luminance component map and chrominance component map of the image to be processed.

[0031] The image to be processed is the original image acquired by the endoscope, such as the current frame during the real-time acquisition process of the endoscope; the luminance component image is the original luminance component image after noise reduction processing.

[0032] It should be understood that the raw images acquired by the endoscope are affected by low illumination, high gain, and optical scattering, with noise mainly superimposed on the brightness information. If noise reduction is performed simultaneously on the RGB three channels, the degree of noise suppression in each channel is difficult to be completely consistent, which can easily lead to an imbalance in the color ratio between pixels, and thus cause tissue color shift. For example, normal mucosa is light pink, but due to algorithm misadjustment, it may appear purplish or yellowish, seriously interfering with doctors' judgment of pathological features such as congestion, bleeding, or cancer.

[0033] Therefore, this embodiment provides the following optional real-time methods for step S1: S1-1: Separate the original luminance component map from the image to be processed and filter it to obtain the filtered luminance map.

[0034] In this embodiment, the problem of color distortion in endoscopic images during noise reduction is solved by converting the image to be processed from the red-green-blue (RGB) color space to the luminance-chromaticity space and separating the original luminance component image. Therefore, the host computer can first complete the color space transformation and decoupling, so that all subsequent noise reduction operations only apply to the original luminance component. , and chromaticity components Then the entire process remains unchanged or essentially unchanged; this mapping process can be represented as:

[0035] Based on the original luminance component map and its corresponding edge region separated in the above embodiments, this embodiment further implements edge-preservation-based noise reduction processing in the luminance domain. Specifically, high-frequency components of random noise are first suppressed, while preserving the integrity of the mucosal edge and texture structure as much as possible. On this basis, the luminance component map from the previous frame is used as a reference luminance map, and based on this reference luminance map, the inter-frame motion ratio of the original luminance component map relative to the reference luminance map is estimated in the low-resolution statistical domain, thereby generating temporal fusion coefficients. The luminance information of the current frame and the previous frame is then fused using a recursive weighting method. Therefore, step S1 further includes: S1-2, obtain the motion ratio between the filtered brightness map and the reference brightness map.

[0036] The reference luminance map represents the luminance component map based on the output of the previous frame.

[0037] As an optional implementation, to suppress the interference of inherent noise in the luminance component image on the motion ratio, the reference luminance image here can also be based on the luminance component image after filtering the previous frame. Therefore, this embodiment introduces adaptive bilateral filtering as a preprocessing mechanism. Of course, other suitable filtering algorithms can be used, and this embodiment does not specifically limit them.

[0038] In practical applications, for the original luminance component map of the previous frame, the host can perform an adaptive bilateral filtering operation to obtain a bilateral base luminance map, which is then used as the reference luminance map. In other words, in this embodiment, the reference luminance map can actually be a bilateral base luminance map. Of course, it can also be understood that the reference luminance map is not limited to the previous frame, but can also be a fused image of the bilateral base luminance maps of the previous N frames, where N≥1. Similarly, when the host obtains the motion ratio between the original luminance component map and the reference luminance map, it does not directly calculate the motion ratio between the original luminance component map and the reference luminance map. Instead, it first performs the same adaptive bilateral filtering operation on the original luminance component map separated from the current image to be processed, thereby obtaining the filtered luminance map, and compares it with the reference luminance map to obtain the inter-frame motion ratio. In other words, the filtered luminance map can actually be another bilateral base luminance map. For ease of description using the adaptive bilateral filtering algorithm as an example, the filtered luminance map is referred to as the bilateral base luminance map of the current frame, and the reference luminance map is referred to as the bilateral base luminance map of the previous frame.

[0039] The expression for adaptive bilateral filtering is as follows:

[0040] In the formula, Indicates the first Frame bilateral substrate brightness map The brightness value at pixel coordinate X. Indicates the first The brightness value at pixel coordinate X in the original brightness component image of the frame. Indicates the first Neighboring pixels in the original luminance component image of the frame The brightness value at the given location, where the spatial kernel and gray-level similarity kernel are used to suppress cross-edge propagation, can be expressed as:

[0041] In the formula, The spatial kernel function is used to characterize the spatial proximity relationship between the current pixel and its neighboring pixels. This represents the spatial distance between the current pixel position and the neighboring pixel positions. This represents the gray-level similarity kernel function, used to characterize the gray-level similarity relationship between the current pixel and its neighboring pixels; This represents the brightness difference between the current pixel and its neighboring pixels, i.e. ; For bilateral filtering neighborhood, This represents the pixel coordinates within the neighborhood that are involved in the filtering calculation. Represents pixel coordinates; parameters (Corresponding to grayscale scale) can be adaptively updated according to the noise level to control the denoising intensity. Parameter (Corresponding spatial domain scale) can be fixed or slightly updated with the scene; The filter kernel size can be an odd number (corresponding to the size of a bilateral filter). (parameters), for example, In this way, a balance can be achieved between real-time performance, noise suppression, and detail preservation.

[0042] For ease of subsequent description, the following text will refer to the first... The frame is called the current frame, the 1st frame. The frame is referred to as the previous frame; furthermore, for the following text containing pixel coordinates... The variables, although they actually represent the corresponding image location The pixel values ​​at each location are all simplified to the description of the image to which they belong, for example, It is simplified as the original luminance component map of the current frame.

[0043] Based on the bilateral basalt brightness maps of the current frame and the previous frame, the host further calculates the absolute difference between the two bilateral basalt brightness maps pixel by pixel within a low-resolution scale, forming a pixel difference image. Next, the proportion of pixels in the pixel difference image whose difference value exceeds the difference threshold is counted, and this proportion is used as the motion ratio. The corresponding expression is:

[0044]

[0045] In the formula, and These represent the low-resolution images of the current frame and the previous frame after dimensionality reduction of the bilateral basis brightness maps, respectively. Represents the total number of pixels in a pixel-differential image. Indicates the first The frame differential threshold.

[0046] As an optional implementation method, the differential threshold Low-frequency adaptive updates (e.g., updates every few frames) can be used to adapt to different noise and motion scenarios, using differential thresholding. The expression is:

[0047] In the formula, This represents the difference threshold of the previous frame, which serves as the historical benchmark for updating the current frame. This represents a learning rate between 0 and 1. This represents the average pixel value in the pixel difference image. This represents the clipping function. This represents the lower and upper limits of the differential threshold.

[0048] It should be noted that bilateral base brightness is used here instead of the original brightness for differential analysis because bilateral filtering has suppressed high-frequency random noise in advance, so that the differential results more realistically reflect the actual movement changes of the tissue without being interfered with by noise; at the same time, using a low-resolution statistical domain (for example, downsampling the image to one-quarter of its original size) can reduce the computational load and improve the real-time performance of the final displayed image.

[0049] S1-3, determine the first weight of the fused reference brightness map based on the motion ratio.

[0050] The first weight is inversely correlated with the motion ratio. During this process, the host computer calculates the first weight using the following expression. The fusion coefficients of the reference brightness map and the filtered brightness map in the time domain are represented as follows:

[0051] In the formula, Represented as the aforementioned motion ratio, The calculation can be performed using the formula mentioned above. The preset motion sensitivity coefficient, Based on the time-domain intensity, This means cropping the results to a preset upper and lower limit range. This expression embodies the principle of strong blending when static and weak blending when moving; that is, when the proportion of motion is low, it indicates that the overall image tends to be static. Increasing the value allows the current frame to incorporate more brightness information from the previous frame, thereby improving the signal-to-noise ratio; when the motion ratio increases, The value is reduced so that the current frame relies more on its own brightness information, thus avoiding motion blur artifacts caused by too much involvement of historical frames.

[0052] Based on the first weight obtained from the above embodiments, step S1 further includes: S1-4: The reference luminance map and the filtered luminance map are fused using the first weight to obtain the denoised luminance component map.

[0053] The study also found that during endoscopic examinations, slight lens shake, operator hand movements, or physiological activities such as the patient's breathing and heartbeat can cause continuous displacement of certain areas in the image (e.g., mucosal boundaries, vascular outlines); these edge areas are precisely the key locations for doctors to determine the nature of lesions. If a strong temporal fusion method suitable for static areas is directly applied to edge areas, the structural information that has been shifted in historical frames will be incorrectly superimposed onto the current frame, forming ghosting or string-like artifacts. Therefore, this embodiment also provides the following optional implementation methods for steps S1-4: S1-4-1 identifies the edge regions in the filtered brightness map.

[0054] When identifying edge regions in the filtered brightness map, as an optional implementation, the host can perform edge detection on the filtered brightness map to obtain the actual edges in the filtered brightness map; and then perform morphological dilation on the actual edges to obtain the edge regions in the filtered brightness map.

[0055] In practical applications, this embodiment performs edge detection and morphological dilation on the filtered brightness map via the host computer to identify edge regions, thereby protecting edge information in the image and reducing the risk of motion blur during temporal fusion. Here, the actual edge can refer to the initial edge line detected on the filtered brightness map by the Canny operator; while morphological dilation refers to the operation of expanding the neighborhood of edge pixels using structuring elements.

[0056] Continuing with the example of the bilateral base brightness map of the current frame, when performing morphological dilation on its initial edges, the host can use structuring elements. By moderately extending the edges, a wider edge area can be obtained. The corresponding expression is:

[0057] In the formula, Represents the bilateral basis brightness map for the current frame. The Canny edge detection operation performed, where For low threshold, Both are set as high thresholds, and together they control the sensitivity of edge extraction; This indicates a morphological dilation operation. As a structural element, it is used to moderately extend the detected edge lines to form an edge protection band with a certain width to accommodate minor alignment errors or tissue deformation.

[0058] Based on the above description of the edge region in the embodiments, steps S1-4 further include: S1-4-2, using the first weight to fuse the non-edge regions of the reference brightness map with the non-edge regions of the filtered brightness map, and using the second weight to fuse the edge regions of the reference brightness map with the edge regions of the filtered brightness map, to obtain the denoised brightness component map.

[0059] The second weight is less than the first weight. Continuing with the example of the bilateral base brightness map of the current frame, the host, based on this edge mask, adjusts the first weight within the edge region. Further reduction to second weight The calculation formula is as follows:

[0060] In the formula, This is the edge reduction factor.

[0061] The first weight obtained based on the above embodiments With further reduction to second weight The host computer applies the first weight to pixels in non-edge regions. Perform recursive weighted fusion, that is:

[0062] For pixels in the edge region, a second weight is applied. Execution fusion, that is:

[0063] This represents the luminance component map of the current frame after fusion operation, after noise reduction. This represents the luminance component image after noise reduction from the previous frame. This represents the bilateral base brightness map of the current frame, or the bilateral base brightness map after exposure micro-compensation correction. Compared with the original bilateral base brightness map, it has undergone inter-frame brightness mean compensation and cropping, thereby reducing inter-frame flicker caused by lens jitter, light source fluctuations or gain jumps while retaining its noise suppression and edge preservation capabilities.

[0064] In this way, by using a regional weighting method, the advantages of temporal denoising are fully utilized in non-edge regions, while the influence of historical frames is actively reduced in edge regions related to mucosal boundaries, vascular contours and diagnostic conclusions, thereby suppressing noise and preventing blurring and ghosting.

[0065] Based on the above embodiment's description of the luminance component map in step S1, please refer to [link to previous document]. Figure 1 Next, for Figure 1 Step S2 will be explained below: S2, determine the distorted pixels from the edge region of the luminance component map.

[0066] Among them, distorted pixels represent pixels whose information is distorted due to noise reduction processing.

[0067] It should be noted that after luminance domain noise reduction processing of the endoscopic image, although some pixels in the edge region change numerically, this change may not reflect the true structural information. Instead, it may cause the loss of subtle textures or blurring of contours required for diagnosis. Therefore, this embodiment also provides the following optional implementation of step S2: S2-1, obtain the target distortion threshold that is positively correlated with the noise intensity of the luminance component map.

[0068] It should be understood that when the endoscope is in a dark cavity or under high magnification, it automatically increases the image gain to enhance brightness, but at the same time, it also significantly amplifies the original noise. In this case, the noise reduction algorithm modifies the edge pixels more significantly. If the fixed threshold used in low-noise scenes is still used, a large number of real structural changes that should be preserved will be misjudged as distortions, resulting in over-restoration and image oscillation. Conversely, in bright scenes with high signal-to-noise ratios, if the threshold is not adjusted accordingly, slight but critical distortions may be missed, making the edges of tiny blood vessels blurry and irrecoverable.

[0069] In other words, the noise intensity of endoscopic images varies greatly under different lighting and equipment gain conditions. If a fixed distortion threshold is used to identify edge-distorted pixels caused by noise reduction, it will lead to inaccurate judgment.

[0070] It should also be understood that in the adaptive bilateral filtering algorithm... The scale parameter corresponding to the grayscale domain is used to control the decay rate of grayscale and brightness similarity weights, and can be adaptively updated according to the noise level to control the denoising intensity. Specifically, The smaller the value, the faster the weight decays across grayscale differences, resulting in a stronger cross-edge suppression effect, making the edges easier to protect, but it may not be sufficient for noise suppression in high-noise scenes. A larger value allows for greater grayscale differences in noise reduction, but may introduce the risk of over-smoothing. Therefore, this embodiment can adaptively update based on noise intensity to dynamically balance noise suppression and detail preservation at different noise levels.

[0071] Based on the above concept, when obtaining a distortion threshold that is positively correlated with the noise intensity of the luminance component map, the host obtains a preset basic distortion threshold; the basic distortion threshold is adjusted using the noise intensity to obtain the target distortion threshold.

[0072] This embodiment can be understood as follows: by combining a preset basic distortion threshold with the real-time estimated noise intensity, a target distortion threshold that adapts to the image noise level is dynamically generated, thereby ensuring that edge pixels that have suffered substantial information loss due to noise reduction processing can be accurately identified under different shooting conditions.

[0073] In practical applications, this basic distortion threshold can be a benchmark reference value calibrated based on typical endoscopic imaging conditions. Continuing with the example of the bilateral base plate brightness map of the current frame, the host estimates the noise intensity corresponding to the current brightness component map based on the statistical characteristics of the current bilateral base plate brightness map; and directly adds the basic distortion threshold to the noise intensity to obtain the final target distortion threshold used for distortion determination. As an optional implementation, due to the scale parameter of the grayscale domain... It can adaptively update to control the denoising intensity according to the noise level, and it can reflect the noise intensity in the image. Therefore, the current denoising intensity can be adjusted accordingly. Scale parameters of grayscale domain in bilateral filtering algorithm for frame image As noise intensity, the corresponding target distortion threshold The specific expression is:

[0074] In the formula, That is, the basic distortion threshold. This is the proportionality coefficient. This represents the scale parameter of the grayscale domain used in the bilateral filtering of the current frame, which can reflect the noise level of the current image.

[0075] Thus, when the endoscope is in a dark area or in a high-gain state, and the image noise is enhanced, The value increases, leading to Automatically raised to avoid misjudging minute pixel shifts caused by noise disturbances as distortions that need to be restored; when the image signal-to-noise ratio is good and the noise is weak, The value decreases. The corresponding reduction allows the host to more sensitively capture subtle but critical structural attenuation, such as slight blurring of the edges of mucosal microvessels.

[0076] S2-2, For the pixel to be analyzed in the edge region, obtain the degree of distortion of the pixel to be analyzed.

[0077] Optionally, the host computer can obtain the pixel residual between the pixel to be analyzed and the corresponding pixel in the original luminance component image; and use the pixel residual as the degree of distortion of the pixel to be analyzed.

[0078] In practical applications, for any pixel to be analyzed located in the edge region of the luminance component image, the host locates the pixel whose spatial position is completely corresponding to its position in the original luminance component image; and calculates the difference in luminance values ​​between these two pixels, i.e., the pixel residual, which is mathematically expressed as:

[0079] In the formula, This represents the original luminance component map. This represents the luminance component image after noise reduction for the current frame. Represents pixel coordinates.

[0080] The host will use the absolute value of this difference The difference is directly used to measure the distortion level of the pixel being analyzed. If the difference is close to zero, it means that the pixel has not been significantly altered by the noise reduction operation, and its structural information can be considered to be basically preserved. If the difference is large, it indicates that the brightness of the pixel has been corrected by smoothing, averaging, or temporal fusion, and the original subtle textures or edge contrast may have been lost.

[0081] S2-3, If the degree of distortion is greater than the distortion threshold, then the pixel to be analyzed is determined to be a distorted pixel.

[0082] In this way, the noise changes in the original luminance component map are adapted to the dynamically changing distortion threshold, and the distorted pixels that need to be repaired are accurately selected.

[0083] Based on the description of the distorted image in step S2 in the above embodiments, the following will continue to discuss... Figure 1 Step S3 will be explained below: S3. Use the reference pixels in the original luminance component map to restore the distorted pixels and obtain the optimized luminance component map.

[0084] The reference pixel is the pixel in the original luminance component image that corresponds to the position of the distorted pixel.

[0085] This embodiment can be understood as follows: by utilizing the reference pixels corresponding to the positions of distorted pixels in the original luminance component image, targeted local compensation is performed on pixels whose information is distorted due to noise reduction processing. In this way, while preserving the overall noise reduction effect, the structural information of edges and fine texture areas that have been weakened by excessive smoothing is restored.

[0086] Specifically, the host computer uses the difference between the reference pixel and the corresponding pixel in the currently denoised luminance component image as the detail residual, and multiplies it by a small recharge intensity coefficient. Then, it is superimposed on the current pixel output value, and the result is cropped to within the effective brightness range of 0 to 255. The corresponding expression is:

[0087] In the formula, Indicates the edge region. Indicates detailed residuals, Indicates the distortion threshold. This includes pixels that need to be refilled. Therefore, the refilling operation in the formula is triggered only when the detail residual amplitude exceeds the adaptive distortion threshold within the edge mask, thereby ensuring that only light compensation is performed at the edge of the distortion, avoiding the misinterpretation of random noise as detail.

[0088] Based on the above embodiment's description of the optimized luminance component map in step S3, please refer to [link to previous document]. Figure 1 Next, we will explain step S4 in the diagram: S4 merges the optimized luminance component map and chrominance component map into a real-time image for display.

[0089] This embodiment can be understood as follows: by recombining the optimized luminance component map, which has undergone detail restoration processing, with the chrominance component map that has not been modified in any way, and by performing an inverse color space transformation, a final color image for real-time clinical display is generated, thereby improving the clarity of the luminance domain structure while strictly maintaining the color characteristics of the original acquired image.

[0090] Specifically, the host computer inputs the optimized luminance component map and chrominance component map into a preset color space mapping relationship. This mapping relationship adopts a linear transformation form, and its inverse transformation formula can be expressed as:

[0091] in, This represents the optimized luminance component map. and Represents the chromaticity component diagram. This is the color space transformation matrix. The bias vector is used; this inverse transformation restores the luminance and chrominance components to a standard red (R), green (G), and blue (B) three-channel image. Since the chrominance component image is not modified, the hue and saturation of the output image are completely inherited from the original image and are not shifted due to luminance optimization.

[0092] It should be noted that during endoscopic examinations, doctors often need to press the freeze button to capture a specific image for detailed observation. However, the image at that moment may be affected by slight tremors of the endoscope, fluctuations in the patient's breathing, local tissue deformation, or instantaneous changes in lighting, resulting in motion blur, exposure deviations, or significant noise in the frozen frame. Directly enhancing this single frame in such cases is limited by the insufficient information in a single frame, easily amplifying noise or introducing artifacts, making it difficult to reliably recover the true structure. Conversely, merging any clear image selected from the entire historical cache may cause structural misalignment due to shifts in perspective or different anatomical positions. Therefore, a pre-set cache queue contains historical images captured by the endoscope, such as... Figure 2 As shown, the endoscopic image optimization method provided in this embodiment further includes: S5, in response to the freeze operation on the real-time image, selects multiple adjacent images from a preset buffer queue.

[0093] Among them, multiple adjacent images are those that are within a preset time window from the real-time image and whose clarity meets preset conditions.

[0094] It should be understood that the preset buffer queue is continuously maintained by the host, and it stores the output frames after real-time noise reduction processing, represented as... The host records the corresponding timestamp for each frame. This ensures a one-to-one correspondence between frames and time points; the length of this buffer queue is set to a fixed value. This is used to limit the computing resources and storage overhead required by the host, and to prevent unlimited accumulation that could lead to response delays or memory overflows.

[0095] Based on this, when the host responds to a freeze operation on a real-time image, it does not randomly select any historical frame, but strictly filters them according to the time dimension, that is, it only selects frames whose timestamps fall near the freeze time and whose time window length is [missing information]. Historical images within the selected range. The number of selected historical images shall not exceed a preset limit. This ensures that the selected images and the frozen frames are highly similar in terms of tissue motion, camera angle, and lighting conditions.

[0096] For historical images within this time window, the host has performed three types of complementary sharpness evaluations on each frame during caching, so that they can be directly used when generating frozen images. The three types of complementary sharpness are explained below: The first type is Laplace variance, which is calculated by measuring the grayscale image. Second-order differential response variance It is used to measure the intensity of high-frequency details in an image and is highly sensitive to changes in fine structures such as mucosal edges and microvessels. The second type is Tenengrad gradient energy, which first processes the grayscale image... Calculate the Sobel horizontal gradient separately and vertical gradient Then, the energy is normalized using the following expression to obtain the saliency representing the overall edge and texture:

[0097] The third category is edge ratio, which extracts the Canny edge map from the grayscale image. Then, calculate The proportion of edge pixels to the entire image is obtained to reflect the density of effective structural information in the image that can be clinically interpreted.

[0098] During this process, the host computer did not use any of the above indicators as a screening criterion alone. Instead, it addressed the problem of drastic fluctuations in indicators caused by reflections, dark areas, partial obstruction, and noise interference that are common in endoscopy scenarios. The host computer then performed robust normalization on the three types of indicator sequences using the median and median absolute deviation (MAD).

[0099] Specifically, the host first calculates the indicator sequence. the median of Median of absolute deviation Then, normalization is performed using the following expression:

[0100] In the formula, To prevent the lower bound constant of the divisor being zero.

[0101] The service seeks to improve the normalized Laplace variance. Tenengrad gradient energy With edge proportion By weight , , The linear weighted combination is used to score frame quality, and the expression is:

[0102] Wherein, the weights satisfy ; Finally, the service areas are rated. Sort from highest to lowest, and select the top... A frame is considered as a neighboring image that meets preset conditions.

[0103] Based on the above embodiments illustrating multiple adjacent images, the following will continue to discuss... Figure 2 Step S2 will be explained below: S6: Select the image to be fused from multiple adjacent images.

[0104] Among them, the images to be fused represent neighboring images whose overlap with the real-time image is higher than a threshold.

[0105] The study also found that although adjacent frames in endoscopic videos are close in time, the same anatomical location may show obvious displacement, deformation or even temporary disappearance in different frames due to the influence of slight shaking of the endoscope, tissue peristalsis, breathing fluctuations or local occlusion. If all adjacent frames are used for fusion without discrimination, the host may forcibly align pixels that do not belong to the same anatomical structure.

[0106] For example, overlapping the location of blood vessels in one frame with the location of mucosal folds in another frame can cause ghosting, double contours, or abnormal textures in the output image. Furthermore, when adjacent images have interference such as reflective flares, blood coverage, or lens fogging, their pixel values ​​differ significantly from the frozen frame, and forcibly participating in the fusion will introduce serious contamination.

[0107] Therefore, this embodiment provides the following optional implementation of step S6: S6-1, For each adjacent image, the adjacent image is pixel-aligned with the real-time image to obtain multiple pixel pairs with spatial correspondence.

[0108] In this embodiment, the host can perform dense optical flow estimation on each adjacent image and the real-time image, and remap according to the obtained displacement field, thereby establishing a one-to-one correspondence between the spatial positions of two frames at the pixel level, forming a large number of pixel pairs with accurate geometric correlation.

[0109] It should be understood that adjacent images and real-time images have a high similarity in terms of organizational motion morphology, thereby reducing the probability of alignment failure due to excessive deformation. Based on this, the host computer performs alignment checks on each adjacent image. With real-time images Estimate the dense displacement fields between them respectively. The displacement field is for each pixel. Provide horizontal offset and vertical offset Based on this displacement field, adjacent images are... Central The pixel value at that location is remapped to its position in the real-time image coordinate system. This generates alignment candidate frames. Its mathematical expression is:

[0110] This dense alignment method can accurately compensate for non-rigid motions (e.g., mucosal folding and stretching, vascular pulsation deformation) and local deformations (e.g., tissue indentation caused by instrument contact) that are common in endoscopic scenes, so that pixels of the same anatomical structure in different frames overlap as much as possible in space.

[0111] Based on the above description of pixel pairs in the embodiments, step S6 further includes: S6-2, Calculate the proportion of similar pixel pairs among multiple pixel pairs, and use it as the similarity proportion in adjacent images.

[0112] Among them, similar pixel pairs represent pixel pairs whose pixel values ​​are within a preset difference range.

[0113] This embodiment can be understood as follows: by using the host to count the proportion of pixels in each adjacent image that are highly consistent with the real-time image in terms of color value after pixel alignment, the actual degree of structural matching between the adjacent image and the real-time image can be quantified.

[0114] It should be understood that a pixel pair is a spatially corresponding combination of pixels obtained by performing dense optical flow estimation and remapping on adjacent images and the real-time image; for each pixel pair, the host calculates the color difference between its two constituent pixels. The norm squared is used to construct the pixel confidence map. Its expression is:

[0115] The host further performs thresholding on the confidence graph, setting a preset confidence threshold. Generate an effective area mask Its expression is:

[0116] In the formula, square brackets represent logical judgments, with a result of 1 (true) or 0 (false); the mask identifies all pixel locations with sufficiently high confidence, i.e., sufficiently small color differences.

[0117] Based on this, the host computer counts the total number of pixels with a value of 1 in the mask and divides it by the total number of pixels in the image. The effective overlap rate of adjacent images is obtained. Its expression is:

[0118] Similar pixel pairs here refer to those that satisfy... The condition is the pixel pair; and the similarity ratio is the effective overlap rate. This is used to characterize the proportion of the adjacent image that can still maintain reliable visual consistency with the real-time image after pixel alignment is completed.

[0119] Based on the above explanation of the similarity ratio of each adjacent image, step S6 further includes: S6-3, based on the similarity ratio with each adjacent image, images with a similarity threshold greater than the threshold are selected as images to be fused.

[0120] In this embodiment, the host compares the similarity ratio of each adjacent image with a preset similarity threshold, and only identifies adjacent images with a similarity ratio greater than the similarity threshold as images to be fused, thereby excluding candidate frames that are severely interfered with.

[0121] In practical applications, the host computer pre-sets a similarity threshold. This threshold is an empirical parameter used to characterize the minimum acceptable overlap standard. It is determined by the similarity ratio of adjacent images. Less than At this point, the host computer determines that there is a large area of ​​content mismatch between the image and the real-time image. Including this image in the fusion process would lead to structural misalignment, artifact overlay, and even misdiagnosis. Therefore, the host computer only merges images that meet the requirements... The adjacent images of the conditions are formally identified as the images to be fused.

[0122] Based on the above description of the images to be fused in step S6, the following will continue to discuss... Figure 2 Step S7 will be explained as follows: S7 merges the real-time image with the image to be merged to obtain a frozen image.

[0123] In practice, it was found that simply selecting neighboring frames for direct fusion, without quantitatively evaluating the matching quality between each frame and the frozen frame, may force misaligned, insufficiently overlapping, or inherently blurry images into the fusion process, thereby reducing the quality of the frozen image. Therefore, this embodiment provides the following optional implementation methods for step S7: S7-1, obtain the sharpness of the image to be fused, the overlap rate with the real-time image, and the alignment penalty coefficient with the real-time image.

[0124] The alignment penalty coefficient represents the degree of alignment between the image to be fused and the real-time image.

[0125] S7-2 maps sharpness, overlap rate, and alignment penalty coefficient to fusion weights for the images to be fused. S7-3, based on the fusion weights of the images to be fused, merge the real-time image and the image to be fused into a frozen image.

[0126] This embodiment can be understood as follows: by performing a fusion operation based on triple quality assessment and pixel-level confidence weighting on the host side, a real-time image and the image to be fused are fused to obtain a frozen image. However, this embodiment does not simply superimpose or average the images. Instead, it calculates the contribution intensity of the image to be fused to the final frozen image pixel by pixel based on the comprehensive performance of the image to be fused in three dimensions: sharpness, spatial overlap reliability, and alignment accuracy. This robustly introduces high-quality details while preserving the main anatomical structure of the original image, and actively suppresses ghosting, misalignment, and pseudo-textures caused by alignment deviations.

[0127] It is important to note that the clinical value of endoscopic frozen images highly depends on their structural accuracy and the reliability of their details. If only a seemingly clear historical image is directly stitched together with a real-time image, pixel misalignment is easily caused by tissue deformation, lens micro-movements, or local occlusion. For example, incorrectly superimposing a blood vessel from a nearby intestinal segment onto the surface of a polyp in the current field of view creates a false boundary, misleading the physician's judgment. Therefore, it is essential to evaluate whether each image to be fused is truly usable; that is, it must not only be sufficiently clear itself but also spatially highly matched with the real-time image and substantially consistent in pixel values. Any significant deficiencies in either aspect should reduce its weight during fusion.

[0128] Specifically, the host computer first acquires the sharpness of the image to be fused, which quantifies the richness of edge and texture details contained in the image itself. Secondly, the host computer calculates the overlap rate between the image to be fused and the real-time image; a higher overlap rate indicates a wider effective area that can be aligned anatomically between the two images, resulting in a more robust fusion foundation. Furthermore, the host computer also acquires the alignment penalty coefficient between the image to be fused and the real-time image. The alignment penalty coefficient is determined by the alignment within the effective region. The mean squared error (MSE) is calculated and normalized.

[0129] In the calculation of the alignment penalty coefficient, the host first calculates the mean square error, expressed as:

[0130] Then, normalization is further performed according to the following expression to obtain the alignment penalty coefficient:

[0131] The smaller the coefficient, the better the consistency of color and brightness between the two images after alignment, and the higher the alignment quality; conversely, the larger the error, the more significant the photometric deviation, which may be due to reflection, shadow changes or local occlusion, even after motion compensation is completed. In this case, its fusion weight should be reduced.

[0132] Finally, the host computer will analyze the above three parameters, including resolution. overlap rate Alignment penalty coefficient Commonly mapped to fusion weights of the images to be fused During the mapping process, the host first combines the three factors into a global score, expressed as:

[0133] In the formula, the exponential decay term ensures that even a slight increase in the alignment error will cause a significant decrease in the weight.

[0134] Then, the host imposes an upper limit constraint on the global score to prevent a particular image from excessively dominating the fusion result due to an abnormally high score in a certain metric. The expression is as follows:

[0135] Based on the above fusion weights The host computer combines the fusion weights with the pixel confidence map. For real-time images Aligned images to be merged Normalized weighted fusion: This involves using real-time images as a baseline and assigning them fixed base weights. For each image to be fused, its fusion weight is used. With corresponding pixel confidence Multiply by the given values ​​to obtain a dynamic weighting coefficient for that pixel location; the final frozen image is obtained using the following expression:

[0136] From the above expression, it is easy to see that at the pixel level, only images to be fused with high confidence (i.e., small color difference) and excellent global scores at a given location are allowed to participate in the weighting. If all images to be fused have low confidence or insufficient scores in a certain region, then that region is entirely dominated by the real-time image, thus ensuring that the output results always have determinism and clinical usability. When all candidate frames are rejected due to excessively low overlap or excessive alignment penalty, the output degenerates into the original real-time image. To avoid blank or abnormal screens.

[0137] In practice, it was also found that endoscopic frozen images often face the following problems in clinical observation: First, uneven illumination inside the cavity leads to dark areas in deep mucosal folds or curved regions, making it difficult to identify key lesion details; second, low-light and high-gain imaging easily introduces noise, especially in the blue channel, which is prone to graininess. Blindly enhancing the image can misinterpret noise as texture, creating false structures; third, doctors rely on the inherent color of tissue to determine benignity or malignancy, but common sharpening algorithms often alter the proportions of red, green, and blue, leading to color cast risks that may mislead diagnosis. Therefore, based on the frozen images obtained from the above embodiments, such as... Figure 3 As shown, the endoscopic image optimization method provided in this embodiment further includes: S8 separates the base image and detail image from the frozen image.

[0138] The base image carries structural and brightness information from the frozen image, while the detail image carries texture and edge information from the frozen image.

[0139] In this embodiment, the host separates the basal image and the detail image from the frozen image; however, this operation is not simply high-frequency or low-frequency filtering of the image, but has the dual goal of preserving clinically critical structural information and suppressing noise interference. The basal image carries the large-scale structure and overall brightness distribution in the frozen image, while the detail image specifically carries fine textures and edge contours.

[0140] It should be understood here that although the endoscopic frozen image has improved the blur problem through the aforementioned fusion processing, it still retains defects such as noise, uneven lighting, and insufficient local contrast. Directly enhancing the entire image could easily lead to underexposure of dark areas and overexposure of bright areas, or misinterpreting noise as texture and amplifying it. Therefore, this embodiment requires a two-stage decomposition of the frozen image in the brightness domain. The first stage uses bilateral filtering to decompose the frozen image... Extracting structural layers It satisfies the expression:

[0141] In the formula, For spatial scale parameters, The grayscale similarity scale parameter is used; this structural layer can suppress random noise while preserving the geometric integrity of the mucosal edge and tissue texture.

[0142] The host computer further calculates the noise residual. And obtain the noise-suppressed image:

[0143] In the formula, This is the noise suppression strength coefficient.

[0144] Based on this noise-suppressed image, continue to refine the noise-suppressed image. Apply edge-preserving smoothing operator (e.g., fast global smoothing or guided filtering) to generate a base image The basal image mainly contains low-frequency illumination information and large-scale anatomical structures, such as the overall outline of the intestinal wall, the orientation of folds, and the trend of background brightness.

[0145] Finally, the host computer obtains the detailed image in residual form, expressed as:

[0146] This detailed image specifically carries high-frequency texture and fine edge information, such as gland openings, microvascular branches, and subtle ridges on the surface of lesions.

[0147] Based on the base image and detail image obtained from the above embodiments, see below. Figure 3 The method also includes: S9, the dark areas in the base image are enhanced and the bright areas are suppressed to obtain the optimized base image.

[0148] In this embodiment, the host obtains an optimized base image by enhancing the brightness of dark areas and suppressing the brightness of bright areas in the base image. This operation does not uniformly brighten or stretch the contrast of the entire image, but dynamically determines whether to enhance, how much to enhance, and whether to limit the enhancement range based on the brightness position of each pixel. This improves the visibility of lesions in dark areas while strictly avoiding overexposure in bright areas, highlight clipping, and the amplification of noise.

[0149] It should be understood that the lighting inside the endoscopic cavity is naturally uneven, resulting in a bright and clear area directly in front of the lens, while large dark areas are formed in the deep folds, bends, or areas obscured by tissue. These areas may hide tiny polyps or early signs of cancer. If global gamma correction is used, although the brightness of dark areas can be improved, it will inevitably lead to severe overexposure of the already clear bright areas, loss of vascular texture, and even the production of halo artifacts. If only local enhancement is used, there is a lack of grasp of the overall lighting trend, which can easily cause brightness discontinuity.

[0150] Therefore, in this embodiment, the host can be based on the base image. luminance component First, use the Gaussian smoothing operator. A smooth illumination term is generated to characterize the slowly changing background illumination distribution in the image, expressed as:

[0151] The host further constructs a specular suppression term, with the expression:

[0152] In the formula, The high-light protection threshold is used to identify and mark areas of strong reflection that are close to saturation (e.g., specular reflection points) to prevent subsequent enhancement from being applied to these areas. The host then incorporates a dark-area sensitive weighting function, causing the weights to approach 1 in dark areas, approach 0 in bright areas, and be forcibly set to zero in highlight areas. The expression is:

[0153] Based on this sensitive weighting function, the host then performs pixel-by-pixel gamma mapping to ensure that only dark areas receive non-linear enhancement, and that the enhancement lower limit in the darkest areas is affected. The constraint, expressed as:

[0154] Based on the above process of light enhancement, the host calculates the enhanced brightness using the following expression:

[0155] In the formula, The result is the gamma transform. This is the global limiting coefficient. To preset the maximum brightness increase, It is a numerical stability constant used to rigidly cap the overall strength.

[0156] Finally, the host image is scaled to the RGB channels according to the brightness ratio to obtain the optimized base image. .

[0157] This makes the dark areas of the base image brighter, while the bright areas must not be overexposed.

[0158] Based on the optimized base image obtained from the above embodiments, see below. Figure 3 Next, we will continue with... Figure 3 Step S10 will be explained as follows: S10: Select the areas to be enhanced in the detail image and optimize the different color channels of the areas to be enhanced to obtain the optimized detail image.

[0159] It should be understood that while the detailed images in endoscopic frozen images have separated texture and edge information, they contain two types of components that cannot be enhanced: one type is the texture of real lesions (such as gland openings and microvascular branches), which needs to be highlighted; the other type is noise residuals or artifacts (such as blue channel speckle under low illumination and bright spots in reflective areas). If these are enhanced together, noise will be mistaken for pathological features, leading to false positives. In addition, relying solely on the energy threshold of the detailed image itself to define the enhancement area is prone to missing weak but crucial textures in smooth tissue areas and mistakenly selecting false bright spots in areas of high reflectivity.

[0160] Therefore, when filtering out the regions to be enhanced in the detail image, the host determines the initial enhancement region in the detail image; the optimized base image is then used to filter out the regions to be enhanced from the initial enhancement region.

[0161] This embodiment does not directly set a fixed threshold on the detail image to define the enhancement range. Instead, it first generates an initial enhancement region based on the detail energy, and then uses the spatial structure information of the optimized base image to perform fine filtering, thereby ensuring that the finally selected enhancement region not only truly reflects the anatomical texture position, but also avoids noise interference and specular artifact regions.

[0162] In practical applications, the host computer first obtains the luminance component based on the noise-suppressed image. A Difference of Gaussians (DoG) energy map is constructed to highlight texture regions with multi-scale structural responses. The expression is as follows:

[0163] The host is further processed by the local mean operator Smoothing yields the energy map and utilize quantiles Robust normalization is obtained Then through a continuous sigmoid function Mapped as a soft mask To suppress false enhancement of reflective areas, it is also necessary to... A specular suppression factor is superimposed. Based on this, the host computer uses the optimized base image. luminance component For the guiding diagram, Perform guided filtering to obtain the mask. This ensures that the mask boundaries strictly conform to the contours of the actual tissue structure, resulting in the final mask. This is a pixel-level indicator map of the area to be enhanced, used to determine the area to be enhanced.

[0164] In practice, it was also found that the diagnostic information value and noise sensitivity of different color channels in endoscopic images vary. The green channel contributes the most to human eye brightness perception and has the highest signal-to-noise ratio in most endoscopic imaging systems, making it the main channel for expressing key structures such as mucosal texture and vascular morphology. The blue channel is easily mixed with speckle noise and electronic interference under low illumination and high gain; if enhanced at the same rate as other channels, it will significantly amplify graininess, creating false textures. The red channel is directly related to tissue blood supply, congestion, and lesion color interpretation (e.g., bleeding, reddened areas); arbitrary enhancement may change its hue ratio, leading to misdiagnosis by doctors. Therefore, uniformly enhancing all three channels will not only fail to focus the most effective visual information but will also simultaneously amplify blue channel noise and distort red channel color, severely compromising diagnostic reliability.

[0165] Therefore, when performing targeted optimization on different color channels of the area to be enhanced to obtain an optimized detail image, the host can fully enhance the green channel of the area to be enhanced, and proportionally enhance the blue channel of the area to be enhanced, while keeping the red channel of the area to be enhanced unchanged, to obtain an optimized detail image.

[0166] Therefore, it can be understood that this embodiment does not apply the same intensity of enhancement to the red, green, and blue channels. Instead, it sets enhancement strategies for each channel based on their information value in clinical diagnosis: the green channel undertakes the main task of sharpness enhancement, the blue channel only makes limited compensation, and the red channel remains completely unchanged, thereby improving texture visibility while avoiding the risks of color cast and false texture.

[0167] In practical applications, the host computer first generates a mask based on the above-described embodiments. Constructing pixel-wise gain The expression is:

[0168] In the formula, This is the detail gain factor, used to control the enhancement intensity.

[0169] Furthermore, detailed images are processed separately according to channel selectivity rules. The three components, for the red channel Simply retain the original value, that is Ensure that the inherent color of the tissue (such as congestion, bleeding, and mucosal color) does not shift due to enhancement; for the green channel Apply full gain, i.e. Because it has the highest weight and best signal-to-noise ratio in human eye brightness perception, it is the most reliable information carrier for expressing key textures such as microvessels and gland openings; for the blue channel only proportionally (in Restricted enhancement is performed, i.e. This is to suppress the risk of the inherent speckle noise and electronic interference being amplified simultaneously under low illumination and high gain conditions.

[0170] Based on the optimized base image and optimized detail image obtained from the above embodiments, the following will continue to... Figure 3 Step S11 will be explained as follows: S11, the optimized base image and the optimized detail image are fused to obtain the optimized frozen image.

[0171] It should be understood that the optimized base image Adaptive brightness enhancement has been implemented in dark areas, brightness suppression in bright areas, and highlight clipping and overall overexposure have been strictly avoided; the optimized image shows detailed images. Then only through the mask Within a precisely defined area to be enhanced, the green channel is fully enhanced, the blue channel is proportionally limited, and the red channel remains completely unchanged. This improves texture clarity while suppressing the risk of hue drift and blue channel noise amplification.

[0172] Therefore, the purpose of fusion is not to readjust, but to physically superimpose semantically complementary information from the two, expressed as:

[0173] In the formula, This indicates that the pixel values ​​are cropped to... The effective display area is determined to prevent brightness overflow or color distortion caused by superposition.

[0174] Based on the same inventive concept as the endoscopic image optimization method provided in this embodiment, this embodiment also provides an endoscopic image optimization device. This device includes at least one software functional module that can be stored in a memory or embedded in an electronic device. The processor in the electronic device executes the executable module stored in the memory. For example, the software functional modules and computer programs included in this device. Please refer to... Figure 4 Functionally, the device may include: Image separation module 11 is used to acquire the luminance component map and chrominance component map of the image to be processed, wherein the image to be processed is the original image acquired by the endoscope, and the luminance component map is the original luminance component map of the original image after noise reduction processing. Pixel recognition module 12 is used to determine distorted pixels from the edge region of the luminance component map, wherein the distorted pixels represent pixels whose information is distorted due to noise reduction processing; Image optimization module 13 is used to restore distorted pixels using reference pixels in the original luminance component image to obtain an optimized luminance component image, wherein the reference pixel is the pixel in the original luminance component image that corresponds to the position of the distorted pixel. Image fusion module 14 is used to fuse the optimized luminance component map and chrominance component map into a real-time image for display.

[0175] In this embodiment, the image separation module 11 is used to implement Figure 1 In step S1, the pixel recognition module 12 is used to implement Figure 1 In step S2, the image optimization module 13 is used to implement... Figure 1 In step S3, the image fusion module 14 is used to implement Figure 1 Step S4 in the above process. Therefore, for a detailed description of each of the above modules, please refer to the specific implementation of the corresponding step.

[0176] Optionally, the pixel recognition module 12 determines the distorted pixels from the edge regions of the luminance component map in the following ways: Obtain the distortion threshold that is positively correlated with the noise intensity of the luminance component map; For the pixel to be analyzed in the edge region, obtain the degree of distortion of the pixel to be analyzed; If the degree of distortion is greater than the distortion threshold, the pixel to be analyzed is identified as a distorted pixel.

[0177] Optionally, the pixel recognition module 12 obtains the distortion threshold that is positively correlated with the noise intensity of the luminance component map in the following ways: Obtain the preset base distortion threshold; The distortion threshold is obtained by adding the base distortion threshold to the noise intensity.

[0178] Optionally, the pixel recognition module 12 obtains the degree of distortion of the pixel to be analyzed in the following ways: Obtain the pixel residual between the pixel to be analyzed and the corresponding pixel in the original luminance component image; The pixel residual is used as the degree of pixel distortion to be analyzed.

[0179] Optionally, the image separation module 11 acquires the brightness component map of the image to be processed in the following ways: The original luminance component map is separated from the image to be processed and filtered to obtain the filtered luminance map; Obtain the motion ratio between the filtered luminance map and the reference luminance map, where the reference luminance map represents the luminance component map output based on the previous frame; The first weight of the fused reference brightness map is determined based on the motion ratio, wherein the first weight is inversely correlated with the motion ratio; The reference luminance map and the filtered luminance map are fused using the first weight to obtain the denoised luminance component map.

[0180] Optionally, the image separation module 11 fuses the reference brightness map and the filtered brightness map using a first weight to obtain the denoised brightness component map, including: Identify the edge regions in the filtered brightness map; The non-edge regions of the reference brightness map and the non-edge regions of the filtered brightness map are fused using a first weight, and the edge regions of the reference brightness map and the edge regions of the filtered brightness map are fused using a second weight to obtain the denoised brightness component map, wherein the second weight is less than the first weight.

[0181] Optionally, the image separation module 11 identifies edge regions in the filtered brightness map by means of: By performing edge detection on the filtered brightness map, the actual edges in the filtered brightness map are obtained; The actual edges are morphologically dilated to obtain the edge regions in the filtered brightness map.

[0182] Optionally, a preset buffer queue stores historical images captured by the endoscope, and the image fusion module 14 is also used for: In response to the freeze operation on the real-time image, multiple adjacent images are selected from a preset cache queue. These multiple adjacent images are images that are within a preset time window from the real-time image and whose clarity meets preset conditions. The image to be fused is selected from multiple adjacent images, where the image to be fused represents an adjacent image whose overlap with the real-time image is higher than a threshold. The real-time image is fused with the image to be fused to obtain a frozen image.

[0183] Optionally, the image fusion module 14 fuses the real-time image with the image to be fused to obtain a frozen image in the following ways: The sharpness of the image to be fused, the overlap rate with the real-time image, and the alignment penalty coefficient with the real-time image are obtained. The alignment penalty coefficient represents the degree of alignment between the image to be fused and the real-time image. Sharpness, overlap rate, and alignment penalty coefficient are mapped to fusion weights for the images to be fused. Based on the fusion weights of the images to be fused, the real-time image and the image to be fused are fused into a frozen image.

[0184] Optionally, the image fusion module 14 selects the image to be fused from multiple adjacent images in the following ways: For each adjacent image, the adjacent image is pixel-aligned with the real-time image to obtain multiple pixel pairs with spatial correspondence. The proportion of similar pixel pairs in multiple pixel pairs is counted and used as the similarity proportion in adjacent images. Here, a similar pixel pair represents a pixel pair whose pixel value is within a preset difference range. Based on the similarity ratio with each adjacent image, images with a similarity threshold greater than a certain threshold are selected as images to be fused.

[0185] Optionally, the image fusion module 14 is also used for: The base image and detail image are separated from the frozen image. The base image carries the structural and brightness information of the frozen image, and the detail image carries the texture and edge information of the frozen image. The base image is optimized by enhancing the brightness of dark areas and suppressing the brightness of bright areas. The regions to be enhanced in the detail image are selected, and different color channels of the regions to be enhanced are optimized to obtain the optimized detail image. The optimized base image and the optimized detail image are fused together to obtain the optimized frozen image.

[0186] Optionally, the image fusion module 14 filters out the regions to be enhanced located in the detail image in the following ways: Identify the initial enhancement region located in the detail image; The optimized base image is used to filter out the region to be enhanced from the initial enhancement region.

[0187] Optionally, the image fusion module 14 optimizes different color channels of the region to be enhanced to obtain an optimized detail image, including: The green channel of the area to be enhanced is fully enhanced, and the blue channel of the area to be enhanced is enhanced proportionally, while the red channel of the area to be enhanced remains unchanged, resulting in an optimized detail image.

[0188] In addition, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0189] It should also be understood that if the above embodiments are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.

[0190] Therefore, this embodiment also provides a storage medium, which is a computer-readable storage medium. The storage medium stores a computer program, which, when executed by a processor, implements the endoscopic image optimization method provided in this embodiment. The storage medium can be any medium capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0191] This embodiment also provides an electronic device for implementing the endoscopic image optimization method. For example... Figure 5 As shown, the electronic device may include a processor 22 and a memory 21. The memory 21 stores a computer program, and the processor implements the endoscopic image optimization method provided in this embodiment by reading and executing the computer program corresponding to the above embodiments stored in the memory 21.

[0192] See also Figure 5 The electronic device also includes a communication unit 23. The memory 21, processor 22 and communication unit 23 are electrically connected to each other directly or indirectly through system bus 24 to realize data transmission or interaction.

[0193] The memory 21 can be an information recording device based on any electronic, magnetic, optical, or other physical principles, used to record execution instructions, data, etc. In some embodiments, the memory 21 can be, but is not limited to, volatile memory, non-volatile memory, memory drive, etc.

[0194] In some embodiments, the volatile memory may be random access memory (RAM); in some embodiments, the non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, etc.; in some embodiments, the storage drive may be a disk drive, solid-state drive, any type of storage disk (such as optical disc, DVD, etc.), or similar storage media, or a combination thereof.

[0195] The communication unit 23 is used to send and receive data over a network. In some embodiments, the network may include a wired network, a wireless network, a fiber optic network, a telecommunications network, an intranet, the Internet, a local area network (LAN), a wide area network (WAN), a wireless local area network (WLAN), a metropolitan area network (MAN), a public switched telephone network (PSTN), a Bluetooth network, a ZigBee network, or a near field communication (NFC) network, or any combination thereof. In some embodiments, the network may include one or more network access points. For example, the network may include wired or wireless network access points, such as base stations and / or network switching nodes, through which one or more components of the service request processing system can connect to the network to exchange data and / or information.

[0196] The processor 22 may be an integrated circuit chip with signal processing capabilities, and may include one or more processing cores (e.g., a single-core processor or a multi-core processor). By way of example only, the processor described above may include a Central Processing Unit (CPU), an Application Specific Integrated Circuit (ASIC), an Application Specific Instruction-set Processor (ASIP), a Graphics Processing Unit (GPU), a Physics Processing Unit (PPU), a Digital Signal Processor (DSP), a Field Programmable Gate Array (FPGA), a Programmable Logic Device (PLD), a controller, a microcontroller unit, a Reduced Instruction Set Computing (RISC) computer, or a microprocessor, or any combination thereof.

[0197] Understandable. Figure 5The structure shown is for illustrative purposes only. Electronic devices may also have more advanced features. Figure 5 Showing more or fewer components, or having with Figure 5 The different configurations shown. Figure 5 The components shown can be implemented using hardware, software, or a combination thereof.

[0198] It should be understood that the apparatus and methods disclosed in the above embodiments can also be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0199] The above descriptions are merely various embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An endoscopic image optimization method, characterized in that, The method includes: The luminance component map and chrominance component map of the image to be processed are obtained, wherein the image to be processed is the original image acquired by the endoscope, and the luminance component map is the original luminance component map of the original image after noise reduction processing. Distorted pixels are determined from the edge regions of the luminance component map, wherein the distorted pixels represent pixels whose information is distorted due to the noise reduction process; The distorted pixel is restored using the reference pixel in the original luminance component image to obtain an optimized luminance component image, wherein the reference pixel is the pixel in the original luminance component image that corresponds to the position of the distorted pixel; The optimized luminance component map and the chrominance component map are fused together to form a real-time image for display.

2. The endoscopic image optimization method according to claim 1, characterized in that, Determining distorted pixels from the edge regions of the luminance component map includes: Obtain the target distortion threshold that is positively correlated with the noise intensity of the luminance component map; For the pixel to be analyzed in the edge region, the degree of distortion of the pixel to be analyzed is obtained; If the degree of distortion is greater than the target distortion threshold, then the pixel to be analyzed is determined to be a distorted pixel.

3. The endoscopic image optimization method according to claim 2, characterized in that, Obtaining a target distortion threshold that is positively correlated with the noise intensity of the luminance component map includes: Obtain the preset base distortion threshold; The target distortion threshold is obtained by adjusting the base distortion threshold using the noise intensity.

4. The endoscopic image optimization method according to claim 2, characterized in that, Obtaining the distortion level of the pixel to be analyzed includes: Obtain the pixel residual between the pixel to be analyzed and the corresponding pixel in the original luminance component image; The pixel residual is used as the degree of distortion of the pixel to be analyzed.

5. The endoscopic image optimization method according to claim 1, characterized in that, Obtain the brightness component map of the image to be processed, including: The original luminance component map is separated from the image to be processed and filtered to obtain a filtered luminance map; Obtain the motion ratio between the filtered brightness map and the reference brightness map, wherein the reference brightness map represents the brightness component map output based on the previous frame; Based on the motion ratio, a first weight is determined for fusing the reference brightness map, wherein the first weight is inversely correlated with the motion ratio; The reference luminance map and the filtered luminance map are fused using the first weight to obtain the denoised luminance component map.

6. The endoscopic image optimization method according to claim 5, characterized in that, The reference luminance map and the filtered luminance map are fused using the first weight to obtain the denoised luminance component map, including: Identify the edge regions in the filtered brightness map; The non-edge regions of the reference luminance map and the non-edge regions of the filtered luminance map are fused using the first weight, and the edge regions of the reference luminance map and the edge regions of the filtered luminance map are fused using the second weight to obtain the denoised luminance component map, wherein the second weight is less than the first weight.

7. The endoscopic image optimization method according to claim 6, characterized in that, Identifying the edge regions in the filtered brightness map includes: By performing edge detection on the filtered brightness map, the actual edges in the filtered brightness map are obtained; The actual edge is morphologically dilated to obtain the edge region in the filtered brightness map.

8. The endoscopic image optimization method according to claim 1, characterized in that, The method further includes: The pre-defined cache queue stores historical images captured by the endoscope. In response to the freeze operation on the real-time image, multiple adjacent images are selected from a preset cache queue, wherein the multiple adjacent images are images that are within a preset time window from the real-time image and whose clarity meets preset conditions. An image to be fused is selected from multiple adjacent images, wherein the image to be fused represents an adjacent image whose overlap with the real-time image is higher than a threshold; The real-time image is fused with the image to be fused to obtain a frozen image.

9. The endoscopic image optimization method according to claim 8, characterized in that, The real-time image is fused with the image to be fused to obtain a frozen image, including: The sharpness of the image to be fused, the overlap rate between the image and the real-time image, and the alignment penalty coefficient between the image and the real-time image are obtained, wherein the alignment penalty coefficient characterizes the degree of alignment between the image to be fused and the real-time image; The sharpness, the overlap rate, and the alignment penalty coefficient are mapped to the fusion weights of the images to be fused. Based on the fusion weights of the images to be fused, the real-time image and the images to be fused are fused into the frozen image.

10. The endoscopic image optimization method according to claim 8, characterized in that, Selecting the image to be fused from the plurality of adjacent images includes: For each of the adjacent images, the adjacent images are pixel-aligned with the real-time image to obtain multiple pixel pairs with spatial correspondence. The proportion of similar pixel pairs among multiple pixel pairs is counted and used as the similarity proportion in the adjacent images, wherein the similar pixel pairs represent pixel pairs whose pixel values ​​are within a preset difference range; Based on the similarity ratio with each of the adjacent images, images with a similarity threshold greater than a certain threshold are selected as the images to be fused.

11. The endoscopic image optimization method according to claim 8, characterized in that, The method further includes: A base image and a detail image are separated from the frozen image, wherein the base image carries structural and brightness information of the frozen image, and the detail image carries texture and edge information of the frozen image; The dark areas of the base image are enhanced in brightness, and the bright areas are suppressed in brightness to obtain an optimized base image. The regions to be enhanced in the detailed image are selected, and the different color channels of the regions to be enhanced are optimized to obtain the optimized detailed image. The optimized base image and the optimized detail image are fused together to obtain the optimized frozen image.

12. The endoscopic image optimization method according to claim 11, characterized in that, Filtering out the regions to be enhanced located in the detailed image includes: Identify the initial enhancement region located in the detailed image; The region to be enhanced is filtered out from the initial enhancement region using the optimized base image.

13. The endoscopic image optimization method according to claim 11, characterized in that, The different color channels of the region to be enhanced are optimized to obtain an optimized detail image, including: The green channel of the region to be enhanced is fully enhanced, and the blue channel of the region to be enhanced is enhanced proportionally, while the red channel of the region to be enhanced remains unchanged, to obtain the optimized detail image.

14. An endoscopic image optimization device, characterized in that, The device includes: The image separation module is used to obtain the luminance component map and chrominance component map of the image to be processed, wherein the image to be processed is the original image acquired by the endoscope, and the luminance component map is the original luminance component map of the original image after noise reduction processing. A pixel recognition module is used to determine distorted pixels from the edge region of the luminance component map, wherein the distorted pixels represent pixels whose information is distorted due to the noise reduction process; An image optimization module is used to restore the distorted pixel using a reference pixel in the original luminance component image to obtain an optimized luminance component image, wherein the reference pixel is the pixel in the original luminance component image that corresponds to the position of the distorted pixel; An image fusion module is used to fuse the optimized luminance component image and the chrominance component image into a real-time image for display.

15. A storage medium, characterized in that, The storage medium stores a computer program that, when executed by a processor, implements the endoscopic image optimization method according to any one of claims 1-13.

16. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory storing a computer program that, when executed by the processor, implements the endoscopic image optimization method according to any one of claims 1-13.