Endoscope image enhancement method based on joint retinex mechanism

By using a combined retinex mechanism and a three-branch full convolutional neural network to separate image components in endoscopic image processing, combined with the denoising module and the CBAM module, the scene generalization performance and suppressing image distortion in low-illumination endoscopic image processing are solved, and a more efficient image enhancement effect is achieved.

CN119991487APending Publication Date: 2025-05-13CHONGQING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202510084647.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

When processing low-illumination endoscopic images, the scene generalization performance and suppression of image distortion are insufficient, resulting in unsatisfactory image enhancement effect.

Method used

Using the endoscopic image enhancement method based on the combined retinex mechanism, the reflection components, illumination components and noise are separated by a three-branch fully convolutional neural network, the noise-free reflectivity is calculated and the illumination adjustment is performed. Combined with the denoising module and the CBAM module, the model is trained through the model total loss to generate more accurate endoscopic enhancement images.

Benefits of technology

It effectively improves the visual quality of low-illumination endoscopic images, enhances the brightness and clarity of the images, improves the doctor's diagnostic accuracy, and optimizes the retention of image details.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991487A_ABST
    Figure CN119991487A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image processing, and particularly relates to an endoscope image enhancement method based on a joint retinex mechanism. The method comprises the following steps: acquiring an original endoscope image and inputting the original endoscope image into a three-branch full convolutional neural network for processing to obtain a reflection component, an illumination component and noise; calculating noiseless reflectivity according to the original endoscope image, the illumination component and the noise; adjusting the illumination component, and performing point multiplication on the noiseless reflectivity and the adjusted illumination component to obtain an initial recovery image; the sampling and denoising module processes the initial recovery image to obtain an endoscope enhanced image; calculating the total loss of the model and adjusting model parameters according to the total loss of the model to obtain a trained endoscope image enhancement model; an endoscope enhanced image can be obtained by using the model of the training number; according to the method, the endoscope image enhancement effect is improved, and the method has a good application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of image processing, and in particular relates to an endoscopic image enhancement method based on a joint retinex mechanism. Background Art

[0002] Endoscopic technology has become one of the key tools for diagnosis and treatment in modern medicine, and is widely used in many medical fields such as the digestive tract, respiratory tract, and gynecology. Nowadays, high-definition endoscopes can not only improve the efficiency of medical diagnosis, but also greatly improve patients' treatment experience and recovery. However, sometimes the light source of the endoscope may not be enough to provide sufficient brightness, or in certain organs or deep tissues, the light may be blocked or scattered, resulting in a decrease in the light received by the endoscope, causing the image to become dim. Therefore, to address this problem, image enhancement technology is needed to improve the brightness and clarity of the image, so as to improve the doctor's diagnostic accuracy.

[0003] There are many mainstream low-light image enhancement technologies. One of them is to use convolutional neural networks to learn the salient features of low light and perform image restoration and enhancement. The use of convolutional neural networks can effectively improve the brightness, contrast and detail retention of images. The Retinex algorithm is a classic image enhancement algorithm. The Retinex theory is a traditional image enhancement algorithm. Its core idea is that the human visual system does not simply perceive color and brightness based on the intensity of the received light, but perceives changes in color and brightness by comparing the light intensity at different locations around a pixel. This comparison usually includes the light intensity in the same area and the light intensity between different areas. Combining neural networks with the retinex mechanism can further improve the visual quality of low-light images. Compared with the algorithm based on histogram equalization, the algorithm based on the retinex theory has obvious improvements in computational efficiency and robustness, but it still has obvious deficiencies in scene generalization performance and suppression of image distortion. Summary of the invention

[0004] In view of the shortcomings of the prior art, the present invention proposes an endoscopic image enhancement method based on a joint retinex mechanism, the method comprising: obtaining an endoscopic image to be processed and inputting the image into a trained endoscopic image enhancement model to obtain an endoscopic enhanced image;

[0005] The training process of the endoscopic image enhancement model includes:

[0006] S1: Obtain the original endoscopic image and input it into a three-branch fully convolutional neural network for processing to obtain the reflection component, illumination component and noise;

[0007] S2: Calculate the noise-free reflectance based on the original endoscopic image, illumination components, and noise;

[0008] S3: Adjust the illumination component, perform dot multiplication of the noise-free reflectance and the adjusted illumination component to obtain an initial restored image;

[0009] S4: The sampling and denoising module processes the initial restored image to obtain an endoscope enhanced image;

[0010] S5: Calculate the total model loss and adjust the model parameters according to the total model loss to obtain a trained endoscopic image enhancement model.

[0011] Preferably, the formula for calculating the noise-free reflectivity is expressed as:

[0012]

[0013] in, represents the noise-free reflectivity of pixel x, I(x) represents the original endoscopic image of pixel x, N(x) represents the noise of pixel x, and S(x) represents the illumination component of pixel x.

[0014] Preferably, the process of the denoising module processing the initial restored image includes:

[0015] Four encoders are used to process the initial restored image to obtain the encoded features; each encoder consists of a convolutional layer, a pooling layer and a CBAM module;

[0016] Four decoders are used to process the encoded features, and the last three decoders are connected to the corresponding first three encoders by skip connections to obtain the decoded features; each decoder consists of an upsampling layer and a convolutional layer;

[0017] The decoded features are processed by back diffusion to obtain the endoscope enhanced image.

[0018] Furthermore, the CBAM module includes a channel attention module and a spatial attention module;

[0019] The data processing process of the channel attention module includes: performing global maximum pooling and global average pooling on the input features in the spatial dimension to obtain two channel pooling features; inputting the two channel pooling features into the multi-layer perceptron for processing to obtain two initial channel attention maps; adding the two initial channel attention maps and normalizing them to obtain the channel attention weight matrix; multiplying the channel attention weight matrix with the input features to obtain the final channel attention map;

[0020] The data processing process of the spatial attention module includes: performing global maximum pooling and global average pooling on the channel attention map along the channel direction to obtain two spatial pooling features; splicing the two spatial pooling features, and processing the splicing result through a convolution layer with a convolution kernel size of 7×7 to obtain a spatial attention weight matrix; multiplying the spatial attention weight matrix with the channel attention map to obtain the fine feature, that is, the output feature of the CBAM module.

[0021] Preferably, the total model loss is a weighted sum of reconstruction loss, texture enhancement loss and noise estimation loss.

[0022] Furthermore, the reconstruction loss is expressed as:

[0023] L r =||I-(R·S+N)||+||SS 0 ||+||RI / S||

[0024] Among them, L r represents the reconstruction loss, I represents the original endoscopic image, R represents the reflection component, S represents the illumination component, and S 0 represents the maximum pixel value, and N represents the noise.

[0025] Furthermore, the texture enhancement loss is expressed as:

[0026]

[0027] Among them, L t represents the texture enhancement loss, w x represents the horizontal weight, w y represents the vertical weight, represents the horizontal gradient, represents the vertical gradient, and S represents the illumination component.

[0028] Furthermore, the noise estimation loss is expressed as:

[0029]

[0030] Among them, L n represents the noise estimation loss, w n represents the horizontal guidance weight, N represents the noise, and w r represents the vertical guidance weight, λ n represents the adjustment factor, represents the horizontal gradient, represents the vertical gradient, R represents the reflection component, S represents the illumination component, and λ represents the regularization coefficient.

[0031] The beneficial effects of the present invention are as follows: the present invention designs an endoscopic image enhancement model to process the original endoscopic image, which combines the neural network with the retinex mechanism to further improve the visual quality of low-light images; the present invention designs a denoising module, and the forward diffusion process gradually adds noise to the input image until the original image is destroyed; the reverse diffusion promotes image reconstruction, retains image details, and optimizes the target. The present invention uses the weighted sum of the reconstruction loss, texture enhancement loss, and noise estimation loss as the total loss of the model to train the model, so that the model generates more accurate endoscopic enhanced images; the present invention can effectively process low-brightness endoscopes and improve the enhancement effect of endoscopic enhanced images. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 It is a structural block diagram of the endoscope image enhancement model in the present invention;

[0033] Figure 2 This is a structural diagram of the denoising module in the present invention;

[0034] Figure 3 It is the structural diagram of the CBAM module in the present invention;

[0035] Figure 4 Endoscopic image enhancement model. DETAILED DESCRIPTION

[0036] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0037] The present invention proposes an endoscopic image enhancement method based on a joint retinex mechanism, the method comprising the following contents:

[0038] The endoscopic image to be processed is obtained and input into the trained endoscopic image enhancement model to obtain an endoscopic enhanced image.

[0039] The training process of the endoscopic image enhancement model includes:

[0040] S1: The original endoscopic image is obtained and input into a three-branch fully convolutional neural network for processing to obtain the reflection component, illumination component and noise.

[0041] like Figure 1As shown, the three branches of the three-branch fully convolutional neural network are used to estimate the reflection component, illumination and noise respectively. The three-branch fully convolutional neural network has the same structure. Preferably, each branch is composed of multiple convolutional layers and SE modules; the last layer adopts different processing methods, and the reflection and illumination components end with a sigmoid layer to ensure that they fall within the range of [0, 1]. The last layer of the noise branch uses a tanh layer to reduce the noise to [-1, 1].

[0042] The three-branch fully convolutional neural network processes the original endoscopic image respectively to obtain the reflection component R, illumination component S and noise N.

[0043] S2: Calculate the noise-free reflectance based on the original endoscopic image, illumination components and noise.

[0044]

[0045] in, represents the noise-free reflectivity of pixel x, I(x) represents the original endoscopic image of pixel x, N(x) represents the noise of pixel x, and S(x) represents the illumination component of pixel x.

[0046] S3: Adjust the illumination component, perform dot multiplication of the noise-free reflectance and the adjusted illumination component to obtain an initial restored image.

[0047] Adjust the lighting components:

[0048]

[0049] in, Represents the adjusted illumination component, and γ is the core parameter for adjusting brightness set in advance.

[0050] Combining the adjusted illumination component and the noise-free reflectance, the final restored result can be calculated as:

[0051]

[0052] in, represents the initial restored image.

[0053] S4: The sampling and denoising module processes the initial restored image to obtain an endoscope enhanced image.

[0054] like Figure 2 As shown, the process of the denoising module processing the initial restored image includes:

[0055] Four encoders are used to process the initial restored image to obtain the encoded features; each encoder consists of a convolutional layer, a pooling layer and a CBAM module;

[0056] Four decoders are used to process the encoded features, and jump connections are used between the last three decoders and the corresponding first three encoders to obtain the decoded features; each decoder consists of an upsampling layer and a convolutional layer.

[0057] Among them, Figure 3 As shown, the CBAM module includes a channel attention module and a spatial attention module:

[0058] The data processing process of the channel attention module includes: performing global maximum pooling and global average pooling on the input features in the spatial dimension to obtain two channel pooling features; inputting the two channel pooling features into the multi-layer perceptron for processing to obtain two initial channel attention maps (size is 1×1×C); adding the two initial channel attention maps and normalizing them to obtain the channel attention weight matrix; multiplying the channel attention weight matrix with the input features to obtain the final channel attention map;

[0059] The data processing process of the spatial attention module includes: performing global maximum pooling and global average pooling on the channel attention map along the channel direction to obtain two spatial pooling features (size is 1×H×W); splicing the two spatial pooling features, and processing the splicing result through a convolution layer with a convolution kernel size of 7×7 to obtain a spatial attention weight matrix; multiplying the spatial attention weight matrix with the channel attention map to obtain the fine feature, that is, the output feature of the CBAM module.

[0060] The decoded features are processed by back diffusion to obtain the endoscope enhanced image.

[0061] The purpose of back diffusion is to recover detail information from the feature map during the decoding process. By gradually applying the back diffusion technique, low-resolution or lost feature information is gradually restored. The core idea of ​​back diffusion is to enhance or repair image features by transferring information from high layers to low layers, or by gradually restoring noise or details in the image. The back diffusion process can use Markov modeling; the decoded features may lose some details or edge information during the upsampling process, and back diffusion makes the final output image or feature map visually clearer or more accurate by learning how to restore this lost information.

[0062] S5: Calculate the total model loss and adjust the model parameters according to the total model loss to obtain a trained endoscopic image enhancement model.

[0063] In order to update the weight value of the decomposition network, a loss function is needed to estimate the current decomposition and guide the network to generate more accurate components. The total loss L of the model designed by the present invention consists of three parts, namely, the weighted sum of reconstruction loss, texture enhancement loss and noise estimation loss, expressed as:

[0064] L=λr L r +λ t L t +λ n L n

[0065] Among them, L t represents the reconstruction loss, L r represents the texture enhancement loss, L n represents the noise estimation loss, λ r ,λ t ,λ n is the weight factor of the corresponding loss component.

[0066] Reconstruction loss:

[0067] The decomposed components of the image must first satisfy the following conditions: I(x) = R(x)·S(x)+N(x); in Retinex theory, the maximum value of the R, G, and B channels is usually used. As an initial estimate of the illumination, the reflectance component is calculated by pixel-by-pixel division between the image and its illumination component. The reconstruction loss of the reflectance component can be expressed as:

[0068] L r =||I-(R·S+N)||+||SS 0 ||+||RI / S||

[0069] Among them, L r represents the reconstruction loss, I represents the original endoscopic image, R represents the reflection component, S represents the illumination component, and S 0 represents the maximum pixel value, N represents the noise; the reconstructed image is (R·S+N). ||X|| represents the calculation of the sum of the absolute values ​​of all items in X.

[0070] Texture enhancement loss:

[0071] In endoscopic images of normal brightness, the illumination component of the surface is usually relatively flat. A piecewise flat illumination component helps to enhance the texture of dark areas. This is because when the brightness of adjacent pixels is close, their contrast will be amplified by the same illumination value, which is distributed in [0, 1]. In order to ensure that the texture is enhanced, the texture enhancement loss is designed as:

[0072]

[0073] Among them, L t represents the texture enhancement loss, w x represents the horizontal weight, w y represents the vertical weight, represents the horizontal gradient, represents the vertical gradient, and S represents the illumination component.

[0074] The horizontal weights are:

[0075]

[0076] The vertical weight is:

[0077]

[0078] Where G represents a Gaussian filter, represents the convolution operator, I g is the grayscale image of the input image.

[0079] Noise estimation loss:

[0080] The illumination component guides the noise estimation loss. In the task of endoscopic underexposed image restoration, the contrast of dark areas is stretched to improve their visibility. But at the same time, the noise hidden in the dark areas will be amplified, so it is necessary to suppress the noise, especially in the dark areas. The illumination component of the image has been decomposed and can be used to guide the image denoising task, and can help the decomposition network focus on estimating the noise in the dark through weighting. The noise estimation loss is designed as:

[0081]

[0082] Among them, L n represents the noise estimation loss, w n represents the horizontal guidance weight, N represents the noise, and w r represents the vertical guidance weight, λ n represents the adjustment factor, represents the horizontal gradient, represents the vertical gradient, R represents the reflection component, S represents the illumination component, and λ represents the regularization coefficient; ||x|| F is the Frobenius norm of the matrix x.

[0083] w r and w n All are light-guided weighted items:

[0084] w n (x) = I(x)

[0085]

[0086] Among them, normalize means minimum-maximum normalization.

[0087] The loss designed by the present invention for noise estimation, i.e., the noise estimation loss, is based on two considerations. First, the range of the noise map needs to be limited. Second, the noise is suppressed by smoothing the reflection component. Unlike illumination smoothing, it focuses on points with small horizontal and vertical gradients, ensuring that real noise points rather than edges are smoothed. In order to estimate noise in the dark, the above two terms are weighted and limited by the illumination component.

[0088] The model parameters are adjusted according to the total loss of the model. When the loss function converges or reaches the maximum number of iterations, the training is stopped, the model parameters are saved, and the trained endoscopic image enhancement model is obtained. The endoscopic image to be processed is obtained and input into the trained endoscopic image enhancement model to obtain an endoscopic enhanced image; Figure 4 shown.

[0089] The above embodiments further illustrate the purpose, technical solutions and advantages of the present invention in detail. It should be understood that the above embodiments are only preferred implementation modes of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made to the present invention within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. An endoscopic image enhancement method based on a joint retinex mechanism, characterized in that: include: Acquire the endoscopic image to be processed and input it into the trained endoscopic image enhancement model to obtain an endoscopic enhanced image; The training process of the endoscopic image enhancement model includes: S1: Obtain the original endoscopic image and input it into a three-branch fully convolutional neural network for processing to obtain the reflection component, illumination component and noise; S2: Calculate the noise-free reflectance based on the original endoscopic image, illumination components, and noise; S3: Adjust the illumination component, perform dot multiplication of the noise-free reflectance and the adjusted illumination component to obtain an initial restored image; S4: The sampling and denoising module processes the initial restored image to obtain an endoscope enhanced image; S5: Calculate the total model loss and adjust the model parameters according to the total model loss to obtain a trained endoscopic image enhancement model.

2. The endoscopic image enhancement method based on the joint retinex mechanism according to claim 1, characterized in that: The formula for calculating the noise-free reflectivity is expressed as: in, represents the noise-free reflectivity of pixel x, I(x) represents the original endoscopic image of pixel x, N(x) represents the noise of pixel x, and S(x) represents the illumination component of pixel x.

3. The endoscopic image enhancement method based on the joint retinex mechanism according to claim 1, characterized in that: The process of the denoising module processing the initial restored image includes: Four encoders are used to process the initial restored image to obtain the encoded features; each encoder consists of a convolutional layer, a pooling layer and a CBAM module; Four decoders are used to process the encoded features, and the last three decoders are connected to the corresponding first three encoders by skip connections to obtain the decoded features; each decoder consists of an upsampling layer and a convolutional layer; The decoded features are processed by back diffusion to obtain the endoscope enhanced image.

4. The endoscopic image enhancement method based on the joint retinex mechanism according to claim 3, characterized in that: The CBAM module includes a channel attention module and a spatial attention module; The data processing process of the channel attention module includes: performing global maximum pooling and global average pooling on the input features in the spatial dimension to obtain two channel pooling features; inputting the two channel pooling features into the multi-layer perceptron for processing to obtain two initial channel attention maps; adding the two initial channel attention maps and normalizing them to obtain the channel attention weight matrix; multiplying the channel attention weight matrix with the input features to obtain the final channel attention map; The data processing process of the spatial attention module includes: performing global maximum pooling and global average pooling on the channel attention map along the channel direction to obtain two spatial pooling features; splicing the two spatial pooling features, and processing the splicing result through a convolution layer with a convolution kernel size of 7×7 to obtain a spatial attention weight matrix; multiplying the spatial attention weight matrix with the channel attention map to obtain the fine feature, that is, the output feature of the CBAM module.

5. The endoscopic image enhancement method based on the joint retinex mechanism according to claim 1, characterized in that: The total loss of the model is the weighted sum of reconstruction loss, texture enhancement loss and noise estimation loss.

6. The endoscopic image enhancement method based on the joint retinex mechanism according to claim 5, characterized in that: The reconstruction loss is expressed as: L r =||I-(R·S+N)||+||S-S0||+||R-I / S|| Among them, L r represents the reconstruction loss, I represents the original endoscopic image, R represents the reflection component, S represents the illumination component, S0 represents the maximum pixel value, and N represents the noise.

7. The endoscopic image enhancement method based on the joint retinex mechanism according to claim 5, characterized in that: The texture enhancement loss is expressed as: Among them, L t represents the texture enhancement loss, w x represents the horizontal weight, w y represents the vertical weight, represents the horizontal gradient, represents the vertical gradient, and S represents the illumination component.

8. The endoscopic image enhancement method based on the joint retinex mechanism according to claim 5, characterized in that: The noise estimation loss is expressed as: Among them, L n represents the noise estimation loss, w n represents the horizontal guidance weight, N represents the noise, and w r represents the vertical guidance weight, λ n represents the adjustment factor, represents the horizontal gradient, represents the vertical gradient, R represents the reflection component, S represents the illumination component, and λ represents the regularization coefficient.

Citation Information

Patent Citations

  • Depth Retinex image enhancement method under weak illumination condition

    CN115205146A

  • Self-adaptive low-light noise image enhancement method and system based on zero-order learning

    CN116579944A

  • Non-supervision-based low-illumination non-contact fingerprint enhancement method and device

    CN118135616A

  • Embedding an input image to a diffusion model

    US20240161462A1

Cited By

  • Automatic artifact detection and restoration method and system based on endoscope image

    CN121392066A