Endoscopic Image Enhancement Method Based on Unsupervised Learning

The integration of classical image enhancement techniques with a no-supervised learning network enhances endoscope images, addressing the limitations of existing methods by improving contrast, clarity, and color richness for better medical diagnosis.

CN113808057BActive Publication Date: 2025-07-15SHANGHAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202110885619.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-03
Publication Date
2025-07-15
Estimated Expiration
2041-08-03

AI Technical Summary

Technical Problem

Existing endoscopic image enhancement methods cannot simultaneously improve the contrast, brightness uniformity, clarity and color nature of images, resulting in a decrease in doctors' diagnostic accuracy.

Method used

The endoscopic image enhancement method based on unsupervised learning is adopted, and the original image is processed through CLAHE, Gamma and LIME image enhancement technology, combined with the unsupervised learning network DerivedFuse for feature extraction, fusion and reconstruction, adaptive nonlinear stretching processing, and finally output the enhanced image.

Benefits of technology

It significantly improves the contrast, clarity and color richness of the endoscopic image, improves the visual effect, and improves the accuracy of clinical diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113808057B_ABST
    Figure CN113808057B_ABST
Patent Text Reader

Abstract

The present invention relates to an endoscope image enhancement method based on unsupervised learning, which includes four steps: (1) preprocessing the original endoscope image dataset, and using three image enhancement techniques to process the original image to obtain three derivative images; (2) converting the original image and its corresponding derivative images to the HSI color space, keeping the H-channel image unchanged, inputting the I-channel images of the derivative images into an unsupervised learning network, and performing deep network model training; (3) obtaining the I-channel image enhancement result according to the training parameters obtained after network training; (4) performing adaptive non-linear stretching processing on the S-channel image of the original image, and converting the HSI color space back to the RGB color space to output the final enhanced image. The unsupervised learning network of the present invention does not require a ground truth as a reference image, and the network training is realized through a non-reference loss function. This method has significantly improved in terms of contrast, clarity, and detail information compared with the existing methods, and has high clinical application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of medical image enhancement and computer network, and particularly relates to an endoscopic image enhancement method based on unsupervised learning. Background Art

[0002] Endoscopic imaging is a diagnostic and medical procedure. Doctors can directly observe the tissue morphology and lesion conditions of human internal organs through an electronic endoscope. Endoscopes have been widely used in the examination, diagnosis, and treatment of the esophageal digestive system such as the stomach and intestines. The quality of endoscopic images directly affects doctors' accurate analysis and diagnosis of diseases. However, due to the limitations of lighting conditions and the complex environment of human internal organs, the images directly obtained by endoscopes often have problems such as weak texture features, uneven brightness, and low contrast, resulting in the lack of some tissue morphological features in the human body. Thereby, it affects the accuracy of doctors' disease analysis and diagnosis. Therefore, the research on endoscopic image enhancement is of great significance for assisting doctors in diagnosis.

[0003] In recent years, researchers have proposed some classic image enhancement methods to improve the quality of endoscopic images. With the rapid development of deep learning, deep convolutional neural networks have gradually become the main driving force in the field of image enhancement. Such methods establish a complex non-linear mapping relationship between low-quality images and high-quality images through the learning of deep convolutional neural networks, so as to achieve the purpose of enhancing low-quality images. Due to the complex characteristics of endoscopic images, the research on endoscopic image enhancement technology mainly focuses on aspects such as lighting adjustment, improvement of contrast and clarity. As a difficult point in image enhancement, existing methods cannot obtain comprehensive performances such as enhanced contrast, uniform brightness, high clarity, and natural color, and the enhancement effect on endoscopic images is not good. Summary of the Invention

[0004] The purpose of the present invention is to propose an endoscopic image enhancement method based on unsupervised learning for the deficiencies of the existing technology. It is a method for enhancing endoscopic images based on Python language, Matlab language, and Pytorch framework. This method not only makes the useful detail information and color information in endoscopic images richer, but also improves the image contrast and clarity.

[0005] To achieve the above purpose, the present invention adopts the following technical scheme:

[0006] An endoscopic image enhancement method based on unsupervised learning, the steps are as follows:

[0007] Step 1: Preprocess the original endoscopic image dataset, and use three image enhancement technologies to process the original image to obtain three derivative images;

[0008] Step 2: Convert the original endoscopic image and its corresponding derived image to the HSI color space. Keep the H-channel image unchanged, and input the I-channel image of the derived image into the unsupervised learning network DerivedFuse for deep network model training;

[0009] Step 3: Obtain the enhanced result of the I-channel image according to the training parameters obtained after network training;

[0010] Step 4: Perform adaptive non-linear stretching on the S-channel image of the original endoscopic image, and convert the HSI color space image back to the RGB color space to output the final enhanced image;

[0011] Preferably, the preprocessing in Step 1 includes the following operations:

[0012] 1-1: Adjust the resolution of the original endoscopic image to 256×256 pixels;

[0013] 1-2: Use three classic image enhancement techniques - CLAHE, Gamma, and LIME to process the original image to obtain the corresponding derived image;

[0014] 1-3: Divide the obtained dataset into a training set, a validation set, and a test set in a ratio of 3:1:1.

[0015] Preferably, the network model training in Step 2 includes the following operations:

[0016] 2-1: Convert the RGB color space to the HSI color space;

[0017] 2-2: The unsupervised learning network DerivedFuse consists of three parts: feature extraction, feature fusion, and reconstruction. Input the I-channel image of the derived image in the dataset into the unsupervised learning network for training, and finally obtain the prediction map;

[0018] 2-3: The objective function is the non-reference loss function, and its value is affected by three components: luminance, contrast, and structure of the input network image;

[0019] 2-4: The network model adopts an optimization algorithm with a learning rate of 0.002 and is trained for a total of 100 epochs.

[0020] Preferably, the enhanced result of the I-channel image in Step 3 includes the following operations:

[0021] 3-1: After 100 epochs of training, obtain the corresponding network model training parameters;

[0022] 3-2: Input the test set into the trained network model to obtain the enhanced I-channel image predicted by the model.

[0023] Preferably, the adaptive non-linear stretching in step 4 includes the following operations:

[0024] 4-1: Calculate the maximum value M(R, G, B), minimum value m(R, G, B), and average value mean(R, G, B) of the R, G, and B color components of the corresponding pixel points in the RGB color space of the original endoscopic image;

[0025] 4-2: Perform adaptive non-linear stretching processing on the S-channel image information S of the original endoscopic image in the HSI color space Original to obtain the S-channel map S with adjusted saturation enhanced ;

[0026] 4-3: Integrate the brightness component obtained in step 3-2 and the saturation component obtained in step 4-2 with the original hue component, and inverse transform it to the RGB color space to obtain the final enhanced image.

[0027] Compared with the prior art, the present invention has the following obvious outstanding substantive features and remarkable advantages:

[0028] 1. The method of the present invention combines the advantages of Gamma images, CLAHE images, and LIME images with deep learning, and proposes an unsupervised derived image fusion network DerivedFuse for fine-detail fusion of derived images. This model can accurately extract and fuse useful features of the derived images without the need for groundtruth.

[0029] 2. Compared with the existing methods, the present invention can make the enhanced image have strong contrast, clear details, rich and natural colors, significantly improving the visual effect of endoscopic images, which is of great significance for clinical applications. Description of the Drawings

[0030] Figure 1 is the program block diagram of the method of the present invention.

[0031] Figure 2 is the overall flowchart of the method of the present invention.

[0032] Figure 3 is the architecture of the unsupervised learning network DerivedFuse of the present invention.

[0033] Figure 4 is the enhanced result diagram of the endoscopic image dataset by the network model trained by the method of the present invention.

[0034] Figure 5The prediction of the endoscopic image dataset by the network model trained by the method of the present invention obtains a comparison graph of the enhanced result graph and the results of multiple existing methods. Detailed implementation manners

[0035] The preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0036] See Figure 1 , an endoscopic image enhancement method based on unsupervised learning, comprising the following operation steps:

[0037] Step 1: Preprocess the original endoscopic image dataset, and process the original image with three image enhancement techniques to obtain three derived graphs, including the following operations:

[0038] 1-1: Adjust the resolution of the original endoscopic image to 256×256 pixels;

[0039] 1-2: Use three classic image enhancement techniques - CLAHE, Gamma, and LIME to process the corresponding derived graphs of the original image;

[0040] 1-3: Divide the obtained dataset into a training set, a validation set, and a test set according to a ratio of 3:1:1.

[0041] Step 2: Convert the original endoscopic image in the RGB color space and its corresponding derived graphs to the HSI color space, keep the H-channel image unchanged, and input the I-channel image of the derived graph into an unsupervised learning network for deep network model training, including the following operations:

[0042] 2-1: Convert from the RGB color space to the HSI color space, and the conversion formula is as follows:

[0043]

[0044] Among them, R, G, and B respectively represent the red, green, and blue color components in the RGB color space, and H, S, and I respectively represent the chrominance component, saturation component, and brightness component in the HSI color space.

[0045] 2-2: The unsupervised learning network DerivedFuse consists of three parts, namely feature extraction, feature fusion, and reconstruction. Input the I-channel image of the derived graph in the dataset into the unsupervised learning network for training, and finally obtain a prediction graph. The complete network structure is as Figure 3 shown:

[0046] The DerivedFuse network model proposed by the present invention mainly consists of three parts: feature extraction, feature fusion, and reconstruction.

[0047] The feature extraction module consists of two layers of partial convolutional layers. The size of the convolutional kernel in the first layer is set to 5×5, and the number of channels is 16. The size of the convolutional kernel in the second layer is 7×7, and the number of channels is 32. The 5×5 convolutional kernel can be used to extract low-level features of the image. The receptive field of the 7×7 convolutional kernel is larger, which is used to capture more complex and abstract information of the image. By extracting feature maps through convolutions of different scales, features in different aspects of the data can be extracted from the three derived image luminance components, improving the network training speed.

[0048] The feature fusion module simply combines each feature map by adding the pixels of each feature map through a fusion layer. The features of L21, L22, and L23 are fused using the fusion layer. The size of the fused image remains unchanged at 256×256 and contains 32 channels.

[0049] For the image reconstruction module, the present invention uses a U-Net model to extract deep features of the image, which includes 7 upsampling layers and downsampling layers respectively. The input to the network is a feature map with 64 channels and a size of 256×256. The output is a luminance map with 1 channel and a size of 256×256. The 7-layer downsampling convolutional layer uses a 4×4 convolutional kernel, with a stride of 2 and padding of 1, and the Leaky LeRU activation function is adopted. The size of the convolutional kernel used in the first six upsampling convolutional layers is 4×4, with a stride of 2 and padding of 1, and the LeRU activation function is adopted. The last deconvolution layer uses the Tanh activation function to generate a final luminance image with more complete detail preservation.

[0050] 2-3: The objective function is a non-reference loss function, and its value is affected by the luminance, contrast, and structure components of the input network image;

[0051] The expression of the objective function is as follows:

[0052]

[0053] where, i f represents the image patch extracted from the same spatial position in the fused image actually output from DerivedFuse, represents the expected fused image patch. is the variance of. is and i f the covariance of. C represents a constant. The calculation formula of is and respectively represent the desired luminance, contrast, and structure components.

[0054] SSIM decomposes any given image patch into three components: luminance, contrast, and structure. Then, for the input image patches i extracted at the same spatial position in the I-channel image of the HSI color space for the derived map n n can be represented by Equation (3):

[0055]

[0056] where n represents the number of derived maps. ||·|| represents the l 2 norm of the vector. represents the average value of i n . l n , c n and s n respectively represent the luminance, contrast, and structure components of i n .

[0057] To obtain a high-contrast fused image, the maximum contrast among three different image patches in the I-channel of the HSI color space is selected as the desired contrast of i f . The detailed expression of is as follows:

[0058]

[0059] where max represents taking the mean value.

[0060] To fuse the texture structure information of the CLAHE-derived map and the LIME-derived map, the detailed expression of the desired structure is as follows:

[0061]

[0062] To obtain a high-luminance fused image, the average luminance of three different image patches in the I-channel of the HSI color space is selected as the desired luminance of i f . The detailed expression of is as follows:

[0063]

[0064] where mean represents taking the mean value.

[0065] 2-4: The network model adopts an optimization algorithm with a learning rate of 0.002 and is trained for a total of 100 epochs.

[0066] Step 3: According to the training parameters obtained after network training, obtain the enhanced result of the I-channel image, including the following operations:

[0067] 3-1: After 100 epochs of training, obtain the corresponding network model training parameters;

[0068] 3-2: Input the test set into the trained network model to obtain the enhanced I-channel image I predicted by the model fused .

[0069] Step 4: Perform adaptive non-linear stretching processing on the S-channel image of the original endoscopic image and convert the HSI color space image back to the RGB color space to output the final enhanced image, including the following operations:

[0070] 4-1: Calculate the maximum value M(R, G, B), minimum value m(R, G, B), and average value me(R, G, B) of the R, G, and B color components of the corresponding pixel points in the RGB color space of the original endoscopic image;

[0071] 4-2: Perform adaptive non-linear stretching processing on the S-channel image information S of the original endoscopic image in the HSI color space Original to obtain the S-channel map S with adjusted saturation enhanced , and the adaptive non-linear stretching function constructed in the present invention is defined as: Among them, M(R, G, B) can be obtained through the function max in matlab, m(R, G, B) can be obtained through the function min in matlab, and me(R, G, B) can be obtained by calculating the sum of the pixel values of each pixel point in the three channels of the image divided by the total number of pixel points in the image;

[0072] 4-3: As Figure 2 shown, integrate the original hue component with the brightness component I obtained in 3-2 fused and the saturation component S obtained in 4-2 enhanced and inverse-transform to the RGB color space to obtain the final enhanced image. The formula for converting to the RGB color space is as follows:

[0073]

[0074] Among them, R, G, and B respectively represent the red, green, and blue color components in the RGB color space, and H, S, and I respectively represent the chromaticity component, saturation component, and brightness component in the HSI color space.

[0075] The method of this embodiment can accurately extract and fuse the useful features of the derivative graph, and can complete the enhancement of endoscopic images without ground truth.

[0076] The method of this embodiment selects a part of images from the public endoscopic datasets Kvasir dataset, Kvasir-SEG, CVC-ClinicDB, ETIS-Larib Polyp DB, CVC-EndoSceneStill, and CVC-ClinicSpec to verify the network efficiency. The endoscopic images are enhanced by using the method of the present invention, and are compared with the histogram equalization algorithm (HE), contrast-limited adaptive histogram equalization (CLAHE), single-scale Retinex algorithm (SSR), multi-scale Retinex algorithm with color restoration (MSRCR), multi-scale Retinex with color protection (MSRCP), Structure-revealing low-light enhancement method based on the Rubost Retinex decomposition model (RRM), method of enhancing images by estimating the illumination map of low-brightness images (LIME), adaptive gamma correction weighted distribution (AGCWD), method proposed by Al-Ameen, and unsupervised learning methods Zero-DCE, Zero-DCE++, EnlightenGAN in terms of enhancement effects. The enhancement effect of the method of the present invention is as Figure 4 shown. It can be seen from Figure 4 the enhancement results that the method of the present invention has significantly improved in terms of contrast, clarity, and saturation, and has high clinical application value. Figure 5 Figure is a comparison of the prediction results of the method of the present invention and existing methods. It can be seen from Figure 5 the comparison results that the method of the present invention has a better enhancement effect on the image texture details when processing endoscopic images, and has the best visual effect.

[0077] The above has described the embodiments of the present invention in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments, and various changes can be made according to the purpose of the invention of the present invention. Any changes, modifications, substitutions, combinations, or simplifications made based on the spirit and principle of the technical solution of the present invention shall be equivalent replacement methods, as long as they meet the invention purpose of the present invention and do not depart from the technical principle and inventive concept of the present invention, they all belong to the protection scope of the present invention.

Claims

1. An endoscopic image enhancement method based on unsupervised learning, characterized in that, The operation steps are as follows: Step 1: Preprocess the original endoscopic image dataset, and use three image enhancement techniques to process the original image to obtain three derived images; Step 2: Convert the original endoscopic image and its corresponding derived images to the HSI color space, keep the H-channel image unchanged, input the I-channel images of the derived images into an unsupervised learning network for deep network model training; Step 3: Obtain the enhanced result of the I-channel image according to the training parameters obtained after network training; Step 4: Perform adaptive non-linear stretching on the S-channel image of the original endoscopic image, and convert the HSI color space image back to the RGB color space to output the final enhanced image; The network model in Step 2 includes the following specific operation steps: 2-1: Convert the RGB color space to the HSI color space; 2-2: An unsupervised learning network, called DerivedFuse, which consists of three parts: feature extraction, feature fusion, and reconstruction; input the I-channel images of the derived images in the dataset into the unsupervised learning network DerivedFuse for training to finally obtain a predicted image; 2-3: The objective function is a non-reference loss function, and its value is affected by three components: luminance, contrast, and structure of the input network image; The expression of the objective function is as follows: where i f represents an image patch extracted from the fused image actually output from DerivedFuse at the same spatial position, represents the desired fused image patch; is the variance of; is and i f the covariance of; C represents a constant; The calculation formula of and represent the desired luminance, contrast, and structure components respectively; SSIM decomposes any given image patch into three components: luminance, contrast, and structure. The input image patches i extracted at the same spatial location in the I-channel image of the HSI color space for the derived map n n can be expressed by Equation (3): Among them, n represents the number of derivative diagrams; ||·|| represents the l 2 norm of the vector; represents the average value of i n ; l n , c n and s n respectively represent the luminance, contrast, and structure components of i n ; To obtain a fused image with high contrast, the maximum contrast of three different image patches in the I channel of the HSI color space is selected as the expected contrast of i f The detailed expression of is as follows: where, max means taking the mean; To integrate the texture structure information of the CLAHE-derived map and the LIME-derived map, the desired structure has the following detailed expression: To obtain a fused image with high brightness, the average luminance of three different image patches in the I channel of the HSI color space is selected as the expected luminance of i f The detailed expression of is as follows: where, mean means taking the mean; 2-4: The network model adopts an optimization algorithm with a learning rate of 0.002 and is trained for a total of 100 epochs.

2. The endoscopic image enhancement method based on unsupervised learning according to claim 1, wherein, The preprocessing in Step 1 includes the following specific operation steps: 1-1: Adjust the resolution of the original endoscopic image to 256×256 pixels; 1-2: Use three classic image enhancement techniques, namely CLAHE, Gamma, and LIME, to process the original image to obtain corresponding derived images; 1-3: Divide the obtained dataset into a training set, a validation set, and a test set in a ratio of 3:1:

1.

3. An endoscopic image enhancement method based on unsupervised learning according to claim 1, characterized in that, The enhanced result of the I-channel image in Step 3 includes the following specific operation steps: 3-1: After 100 epochs of training, obtain the corresponding network model training parameters; 3-2: Input the test set into the trained network model to obtain the I-channel enhanced image predicted by the model.

4. A method for endoscopic image enhancement based on unsupervised learning according to claim 1, characterized in that The adaptive non-linear stretching in Step 4 includes the following specific operation steps: 4-1: Calculate the maximum value, minimum value, and average value of the R, G, and B color components of the corresponding pixel points of the original endoscopic image in the RGB color space; 4-2: Perform adaptive non-linear stretching on the S-channel image information of the original endoscopic image in the HSI color space to obtain an S-channel map with adjusted saturation; 4-3: Integrate the brightness component obtained in Step 3-2 and the saturation component obtained in Step 4-2 with the original hue component, and inverse-convert it to the RGB color space to obtain the final enhanced image.