Low-light Digestive Endoscope Image Enhancement Method, Device and Equipment
By normalizing the low-light digestion endoscopic images and using the image enhancement network and denoising network for image enhancement and denoising, the recognition error and noise problems in low-light image processing are solved, and the image quality is improved.
Patent Information
- Application Number
- CN202510124877.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-27
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-01-27
AI Technical Summary
The prior art has problems such as image recognition errors, difficulty in automatically distinguishing image brightness, and inability to effectively process image noise when processing low-light digestion endoscopic images.
By normalizing the initial target image and feature extraction, the preset image enhancement network recognizes brightness features and generates an enhancement description matrix. Then, the feature vectors are denoised using a convolutional layer and an image denoising network, and image enhancement is performed according to the brightness characteristics through the image conversion model.
It effectively improves the quality of low-light digestion endoscopic images, reduces image recognition errors, and processes noise in the image, improving the overall brightness and clarity of the image.
Smart Images

Figure CN119559087B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular, to a low-light digestive endoscope image enhancement method, device, and equipment. Background Art
[0002] In the technical field of image processing and recognition, due to the influence of the environmental structure and environmental brightness during the image acquisition process, some low-light images are often collected. These low-light images often affect subsequent image recognition and cause errors in image recognition.
[0003] For example, the inside of the human body is completely dark. When performing gastroscopy and colonoscopy examinations, a cold light source needs to be used for illumination, and different parts of the digestive tract are examined in the order of light entry. However, due to factors such as the complexity of the digestive tract structure, the technical limitations of the cold light source, and inappropriate endoscopic parameter configurations, the brightness of digestive endoscope images is often too dark, which greatly affects the examination effect and the diagnosis of digestive tract diseases.
[0004] Currently, traditional image enhancement methods based on supervised learning use image pairs composed of low-light images and high-light images to construct a training set for training an AI model. Such methods have the following defects: 1) The requirements for the image pairs in this training set are very high. The content of the low-light images and high-light images used in the training set must be exactly the same and the positions of the pixel points must correspond exactly. However, it is difficult to collect such image pairs of this type of digestive endoscope. Therefore, the supervised learning solution is difficult to implement in actual implementation; 2) Such models cannot automatically distinguish whether the image brightness is too low and will process images with good brightness without discrimination, resulting in too high image brightness; 3) Such models only process the image brightness and cannot process the noise in the image. Summary of the Invention
[0005] In view of the above problems, the present invention provides a low-light digestive endoscope image enhancement method, device, and equipment to overcome or partially overcome the above problems.
[0006] In one aspect of the present invention, a low-light digestive endoscope image enhancement method is provided, and the method includes:
[0007] Normalize the pixel values of each pixel point in the initial target image to be enhanced to obtain a target image;
[0008] Based on a preset image enhancement network, perform feature extraction on the target image to identify the brightness features of each pixel point in the target image, and obtain an enhancement description matrix, where the enhancement description matrix is a two-dimensional matrix with the same scale as the target image;
[0009] The convolutional layer is respectively used for the target image and the enhanced description matrix to obtain feature vectors, the feature vectors are fused to obtain a fused image feature vector, the fused image feature vector is subjected to feature extraction, and the extracted eigenvalues are subjected to a denoising operation to obtain a denoised image feature vector;
[0010] Based on the enhanced description matrix, the denoised image feature vector is subjected to image conversion, so that the eigenvalues of the denoised image feature vector at each pixel point are enhanced to different degrees according to the brightness characteristics of the corresponding pixel points, and an enhanced image feature vector is obtained;
[0011] The eigenvalues of each pixel point of the enhanced image feature vector are subjected to a denormalization operation to obtain the target enhanced image of the initial target image.
[0012] Further, the image conversion of the denoised image feature vector based on the enhanced description matrix includes:
[0013] The preset image conversion model is used to implement the image conversion of the denoised image feature vector based on the enhanced description matrix. The image conversion model is expressed as:
[0014]
[0015] where p refers to the pixel point in the image, and Φ L (p) refers to the eigenvalue of the enhanced description matrix at the pixel point p, and L D (p) refers to the eigenvalue of the denoised image feature vector at the pixel point p, and L E (p) refers to the eigenvalue of the calculated enhanced image feature vector at the pixel point p.
[0016] Further, the image enhancement network is used to complete the following operations:
[0017] The target image is subjected to hierarchical iterative feature extraction through a plurality of sequentially connected feature extraction modules to obtain feature matrices of multiple channels for representing brightness features;
[0018] The feature matrices of multiple channels are subjected to a convolution operation to obtain a two-dimensional matrix of one channel, and the obtained two-dimensional matrix is used as the enhanced description matrix.
[0019] Further, before the feature extraction of the target image based on the preset image enhancement network to identify the brightness features of each pixel point in the target image, the method further includes:
[0020] Based on the total loss function of the preset image enhancement network, the image enhancement network is trained; the total loss function of the image enhancement network includes a self-supervised loss function with a preset first weight coefficient, and the self-supervised loss function is expressed as:
[0021]
[0022] Among them, 0 < the first weight coefficient ≤ 1, Loss s is the self-supervised loss, L refers to the low-light image, and F M refers to the image enhancement network, and F M (L) is the result output after the low-light image is processed by the image enhancement network. ω is a custom parameter for self-supervised learning, 0 < ω < 1, and Lω is the calculation result after optimizing the normalized pixel value of each pixel point in L using the parameter ω. F M (Lω) is the result output after Lω is processed by the image enhancement network. The calculation formula of Lω is as follows:
[0023]
[0024] Among them, L(p) is the magnitude of the normalized pixel value of a single pixel point of the low-light image.
[0025] Furthermore, the total loss function of the image enhancement network also includes a low-light image enhancement loss function with a preset second weight coefficient, and the low-light image enhancement loss function is expressed as:
[0026]
[0027] Among them, 0 < the second weight coefficient ≤ 1, Loss C is the low-light image enhancement loss, and L M = F M (L), and F M (L M )is the result output after the low-light image is processed by the image enhancement network and then processed by the image enhancement network for the second time on its output result. E is a matrix with the same scale as the enhancement description matrix and all elements being 1.
[0028] Furthermore, the total loss function of the image enhancement network also includes a high-light image enhancement loss function with a preset third weight coefficient, and the high-light image enhancement loss function is expressed as:
[0029]
[0030] Among them, 0 < the third weight coefficient ≤ 1, Loss H is the high-light image enhancement loss, H is the high-light image, and F M (H) is the result output after the high-light image is processed by the image enhancement network.
[0031] Further, obtaining feature vectors by respectively using convolutional layers for the target image and the enhanced description matrix, performing feature fusion on the feature vectors to obtain a fused image feature vector, performing feature extraction on the fused image feature vector, and performing a denoising operation on the extracted eigenvalues to obtain a denoised image feature vector includes:
[0032] Based on a preset image denoising network, the following operations are completed for the target image and the enhanced description matrix input to the image denoising network:
[0033] Performing a convolution operation on the target image to obtain a first convolution input image feature vector;
[0034] Performing a convolution operation on the enhanced description matrix to obtain a second convolution enhanced description feature vector;
[0035] Performing a concatenation operation on the first convolution input image feature vector and the second convolution enhanced description feature vector to obtain a fused image feature vector of the target image and the enhanced description matrix;
[0036] Performing a feature extraction operation and a soft threshold denoising operation on the fused image feature vector respectively through at least one feature extraction module and at least one soft threshold denoising module to obtain an initial denoised image feature vector; wherein, the soft threshold denoising module is used to perform the following operations:
[0037]
[0038] Wherein, v is the eigenvalue input to the soft threshold denoising module, δ is a preset soft threshold parameter, when v > 0, Sign(v) = 1, and when v < 0, Sign(v) = -1;
[0039] Performing a convolution operation on the initial denoised image feature vector to obtain at least one feature matrix with the same number of channels and scale as the target image, and taking the obtained at least one feature matrix as the denoised image feature vector.
[0040] Further, before respectively using convolutional layers to calculate and form feature vectors for the target image and the enhanced description matrix, the method further includes:
[0041] Training the model of the image denoising network based on the total loss function of the preset image denoising network;
[0042] The total loss function of the image denoising network is obtained by weighted addition of the low-light image denoising loss and the high-light image denoising loss; wherein,
[0043] The low-light image denoising loss function is expressed as:
[0044]
[0045] Among them, Loss D (L) refers to the low-light image denoising loss, and F D (L) is the result output after the low-light image is processed by the image denoising network, and L refers to the low-light image;
[0046] The high-light image denoising loss is expressed as:
[0047]
[0048] Among them, Loss D (H) refers to the high-light image denoising loss, and F D (H) is the result output after the high-light image is processed by the image denoising network, and H is the high-light image.
[0049] Furthermore, the feature extraction module is used to perform the following operations:
[0050] Perform hierarchical iterative convolutional feature extraction on the image features input to the feature extraction module through multiple feature extraction convolutional layers;
[0051] During the convolutional feature extraction process, adaptively adjust the weights of the channel features through a preset channel attention module to increase the weights of the channels used to represent brightness;
[0052] Perform element-wise addition on the image features obtained after performing convolutional feature extraction and the image features input to the feature extraction module to obtain the output features of the feature extraction module.
[0053] Another aspect of the present invention provides a low-light digestive endoscope image enhancement device, and the device includes:
[0054] An image processing module for normalizing the pixel values of each pixel point in the initial target image to be enhanced to obtain a target image;
[0055] An image enhancement module for performing feature extraction on the target image based on a preset image enhancement network to identify the brightness features of each pixel point in the target image, and obtaining an enhancement description matrix, where the enhancement description matrix is a two-dimensional matrix with the same scale as the target image;
[0056] An image denoising module for respectively using convolutional layers on the target image and the enhancement description matrix to obtain feature vectors, performing feature fusion on the feature vectors to obtain a fused image feature vector, performing feature extraction on the fused image feature vector, and performing a denoising operation on the extracted eigenvalues to obtain a denoised image feature vector;
[0057] An image conversion module, configured to perform image conversion on the denoised image feature vector based on the enhanced description matrix, so as to enhance the eigenvalues of the denoised image feature vector at each pixel point to different degrees according to the brightness feature of the corresponding pixel point, and obtain an enhanced image feature vector;
[0058] The image processing module is further configured to perform a denormalization operation on the eigenvalues of each pixel point of the enhanced image feature vector to obtain a target enhanced image of the initial target image.
[0059] On the other hand, the present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor; when the computer program is executed by the processor, the low-light digestive endoscope image enhancement method described above is implemented.
[0060] A low-light digestive endoscope image enhancement method, device and equipment provided by the present invention perform normalization processing on the pixel values of each pixel point in the initial target image to be enhanced to obtain a target image, and perform feature extraction on the target image to identify the brightness features of each pixel point in the target image, and obtain an enhanced description matrix; at the same time, the enhanced description matrix and the target image are respectively used with a convolutional layer to obtain feature vectors, the feature vectors are subjected to image fusion to obtain a fused image feature vector, and feature extraction is performed on the fused image feature vector, and the extracted eigenvalues are subjected to a denoising operation to obtain a denoised image feature vector; finally, image conversion is performed on the denoised image feature vector based on the enhanced description matrix to obtain an enhanced image feature vector. Since the enhanced description matrix can reflect the brightness features of each pixel point, when performing image conversion on the denoised image, the eigenvalues of the denoised image feature vector of the pixel points with lower brightness can be enhanced to a higher degree, thereby achieving the purpose of improving the image quality and enhancing the low-light image at the same time.
[0061] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the following specifically describes the embodiments of the present invention. Description of the Drawings
[0062] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. In the drawings:
[0063] Figure 1 It is a flowchart of a low-light digestive endoscope image enhancement method according to an embodiment of the present invention;
[0064] Figure 2It is a network architecture diagram of a low-light digestive endoscope image enhancement model according to an embodiment of the present invention;
[0065] Figure 3 It is a comparison diagram of enhanced images obtained by using the low-light digestive endoscope image enhancement method of the present invention; among them,
[0066] (a) is a low-light endoscope image of the stomach;
[0067] (b) is the enhanced endoscope image of the stomach;
[0068] (c) is a low-light endoscope image of the intestine;
[0069] (d) is the enhanced endoscope image of the colon;
[0070] Figure 4 It is a schematic structural diagram of a low-light digestive endoscope image enhancement device according to an embodiment of the present invention. Detailed implementation manners
[0071] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0072] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted in an idealized or overly formal sense unless specifically defined.
[0073] Embodiment 1
[0074] The embodiment of the present invention provides a low-light digestive endoscope image enhancement method, as Figure 1 shown, this low-light image enhancement method includes the following steps:
[0075] S1. Normalize the pixel values of each pixel point in the initial target image to be enhanced to obtain a target image;
[0076] It should be noted that in the embodiments of the present invention, the normalization process of the actual pixel values of each pixel point of the initial target image is specifically to divide the pixel values of each pixel point of the initial target image by 255, so that the normalized pixel values of each pixel point of the obtained target image are floating-point numbers greater than 0 and less than 1. It should be noted here that the Figure 2 specific embodiment of the present invention is a network model built with an RGB image as a specific embodiment. At this time, the initial target image has three channels, that is, each pixel point has pixel values of three colors. When performing the normalization operation, it is necessary to perform the normalization operation on the three pixel values of each pixel point. It can be understood that according to different application scenarios, the initial target images used are also different. The present invention is also applicable to the enhancement of other target images with different numbers of channels.
[0077] S2. Based on a preset image enhancement network, perform feature extraction on the target image to identify the brightness features of each pixel point in the target image, and obtain an enhanced description matrix, where the enhanced description matrix is a two-dimensional matrix with the same scale as the target image;
[0078] In a specific embodiment of the present invention, the identification of the brightness features of each pixel point in the target image can be specifically that for low-light pixel points, they are identified as values greater than 0 and less than 1. After a certain low-light pixel point is processed by the image enhancement network, the feature value of the enhanced description matrix at the corresponding pixel point is identified as 0.75, and for high-light (normal light) pixel points, they are identified as 1.
[0079] S3. Use convolutional layers on the target image and the enhanced description matrix respectively to obtain feature vectors, perform feature fusion on the feature vectors to obtain a fused image feature vector, perform feature extraction on the fused image feature vector, and perform a denoising operation on the extracted feature values to obtain a denoised image feature vector;
[0080] In a specific embodiment of the present invention, a cocat module is used to perform feature fusion on the feature vectors obtained after performing convolutional operations on the target image and the enhanced description matrix, further improving the denoising effect. And perform soft-threshold denoising on the feature values in the network parameters, compare the amplitude of the signal, when the amplitude of the signal is less than the set threshold, set the signal to zero; when the amplitude of the signal is greater than or equal to the threshold, retain the original signal, which can effectively suppress the noise in the signal while retaining the main features of the signal.
[0081] S4. Based on the enhanced description matrix, perform image conversion on the denoised image feature vector to enhance the feature values of the denoised image feature vector at each pixel point to different degrees according to the brightness features of the corresponding pixel points, and obtain an enhanced image feature vector;
[0082] In the embodiment of the present invention, the image conversion of the denoised image feature vector based on the enhanced description matrix mainly enhances the eigenvalues of the denoised image feature vector at low-light pixels to different degrees according to the degree of low light. For the eigenvalues of high-light (normal light) pixels, no further enhancement is required. Thus, an enhanced image feature vector with high light at each pixel is obtained. It should be noted here that in a specific embodiment of the present invention, for an RGB image, the denoised image feature vector has eigenvalue of three different channels at each pixel, and the eigenvalues in each channel need to be image-converted by the eigenvalues of the enhanced description matrix at the same pixel, so that the eigenvalues of the denoised image feature vector in three different channels at each pixel are enhanced.
[0083] S5. Perform denormalization on the eigenvalues of each pixel of the enhanced image feature vector to obtain the target enhanced image of the initial target image.
[0084] In the embodiment of the present invention, denormalization also needs to be performed on the enhanced image feature vector, that is, the eigenvalues of each pixel of the enhanced image feature vector are all multiplied by 255 to obtain a real image. It should be noted here that for an RGB image, the enhanced image feature vector has a feature matrix of three different channels at each pixel, and the eigenvalues in each channel need to perform denormalization.
[0085] While enhancing the low-light image, the present invention can also avoid re-enhancing the high-light image, effectively ensuring the picture quality. At the same time, the present invention also processes the noise in the image, further improving the image quality.
[0086] Figure 2 The network architecture diagram of a low-light digestive endoscope image enhancement model according to an embodiment of the present invention is shown, using the image enhancement network F M Obtain the enhanced description matrix of the target image, and use the image denoising network F D Denoise the target image and perform image conversion using the image conversion model.
[0087] Further, the image enhancement network in the embodiment of the present invention is also called F M network, and the result of the enhanced description matrix output by the F M network is represented as Φ L :
[0088] (1)
[0089] The enhanced description matrix can perform enhancements to different degrees according to different regions of the image. The value of the enhanced description matrix in the low-light region is a non-1 value, while the value of the enhanced description matrix in the high-light region is 1. Given that the size of an input image is 512*512, the size of the generated enhanced description matrix is also 512*512.
[0090] Specifically, the image enhancement network in the embodiments of the present invention is used to perform the following operations: performing hierarchical iterative feature extraction on the target image through a plurality of sequentially connected feature extraction modules to obtain a feature matrix of multiple channels representing luminance features; performing a convolution operation on the feature matrix of multiple channels to obtain a two-dimensional matrix of one channel, and using the obtained two-dimensional matrix as the enhanced description matrix.
[0091] Further, the feature extraction module in the embodiments of the present invention is used to perform the following operations: performing hierarchical iterative convolution feature extraction on the image features input to the feature extraction module through a plurality of feature extraction convolutional layers; during the convolution feature extraction process, adaptively adjusting the weights of the channel features through a preset channel attention module to increase the weights of the channels used to represent luminance; performing an element-wise addition operation on the image features obtained after performing the convolution feature extraction and the image features input to the feature extraction module to obtain the output features of the feature extraction module.
[0092] Specifically, referring to Figure 2 , in a specific embodiment of the present invention, the image enhancement network F M is composed of 3 feature extraction modules and 1 3*3 convolutional layer. The structures of the 3 feature extraction modules are exactly the same, including a plurality of convolutional layers with different numbers of channels and 1 ECA (Efficient Channel Attention Module) channel attention module. Each convolutional layer of the feature extraction module uses the ReLU activation function. The ECA channel attention module adaptively adjusts the weights of the channel features, enabling the network to better focus on important features and suppress unimportant features.
[0093] Specifically, in each feature extraction module, both the first convolutional layer and the second convolutional layer have 32 channels, that is, 32 3*3 convolutional kernels are used to extract features from the input features, and the thickness of each convolutional kernel is the same as the number of channels input to this convolutional kernel, and the ReLU activation function is used for non-linear transformation, and finally 32 different feature matrices are output; the output features of the second convolutional layer are input into the ECA (Efficient Channel Attention Module) channel attention module, and at this time the ECA channel attention module adaptively adjusts the weights of the channel features, increases the weights of the channels that can better reflect brightness, and reduces the weights of the channels that are insensitive to brightness; the third convolutional layer uses 4 1*1 convolutional kernels to extract features from the input features, and the thickness of each convolutional kernel is the same as the number of channels input to this convolutional kernel, forming 4 different feature matrices; the fourth convolutional layer uses 32 1*1 convolutional kernels to extract features from the input features, forming 32 different feature matrices; finally, the output features of the fourth convolutional kernel and the features input to the feature extraction module are subjected to an element-wise addition operation to obtain the output features of the feature extraction module. Among them, the residual network structure is adopted in the feature extraction module to prevent the gradient disappearance of the multi-layer network.
[0094] Furthermore, the image enhancement network F of the present invention M uses three feature extraction modules to perform hierarchical iterative feature extraction, and uses a 3*3 convolutional layer to reduce the output features of the third feature extraction module to 1 channel, so that the finally obtained feature matrix is an enhanced description matrix with the same size as the original image and 1 channel. For example, if the original image is 512*512*3, the output enhanced description matrix is 512*512*1.
[0095] Further, in step S3 of the embodiment of the present invention, a convolutional layer is used for the target image and the enhanced description matrix respectively to obtain feature vectors, the feature vectors are subjected to feature fusion to obtain a fused image feature vector, the fused image feature vector is subjected to feature extraction, and a denoising operation is performed on the extracted eigenvalues to obtain a denoised image feature vector, including: based on a preset image denoising network, the following operations are completed on the target image and the enhanced description matrix input to the image denoising network: performing a convolutional operation on the target image to obtain a first convolutional input image feature vector; performing a convolutional operation on the enhanced description matrix to obtain a second convolutional enhanced description feature vector; performing a splicing operation on the first convolutional input image feature vector and the second convolutional enhanced description feature vector to obtain a fused image feature vector of the target image and the enhanced description matrix; respectively performing a feature extraction operation and a soft threshold denoising operation on the fused image feature vector through at least one feature extraction module and at least one soft threshold denoising module to obtain an initial denoised image feature vector; performing a convolutional operation on the initial denoised image feature vector to obtain at least one feature matrix having the same number of channels and scale as the target image, and taking the obtained at least one feature matrix as the denoised image feature vector.
[0096] Specifically, referring to Figure 2 , in a specific embodiment of the present invention, the image denoising network F D consists of two 3×3 convolutional layers at the feature input end, and a feature splicing module connecting the two 3×3 convolutional layers performs feature fusion of the enhanced description matrix and the target image. Through sequentially connected feature extraction modules, soft threshold denoising modules, and feature extraction modules, feature extraction and denoising are performed on the fused image feature vector. Finally, a 3×3 convolutional layer is used to reduce the number of channels to the same as the number of channels of the target image for the denoised image feature vector. The structure of the feature extraction module of the image denoising network is the same as the structure of the feature extraction module in the image enhancement network.
[0097] Among them, the soft threshold denoising module is used to perform the following operations:
[0098] (2)
[0099] Among them, v is the eigenvalue input to the soft threshold denoising module, that is, the eigenvalue in each feature matrix input from the first feature extraction module of the image denoising network F D , δ is a preset soft threshold parameter. When v > 0, Sign(v) = 1, and when v < 0, Sign(v) = -1.
[0100] Among them, the preset soft threshold parameter δ can be set as needed. It can be known from the calculation model of the soft threshold denoising module that when the absolute value of the eigenvalue (which can also be called the signal value) in the feature matrix is greater than the soft threshold parameter, the eigenvalue is retained; if it is less than or equal to the soft threshold parameter, it proves that the eigenvalue is noise, and the eigenvalue is set to zero. Therefore, by adjusting the threshold size, the noise in the signal can be effectively suppressed while retaining the main features of the signal.
[0101] Furthermore, in step S4 of the embodiment of the present invention, the image conversion of the denoised image feature vector based on the enhanced description matrix includes: implementing the image conversion of the denoised image feature vector based on the enhanced description matrix by using a preset image conversion model, and the image conversion model is expressed as:
[0102] (3)
[0103] where p refers to the pixel point in the image, Φ L (p) refers to the eigenvalue of the enhanced description matrix at the pixel point p, and L D (p) refers to the eigenvalue of the denoised image feature vector at the pixel point p, and L E (p) refers to the eigenvalue of the calculated enhanced image feature vector at the pixel point p.
[0104] Specifically, at the same pixel point p, the eigenvalue of the denoised image feature vector at the pixel point p is subjected to a power exponent operation using the eigenvalue of the enhanced description matrix at the pixel point p to obtain the eigenvalue of the enhanced image feature vector at the pixel point p.
[0105] For example, in a low-brightness image area, the corresponding numerical range of the enhanced description matrix is greater than 0 and less than 1, then the image eigenvalue will increase. For example L D (p) the eigenvalue of a certain pixel point is 0.3, Φ L (p) and the corresponding value is 0.75, 0.3 0.75 ≈ 0.4053 ; in an over-bright image area, the corresponding value of the enhanced description matrix is greater than 1, then the image eigenvalue will decrease. For example L D (p) the value of a certain pixel is 0.9, Φ L (p) and the corresponding value is 1.25, 0.9 1.25 ≈ 0.8766; This value is not within the scope of the present invention. By default, all images are low-brightness images, and there is no over-bright situation. In the region of images with moderate brightness, the value corresponding to the enhancement description matrix is equal to 1, and then the image eigenvalue remains unchanged. For example L D (p) the eigenvalue of a certain pixel point of (p) is 0.5, Φ L (p) and the corresponding value is 1, 0.5 1 = 0.5。
[0106] It should be noted here that the number of channels of the enhanced image feature vector is the same as that of the denoised image feature vector. When the denoised image feature vector has three eigenvalues at pixel point p, the eigenvalues of the enhanced description matrix at pixel point p are respectively used for power exponent operations on them, and finally the three eigenvalues of the enhanced image feature vector at pixel point p are obtained.
[0107] In the embodiment of the present invention, before using the image enhancement network to perform image enhancement on the target image, it is necessary to first iteratively optimize and train the image enhancement network using a sample data set containing low-light images and high-light images, calculate the total loss after each iteration using the total loss function of the image enhancement network, and continuously adjust the model parameters in the image enhancement network according to the total loss after each iteration until the total loss of the image enhancement network converges, so that the image enhancement network has a good image enhancement effect.
[0108] Furthermore, before performing image enhancement on the target image in the embodiment of the present invention, the model also needs to be trained. The method further includes:
[0109] S01. Construct a training set, where the training set includes a plurality of low-light images under low-light conditions and a plurality of high-light images under well-illuminated conditions;
[0110] In the embodiment of the present invention, constructing the training set specifically means collecting two sets of digestive endoscope images with the same number. One set is images L containing images under low-light conditions and the other set is images H containing images under well-illuminated conditions. The contents of images L and images H can be different. It should be noted that the images L and H in the embodiment of the present invention have been normalized, so that each pixel point in the image is a normalized pixel value greater than 0 and less than 1.
[0111] Specifically, for digestive endoscope images, count the shooting parts of the digestive endoscope images. In order to ensure the balance of the number of the training set, try to keep the consistency of the number of images in different digestive tract parts, delete duplicate images or supplement images of parts with fewer numbers. After collecting the images, images L and H are placed in two folders respectively for distinction.
[0112] Fix the model parameters in the fixed-image denoising network, perform iterative model training on the image enhancement network based on the training set, and calculate the total image enhancement loss after each iterative training according to the total loss function of the image enhancement network. Adjust the model parameters in the image enhancement network according to the total image enhancement loss until the total image enhancement loss reaches the expected loss threshold, and then stop the iterative training;
[0113] Fix the model parameters in the image enhancement network, perform iterative model training on the image denoising network based on the training set, and calculate the total image denoising loss after each iterative training according to the total loss function of the image denoising network. Adjust the model parameters in the image denoising network according to the total image denoising loss until the total image denoising loss reaches the expected loss threshold, and then stop the iterative training.
[0114] The present invention separately trains the image enhancement network and the image denoising network. It does not require manually labeled data for training and uses the characteristics of the data itself as the supervision signal, enabling the model to self-learn and extract useful features.
[0115] Furthermore, the total loss function of the image enhancement network provided by the embodiment of the present invention includes a self-supervised loss function with a preset first weight coefficient, and the self-supervised loss function is expressed as:
[0116] (4)
[0117] where 0 < the first weight coefficient ≤ 1, Loss s is the self-supervised loss, L refers to the low-light image, F M refers to the image enhancement network, F M (L) is the result output after the low-light image is processed by the image enhancement network, ω is a custom parameter for self-supervised learning, 0 < ω < 1, Lω is the calculation result after optimizing the normalized pixel value of each pixel point in L using the parameter ω, and F M (Lω) is the result output after Lω is processed by the image enhancement network. The calculation formula of Lω is as follows:
[0118] (5)
[0119] where L(p) is the magnitude of the normalized pixel value of a single pixel point in the low-light image;
[0120] For example, when the parameter ω = 0.75 and the P value of a certain pixel point in the image L is 0.8, after calculation by this formula, the value of this pixel point becomes 0.8 0.75≈0.8458. Through the processing of parameter ω, the normalized pixel values of the image will increase, indicating that its brightness will be enhanced. This Lω is for pixel value amplification to enhance the overall brightness of the image.
[0121] Since the calculation of the self-supervised loss Loss(L, ω) is based on the image L and the parameter ω, rather than two pixel-by-pixel corresponding images, the training set can use unpaired low-light digestive endoscope images and high-light digestive endoscope images, rather than necessarily using pairs of digestive endoscope images with exactly the same content.
[0122] Furthermore, to ensure the consistency of the low-light image enhancement effect, a low-light image enhancement loss is designed. This loss ensures that for the result of the image L after being processed by the F M network, if it is processed by the F M network again, the output enhancement Map is a matrix with all elements being 1, so that the image will no longer be enhanced.
[0123] Based on the above design idea, the total loss function of the image enhancement network can also include a low-light image enhancement loss function with a preset second weight coefficient, and the low-light image enhancement loss function is expressed as:
[0124] (6)
[0125] where 0 < the second weight coefficient ≤ 1, Loss C is the low-light image enhancement loss, L M = F M (L), F M (L M ) is the result of the second use of the image enhancement network for processing the output result of the low-light image after being processed by the image enhancement network, and E is a matrix with the same scale as the enhancement description matrix and all elements being 1;
[0126] Furthermore, since the generated enhancement description matrix of F M (L M ) has all elements being 1, when using F M (L M ) for the low-light image L, any pixel value of the image remains unchanged. For example, if the normalized pixel value of the image is 0.6, after being processed by F M (L M ), 0.6 1 = 0.6. In this way, it is ensured that for the result of the image L after being processed by the F M network, if it is processed by the F M network again, its brightness remains unchanged, effectively avoiding the secondary enhancement of low-light digestive endoscope images.
[0127] Furthermore, to ensure the consistency of the high-brightness image enhancement effect, the present invention designs a high-brightness image enhancement loss. This loss ensures that for image H, when using F M the enhanced Map output by the network is a matrix with all elements being 1, such that the high-brightness image is not enhanced.
[0128] Based on the above design concept, the total loss function of the image enhancement network may further include a high-brightness image enhancement loss function with a preset third weight coefficient, and the high-brightness image enhancement loss function is expressed as:
[0129] (7)
[0130] where 0 < the third weight coefficient ≤ 1, Loss H is the high-brightness image enhancement loss, H is the high-brightness image, and F M (H) is the result output after the high-brightness image is processed by the image enhancement network.
[0131] Since all elements of F M (H) are 1, when at F M (H), any pixel value of the high-brightness image remains unchanged. For example, if the image pixel value is 0.9, after the enhancement process, 0.9 1 = 0.9. In this way, it is ensured that the result of image H after being processed by the F M network has its brightness unchanged, effectively avoiding the enhancement of high-brightness digestive endoscope images.
[0132] In a specific embodiment of the present invention, the total loss function of the image enhancement network is expressed as:
[0133] (8)
[0134] where Loss s refers to the self-supervised loss, Loss C refers to the low-brightness image enhancement loss, Loss H refers to the high-brightness image enhancement loss, α and β represent weight parameters and can be adjusted according to the effect. In a specific embodiment of the present invention, both α and β are 0.02;
[0135] The present invention uses the characteristics of the data itself as the supervision signal, enabling the model to self-learn and extract useful features. Therefore, it is not necessary to use exactly the same low-brightness images and high-brightness images for model training, simplifying the process of model training.
[0136] Further, before calculating the feature vectors by using the convolutional layer for the target image and the enhanced description matrix respectively, the method further includes: training the model of the image denoising network based on the total loss function of the preset image denoising network.
[0137] In the embodiment of the present invention, the total loss function of the image denoising network is obtained by weighted addition of the low-light image denoising loss and the high-light image denoising loss; the total loss function of the image denoising network is expressed as:
[0138] (9)
[0139] wherein, Loss D (L) refers to the low-light image denoising loss, Loss D (H) refers to the high-light image denoising loss, L refers to the low-light image, H refers to the high-light image, and k and z represent weight parameters.
[0140] Further, in order to ensure that the details of the original image content are not lost after denoising the low-light image, the low-light image denoising loss is designed. The low-light image denoising loss is expressed as:
[0141] (10)
[0142] wherein, Loss D (L) refers to the low-light image denoising loss, F D (L) is the result output after the low-light image is processed by the image denoising network.
[0143] Further, in order to ensure that the details of the original image content are not lost after denoising the high-light image, the high-light image denoising loss is designed. The high-light image denoising loss is expressed as:
[0144] (11)
[0145] wherein, Loss D (H) refers to the high-light image denoising loss, F D (H) is the result output after the high-light image is processed by the image denoising network.
[0146] Further, in order to verify the feasibility of this method, in a specific embodiment of the present invention, 4000 digestive endoscope images are used for experiments, wherein the training set, the validation set and the test set are divided according to the ratio of 70%, 20%, 10%. Among them, the training set adopts a group of 1400 low-light digestive endoscope images and a group of 1400 high-light digestive endoscope images, and the two groups of images are non-matching image pairs. The model training is divided into two stages. In the first stage, only F MThe network fixes the network weight parameters of this module after training is completed, and then trains F D the network. After the two-stage training of the model is completed, the entire network model is used for experiments. Some experimental results are as Figure 3 shown, where Figure 3 in (a) is a low-light endoscopic image of the stomach, Figure 3 in (b) is the enhanced endoscopic image of the stomach, Figure 3 in (c) is a low-light endoscopic image of the colon, Figure 3 in (d) is the enhanced endoscopic image of the colon. This method has obvious enhancement effect on the low-light images of digestive endoscopes, retains the original image details, and effectively improves the image brightness.
[0147] The present invention establishes a method for enhancing low-light images of digestive endoscopes based on self-supervised learning, and completes image brightness enhancement and denoising through the network model shown in Figure 2 . This model has the following advantages: 1) The training set uses unpaired low-light digestive endoscopic images and high-light digestive endoscopic images, which are very easy to collect; 2) It effectively avoids enhancing high-light digestive endoscopic images and secondarily enhancing low-light digestive endoscopic images; 3) It effectively processes the noise in the images and improves the image quality. The present invention helps to improve the quality of digestive endoscopic images, can effectively improve the effect of digestive endoscopic examinations, and has high clinical and scientific research application values.
[0148] For the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present invention are not limited by the described action sequence, because according to the embodiments of the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for the embodiments of the present invention.
[0149] Embodiment 2
[0150] Figure 3 Schematically shows a structural schematic diagram of a device for enhancing low-light digestive endoscopic images provided by an embodiment of the present invention. Referring to Figure 4 , a device for enhancing low-light digestive endoscopic images according to an embodiment of the present invention specifically includes an image processing module 401, an image enhancement module 402, an image denoising module 403, and an image conversion module 404, where:
[0151] The image processing module 401 is configured to perform normalization processing on the pixel values of each pixel point in the initial target image to be enhanced to obtain a target image;
[0152] An image enhancement module 402, configured to perform feature extraction on a target image based on a preset image enhancement network to identify the brightness features of each pixel point in the target image, and obtain an enhanced description matrix, where the enhanced description matrix is a two-dimensional matrix with the same scale as the target image;
[0153] An image denoising module 403, configured to respectively use a convolutional layer on the target image and the enhanced description matrix to obtain feature vectors, perform feature fusion on the feature vectors to obtain a fused image feature vector, perform feature extraction on the fused image feature vector, and perform a denoising operation on the extracted eigenvalues to obtain a denoised image feature vector;
[0154] An image conversion module 404, configured to perform image conversion on the denoised image feature vector based on the enhanced description matrix, so as to enhance the eigenvalues of the denoised image feature vector at each pixel point to different degrees according to the brightness features of the corresponding pixel points, and obtain an enhanced image feature vector;
[0155] The image processing module 401 is further configured to perform a denormalization operation on the eigenvalues of each pixel point of the enhanced image feature vector to obtain a target enhanced image of the initial target image.
[0156] Further, the image conversion module 404 is specifically configured to implement image conversion on the denoised image feature vector based on the enhanced description matrix by using a preset image conversion model, and the image conversion model is expressed as:
[0157]
[0158] Further, the image enhancement module 402 is specifically configured to perform corresponding image enhancement operations on the input of the image enhancement network based on the preset image enhancement network, where the operation process of the image enhancement network is the same as that in the method embodiment and will not be elaborated here.
[0159] Further, the embodiment of the present invention further includes a network training module, configured to perform model training on the image enhancement network based on the total loss function of the preset image enhancement network before the image enhancement module 402 performs feature extraction on the target image to identify the brightness features of each pixel point in the target image. The total loss function of the image enhancement network is the same as that in the method embodiment and will not be elaborated here.
[0160] Further, the image denoising module 403 is specifically configured to perform corresponding image denoising operations on the target image and the enhanced description matrix input to the image denoising network based on the preset image denoising network, where the operation process of the image denoising network is the same as that in the method embodiment and will not be elaborated here.
[0161] Further, the network training module of the embodiment of the present invention is further configured to train the model of the image denoising network based on the total loss function of the preset image denoising network before the image denoising module 403 performs feature fusion on the target image and the enhanced description matrix to obtain a fused image feature vector. The total loss function of the image denoising network is the same as that in the method embodiment and will not be elaborated here;
[0162] A low-light digestive endoscope image enhancement method and device provided by an embodiment of the present invention normalize the pixel values of each pixel point in an initial target image to be enhanced to obtain a target image, and perform feature extraction on the target image to identify the brightness features of each pixel point in the target image, obtaining an enhanced description matrix; at the same time, a convolutional layer is respectively used for the enhanced description matrix and the target image to obtain feature vectors, the feature vectors are subjected to image fusion to obtain a fused image feature vector, feature extraction is performed on the fused image feature vector, and a denoising operation is performed on the extracted eigenvalue to obtain a denoised image feature vector; finally, an image transformation is performed on the denoised image feature vector based on the enhanced description matrix to obtain an enhanced image feature vector. Since the enhanced description matrix can reflect the brightness features of each pixel point, when performing image transformation on the denoised image, the eigenvalue of the denoised image feature vector of the pixel point with a lower brightness can be enhanced to a higher degree, thereby achieving the purpose of improving the image quality and enhancing the low-light image at the same time.
[0163] Embodiment III
[0164] An embodiment of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps in the above-mentioned method embodiments for enhancing low-light digestive endoscope images are implemented, such as Figure 1 the steps S1 - S5 shown. Alternatively, when the processor executes the computer program, the functions of each module / unit in the above-mentioned embodiments of the low-light image enhancement device are implemented, such as Figure 4 the image processing module 401, the image enhancement module 402, the image denoising module 403, and the image transformation module 404 shown.
[0165] In addition, those skilled in the art can understand that although some embodiments herein include certain features included in other embodiments rather than other features, the combination of the features of different embodiments means that it is within the scope of the present invention and forms different embodiments. For example, any one of the claimed embodiments can be used in any combination.
[0166] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A low-light digestive endoscope image enhancement method, characterized in that: The method comprises: Normalizing the pixel value of each pixel in the initial target image to be enhanced to obtain the target image; Based on a preset image enhancement network, feature extraction is performed on the target image to identify the brightness features of each pixel in the target image, and an enhancement description matrix is obtained, wherein the enhancement description matrix is a two-dimensional matrix with the same scale as the target image; The image enhancement network is used to complete the following operations: Iteratively extracting features of the target image through a plurality of feature extraction modules connected in sequence, and obtaining a feature matrix of a plurality of channels for representing brightness features; Perform convolution operation on the feature matrices of multiple channels to obtain a two-dimensional matrix of one channel, and use the obtained two-dimensional matrix as an enhanced description matrix; Using convolutional layers to calculate feature vectors for the target image and the enhanced description matrix respectively, performing feature fusion on the feature vectors to obtain a fused image feature vector, performing feature extraction on the fused image feature vector, and performing a denoising operation on the extracted feature values to obtain a denoised image feature vector; Performing image transformation on the denoised image feature vector based on the enhanced description matrix, so as to enhance the feature value of the denoised image feature vector at each pixel point to different degrees according to the brightness characteristics of the corresponding pixel point, thereby obtaining an enhanced image feature vector; The image conversion of the denoised image feature vector based on the enhanced description matrix includes: The preset image conversion model is used to realize the image conversion of the denoised image feature vector based on the enhanced description matrix. The image conversion model is expressed as: Among them, p refers to the pixel point in the image, Φ L (p) refers to the eigenvalue of the enhanced description matrix at pixel p, L D (p) refers to the eigenvalue of the denoised image feature vector at pixel p, L E (p) refers to the calculated eigenvalue of the enhanced image feature vector at pixel p; A denormalization operation is performed on the eigenvalues of each pixel point of the enhanced image eigenvector to obtain a target enhanced image of the initial target image.
2. The method according to claim 1, characterized in that Before extracting features of the target image based on a preset image enhancement network to identify brightness features of each pixel in the target image, the method further includes: The image enhancement network is model trained based on a preset total loss function of the image enhancement network, wherein the total loss function of the image enhancement network includes a self-supervised loss function of a preset first weight coefficient, and the self-supervised loss function is expressed as: Among them, 0<first weight coefficient≤1, Loss s is the self-supervised loss, L refers to low-light images, and F M refers to the image enhancement network, F M (L) is the output result of the low-light image after being processed by the image enhancement network, ɷ is a custom parameter for self-supervised learning, 0<ɷ<1, Lɷ is the calculation result after optimizing the normalized pixel value of each pixel in L using the parameter ɷ, and F M (Lɷ) is the output result of Lɷ after being processed by the image enhancement network. The calculation formula of Lɷ is as follows: Among them, L(p) is the normalized pixel value of a single pixel in the low-light image.
3. The method according to claim 2, characterized in that The total loss function of the image enhancement network also includes a low-light image enhancement loss function of a preset second weight coefficient, and the low-light image enhancement loss function is expressed as: Among them, 0<second weight coefficient≤1, Loss C is the low-light image enhancement loss, L M = F M (L), F M (L M ) is the result of the second processing of the low-light image by the image enhancement network after the image enhancement network has processed the output result. E is a matrix with the same scale as the enhancement description matrix and all elements are 1.
4. The method according to claim 2 or 3, characterized in that: The total loss function of the image enhancement network also includes a high-light image enhancement loss function of a preset third weight coefficient, and the high-light image enhancement loss function is expressed as: Among them, 0<third weight coefficient≤1, Loss H is the highlight image enhancement loss, H is the highlight image, and F M (H) is the output result of the highlight image after being processed by the image enhancement network.
5. The method according to claim 1, characterized in that: The method of using convolution layers to obtain feature vectors for the target image and the enhanced description matrix respectively, performing feature fusion on the feature vectors to obtain a fused image feature vector, performing feature extraction on the fused image feature vector, and performing a denoising operation on the extracted feature values to obtain a denoised image feature vector comprises: Based on the preset image denoising network, the following operations are performed on the target image and the enhanced description matrix input into the image denoising network: Performing a convolution operation on the target image to obtain a first convolution input image feature vector; Performing a convolution operation on the enhanced description matrix to obtain a second convolution enhanced description feature vector; Performing a concatenation operation on the first convolution input image feature vector and the second convolution enhancement description feature vector to obtain a fused image feature vector of the target image and the enhancement description matrix; At least one feature extraction module and at least one soft threshold denoising module respectively perform feature extraction operations and soft threshold denoising operations on the fused image feature vector to obtain an initial denoised image feature vector; wherein the soft threshold denoising module is used to perform the following operations: Wherein, v is the characteristic value input to the soft threshold denoising module, δ is the preset soft threshold parameter, when v is greater than 0, Sign(v) = 1, when v is less than 0, Sign(v) = -1; A convolution operation is performed on the initial denoised image feature vector to obtain at least one feature matrix having the same number of channels and scale as the target image, and the obtained at least one feature matrix is used as the denoised image feature vector.
6. The method according to claim 5, characterized in that Before respectively using convolutional layers to calculate the target image and the enhanced description matrix to form feature vectors, the method further includes: Performing model training on the image denoising network based on a preset total loss function of the image denoising network; The total loss function of the image denoising network is obtained by weighted addition of the low-light image denoising loss and the high-light image denoising loss; wherein, The low-light image denoising loss function is expressed as: Among them, Loss D (L) refers to the low-light image denoising loss, F D (L) is the output result after the low-light image is processed by the image denoising network, L refers to the low-light image; The high light image denoising loss is expressed as: Among them, Loss D (H) refers to the denoising loss of high-light images, F D (H) is the output result after the highlight image is processed by the image denoising network, and H is the highlight image.
7. The method according to claim 1 or 5, characterized in that: The feature extraction module is used to complete the following operations: Performing iterative convolution feature extraction on the image features of the input feature extraction module through multiple feature extraction convolution layers; In the process of convolutional feature extraction, the weight of channel features is adaptively adjusted through the preset channel attention module to increase the weight of the channel used to characterize brightness; The image features obtained after performing convolution feature extraction and the image features input to the feature extraction module are element-wise added to obtain the output features of the feature extraction module.
8. A low-light digestive endoscope image enhancement device, characterized in that: The device comprises: An image processing module is used to normalize the pixel value of each pixel in the initial target image to be enhanced to obtain a target image; An image enhancement module is used to extract features of a target image based on a preset image enhancement network to identify brightness features of each pixel in the target image, and obtain an enhancement description matrix, wherein the enhancement description matrix is a two-dimensional matrix with the same scale as the target image; The image enhancement module is specifically used to perform iterative feature extraction on the target image through multiple feature extraction modules connected in sequence to obtain feature matrices of multiple channels for representing brightness features; perform convolution operations on the feature matrices of multiple channels to obtain a two-dimensional matrix of one channel, and use the obtained two-dimensional matrix as an enhancement description matrix; An image denoising module, used to obtain feature vectors using convolution layers for the target image and the enhanced description matrix respectively, perform feature fusion on the feature vectors to obtain a fused image feature vector, perform feature extraction on the fused image feature vector, and perform denoising operation on the extracted feature values to obtain a denoised image feature vector; An image conversion module is used to perform image conversion on the denoised image feature vector based on the enhancement description matrix, so as to enhance the feature value of the denoised image feature vector at each pixel point to different degrees according to the brightness characteristics of the corresponding pixel point, thereby obtaining an enhanced image feature vector; The image conversion module is specifically used to implement image conversion of denoised image feature vectors based on the enhanced description matrix using a preset image conversion model. The image conversion model is expressed as: Among them, p refers to the pixel point in the image, Φ L (p) refers to the eigenvalue of the enhanced description matrix at pixel p, L D (p) refers to the eigenvalue of the denoised image feature vector at pixel p, L E (p) refers to the calculated eigenvalue of the enhanced image feature vector at pixel p; The image processing module is further used to perform a denormalization operation on the eigenvalues of each pixel point of the enhanced image eigenvector to obtain a target enhanced image of the initial target image.
9. A computer device, characterized in that: It includes a memory, a processor, and a computer program stored in the memory and executable on the processor; when the computer program is executed by the processor, the low-light digestive endoscopy image enhancement method as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Image target detection method and system based on two-way deep residual network
CN114241340A
Weak light image enhancement method and system based on illumination decomposition
CN119205541A