An endoscopic image enhancement method and system

By using a lightweight end-to-end network for local perception and global color mapping, the problem of uneven illumination and color distortion in endoscopic images under complex lighting conditions is solved, improving image quality and diagnostic accuracy, and making it suitable for portable endoscopic terminals.

CN121213368BActive Publication Date: 2026-07-24WUHAN INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
WUHAN INST OF TECH
Filing Date
2025-09-29
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing endoscopic image enhancement algorithms struggle to achieve a balance in complex lighting scenarios where underexposure and overexposure coexist, and they are computationally intensive, leading to uneven image illumination, color distortion, and loss of detail, which affects diagnostic accuracy.

Method used

We employ an end-to-end lightweight network, combining a local perception network, a global color mapping network, and a fusion network. By optimizing the model through the total loss function, we achieve local detail enhancement and global color correction while reducing computational complexity.

Benefits of technology

It improves the visual consistency and diagnostic reliability of gastrointestinal endoscopy images, reduces hardware costs, is suitable for portable endoscopy terminals, and provides high-quality image data to support AI-assisted diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121213368B_ABST
    Figure CN121213368B_ABST
Patent Text Reader

Abstract

The application provides an endoscope image enhancement method and system, and relates to the technical field of medical image processing; the method comprises the following steps: constructing an endoscope image enhancement model composed of a local perception network, a global color mapping network and a fusion network; extracting details of a low-quality image through the local perception network; learning colors of a high-quality image through the global color mapping network; integrating double-path features through the fusion network and outputting an endoscope enhanced image; generating a total loss through a total loss function to optimize the endoscope image enhancement model, so as to output an endoscope image with high definition and real colors. An end-to-end lightweight network is adopted, and unified enhancement of an extreme light coexistence scene is completed at one time, so that the visual consistency and diagnostic reliability of a gastrointestinal endoscope image are significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical image processing technology, specifically to an endoscopic image enhancement method and system. Background Technology

[0002] Electronic endoscopy has become an indispensable tool in modern minimally invasive diagnosis and treatment. However, due to the narrowness of human cavities, the refraction or scattering of intracavitary media (such as blood, mucus, and exudate), and the fact that fiber optic light sources can only be introduced at a fixed angle, the imaging environment is characterized by suboptimal lighting posture, high surface humidity, and local occlusion. The direct consequence is severe uneven illumination in single-frame images, with some areas overexposed due to specular reflection or close-range illumination, resulting in loss of texture details; and other areas underexposed due to occlusion or tissue curvature, leading to a low signal-to-noise ratio. The compressed dynamic range makes it difficult to identify lesion boundaries and microvascular morphology, significantly affecting the accuracy of clinical diagnosis.

[0003] Existing enhancement algorithms are mostly designed for low-light scenes, improving visibility in dark areas by increasing overall brightness, but they neglect the special case where endoscopic images also contain overexposed areas. If such methods are applied directly, overexposed pixels will further exceed the sensor's effective dynamic range, causing originally discernible textures to be excessively brightened, resulting in further information loss. Therefore, there is an urgent need for an endoscopic image enhancement technology to meet the pressing demand for high-quality images in precision medicine. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide an endoscopic image enhancement method and system that addresses the shortcomings of the prior art.

[0005] The technical solution of the present invention to solve the above-mentioned technical problems is as follows: An endoscopic image enhancement method, comprising the following steps: An endoscopic image enhancement model is constructed, which includes a local perception network, a global color mapping network, and a fusion network. Import endoscopic image data, which includes initial low-quality images and initial high-quality images. Preprocess the endoscopic image data to obtain low-quality images and high-quality images. The low-quality image is enhanced by the local perceptual network to obtain a low-resolution feature map. The high-quality image is then color-mapped by the global color mapping network to obtain a high-resolution feature map. Finally, the low-resolution feature map and the high-resolution feature map are enhanced by the fusion network to obtain an enhanced endoscopic image. A total loss function is constructed, and the total loss value is calculated on the imported target enhanced image and the endoscope enhanced image using the total loss function. The endoscope image enhancement model is then optimized based on the total loss value to obtain the optimal endoscope image enhancement model. The endoscopic image enhancement model is used to predict the endoscopic image dataset to obtain the endoscopic enhanced image dataset.

[0006] Another technical solution of the present invention to solve the above-mentioned technical problems is as follows: An endoscopic image enhancement system, comprising: A construction unit is used to construct an endoscope image enhancement model, which includes a local perception network, a global color mapping network, and a fusion network. A preprocessing unit is used to import endoscopic image data, which includes initial low-quality images and initial high-quality images, and to preprocess the endoscopic image data to obtain low-quality images and high-quality images. The training unit is used to perform feature enhancement processing on the low-quality image through the local perception network to obtain a low-resolution feature map, perform color mapping processing on the high-quality image through the global color mapping network to obtain a high-resolution feature map, and perform enhancement processing on the low-resolution feature map and the high-resolution feature map through the fusion network to obtain an endoscopy enhanced image. An optimization unit is used to construct a total loss function, calculate the total loss value by applying the total loss function to the imported target enhanced image and the endoscope enhanced image, and optimize the endoscope image enhancement model based on the total loss value to obtain the optimal endoscope image enhancement model. The application unit is used to predict the predicted endoscope image dataset using the endoscope image enhancement model to obtain the endoscope enhanced image dataset.

[0007] The beneficial effects of this invention are as follows: A collaborative transformation framework integrates local and global pixel transformations, employing a lightweight local perception network and an improved attention mechanism to enhance feature extraction capabilities, resulting in a low-resolution image with enhanced local details. Simultaneously, an adaptive global color mapping network is used to achieve effective global contrast, adhering to the lightweight principle. The image is then scaled to a high-resolution space after operation at low resolution, yielding a high-resolution output image with global color details. In the fusion network, deformable convolutions are used to extract feature maps from two images, and a multi-head attention mechanism is employed to fuse the images. This allows for dynamic channel calibration, improving the model's sensitivity to key features. Learnable offsets dynamically adjust the convolution kernel sampling position, accurately adapting to complex structural deformations in endoscopic images. This improves model performance while reducing computational cost, solving problems such as uneven illumination, color distortion, and loss of detail in gastrointestinal endoscopy images. Attached Figure Description

[0008] Figure 1 A flowchart of an endoscopic image enhancement method provided in an embodiment of the present invention; Figure 2 This is a structural diagram of a local sensing network provided in an embodiment of the present invention; Figure 3 This is a structural diagram of the fusion network provided in an embodiment of the present invention; Figure 4 This is a structural diagram of the deformable convolution module provided in an embodiment of the present invention; Figure 5 This is a block diagram of an endoscopic image enhancement system provided in an embodiment of the present invention. Detailed Implementation

[0009] The principles and features of the present invention are described below with reference to the accompanying drawings. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.

[0010] In existing research, an algorithm for electronic endoscope images can address the high noise problem in low-light areas, solving the problems of uneven illumination, blurred details, and noise in low-light areas in current electronic endoscope images.

[0011] An enhancement algorithm based on Retinex theory (i.e., the characteristics of the human visual system) can perform illumination correction and detail stretching separately to improve contrast and solve the problem of overly dark endoscopic images. An image enhancement method based on LSMPEC (Laplacian Self-Attention Multi-Pyramid Enhanced Contrast) uses a pyramid structure to extract and fuse features from images of different resolutions while preserving global color and local details. However, this multi-scale pyramid method has high computational complexity, and the model sometimes produces incorrect tones. A unified illumination correction method (ENdoUIC) specifically designed for capsule endoscopy images can navigate different regions in the parameter space, allowing for customized adjustments to address the different challenges posed by images with excessively strong or weak illumination; however, it lacks a certain degree of adaptability.

[0012] In summary, existing technologies still face two major bottlenecks: first, their processing capabilities are limited, making it difficult to achieve a balance in complex lighting scenarios where underexposure and overexposure coexist; second, the models require a large amount of computation and consume a lot of computing resources.

[0013] like Figure 1 As shown in the figure, an endoscopic image enhancement method provided by an embodiment of the present invention includes the following steps: An endoscopic image enhancement model is constructed, which includes a local perception network, a global color mapping network, and a fusion network. Import endoscopic image data, which includes initial low-quality images and initial high-quality images. Preprocess the endoscopic image data to obtain low-quality images and high-quality images. The low-quality image is enhanced by the local perceptual network to obtain a low-resolution feature map. The high-quality image is then color-mapped by the global color mapping network to obtain a high-resolution feature map. Finally, the low-resolution feature map and the high-resolution feature map are enhanced by the fusion network to obtain an enhanced endoscopic image. A total loss function is constructed, and the total loss value is calculated on the imported target enhanced image and the endoscope enhanced image using the total loss function. The endoscope image enhancement model is then optimized based on the total loss value to obtain the optimal endoscope image enhancement model. The endoscopic image enhancement model is used to predict the endoscopic image dataset to obtain the endoscopic enhanced image dataset.

[0014] It should be understood that the endoscopic image data can be an endoscopic image dataset, which includes multiple sets of endoscopic images, each set of endoscopic images being a one-to-one correspondence of an initial low-quality image and an initial high-quality image.

[0015] In this embodiment, an end-to-end lightweight network is employed to uniformly enhance images in extreme lighting scenarios (where local overexposure and underexposure coexist), avoiding color banding and detail loss caused by traditional step-by-step processing. This significantly improves the visual consistency and diagnostic reliability of gastrointestinal endoscopic images. Through dynamic illumination compensation (i.e., a local perception network) and a color mapping mechanism (i.e., a global color mapping network), overall brightness, contrast, and color cast are simultaneously corrected, making key structures such as mucosal microvessels and lesion boundaries clearer and colors more realistic, providing high-quality data for subsequent AI-assisted diagnosis. Efficient feature extraction and reconstruction are achieved within a single-scale feature space, eliminating redundant computations from traditional multi-scale pyramids. The overall parameter count is less than or equal to 1M, and the inference latency is less than or equal to 5ms, allowing direct deployment on portable endoscopic terminals and reducing hardware costs.

[0016] Preferably, the preprocessing of the endoscopic image data to obtain low-quality and high-quality images includes: The endoscopic image data is subjected to bilinear interpolation to obtain interpolated endoscopic image data. Data augmentation is then performed on the interpolated endoscopic image data to obtain low-quality and high-quality images.

[0017] Specifically, bilinear interpolation is performed on the initial low-quality image and the initial high-quality image that correspond one-to-one, reducing or enlarging them to a uniform size. Data augmentation is then performed on the interpolated initial low-quality image and the initial high-quality image respectively. The images are then randomly flipped and rotated (including 90-degree, 180-degree, 270-degree rotation or transpose) according to random parameters to obtain the low-quality image and the high-quality image.

[0018] In this embodiment, each group of images in the endoscopic image dataset is first reduced or enlarged to the same size to maintain the smoothness of the image content and avoid artificial artifacts such as jagged edges. Random data augmentation is then applied to images with corresponding relationships in the dataset to ensure that the spatial transformations of each pair of images are completely consistent, without destroying the correspondence between images and improving the model's generalization ability.

[0019] Preferably, such as Figure 2 As shown, the step of performing feature enhancement processing on the low-quality image through the local perceptual network to obtain a low-resolution feature map includes: The local sensing network includes an illumination compensation module, an autoencoder and decoder module, and a residual self-attention module; The low-quality image is adjusted in brightness using the illumination compensation module to obtain a low-quality compensated image. The low-quality compensated image is encoded by the autoencoder and decoder module to obtain a downsampled compensated image, and the downsampled compensated image is decoded to obtain an initial low-resolution image. The initial low-resolution image is weighted by the residual self-attention module to obtain a low-resolution feature map.

[0020] Before the illumination compensation module processing step, the method further includes: performing a convolution operation on the low-quality image through a convolutional layer to increase the number of image channels from 3 to 8.

[0021] Specifically, the encoder includes two encoding layers, each consisting of a 1×1 convolutional layer, a depthwise separable convolutional layer, and a 1×1 convolutional layer. The number of image channels output by the first encoding layer is increased to 16, and the number of image channels output by the second encoding layer is increased to 32. The decoder includes two decoding layers, each consisting of a transposed convolutional layer, a 1×1 convolutional layer, a depthwise separable convolutional layer, and a 1×1 convolutional layer, progressively restoring spatial resolution and the number of channels. A Leaky ReLU activation function with a negative slope is used to improve gradient flow, accelerate training, and enhance the model's training performance.

[0022] In this embodiment, pixel-by-pixel conversion is employed to adjust pixels according to the local environment and optimize local details. A lighting compensation module normalizes and compensates for lighting conditions, reducing feature differences between different exposures. A simple autoencoder and decoder structure accelerates training and improves model performance. A residual self-attention algorithm further enhances the model's performance and adaptive feature selection capabilities.

[0023] Preferably, the step of adjusting the brightness of the low-quality image through the illumination compensation module to obtain a low-quality compensated image includes: The illumination compensation module includes an exposure normalization layer, a compensation layer, and an attention network. The attention network includes a pooling layer, a fully connected layer, and an activation function layer. The exposure features of the low-quality image are normalized using the exposure normalization layer to obtain a low-quality exposure image. The instance normalization operation is as follows: , in, Indicates a low-quality exposed image, IN ( ) represents instance normalization, and F represents a low-quality image. This represents the average value calculated for each channel and sample across the spatial dimension. This represents the standard deviation of each channel relative to the sample in the spatial dimension, where γ and β represent the learning parameters; The unexposed features of the low-quality image are integrated using the compensation layer to obtain an exposure feature image and a compensation feature image. The expression for the integration calculation is as follows: , , in, Represents the exposure feature image, Indicates a low-quality exposed image. Indicates the exposure correlation coefficient. Represents the compensated feature image. This represents the non-exposure features of a low-quality image. Represents the compensation correlation coefficient, ( () indicates element-wise multiplication; it should be understood that low-quality exposed images are normalized features, while unexposed features of low-quality images are unnormalized features of low-quality images.

[0024] The exposure feature image and the compensation feature image are pooled by a first pooling layer, the pooled exposure feature image and the compensation feature image are connected by a first fully connected layer, and the connected exposure feature image and the compensation feature image are classified by a first activation function layer to obtain a low-quality compensation image.

[0025] Specifically, illumination normalization and compensation are achieved through an illumination compensation module (ENC). This module comprises two parts: an exposure normalization part and a compensation part. The exposure normalization part maps various exposure features to an exposure-invariant feature space, using instance normalization to coarsely align the features. For the input low-quality image F, instance normalization is performed using the following formula: .

[0026] The compensation component is used to compute the normalized features in the spatial dimension. The correlation coefficient A between the unnormalized feature F and the unnormalized feature F The characteristic map of spatial interaction is obtained. and .

[0027] An attention weight graph is generated using an attention network, which includes one pooling layer and two fully connected layers. The attention weights are adaptively adjusted. and The features are jointly weighted, and the final output is a low-quality compensated image. The attention network processing expression is: , Where Sigmoid represents the activation function, FC represents the fully connected layer, and GAP represents average pooling.

[0028] In this embodiment, since the input image contains various levels of illumination, in order to reduce the differences between image groups, achieve unified color correction and brightness adjustment, reduce feature differences between different exposures, integrate the unnormalized part to compensate for the information removed after normalization, and integrate the exposure feature image and the compensation feature image based on attention weights to reduce the differences in illumination between images.

[0029] Preferably, the step of weighting the initial low-resolution image through the residual self-attention module to obtain a low-resolution feature map includes: The residual self-attention module includes a pooling layer, multiple fully connected layers, and multiple normalization layers; The initial low-resolution image is linearly transformed using a second fully connected layer to obtain a value vector. A second pooling layer then performs spatial downsampling on the initial low-resolution image. The downsampled initial low-resolution image is then concatenated using a second fully connected layer to obtain attention weights. These attention weights are then normalized using a first normalization layer and classified using a second normalization layer to obtain normalized attention weights. Regularized attention weights are then obtained. The value vector and the regularized attention weights are weighted and summed to obtain a weighted value vector. A third fully connected layer performs a linear transformation on the weighted value vector to obtain a low-resolution feature map, which is then regularized to obtain a regularized low-resolution feature map. Finally, the initial low-resolution image and the regularized low-resolution feature map are residually concatenated and summed to obtain the final low-resolution feature map.

[0030] Specifically, the input x (i.e., the initial low-resolution image) is saved for residual connections in the final processing. A linear transformation is performed on the input x to obtain a value vector, including: projecting the input x into a value tensor v through a fully connected layer and performing local window expansion. The expansion operation (i.e., sliding window extraction of local blocks) is used to expand the value tensor v into local windows; specifically, a window size k is set, and the pixels within each window are flattened into a vector, resulting in L k×k local windows. Average pooling is used for spatial downsampling of the input x, and an attention matrix is ​​obtained through a fully connected layer to generate attention weights. Attention relationships between all locations within its local window are predicted for each location. The attention weights are normalized by scaling the attention matrix and applying a softmax function along the last dimension to obtain normalized attention weights, followed by dropout regularization. The regularized attention weights are then weighted and summed with the value vector, and the weighted aggregation result (i.e., the expanded local window) is folded back into a feature map through a folding operation. Finally, the folded features are projected through a fully connected layer (i.e., the channel information is mixed), and dropout regularization is performed. The projection result is added to the residual connection to obtain the module output (i.e., the low-resolution feature map). ).

[0031] In this embodiment, the residual structure is used as a conservative adjustment strategy, which can limit the deviation between the model output and the input, avoid oversensitivity to outliers, and thus adapt to the complex and varied lighting conditions in real-world scenarios, so that the model can remain robust to unidentified lighting distributions.

[0032] Preferably, the step of performing color mapping processing on the high-quality image through the global color mapping network to obtain a high-resolution feature map includes: The global color mapping network includes an adaptive downsampling module and a color enhancement module; The high-quality image is sampled using the adaptive downsampling module to obtain a resampled image. The color enhancement module performs global color enhancement processing on the resampled image to obtain a high-resolution feature map.

[0033] In this embodiment, adaptive sampling preserves more color information during downsampling, improving data quality. Color mapping is then used to correct the global color of the adaptively downsampled image, resulting in a corrected high-resolution feature map. The output image is regenerated in high-resolution space using a high-quality input image and sampling grid, thus reducing computational complexity while achieving a clear, high-resolution global enhancement result.

[0034] Preferably, the step of sampling the high-quality image through the adaptive downsampling module to obtain a resampled image includes: The adaptive downsampling module includes a lightweight backbone network, an interval prediction layer, a normalization layer, and a sampling construction layer; The high-quality image is downsampled using the lightweight backbone network to obtain a multi-dimensional global feature vector. The interval prediction layer then predicts the interval of the global feature vector to obtain multiple sampling intervals. A third normalization layer normalizes these multiple sampling intervals to obtain multiple coordinate points. The sampling construction layer generates a sampling grid based on these coordinate points, and bilinear interpolation is performed on the high-quality image using the sampling grid to obtain a resampled image.

[0035] Specifically, the input image (i.e., the high-quality image) is first used as the input to the lightweight backbone network (LightBackbone), and after downsampling, a 512-dimensional feature vector (i.e., the global feature vector) is obtained as the output. The interval prediction layer predicts the interval distribution of the non-uniform sampling grid on the output feature vector. The generated interval is normalized by softmax through the normalization layer, and then the cumulative sum is calculated to obtain the normalized coordinate points. Finally, the sampling grid is generated based on these coordinate points, and the high-quality image is resampled using bilinear interpolation.

[0036] In this embodiment, adaptive sampling is used to retain more color information during downsampling, thereby improving data quality.

[0037] Preferably, the step of performing global color enhancement processing on the resampled image through the color enhancement module to obtain a high-resolution feature map includes: The color enhancement module includes a lookup table construction layer, a convolutional neural network, a linear layer, a weighted fusion layer, and an enhancement processing layer; Multiple three-dimensional lookup tables are constructed through the lookup table construction layer. Features are extracted from the multiple three-dimensional lookup tables by the convolutional neural network to obtain multiple resampled features. The multiple resampled features are linearly mapped by the linear layer to obtain multiple prediction weights. The multiple prediction weights are weighted with the multiple three-dimensional lookup tables by the weighted fusion layer to obtain the optimal three-dimensional lookup table. The enhancement processing layer performs global mapping of the colors of the resampled image according to the optimal three-dimensional lookup table to obtain a high-resolution feature map.

[0038] Specifically, establishing a 3D lookup table (3D LUT) involves: first, creating three basic lookup table LUTs; initializing the first 3D LUT with an identity transformation (i.e., input equals output); and initializing the remaining 3D LUTs with zero matrices (i.e., vectors to be learned); extracting feature vectors from each 3D lookup table using a convolutional neural network (CNN); transforming the feature vectors of multiple lookup tables into prediction weights using a linear layer; and finally, achieving weighted fusion through matrix multiplication to obtain a weighted 3D LUT; mapping the weighted 3D LUT (i.e., the optimal 3D lookup table) onto a resampled image to obtain a high-resolution global color enhancement image (i.e., a high-resolution feature map); the calculation expression for the weighted 3D LUT is as follows: , in, This represents the optimal three-dimensional lookup table. This represents the i-th set of prediction weights learned by the CNN. This represents the i-th 3D LUT. Represents n sets of 3D LUTs, ( () indicates element-wise multiplication.

[0039] To be understood, the basic principle of 3D LUT (i.e., three-dimensional lookup table) is as follows: transform the three channels (R, G, B) in RGB space respectively to create a 17×17×17 cubic grid. Each grid point stores an RGB output value. For the RGB value of the input pixel, find the 8 nearest vertices (cube cells) in the three-dimensional grid and use trilinear interpolation to calculate the final output color.

[0040] In this embodiment, an adaptive multi-core 3D LUT transformation framework is used to overcome the shortcomings of traditional pixel-by-pixel methods in terms of global color consistency. 3D LUTs can flexibly express nonlinear mappings, transforming the color, brightness, and contrast of an image. Multiple sets of 3D LUTs are used to handle different lighting conditions, and an adaptive method is employed to fuse these multiple sets of 3D LUTs, thus solving the problem of large differences in image brightness and contrast.

[0041] Preferably, such as Figure 3 As shown, the step of enhancing the low-resolution feature map and the high-resolution feature map through the fusion network to obtain the enhanced endoscopic image includes: The fusion network includes activation function layers, multi-head attention mechanism modules, gated feedforward networks, convolutional layers, multiple deformable convolutional modules, multiple normalization layers, and multiple residual blocks; The low-resolution feature map is extracted by the first deformable convolution module to obtain low-resolution weight features, and the high-resolution feature map is extracted by the second deformable convolution module to obtain high-resolution weight features. The low-resolution weight features are nonlinearly transformed by a second activation function layer, and the high-resolution weight features are nonlinearly transformed by a third activation function layer. The low-resolution weight features after nonlinear transformation are normalized by a fourth normalization layer to obtain low-resolution intermediate features, which are then input into the multi-head attention mechanism module as query vectors. The high-resolution weight features after nonlinear transformation are normalized by a fifth normalization layer to obtain high-resolution intermediate features, which are then input into the multi-head attention mechanism module as key vectors and value vectors. The dimensions of the key vector and the query vector are unified. The key vector, value vector, and query vector are fused through the multi-head attention mechanism module to obtain an initial fused feature. The low-resolution feature map is then added element-wise to the initial fused feature through the first residual block to obtain the fused feature. The number of channels of the fused feature is expanded by the gated feedforward network, and the fused feature is optimized and adjusted based on the multiple channel numbers to obtain a multi-dimensional fused feature; The multidimensional fusion features are compressed by the first convolutional layer to obtain key features. The key features are then added element-wise to the fusion features by the second residual block to obtain the endoscopic enhanced image.

[0042] Specifically, a deformable convolutional module dynamically extracts weighted feature maps from two input images (i.e., a low-resolution feature map and a high-resolution feature map). The GELU activation function is applied to balance the aggressive activation of ReLU with the smoothness of the Sigmoid activation function. Then, LayerNorm with a bias layer is used for layer normalization to stabilize the training process.

[0043] Two feature maps are fused using the Transformer's multi-head cross-attention mechanism. The two normalized features are input into the multi-head cross-attention module, where the high-resolution intermediate features... Low-resolution intermediate features serve as key vectors (K) and value vectors (V). The key vector is used as the query vector (Q), and the key vector is adjusted to the size of the query vector through adaptive pooling to achieve spatial alignment; the input feature maps are fused through a multi-head attention mechanism. and Then, residual connections are performed to obtain fused features. .right The feature map is processed using a gated feedforward (FFN) network, including: First, 1×1 convolutions are used to expand the number of channels, activating more possibilities for feature combinations; second, the gating mechanism of the feedforward network, GELU activation, controls the information flow, suppresses noisy channels, and enhances the lesion signal. Then, feature compression is performed, including: first, 1×1 convolutions are used to restore the feature map to its original dimensions to obtain key features. Finally, residual connections are performed to output the enhanced endoscopic image (represented as...). ), that is, to obtain the fused output.

[0044] In this embodiment, residual connections are used to preserve the original information and avoid the loss of important features during the fusion process; a feedforward network (FFN) is used to enhance the model's expressive power and compensate for local details that the attention mechanism may ignore, so as to obtain better feature representation.

[0045] Preferably, such as Figure 4 As shown, the step of extracting features from the low-resolution feature map using the first deformable convolution module to obtain low-resolution weighted features includes: The deformable convolutional module includes a sampling layer, a pooling layer, an activation function layer, a residual block, and multiple convolutional layers; The low-resolution feature map is convolved by a second convolutional layer to obtain an offset. The position of each sampling point in the low-resolution feature map is calculated based on the offset to obtain multiple sampling points. The sampling layer selects sample groups for each sampling point according to the sampling rules to obtain multiple sample groups, and calculates the bilinear weights corresponding to the multiple sample groups. Based on each bilinear weight, the corresponding four-corner feature values ​​are extracted from the low-resolution feature map. The multiple four-corner feature values ​​are weighted and summed to obtain the final sampling value. The low-resolution feature map is sampled according to the final sampling value to obtain deformed convolution features. The deformed convolutional features are transformed and compressed by the third pooling layer to obtain compressed convolutional features. The compressed convolutional features are then convolved by the third convolutional layer to obtain local cross-channel features. The local cross-channel features are weighted by the fourth activation function layer to obtain weighted features. The deformed convolutional features and the weighted features are then multiplied element-wise by the third residual block to obtain low-resolution weighted features.

[0046] Specifically, the low-resolution feature map is first convolved using a 3×3 ordinary convolution to generate the offset. The position (p) of each sampling point is calculated based on the offset, including the initial position. Add the relative positions inside the convolution kernel plus offset Then, bilinear sampling is performed, which involves calculating the four nearest integer coordinates for each sampling point (top left, top right, bottom left, and bottom right) and calculating bilinear weights. Based on the integer coordinates of the sampling points, the four corner feature values ​​are extracted from the input feature map, and a weighted sum is obtained to obtain the final sampled value. Finally, the shape is reconstructed by unfolding the N feature points at each position according to the arrangement of the convolution kernels. Channel-level weights are adjusted on the convolution output through a channel attention component. Global average pooling (GAP) is used to transform the feature map size to fuse contextual information. Then, the feature map is compressed, and an adaptive convolution kernel size is calculated. The compressed feature map is then subjected to one-dimensional convolution to obtain local cross-channel information. A sigmoid activation function is then used to remap the weights between 0 and 1. Finally, the channel weights are multiplied channel-by-channel with the original input to achieve attention enhancement. For points... The output formula for this convolution is as follows: , Where n represents the n sampling points, This represents the convolution weight coefficients at the corresponding point positions. This represents the fixed offset of the convolution kernel. This represents the offset used in dynamic learning. This indicates that the pixel value is calculated using bilinear interpolation to obtain the floating-point coordinates. () indicates element-wise multiplication.

[0047] It should be understood that the steps for extracting features from the high-resolution feature map using the second deformable convolution module to obtain high-resolution weighted features are the same as this step, and will not be repeated here.

[0048] In this embodiment, deformable convolution generates weights to enhance the contribution of important locations (such as edges and high gradient regions). Channel attention is added to further weight the channel dimensions, highlighting the feature channels related to brightness. For low-light areas, deformable convolution more accurately enhances the details of dark areas and avoids smoothing noise.

[0049] Preferably, a total loss function is constructed, and the expression for the total loss function is: , in, This represents the total loss, based on experience. Set to 0.5. Set it to 0.1.

[0050] A total loss function is constructed, and the total loss value is obtained by calculating the total loss value using the imported target enhanced image and the endoscopic enhanced image through the total loss function, including: Initialize the training model instance by creating an exponential moving average (EMA) model to optimize and update parameters during the training of the endoscopic image enhancement model, setting the decay coefficient to 0.999 to smooth parameter updates. Initialize the loss function using mean absolute error loss to maintain basic consistency between the output and the high-resolution target image. Provides pixel-level guidance, directly calculating the pixel differences between the generated and target images, which helps maintain the overall brightness and color consistency of the generated image with the target image, and reduces the mean absolute error loss. The expression is: , in, Indicates the number of pixels. This indicates an enhanced endoscopic image. This represents the enhanced image of the target.

[0051] Using mean absolute error loss The base loss ensures minimal pixel differences between the generated and target images. The model not only needs to generate pixel-perfect images but also structurally identical to the target image. To make the images appear more visually appealing to the human eye, SSIM loss is used to measure image similarity in terms of brightness, contrast, and structure, emphasizing local structural similarity to optimize the sharpness and detail of the generated images. Structural similarity loss is also employed. The expression is: , in, Indicates batch size. The closer to 0, the smaller the loss, meaning the closer the predicted endoscopic enhanced image is to the target enhanced image.

[0052] In addition, the model involves multi-stage image processing and high-resolution image generation, incorporating perceptual loss to optimize feature space similarity and improve the overall perceptual quality of the generated images. (Perceptual Loss) The expression is: , in, This represents the enhanced endoscopic image of layer l. This represents the target enhancement image of layer l.

[0053] Perceived loss Optimizing the model by comparing the similarity between the generated image (i.e., the endoscopic-enhanced image) and the target image (i.e., the target-enhanced image) in a high-level feature space, rather than focusing solely on pixel-level differences, more closely resembles how the human visual system perceives images. By focusing on the global structure, texture, and semantic information of the image, it helps avoid optimizing only the generated image. The problem of excessive smoothing that occurs when there is loss or SSIM loss.

[0054] The endoscopic image enhancement model is optimized based on the total loss value to obtain the optimal endoscopic image enhancement model, including: The optimizer uses the Adam algorithm to update the parameters of the generator network and uses the cosine annealing algorithm to restart the learning rate scheduler to adjust the learning rate.

[0055] During the training loop, each iteration executes the following five steps: Step 1: Data loading: Acquire batch data, transferring low-quality and high-quality images to the computing device; Step 2: Forward computation: After clearing the gradients, input the low-quality image into the generator network to obtain the output image; Step 3: Loss calculation: Calculate the three types of loss in parallel, with the perceptual loss performing feature matching in a specified convolutional layer; Step 4: Backpropagation: Perform backpropagation on the weighted total loss, and update the network parameters using the Adam optimizer; Step 5: EMA update: Update the EMA model parameters with a decay rate of 0.999.

[0056] In this embodiment, the total loss value is calculated using the total loss function, and the network parameters are iteratively updated by minimizing the multi-objective loss function to obtain the final optimal endoscopic image enhancement model.

[0057] Preferably, a preset test set is used as input to the optimal endoscopic image enhancement model to obtain the final enhancement result, specifically: Inference is performed using the EMA model. In gradient computation disabled mode, low-quality images and high-quality images (i.e., the test set) are processed by the generator to obtain the final output, which returns the enhanced endoscopic images (i.e., the endoscopic enhanced image dataset).

[0058] like Figure 5 As shown, an embodiment of the present invention provides an endoscopic image enhancement system, comprising: A construction unit is used to construct an endoscope image enhancement model, which includes a local perception network, a global color mapping network, and a fusion network. A preprocessing unit is used to import endoscopic image data, which includes initial low-quality images and initial high-quality images, and to preprocess the endoscopic image data to obtain low-quality images and high-quality images. The training unit is used to perform feature enhancement processing on the low-quality image through the local perception network to obtain a low-resolution feature map, perform color mapping processing on the high-quality image through the global color mapping network to obtain a high-resolution feature map, and perform enhancement processing on the low-resolution feature map and the high-resolution feature map through the fusion network to obtain an endoscopy enhanced image. An optimization unit is used to construct a total loss function, calculate the total loss value by applying the total loss function to the imported target enhanced image and the endoscope enhanced image, and optimize the endoscope image enhancement model based on the total loss value to obtain the optimal endoscope image enhancement model. The application unit is used to predict the predicted endoscope image dataset using the endoscope image enhancement model to obtain the endoscope enhanced image dataset.

[0059] The above-mentioned endoscopic image enhancement system can be found in the implementation details and beneficial effects of the endoscopic image enhancement method described above, which will not be repeated here.

[0060] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0061] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the system and unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0062] In the several embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.

[0063] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of the present invention, depending on actual needs.

[0064] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. An endoscopic image enhancement method, characterized in that, Includes the following steps: An endoscopic image enhancement model is constructed, which includes a local perception network, a global color mapping network, and a fusion network. Import endoscopic image data, which includes initial low-quality images and initial high-quality images. Preprocess the endoscopic image data to obtain low-quality images and high-quality images. The low-quality image is enhanced by the local perceptual network to obtain a low-resolution feature map. The high-quality image is then color-mapped by the global color mapping network to obtain a high-resolution feature map. Finally, the low-resolution feature map and the high-resolution feature map are enhanced by the fusion network to obtain an enhanced endoscopic image. A total loss function is constructed, and the total loss value is calculated on the imported target enhanced image and the endoscope enhanced image using the total loss function. The endoscope image enhancement model is then optimized based on the total loss value to obtain the optimal endoscope image enhancement model. The optimal endoscopic image enhancement model is used to predict the endoscopic image dataset to obtain the endoscopic enhanced image dataset. The step of enhancing the low-resolution feature map and the high-resolution feature map through the fusion network to obtain an enhanced endoscopic image includes: The fusion network includes activation function layers, multi-head attention mechanism modules, gated feedforward networks, convolutional layers, multiple deformable convolutional modules, multiple normalization layers, and multiple residual blocks; The low-resolution feature map is extracted by the first deformable convolution module to obtain low-resolution weight features, and the high-resolution feature map is extracted by the second deformable convolution module to obtain high-resolution weight features. The low-resolution weight features are nonlinearly transformed by a second activation function layer, and the high-resolution weight features are nonlinearly transformed by a third activation function layer. The low-resolution weight features after nonlinear transformation are normalized by a fourth normalization layer to obtain low-resolution intermediate features, which are then input into the multi-head attention mechanism module as query vectors. The high-resolution weight features after nonlinear transformation are normalized by a fifth normalization layer to obtain high-resolution intermediate features, which are then input into the multi-head attention mechanism module as key vectors and value vectors. The dimensions of the key vector and the query vector are unified. The key vector, value vector, and query vector are fused through the multi-head attention mechanism module to obtain an initial fused feature. The low-resolution feature map is then added element-wise to the initial fused feature through the first residual block to obtain the fused feature. The number of channels of the fused feature is expanded by the gated feedforward network, and the fused feature is optimized and adjusted based on the multiple channel numbers to obtain a multi-dimensional fused feature; The multidimensional fusion features are compressed by the first convolutional layer to obtain key features. The key features are then added element-wise to the fusion features by the second residual block to obtain the endoscopic enhanced image.

2. The endoscopic image enhancement method according to claim 1, characterized in that, The step of performing feature enhancement processing on the low-quality image through the local perceptron to obtain a low-resolution feature map includes: The local sensing network includes an illumination compensation module, an autoencoder and decoder module, and a residual self-attention module; The low-quality image is adjusted in brightness using the illumination compensation module to obtain a low-quality compensated image. The low-quality compensated image is encoded by the autoencoder and decoder module to obtain a downsampled compensated image, and the downsampled compensated image is decoded to obtain an initial low-resolution image. The initial low-resolution image is weighted by the residual self-attention module to obtain a low-resolution feature map.

3. The endoscopic image enhancement method according to claim 2, characterized in that, The step of adjusting the brightness of the low-quality image through the illumination compensation module to obtain a low-quality compensated image includes: The illumination compensation module includes an exposure normalization layer, a compensation layer, and an attention network. The attention network includes a pooling layer, a fully connected layer, and an activation function layer. The exposure features of the low-quality image are normalized using the exposure normalization layer to obtain a low-quality exposure image. The instance normalization operation is as follows: , in, Indicates a low-quality exposed image, IN ( ) indicates instance normalization, F Indicates a low-quality image. This represents the average value calculated for each channel and sample across the spatial dimension. This represents the standard deviation of each channel relative to the sample in the spatial dimension. γ and β Indicates the learning parameters; The compensation layer performs integral calculations on the unexposed features of the low-quality image to obtain an exposure feature image and a compensation feature image. The expression for the integral calculation of the exposure feature image is as follows: , The expression for calculating the integral of the compensated feature image is: , in, Represents the exposure feature image, This indicates the exposure characteristics of a low-quality exposed image. Indicates the exposure correlation coefficient. Represents the compensated feature image. This represents the non-exposure features of a low-quality image. Represents the compensation correlation coefficient, ( () indicates element-wise multiplication; The exposure feature image and the compensation feature image are pooled by a first pooling layer, the pooled exposure feature image and the compensation feature image are connected by a first fully connected layer, and the connected exposure feature image and the compensation feature image are classified by a first activation function layer to obtain a low-quality compensation image.

4. The endoscopic image enhancement method according to claim 2, characterized in that, The step of weighting the initial low-resolution image through the residual self-attention module to obtain a low-resolution feature map includes: The residual self-attention module includes a pooling layer, multiple fully connected layers, and multiple normalization layers; The initial low-resolution image is linearly transformed using a second fully connected layer to obtain a value vector. A second pooling layer then performs spatial downsampling on the initial low-resolution image. The downsampled initial low-resolution image is then concatenated using a second fully connected layer to obtain attention weights. These attention weights are then normalized using a first normalization layer and classified using a second normalization layer to obtain normalized attention weights. Regularized attention weights are then obtained. The value vector and the regularized attention weights are weighted and summed to obtain a weighted value vector. A third fully connected layer performs a linear transformation on the weighted value vector to obtain a low-resolution feature map, which is then regularized to obtain a regularized low-resolution feature map. Finally, the initial low-resolution image and the regularized low-resolution feature map are residually concatenated and summed to obtain the final low-resolution feature map.

5. The endoscopic image enhancement method according to claim 1, characterized in that, The step of performing color mapping processing on the high-quality image through the global color mapping network to obtain a high-resolution feature map includes: The global color mapping network includes an adaptive downsampling module and a color enhancement module; The high-quality image is sampled using the adaptive downsampling module to obtain a resampled image. The color enhancement module performs global color enhancement processing on the resampled image to obtain a high-resolution feature map.

6. The endoscopic image enhancement method according to claim 5, characterized in that, The step of sampling the high-quality image through the adaptive downsampling module to obtain a resampled image includes: The adaptive downsampling module includes a lightweight backbone network, an interval prediction layer, a normalization layer, and a sampling construction layer; The high-quality image is downsampled using the lightweight backbone network to obtain a multi-dimensional global feature vector. The interval prediction layer then predicts the interval of the global feature vector to obtain multiple sampling intervals. A third normalization layer normalizes these multiple sampling intervals to obtain multiple coordinate points. The sampling construction layer generates a sampling grid based on these coordinate points, and bilinear interpolation is performed on the high-quality image using the sampling grid to obtain a resampled image.

7. The endoscopic image enhancement method according to claim 5, characterized in that, The step of performing global color enhancement processing on the resampled image through the color enhancement module to obtain a high-resolution feature map includes: The color enhancement module includes a lookup table construction layer, a convolutional neural network, a linear layer, a weighted fusion layer, and an enhancement processing layer; Multiple three-dimensional lookup tables are constructed through the lookup table construction layer. Features are extracted from the multiple three-dimensional lookup tables by the convolutional neural network to obtain multiple resampled features. The multiple resampled features are linearly mapped by the linear layer to obtain multiple prediction weights. The multiple prediction weights are weighted with the multiple three-dimensional lookup tables by the weighted fusion layer to obtain the optimal three-dimensional lookup table. The enhancement processing layer performs global mapping of the colors of the resampled image according to the optimal three-dimensional lookup table to obtain a high-resolution feature map.

8. The endoscopic image enhancement method according to claim 1, characterized in that, The step of extracting features from the low-resolution feature map using the first deformable convolution module to obtain low-resolution weighted features includes: The deformable convolutional module includes a sampling layer, a pooling layer, an activation function layer, a residual block, and multiple convolutional layers; The low-resolution feature map is convolved by a second convolutional layer to obtain an offset. The position of each sampling point in the low-resolution feature map is calculated based on the offset to obtain multiple sampling points. The sampling layer selects sample groups for each sampling point according to the sampling rules to obtain multiple sample groups, and calculates the bilinear weights corresponding to the multiple sample groups. Based on each bilinear weight, the corresponding four-corner feature values ​​are extracted from the low-resolution feature map. The multiple four-corner feature values ​​are weighted and summed to obtain the final sampling value. The low-resolution feature map is sampled according to the final sampling value to obtain deformed convolution features. The deformed convolutional features are transformed and compressed by the third pooling layer to obtain compressed convolutional features. The compressed convolutional features are then convolved by the third convolutional layer to obtain local cross-channel features. The local cross-channel features are weighted by the fourth activation function layer to obtain weighted features. The deformed convolutional features and the weighted features are then multiplied element-wise by the third residual block to obtain low-resolution weighted features.

9. An endoscopic image enhancement system, characterized in that, include: A construction unit is used to construct an endoscope image enhancement model, which includes a local perception network, a global color mapping network, and a fusion network. A preprocessing unit is used to import endoscopic image data, which includes initial low-quality images and initial high-quality images, and to preprocess the endoscopic image data to obtain low-quality images and high-quality images. The training unit is used to perform feature enhancement processing on the low-quality image through the local perception network to obtain a low-resolution feature map, perform color mapping processing on the high-quality image through the global color mapping network to obtain a high-resolution feature map, and perform enhancement processing on the low-resolution feature map and the high-resolution feature map through the fusion network to obtain an endoscopy enhanced image. An optimization unit is used to construct a total loss function, calculate the total loss value by applying the total loss function to the imported target enhanced image and the endoscope enhanced image, and optimize the endoscope image enhancement model based on the total loss value to obtain the optimal endoscope image enhancement model. The application unit is used to predict the predicted endoscope image dataset using the optimal endoscope image enhancement model to obtain the endoscope enhanced image dataset; The step of enhancing the low-resolution feature map and the high-resolution feature map through the fusion network to obtain an enhanced endoscopic image includes: The fusion network includes activation function layers, multi-head attention mechanism modules, gated feedforward networks, convolutional layers, multiple deformable convolutional modules, multiple normalization layers, and multiple residual blocks; The low-resolution feature map is extracted by the first deformable convolution module to obtain low-resolution weight features, and the high-resolution feature map is extracted by the second deformable convolution module to obtain high-resolution weight features. The low-resolution weight features are nonlinearly transformed by a second activation function layer, and the high-resolution weight features are nonlinearly transformed by a third activation function layer. The low-resolution weight features after nonlinear transformation are normalized by a fourth normalization layer to obtain low-resolution intermediate features, which are then input into the multi-head attention mechanism module as query vectors. The high-resolution weight features after nonlinear transformation are normalized by a fifth normalization layer to obtain high-resolution intermediate features, which are then input into the multi-head attention mechanism module as key vectors and value vectors. The dimensions of the key vector and the query vector are unified. The key vector, value vector, and query vector are fused through the multi-head attention mechanism module to obtain an initial fused feature. The low-resolution feature map is then added element-wise to the initial fused feature through the first residual block to obtain the fused feature. The number of channels of the fused feature is expanded by the gated feedforward network, and the fused feature is optimized and adjusted based on the multiple channel numbers to obtain a multi-dimensional fused feature; The multidimensional fusion features are compressed by the first convolutional layer to obtain key features. The key features are then added element-wise to the fusion features by the second residual block to obtain the endoscopic enhanced image.

Citation Information

Patent Citations

  • Reference-free low-illumination endoscope image quality enhancement method and system and related components

    CN114119422A

  • Intestinal endoscope image enhancement method based on image fusion

    CN116188340A