Endoscope image quality evaluation and enhancement method based on cloud features

CN122550433APending Publication Date: 2026-08-11HARBIN INST OF TECH AT WEIHAI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-14
Publication Date
2026-08-11

AI Technical Summary

Benefits of technology

[0026]1、本发明通过提取局部区域对应的期望Ex、熵En及超熵He,实现了对内镜影像局部退化特征的有效表征,能够提高对模糊、噪声及光照变化区域的识别能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122550433A_ABST
    Figure CN122550433A_ABST
Patent Text Reader

Abstract

This invention discloses a cloud feature-based method for endoscopic image quality assessment and enhancement. The method effectively characterizes local degradation features of endoscopic images by extracting the expected value (Ex), entropy (En), and hyperentropy (He) corresponding to local regions. Uncertainty weights are generated based on entropy (En) and hyperentropy (He), and these weights are used to modulate the enhancement decoding features, enabling the enhancement process to adaptively adjust according to the degree of degradation in different regions, thereby reducing over-enhancement, loss of detail, and artifact enhancement in local areas. Secondary cloud modeling and refined enhancement of the initially enhanced image are performed using a cloud atlas inferrer, further optimizing image details and suppressing residual degradation, thus improving the precision and robustness of the final enhancement result. By generating a pixel-level confidence map, the reliability of the enhancement result is visualized, providing doctors with auxiliary reference information on the reliability of the enhanced region.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of medical image processing and artificial intelligence technology, and relates to a method for endoscopic image quality assessment and enhancement. Background Technology

[0002] Endoscopic technology, as an important tool for screening and diagnosing gastrointestinal diseases, is widely used in clinical practice. However, during actual examinations, due to factors such as the movement of the endoscopic equipment, patient respiration and peristalsis, changes in lighting conditions, and the complex internal environment, the acquired endoscopic images often suffer from various quality degradation problems, including motion blur, uneven lighting, tissue reflection, and mucus obscuring. These problems not only reduce image readability but may also lead to difficulties in identifying lesion areas, thereby increasing the risk of missed diagnoses and misdiagnoses. Therefore, high-quality evaluation and adaptive enhancement of endoscopic images are of great significance.

[0003] Existing image quality assessment methods mainly include those based on subjective evaluation metrics and those based on objective evaluation metrics, such as peak signal-to-noise ratio (PSNR) and structural similarity. These methods typically only output a single numerical value, making it difficult to describe the spatial distribution of degradation levels in different regions of the image, let alone characterize the randomness and uncertainty inherent in the degradation process. Furthermore, these methods are mostly used as post-hoc evaluation tools and cannot provide effective guidance for subsequent image enhancement.

[0004] In image enhancement, traditional methods such as histogram equalization and the Retinex algorithm typically rely on fixed enhancement strategies, making it difficult to handle complex and varied degradation patterns and prone to over-enhancement or artifacts. In recent years, deep learning-based image enhancement methods (such as convolutional neural networks and generative adversarial networks) have improved enhancement results to some extent, but they generally suffer from the following shortcomings: First, they lack explicit modeling of image uncertainty, failing to distinguish the degree of degradation and its random fluctuation characteristics in different regions, resulting in a lack of targeted enhancement strategies; second, most methods rely on pixel-level loss functions for training, making it difficult to constrain the enhancement results from the perspective of uncertainty; and third, the enhancement results lack credibility expression, making it difficult to provide interpretable auxiliary information for clinicians.

[0005] Therefore, there is an urgent need for a new technical solution that can uniformly quantify the complex uncertainties with spatial distribution characteristics in endoscopic images, effectively embed this uncertainty information into the image enhancement process, and provide a credibility evaluation of the enhancement results to improve the stability of the enhancement effect. Summary of the Invention

[0006] To address the aforementioned technical problems, this invention provides a cloud-feature-based method for endoscopic image quality assessment and enhancement. This method is applicable to various endoscopic image enhancement scenarios and has significant engineering application and medical diagnostic value.

[0007] The objective of this invention is achieved through the following technical solution:

[0008] A method for endoscopic image quality assessment and enhancement based on cloud features includes the following steps:

[0009] Step S1: Obtain the endoscopic image to be enhanced, and perform local adaptive sampling on the endoscopic image to obtain multiple local sampling windows;

[0010] Step S2: Calculate the local quality index based on the pixel information in each local sampling window, and use the inverse cloud generator to extract cloud features from the local quality index to obtain the cloud feature parameters of the corresponding local area.

[0011] Step S3: Map the cloud feature parameters corresponding to each local region to cloud feature maps according to their spatial location, and fuse the cloud feature maps with the latent feature maps obtained by encoding the endoscopic images to generate cloud-guided feature maps.

[0012] Step S4: Input the cloud-guided feature map into the enhancement network for image enhancement. Dynamically generate enhancement parameters based on the entropy En and hyperentropy He corresponding to the local region, and generate uncertainty weights based on the entropy En and hyperentropy He to perform weighted modulation on the decoded features in the enhancement network.

[0013] Step S5: Extract secondary cloud features again based on the enhanced preliminary image, and generate feature adjustment signals according to the difference between the secondary cloud features and the preset expected features to dynamically calibrate the features of the enhancement network and output the adjusted final enhanced image.

[0014] Step S6: Generate a pixel-level confidence map based on the entropy En and hyperentropy He of the enhanced image, and output the final enhanced image and the corresponding confidence map.

[0015] A cloud-feature-based endoscopic image quality assessment and enhancement system implementing the above method includes an image acquisition module, a cloud feature extraction module, a cloud-guided enhancement coding network, a feature fusion module, a cloud-guided enhancement decoding network, a cloud atlas inferrer, and an enhancement output module, wherein:

[0016] The image acquisition module is used to acquire the endoscopic images to be processed.

[0017] The cloud feature extraction module is used to extract the expected value Ex, entropy En, and hyperentropy He corresponding to the local region.

[0018] The cloud-guided augmented coding network is used to extract high-dimensional latent feature maps from raw endoscopic images;

[0019] The feature fusion module is used to fuse the cloud feature map with the potential feature map to generate a cloud-guided feature map;

[0020] The cloud-guided enhancement decoding network is used to generate uncertainty weights based on entropy En and hyperentropy He, and to perform weighted modulation on the decoding features to output a preliminary enhanced image;

[0021] The cloud map inferrer is used to perform secondary cloud modeling and fine enhancement on the preliminary enhanced image, and generate the final enhanced image and the corresponding pixel-level confidence map.

[0022] The enhanced output module is used to output the final enhanced image and the corresponding pixel-level confidence map.

[0023] An endoscopic image enhancement device includes a processor, a memory, and a program stored in the memory and executable on the processor; when the program is executed by the processor, it implements the aforementioned cloud-feature-based endoscopic image quality assessment and enhancement method.

[0024] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described cloud-feature-based endoscopic image quality assessment and enhancement method.

[0025] Compared with the prior art, the present invention has the following advantages:

[0026] 1. This invention effectively characterizes the local degradation features of endoscopic images by extracting the expected value Ex, entropy En, and hyperentropy He corresponding to local regions, thereby improving the ability to identify blurred, noisy, and illumination-varying regions.

[0027] 2. This invention generates uncertain weights based on entropy En and hyperentropy He, and uses these uncertain weights to modulate the enhanced decoding features, enabling the enhancement process to adaptively adjust according to the degree of degradation in different regions, thereby reducing over-enhancement in local regions, loss of detail, and artifact enhancement.

[0028] 3. This invention performs secondary cloud modeling and fine enhancement on the preliminary enhanced image through a cloud map inferrer, which can further optimize image details and suppress residual degradation, thereby improving the precision and robustness of the final enhancement result.

[0029] 4. By generating a pixel-level confidence map, this invention enables a visual representation of the credibility of the enhancement results, providing doctors with auxiliary reference information on the reliability of the enhancement area. Attached Figure Description

[0030] Figure 1A schematic diagram of a cloud model and cloud characteristics;

[0031] Figure 2 It is a single-condition uncertainty inferrer;

[0032] Figure 3 This is the overall processing flowchart for Example 1;

[0033] Figure 4 This is a schematic diagram of the Cloud Feature Extraction Module (CFE).

[0034] Figure 5 A schematic diagram of cloud-guided enhanced deep networks;

[0035] Figure 6 This is a schematic diagram of upsampling and downsampling;

[0036] Figure 7 This is a schematic diagram of the cloud attention mechanism;

[0037] Figure 8 This is a schematic diagram of the Cloud Map Inferrer (CAID) structure. Detailed Implementation

[0038] The technical solution of the present invention will be further described below with reference to the accompanying drawings, but it is not limited thereto. Any modifications or equivalent substitutions to the technical solution of the present invention that do not depart from the spirit and scope of the technical solution of the present invention should be covered within the protection scope of the present invention.

[0039] This invention provides a method for endoscopic image quality assessment and enhancement based on cloud features. The method samples local regions of the endoscopic image and extracts corresponding expected value (Ex), entropy (En), and hyperentropy (He) to generate cloud feature parameters. These cloud feature parameters are then mapped to generate a cloud feature map. Uncertainty weights are generated based on the cloud feature map to weight the features in the enhancement network. Simultaneously, a feature adjustment signal is generated based on the difference in cloud feature parameters before and after enhancement, outputting an enhanced image and its corresponding confidence level. Specifically, the method includes the following steps:

[0040] Step S1: Obtain the endoscopic image to be enhanced, and perform local adaptive sampling on the endoscopic image to obtain multiple local sampling windows.

[0041] In this step, local region adaptive sampling uses deformable convolution to generate a sampling offset field, and dynamically adjusts the sampling position of the convolution kernel according to the sampling offset field, so that the sampling window can be dynamically adjusted according to the local structure of the image to adapt to tissue edges, reflective areas and blurred areas in the endoscopic image.

[0042] Step S2: Calculate the local quality index based on the pixel information in each local sampling window, and use the inverse cloud generator to extract cloud features from the local quality index to obtain the cloud feature parameters of the corresponding local area.

[0043] In this step, the local quality index is calculated based on the pixel gradient information within the local window. The local quality index includes at least one of the following: Sobel gradient magnitude; Laplacian response value; local brightness contrast; and high-frequency texture response value.

[0044] In this step, the cloud feature parameters include at least: expectation Ex, used to characterize the quality benchmark of the local region; entropy En, used to characterize the fuzziness of the local region; and hyperentropy He, used to characterize the random discreteness of the local region.

[0045] In this step, the reverse cloud generator is used to convert local quality indicators into cloud feature parameters, including expectation Ex, entropy En, and hyperentropy He.

[0046] The cloud model is an uncertainty quantification model used to convert between qualitative concepts and quantitative values. The cloud model converts qualitative concepts... In the sample space The distribution of data on a cloud is called a cloud. This expresses a qualitative concept. A cloud is composed of multiple cloud droplets. Each cloud droplet... Both are qualitative concepts. Mapping to sample space One point is the concrete realization of the semantic value of a qualitative concept in terms of quantity. This realization carries uncertainty, and the cloud model simultaneously provides cloud droplets. and its descriptive qualitative concepts The degree of certainty, i.e. Cloud Drop With certainty The relationship curve becomes cloud The expected curve is shown in equation (1).

[0047] (1)

[0048] like Figure 1 The cloud model comprehensively reflects the overall characteristics of a concept through expected value Ex, entropy En, and hyperentropy He.

[0049] Expected Ex represents the sample space The most representative of qualitative concepts The value, i.e., the qualitative concept The center of gravity of the cloud. Entropy En reflects both the sample space and the density of the cloud. The concept of being qualitative The acceptable range, i.e., the ambiguity; also reflects the sample space. The elements in the text represent qualitative concepts. The probability, or randomness, of cloud entropy. Hyperentropy He is a measure of the uncertainty of entropy, representing the degree of dispersion and thickness of a cloud.

[0050] The cloud model utilizes a cloud generator (CG) to generate cloud bodies and convert cloud bodies into cloud features. The inverse cloud generator is used to convert a certain number of cloud droplets into cloud feature values; its input is... Data points The output is cloud features (Ex, En, He).

[0051] In this step, the reverse cloud generator is implemented using the mean method, and its calculation is as shown in equation (2).

[0052] (2)

[0053] Through the aforementioned reverse cloud generator, the quality index of each local window is converted into corresponding cloud feature parameters (Ex, En, He), thereby achieving a unified quantitative characterization of the degree of degradation and uncertainty of local areas of endoscopic images.

[0054] Step S3: Map the cloud feature parameters corresponding to each local region to cloud feature maps according to their spatial location, and fuse the cloud feature maps with the latent feature maps obtained by encoding the endoscopic images to generate cloud-guided feature maps.

[0055] In this step, the fusion includes: aligning the cloud feature map and the latent feature map in terms of spatial dimensions; concatenating the channels of the aligned feature maps; and generating the fused cloud-guided feature map through convolutional mapping.

[0056] Step S4: Input the cloud-guided feature map into the enhancement network for image enhancement. Dynamically generate enhancement parameters based on the entropy En and hyperentropy He corresponding to the local region, and generate uncertainty weights based on the entropy En and hyperentropy He to perform weighted modulation on the decoded features in the enhancement network.

[0057] In this step, the enhancement parameters include at least one of the following: convolution weights; feature response gain; attention weights; and enhancement intensity coefficients.

[0058] In this step, the uncertainty weights are generated based at least on the entropy En and the super-entropy He, and are used to weight and modulate the local features in the augmented network. Specifically, this includes: inputting the entropy En and the super-entropy He into the convolutional mapping layer to generate an uncertainty heatmap, and using the uncertainty heatmap to weight the decoded features in the augmented network element by element.

[0059] Step S5: Perform secondary cloud modeling and refinement enhancement on the preliminary enhanced image using a cloud map inferrer: Extract secondary cloud features again based on the enhanced preliminary enhanced image, and generate feature adjustment signals according to the difference between the secondary cloud features and the preset expected features, so as to dynamically calibrate the features of the enhancement network and output the adjusted final enhanced image.

[0060] Optionally, the enhanced cloud features can be fed back to the input of the enhanced coding network or the decoding network to form a feedback adjustment mechanism for the enhancement results, thereby performing closed-loop optimization of the enhanced network.

[0061] In this step, the secondary cloud modeling process is implemented based on a conditional cloud generator. The conditional cloud generator includes an X-cloud generator (XCG) and a Y-cloud generator (YCG), where the X-cloud generator is used to model the cloud based on given input values. The Y-conditional cloud generator generates cloud droplets under given deterministic conditions. Cloud droplets are generated under certain conditions. For example... Figure 2 As shown, by connecting the X-condition cloud generator and the Y-condition cloud generator in series, a single-condition uncertainty inferrer is constructed to realize the uncertainty mapping from input to output.

[0062] In the cloud map inferrer, secondary cloud features are extracted, and the same inverse cloud generator as in step S2 is used to extract cloud feature maps again from the initially enhanced image. The refinement network uses a conditional cloud generator to generate feature adjustment signals based on the difference between the secondary cloud features and the preset desired features, dynamically calibrating the features of the enhancement network to further optimize image details and suppress residual degradation.

[0063] In this step, the generation of the feature adjustment signal is optimized based on the cloud feature difference loss function. During the training phase, the system's total loss function consists of the following three components:

[0064] (1) Pixel loss

[0065] Pixel loss is used to measure the pixel-level reconstruction error between the enhanced image and the high-quality reference image. It is a weighted combination of L1 loss and SSIM loss, and its calculation method is as follows:

[0066] (3)

[0067] in, To enhance the image, For reference ground truth image, These are the weighting coefficients. For L1 loss function, This is the structural similarity loss function. In the training process of this example, the initially enhanced image is... The final cloud-enhanced image is .

[0068] (2) Cloud feature difference loss

[0069] The cloud feature difference loss is used to constrain the cloud features of the final enhanced image to converge towards the cloud feature distribution of the ideal reference image. Ideally, a high-quality image should have low entropy and low hyperentropy, and its calculation method is as follows:

[0070] (4)

[0071] in, For local block indexes, The total number of blocks, and The enhanced image is the first The entropy and hyperentropy of a local block, and These are the weighting coefficients.

[0072] (3) Minimize the regularization term for uncertainty

[0073] The uncertainty minimization regularization term directly affects the final enhanced image, prompting the network to reduce the dispersion and randomness of the quality distribution in the enhanced image. Its calculation method is as follows:

[0074] (5)

[0075] in, It iterates through pixel positions. Total number of pixels and These are pixel-level entropy and hyper-entropy, respectively.

[0076] The total loss function is defined as the weighted sum of the three losses mentioned above:

[0077] (6)

[0078] in, , and These represent the weighting coefficients for each loss term. The entire system employs an end-to-end joint optimization approach, with all modules updating their parameters synchronously within the same training framework.

[0079] Step S6: Generate a pixel-level confidence map based on the entropy En and hyperentropy He of the enhanced image, and output the final enhanced image and the corresponding confidence map.

[0080] This invention also provides a cloud feature-based endoscopic image quality assessment and enhancement system for implementing the above method. The system includes: an image acquisition module, a cloud feature extraction module, a cloud-guided enhancement coding network, a feature fusion module, a cloud-guided enhancement decoding network, a cloud atlas inferrer, and an enhancement output module, wherein:

[0081] The image acquisition module is used to acquire the endoscopic image to be processed, corresponding to step S1;

[0082] The cloud feature extraction module is used to extract the expected value Ex, entropy En and hyperentropy He corresponding to the local region, and output a cloud feature map that matches the spatial size of the encoder feature map, corresponding to step S2;

[0083] The cloud-guided augmented coding network is used to extract high-dimensional latent feature maps from the original endoscopic images, corresponding to the part encoded in step S3.

[0084] The feature fusion module is used to fuse the cloud feature map with the potential feature map to generate a cloud-guided feature map, corresponding to step S3;

[0085] The cloud-guided enhancement decoding network is used to generate uncertainty weights based on entropy En and hyperentropy He, and to perform weighted modulation on the decoding features to output a preliminary enhanced image, corresponding to step S4;

[0086] The cloud map inferrer is used to perform secondary cloud modeling and fine enhancement on the preliminary enhanced image, and generate the final enhanced image and the corresponding pixel-level confidence map, corresponding to steps S5 and S6.

[0087] The enhanced output module is used to output the final enhanced image and the corresponding pixel-level confidence map, corresponding to step S6.

[0088] The present invention also provides an endoscopic image enhancement device, the device including a processor, a memory, and a program stored in the memory and executable on the processor; when the program is executed by the processor, it implements the above-mentioned cloud feature-based endoscopic image quality assessment and enhancement method.

[0089] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method for endoscopic image quality assessment and enhancement based on cloud features.

[0090] Example:

[0091] This embodiment uses the enhancement of endoscopic images, particularly static colonoscopy images, as a typical application scenario to systematically illustrate the specific implementation scheme of the cloud feature-based endoscopic image quality assessment and enhancement method proposed in this invention. This method aims to improve the clarity of endoscopic images, enhance local structural details, and output a confidence level indicator of the enhancement results to assist in clinical diagnosis.

[0092] 1. Overall Architecture

[0093] In this embodiment, the system adopts a modular design, and the overall architecture mainly consists of the following seven modules:

[0094] (1) Image acquisition module: used to acquire endoscopic images to be processed.

[0095] (2) Cloud Feature Extractor (CFE): Used to quantify the sources of uncertainty in endoscopic images (such as motion blur, uneven illumination, etc.) into cloud feature parameters.

[0096] (3) Cloud-guided Enhancement Encoder (CGEE): used to extract high-dimensional latent feature maps from the original image.

[0097] (4) Feature fusion module: used to fuse cloud feature maps with potential feature maps to generate cloud-guided feature maps.

[0098] (5) Cloud-guided Enhancement Decoder (CGED): used to decode and reconstruct the fused features to generate a preliminary enhanced image.

[0099] (6) Cloud Atlas Inference Device (CAID): Used to perform secondary cloud modeling and fine enhancement on the initial enhanced image, and generate the final enhanced image and pixel-level confidence map.

[0100] (7) Enhanced output module: used to output the final enhanced image and the corresponding pixel-level confidence map.

[0101] like Figure 3 As shown, the overall processing flow is as follows: After the original endoscopic image is input into the system, it is sent to the cloud feature extraction module to generate a cloud feature map, and to the cloud-guided enhancement coding network to obtain a latent feature map. The two are then fused by the feature fusion module to generate a cloud-guided feature map rich in uncertainty information. This feature map is subsequently sent to the cloud-guided enhancement decoding network, where it undergoes adaptive enhancement through a cloud attention mechanism to obtain a preliminary enhanced image. Finally, the cloud map inferrer performs secondary cloud modeling and refinement on the preliminary enhanced image, outputting the final cloud-enhanced image and its corresponding pixel-level confidence map in parallel.

[0102] Through the above process, this embodiment realizes integrated processing of endoscopic images from quality assessment, adaptive enhancement to result credibility expression, providing high-quality image data and reliable auxiliary reference information for clinical diagnosis.

[0103] 2. Detailed System Structure

[0104] The system in this embodiment mainly includes the following functional modules: image acquisition module, cloud feature extraction module, cloud-guided enhanced coding network, feature fusion module, cloud-guided enhanced decoding network, cloud map inferrer, and enhanced output module. The structure and function of each module are described in detail below.

[0105] (1) Image acquisition module

[0106] The image acquisition module is used to acquire endoscopic images to be processed. In this embodiment, the input is a single static colonoscopy image, using a standard RGB three-channel format, with a uniform image size of 512×512 pixels. This module also supports frame-by-frame input of video frame sequences. The image acquisition module directly transmits the acquired raw images to the cloud feature extraction module and the cloud-guided augmented coding network for subsequent processing.

[0107] (2) Cloud Feature Extraction Module (CFE)

[0108] like Figure 4 As shown, the cloud feature extraction module is used to quantify the sources of uncertainty in endoscopic images, such as motion blur, uneven illumination, tissue reflection and mucus interference, into cloud feature parameters, namely expectation Ex, entropy En and hyperentropy He, and output a cloud feature map that matches the spatial size of the encoder feature map.

[0109] The workflow of this module includes the following steps:

[0110] First, deformable convolution is used to perform local region adaptive sampling of the image, generating multiple sampling windows. The deformable convolution generates a sampling position offset field through an independent convolutional layer, enabling the sampling windows to be dynamically adjusted according to the local structure of the image, thereby focusing on edges, reflective and blurred areas with significant uncertainty.

[0111] Secondly, local quality indices are calculated based on pixel information within each sampling window. These quality indices reflect the degree of degradation in local image regions. In this embodiment, the gradient energy method is used, employing the Sobel operator to calculate the gradient magnitude and normalizing it to the [0,1] interval.

[0112] Finally, a reverse cloud generator is used to extract cloud features from local quality indicators, obtaining cloud feature parameters for each local window, including: the expectation Ex representing the quality benchmark of the local region, the entropy En representing the degree of ambiguity, and the hyperentropy He representing random discreteness. The cloud feature parameters of all local windows are arranged into three channels according to their spatial location, resulting in a cloud feature map that matches the spatial size of the encoder output feature map.

[0113] (3) Cloud-guided augmented coding network (CGEE)

[0114] The cloud-guided augmented coding network is used to extract high-dimensional latent feature maps from raw endoscopic images while preserving spatial structure information. In this embodiment, the network employs a convolutional encoder structure, comprising four sequentially connected downsampling stages. Each downsampling stage consists of two 3×3 convolutional layers with batch normalization and ReLU activation functions, and a 2×2 max-pooling layer. The encoder's initial channel count is set to 64, doubling with each downsampling stage, ultimately outputting a latent feature map with 256 channels. Furthermore, the coding network retains the feature maps preceding each downsampling stage as skip connections for use by subsequent decoding networks.

[0115] (4) Feature fusion module

[0116] The feature fusion module is used to fuse the cloud feature map and the latent feature map to generate a cloud-guided feature map. The fusion process includes: first, aligning the spatial dimensions of the cloud feature map and the latent feature map; if the dimensions are inconsistent, bilinear interpolation is used for adjustment; then, concatenating the two aligned feature maps by channel; finally, compressing the number of channels of the concatenated feature map back to 256 through a convolutional layer, and applying the ReLU activation function to output a cloud-guided feature map rich in uncertainty information. This cloud-guided feature map serves as both the final output of the encoder and the conditional information passed to the decoder.

[0117] (5) Cloud-guided Enhanced Decoding Network (CGED)

[0118] like Figure 5 As shown, the cloud-guided enhancement decoding network is used to decode and reconstruct the cloud-guided feature map, outputting a preliminary enhanced image. In this embodiment, the decoding network adopts a convolutional decoder structure symmetrical to the encoder, containing four sequentially connected upsampling stages. Figure 6 As shown, each upsampling stage first uses transposed convolution for upsampling, halving the number of output channels. Then, it is concatenated with the skip connection feature map from the corresponding layer of the encoder, and then passed through two convolutional layers with batch normalization and ReLU activation functions.

[0119] In each upsampling stage, this embodiment introduces a cloud attention mechanism. For example... Figure 7 As shown, the cloud attention mechanism generates uncertainty weights based on the entropy En and hyperentropy He in the cloud feature parameters, and performs element-wise weighted modulation on the decoded features. Specifically, the mechanism separates the entropy En and hyperentropy He channels from the cloud-guided feature map, inputs them into the convolutional mapping layer to generate a single-channel uncertainty heatmap, and upsamples it to the size of the current decoded feature map. Then, the uncertainty heatmap is multiplied element-wise with the current decoded feature map, enabling the enhancement process to adaptively enhance regions with high uncertainty such as blur, noise, and illumination changes, thereby reducing over-enhancement and artifacts.

[0120] The final layer of the decoding network employs a convolutional layer and the Tanh activation function, outputting a preliminary enhanced image with pixel values ​​normalized to [-1, 1]. This preliminary enhanced image is then fed into a cloud map inferrer for further refinement.

[0121] (6) Cloud Map Inferrer (CAID)

[0122] like Figure 8 As shown, the cloud map inferrer is used to perform secondary cloud modeling and fine enhancement on the initially enhanced image, ultimately outputting a high-confidence final enhanced image and its corresponding pixel-level confidence map. The inferrer is composed of the following three sub-modules connected in series:

[0123] a) Secondary Cloud Feature Extraction Submodule: This submodule has the same structure as the aforementioned cloud feature extraction module and is used to extract cloud feature maps again from the initially enhanced image. The original image resolution is maintained here for subsequent pixel-level confidence calculations.

[0124] b) Refined Network: The refined network is a lightweight convolutional network used to refine the initially enhanced image. In this embodiment, the network consists of three convolutional layers. The input is a channel concatenation of the initially enhanced image and the secondary cloud feature map, and the output is the final cloud-enhanced image. During the training phase, the refined network generates a feature adjustment signal based on the difference between the secondary cloud features and the preset desired features to dynamically calibrate the features of the enhancement network, thereby further optimizing image details and suppressing residual degradation.

[0125] c) Confidence Map Generator: The confidence map generator calculates confidence values ​​pixel-by-pixel based on the entropy En and hyperentropy He of the final enhanced image, generating a pixel-level confidence map. The confidence calculation formula makes the value closer to 1, indicating that the pixel quality is more reliable. Optionally, the confidence map can be Gaussian smoothed to eliminate isolated noise points.

[0126] Finally, the cloud map inferrer outputs cloud-enhanced images and corresponding confidence maps in parallel. The cloud-enhanced images can be used for clinical diagnosis, and the confidence maps can be overlaid into heat maps to assist doctors in assessing the confidence level of each region in the image.

[0127] (7) Enhanced output module

[0128] The enhanced output module is used to output the final enhanced image and the corresponding pixel-level confidence map.

[0129] System overall reasoning process

[0130] During the inference phase, the system strictly follows the following data flow sequence: the original endoscopic images are simultaneously fed into the cloud feature extraction module and the cloud-guided enhancement coding network, which output cloud feature maps and latent feature maps, respectively; the two are fused by the feature integration module to generate a cloud-guided feature map; the cloud-guided feature map is fed into the cloud-guided enhancement decoding network, which performs adaptive enhancement through the cloud attention mechanism, and outputs a preliminary enhanced image; the preliminary enhanced image is then subjected to secondary modeling and refinement by the cloud map inferrer, and finally outputs the cloud-enhanced image and its corresponding confidence map.

[0131] 3. Key parameter settings

[0132] This embodiment uses endoscopic images (static colonoscopy images) as the application object, and the specific parameters of each module have been determined based on actual experiments. The following description follows the order of dataset construction, cloud feature extraction module parameters, network structure parameters, training parameters, and loss function configuration.

[0133] 3.1 Dataset Construction and Preprocessing

[0134] In this embodiment, the publicly available endoscopy image datasets PolypsSet and GASTROLAB are selected as the basic data sources. To construct paired training data, low-quality-high-quality image pairs are generated in the following way: using high-quality original endoscopy images as the reference ground value, corresponding low-quality images are generated by simulating common quality degradation processes in the endoscopy imaging process. The degradation process includes: (1) Gaussian blur to simulate motion blur, with the kernel size set to 5×5 and the standard deviation uniformly sampled from the interval [1, 5]; (2) random brightness adjustment to simulate uneven illumination, with brightness multiplied by a coefficient and the coefficient uniformly sampled from the interval [0.7, 1.3]; (3) Gaussian noise to simulate sensor noise, with the noise standard deviation uniformly sampled from the interval [0, 10]; (4) random contrast adjustment to simulate tissue reflection and mucus interference, with contrast multiplied by a coefficient and the coefficient uniformly sampled from the interval [0.5, 1.5]. A total of 8000 pairs of training samples, 1000 pairs of validation samples, and 500 pairs of test samples are generated. All images are stored in a uniform size of 512×512 pixels, using an 8-bit RGB three-channel format.

[0135] 3.2 Cloud Feature Extraction Module Parameters

[0136] The cloud feature extraction module employs deformable convolution for adaptive local sampling. The convolution kernel size is set to 3×3, and the offset field is generated through a separate convolutional layer with 9 sampling points. To maintain spatial size matching with the encoder output feature map, the downsampling stride is set to 2, meaning a local window is extracted every 2 pixels. The window stride is consistent with the downsampling factor (when the input size is 512, the output feature map size is 256×256). Each local window covers a 3×3 pixel region, containing 9 pixels. The quality metric uses a local gradient energy scheme: first, the Sobel gradient magnitude is calculated, normalized, and then mapped to the [0,1] interval using the Sigmoid function. In the inverse cloud generator algorithm, the entropy correction coefficient is taken as... The calculation of hyperentropy introduces a minimum constant ε = 1 × 10⁻⁶. -8 To prevent numerical instability, the formula for calculating hyperentropy He is as follows: ,in For sample variance, The entropy is represented by the output cloud feature map channels: expectation Ex, entropy En, and hyperentropy He, each occupying one channel.

[0137] 3.3 Network Structure Parameters

[0138] Cloud-Guided Augmented Encoding Network (CGEE): The encoder consists of four downsampling stages, each consisting of two 3×3 convolutional layers (with BatchNorm and ReLU) and a 2×2 max-pooling layer. The initial number of channels is 64, doubling with each downsampling stage, resulting in a final output of 256 channels. The spatial size of the latent feature map output by the encoder is 256×256.

[0139] Feature integration module: The cloud feature map (3 channels) and the latent feature map (256 channels) are concatenated. The input has 259 channels, which are mapped to a 256-channel output through a convolutional layer. The ReLU activation function is then used to generate the cloud-guided feature map.

[0140] The Cloud-Guided Enhanced Decoding Network (CGED) consists of four upsampling stages, symmetrical to the encoder. Each stage first upsamples using a transposed convolution (halving the number of output channels), then concatenates the channels with the skip connection feature map from the corresponding layer of the encoder, and then passes through two convolutional layers (with BatchNorm and ReLU). The cloud attention mechanism extracts the entropy En and hyperentropy He channels from the cloud-guided feature map, inputs them to the convolutional mapping layer to generate a single-channel uncertainty heatmap, upsamples it to the size of the current decoded feature map, and multiplies it element-wise with the decoded feature map. The final layer of the decoding network uses a convolutional layer and the Tanh activation function to output a preliminary enhanced image with pixel values ​​normalized to [-1, 1].

[0141] The Cloud Map Inference Module (CAID) uses the same parameter configuration as the secondary cloud feature extraction submodule as the CFE. The refined network is a three-layer convolutional network with 64, 32, and 3 channels per layer, followed by BatchNorm and ReLU (except for the output layer). The confidence map generator uses the following parameters: scale factor of 0.1, Gaussian smoothing kernel size of 5, and standard deviation of 1.

[0142] 3.4 Training Parameters and Loss Function Configuration

[0143] The network employs an end-to-end joint training approach. The optimizer chosen is Adam, with β1 = 0.9 and β2 = 0.999, where β1 and β2 are the first and second order momentum decay coefficients of the Adam optimizer, respectively. The initial learning rate is set to 1 × 10⁻⁶. -4 The cosine annealing scheduling strategy is used to gradually decay the loss to zero. The training batch size is 16, and the total number of training rounds is 200. In the loss function calculation formula (3), α=0.8, and the weight coefficients in formulas (4) and (6) are set as follows: =1、 =0.1、 =0.01. All modules (the parameters of the cloud feature extraction module can be fixed or fine-tuned, the feature fusion module, the cloud-guided augmented coding network, the cloud-guided augmented decoding network, and the cloud map inferrer) update their parameters synchronously under the same training framework.

[0144] 3.5 Setting and Implementing Preset Expected Features

[0145] In this method, "preset expected features" refers to the cloud feature parameters corresponding to an ideal high-quality endoscopic image, including expected value Ex, entropy En, and hyperentropy He. The specific setting method is as follows:

[0146] Expected value Ex: Set to 1.0, indicating that the quality benchmark of the local region of the ideal high-quality image is the highest value (after normalization).

[0147] Entropy En: Set to 0.1, indicating that an ideal high-quality image should have extremely low blur (close to sharp).

[0148] Hyperentropy He: Set to 0.05, indicating that the quality distribution of an ideal high-quality image should have extremely low random discreteness (high stability).

[0149] In the cloud map inferrer, the secondary cloud features (Ex2, En2, He2) extracted again from the initially enhanced image are compared with the aforementioned preset expected features (Ex). gt =1.0, En gt =0.1, He gt=0.05) Calculate the difference pixel by pixel to generate a feature adjustment signal. This signal is then used by a lightweight refinement network (a three-layer convolutional network) to refine the initial enhanced image, making the cloud features of the final enhanced image closer to the preset desired features.

[0150] During the training phase, the difference is constrained by the cloud feature difference loss function (Formula (4)); during the inference phase, the difference is directly used as the input condition for refining the network, and adaptive calibration can be achieved without additional labeled data.

[0151] 4. Examples of Implementation Results

[0152] To verify the effectiveness of the technical solution of the present invention, this embodiment selects three representative image enhancement methods as comparison baselines and performs quantitative and qualitative evaluations on synthetic and real datasets, respectively.

[0153] 4.1 Comparative Experiment Design

[0154] This embodiment selects the following three mainstream image enhancement methods as comparison baselines:

[0155] (1) CLAHE (Contrast Limiting Adaptive Histogram Equalization): As a representative method of traditional image enhancement, the parameters are set to block size 8×8, contrast limiting coefficient is 2.0, and histogram distribution adopts uniformity.

[0156] (2) Standard U-Net: It adopts the same encoder-decoder skeleton as the present invention, but does not include the cloud feature extraction module and cloud attention mechanism. The loss function only uses pixel-level L1 loss. This baseline is used to verify the gain effect brought by the cloud model driving mechanism.

[0157] (3) GAN (Generative Adversarial Network, using CycleGAN architecture): As a representative of deep learning-based image enhancement methods, it is provided with the same paired training data as the present invention in this comparative experiment to ensure fair comparison.

[0158] All comparison methods were trained or had their parameters tuned on the same training dataset, and the test set remained consistent.

[0159] 4.2 Quantitative Evaluation Indicators

[0160] This embodiment uses the following three indicators to quantitatively evaluate the enhancement effect of each method:

[0161] PSNR (Peak Signal-to-Noise Ratio): Measures the fidelity of a signal by calculating the pixel-level error between the enhanced image and the reference image. It is measured in decibels (dB), and a higher value indicates better image quality.

[0162] SSIM (Structural Similarity): Evaluates the similarity of images in terms of brightness, contrast, and structure. Its value ranges from -1 to 1, where 1 indicates that the two images are completely identical and more consistent with the perceptual characteristics of the human visual system.

[0163] NIQE (Natural Image Quality Assessment): A no-reference image quality assessment metric that evaluates image quality by calculating the Mahalanobis distance between the statistical features of the image to be evaluated and the statistical model of the natural image. The lower the score, the closer the image is to a natural and distortion-free state, and the higher the quality.

[0164] 4.3 Experimental Results on Synthetic Datasets

[0165] Quantitative evaluations were performed on a synthetic dataset containing 500 pairs of test samples. The comparison results of each method on PSNR and SSIM metrics are shown in Table 1.

[0166] Table 1. Quantitative comparison of various methods on synthetic datasets

[0167]

[0168] Experimental results show that the method of this invention outperforms the comparative methods on all evaluation metrics. Compared with the standard U-Net, this invention improves PSNR by 0.87 dB and SSIM by approximately 0.04. This gain is mainly attributed to the adaptive enhancement capabilities of the cloud feature extraction module and cloud attention mechanism for uncertain regions, as well as the effective suppression of entropy and hyperentropy by the cloud feature difference loss. Compared with the traditional CLAHE method, this invention improves PSNR by 2.08 dB and SSIM by 0.05, fully demonstrating the significant advantages of deep learning combined with cloud model uncertainty modeling in medical image enhancement tasks.

[0169] 4.4 Experimental Results on Real Datasets

[0170] The generalization ability of each method was further evaluated on the GASTROLAB real endoscopic image dataset. Since real data lacks a true reference, the No-Reference Measure (NIQE) was primarily used for quantitative evaluation. Table 2 shows a comparison of the NIQE scores of each method on the real dataset.

[0171] Table 2 Comparison of NIQE scores for different methods on real datasets

[0172]

[0173] Experimental results show that the method of this invention achieved the lowest NIQE score of 4.23 on real endoscopic images, a reduction of 1.45 compared to the original image and approximately 0.34 compared to the standard U-Net. This result verifies the effectiveness and generalization ability of the method in real clinical scenarios. Furthermore, the pixel-level confidence map generated by this invention can effectively indicate the confidence level of the enhanced region, providing physicians with visual auxiliary reference information on the reliability of the enhancement results, thus aiding clinical decision-making.

[0174] 4.5 Summary of Implementation Results

[0175] The above experiments lead to the following conclusions:

[0176] First, on synthetic datasets, the method of this invention significantly outperforms comparative methods such as CLAHE, standard U-Net, and GAN in both PSNR and SSIM metrics, verifying the effect of cloud feature-based adaptive enhancement mechanism on improving image quality.

[0177] Second, on real endoscopic image datasets, the method of this invention achieved the lowest NIQE score, demonstrating its effectiveness and generalization ability in real clinical scenarios.

[0178] Third, the pixel-level confidence map output by the method of the present invention can provide doctors with intuitive indications of enhanced regional confidence, and has practical value in assisting clinical diagnosis.

[0179] In summary, this embodiment fully verifies the feasibility, effectiveness, and superiority of the technical solution of the present invention.

Claims

1. A cloud feature-based endoscopic image quality assessment and enhancement method, characterized in that The method includes the following steps: Step S1: Obtain the endoscopic image to be enhanced, and perform local adaptive sampling on the endoscopic image to obtain multiple local sampling windows; Step S2: Calculate the local quality index based on the pixel information in each local sampling window, and use the inverse cloud generator to extract cloud features from the local quality index to obtain the cloud feature parameters of the corresponding local area. Step S3: Map the cloud feature parameters corresponding to each local region to cloud feature maps according to their spatial location, and fuse the cloud feature maps with the latent feature maps obtained by encoding the endoscopic images to generate cloud-guided feature maps. Step S4: Input the cloud-guided feature map into the enhancement network for image enhancement. Dynamically generate enhancement parameters based on the entropy En and hyperentropy He corresponding to the local region, and generate uncertainty weights based on the entropy En and hyperentropy He to perform weighted modulation on the decoded features in the enhancement network. Step S5: Extract secondary cloud features again based on the enhanced preliminary image, and generate feature adjustment signals according to the difference between the secondary cloud features and the preset expected features to dynamically calibrate the features of the enhancement network and output the adjusted final enhanced image. Step S6: Generate a pixel-level confidence map based on the entropy En and hyperentropy He of the enhanced image, and output the final enhanced image and the corresponding confidence map.

2. The cloud feature based endoscopic video quality assessment and augmentation method of claim 1, wherein In step S1, local region adaptive sampling uses deformable convolution to generate a sampling offset field, and dynamically adjusts the sampling position of the convolution kernel according to the sampling offset field, so that the sampling window can be dynamically adjusted according to the local structure of the image to adapt to tissue edges, reflective areas and blurred areas in the endoscopic image.

3. The method for endoscopic image quality assessment and enhancement based on cloud features according to claim 1, characterized in that... In step S2, the local quality index is calculated based on pixel gradient information within the local window, and the local quality index includes at least one of the following: Sobel gradient magnitude; Laplacian response value; local brightness contrast; high-frequency texture response value; The cloud feature parameters include at least: expectation Ex, used to characterize the quality benchmark of the local region; entropy En, used to characterize the fuzziness of the local region; and hyperentropy He, used to characterize the random discreteness of the local region.

4. The cloud feature based endoscopic video quality assessment and augmentation method of claim 1, wherein In step S3, the fusion includes: aligning the cloud feature map and the latent feature map in terms of spatial dimensions; performing channel stitching on the aligned feature maps; and generating a fused cloud-guided feature map through convolutional mapping.

5. The cloud feature based endoscopic video quality assessment and augmentation method of claim 1, wherein In step S4, the enhancement parameters include at least one of the following: convolution weights, feature response gains, attention weights, and enhancement intensity coefficients; The uncertainty weights are generated based at least on entropy En and hyperentropy He, and are used to weight and modulate local features in the augmentation network. Specifically, the entropy En and hyperentropy He are input into the convolutional mapping layer to generate an uncertainty heatmap, and the uncertainty heatmap is used to weight the decoded features in the augmentation network element by element.

6. The cloud feature based endoscopic video quality assessment and augmentation method of claim 1, wherein In step S5, the cloud map inferrer performs secondary cloud modeling and fine enhancement on the preliminary enhanced image, and feeds back the enhanced cloud features to the input of the enhancement coding network or decoding network to form a feedback adjustment mechanism for the enhancement result, thereby performing closed-loop optimization of the enhancement network.

7. The cloud feature based endoscopic video quality assessment and augmentation method of claim 1, wherein In step S5, the generation of the feature adjustment signal is optimized based on the cloud feature difference loss function. During the training phase, the total loss function of the system consists of the following three components: (1) Pixel loss: wherein, to enhance the image, to refer to a ground truth image, to be a weight coefficient, to be an LI loss function, to be a structural similarity loss function; (2) Cloud feature difference loss: in, For local block indexes, The total number of blocks, and The enhanced image is the first The entropy and hyperentropy of a local block, and These are the weighting coefficients; (3) Uncertainty minimization regularization term: wherein, is the number of pixels traversed, is the total number of pixels, and are the pixel-level entropy and hyperentropy, respectively; The total loss function is defined as the weighted sum of the three losses mentioned above: wherein, , and are weight coefficients for each loss, respectively.

8. A cloud-based endoscopic video quality assessment and enhancement system implementing the method of any one of claims 1-7, characterized in that The system includes an image acquisition module, a cloud feature extraction module, a cloud-guided enhanced coding network, a feature fusion module, a cloud-guided enhanced decoding network, a cloud map inferrer, and an enhanced output module, wherein: The image acquisition module is used to acquire the endoscopic images to be processed. The cloud feature extraction module is used to extract the expected value Ex, entropy En, and hyperentropy He corresponding to the local region. The cloud-guided augmented coding network is used to extract high-dimensional latent feature maps from raw endoscopic images; The feature fusion module is used to fuse the cloud feature map with the potential feature map to generate a cloud-guided feature map; The cloud-guided enhancement decoding network is used to generate uncertainty weights based on entropy En and hyperentropy He, and to perform weighted modulation on the decoding features to output a preliminary enhanced image; The cloud map inferrer is used to perform secondary cloud modeling and fine enhancement on the preliminary enhanced image, and generate the final enhanced image and the corresponding pixel-level confidence map. The enhanced output module is used to output the final enhanced image and the corresponding pixel-level confidence map.

9. An endoscopic image enhancement device comprising a processor, a memory, and a program stored in the memory and executable on the processor; characterized in that When the program is executed by the processor, it implements the cloud-feature-based endoscopic image quality assessment and enhancement method according to any one of claims 1-7.

10. A computer readable storage medium having stored thereon a computer program, characterized in that When the program is executed by the processor, it implements the cloud-feature-based endoscopic image quality assessment and enhancement method according to any one of claims 1-7.