Deep learning-based capsule endoscopic image super-resolution reconstruction method

By designing data enhancement and labeling mechanisms for pollution and occlusion scenarios, optimizing the ESRGAN model with a lightweight structure and attention mechanism, and designing a medical structure-aware loss function, the clarity and structural fidelity issues of capsule endoscopy images are solved, achieving efficient lesion visualization.

CN120807289APending Publication Date: 2025-10-17WUXI FUSHENG SMART MEDICAL TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510911828.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing technologies find it difficult to achieve clarity, structural fidelity, pollution adaptability and deployment feasibility in capsule endoscopy images. Traditional methods lack detail generation and medical perception capabilities, while deep learning methods have complex parameters and poor data generalization capabilities.

Method used

A data enhancement and labeling mechanism designed for pollution and occlusion scenarios is introduced, the ESRGAN model is optimized by combining lightweight structure and attention mechanism, and a loss function for medical structure is designed to improve the lesion visualization capability.

Benefits of technology

It effectively improves the model's ability to focus on key medical information, reduces computing resource requirements, and significantly enhances the visualization of lesion areas, making it suitable for deployment on capsule endoscopy equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120807289A_ABST
    Figure CN120807289A_ABST
Patent Text Reader

Abstract

The invention discloses a capsule endoscopic image super-resolution reconstruction method based on deep learning. The method comprises the steps of designing a data enhancement and label mechanism for pollution shielding; optimizing the ESRGAN model in combination with a lightweight structure and an attention mechanism; generating a structure sensing label; and training the model. The invention provides an image super-resolution reconstruction method fused with medical structure perception in order to solve the problems that a capsule endoscopic image is low in resolution, serious in pollution and shielding, difficult in structural detail recognition and the like in an actual clinical environment. According to the method, a pollution simulation data enhancement mechanism and a structural mask label system are constructed, so that the attention capability of the model on medical key information is effectively improved; a super-division network architecture combining lightweight improvement and attention mechanism optimization is adopted, so that the reconstruction quality is ensured, and meanwhile, the computing resource demand is greatly reduced; and a structure weighting loss function and an edge perception loss strategy are designed, so that the visual performance of the focus area is obviously enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to a deep learning-based capsule endoscopy image super-resolution reconstruction method and belongs to the technical field of image processing and medical artificial intelligence. BACKGROUND

[0002] With the development of intelligent medical treatment and non-invasive examination technology, a capsule endoscope has become an important means for detecting diseases in the digestive tract such as the small intestine and the colon. Due to the advantages of non-invasiveness, no need for intubation and convenience for patients, the capsule endoscope is widely used in the screening of digestive tract bleeding, intestinal polyps and Crohn's disease. However, the capsule endoscope device is usually subject to hardware limitations such as power consumption, bandwidth and storage space, and the image resolution collected by the capsule endoscope is low (for example, 480x480 or smaller), which often accompanies the following problems in actual use:

[0003] Image blur: due to the uncontrolled movement of the capsule in the intestinal tract, the image often appears motion blur or inaccurate focusing;

[0004] Contamination obstruction: foam, mucus, chyme and other contaminants often obstruct the lens, resulting in a decrease in image quality;

[0005] Limited resolution: in order to reduce data transmission and power consumption, the image resolution is usually much lower than that of a traditional endoscope, which is difficult to meet the needs of detailed diagnosis;

[0006] Image instability: the image brightness and color change greatly, and the exposure is uneven, which affects the performance of subsequent image analysis algorithms.

[0007] Therefore, improving the clarity and structural details of the capsule endoscopy image is of great significance for improving the lesion recognition rate and enhancing the performance of the computer-aided diagnosis system.

[0008] At present, image super-resolution (Super-Resolution) technology mainly includes two categories: traditional image magnification methods and deep learning-based methods. Although they have been widely used in natural image processing and other fields, direct application to capsule endoscopy images still faces many challenges:

[0009] 1. Limitations of traditional image interpolation and reconstruction methods

[0010] Traditional methods such as nearest neighbor interpolation (Nearest Neighbor), bilinear interpolation (Bilinear) and bicubic interpolation (Bicubic) are based on the spatial relationship between image pixels for estimation. They are simple to operate, fast and have certain practical value in early image magnification applications.

[0011] In addition, there are also some image reconstruction methods based on signal theory and prior modeling, such as sparse representation, dictionary learning, Bayesian reasoning, graphical models, and non-local means. These methods reconstruct high-resolution details through image patch matching and prior learning. However, these methods generally have the following problems in capsule endoscopy images:

[0012] Lack of detail generation capability: Unable to effectively restore small but important structural features in medical images, such as intestinal mucosal texture, microvascular structure, or early lesion edges;

[0013] Oversmoothing and artifacts: Interpolation and reconstruction methods often produce blurry or spurious structures that obscure true tissue morphology;

[0014] Poor robustness to contamination: Foam, mucus, chyme, and other pollutants often block the lens, making traditional methods unable to make intelligent judgments and may even amplify the contamination.

[0015] Lack of contextual understanding: Such methods lack the ability to model image semantics at the hierarchical level and are unable to distinguish key information in complex backgrounds.

[0016] Computational efficiency or deployment limitations: Some traditional reconstruction methods, such as sparse coding, have high computational complexity and are not suitable for resource-constrained devices such as capsules.

[0017] 2. Limitations of Super-resolution Methods Based on Deep Learning

[0018] In recent years, deep learning methods such as convolutional neural networks (CNNs) and generative adversarial networks (GANs) have achieved remarkable results in image super-resolution. Representative methods include SRCNN, VDSR, SRGAN, ESRGAN, and Real-ESRGAN. ESRGAN utilizes a combination of residual data blocks (RRDBs), perceptual loss, and adversarial loss to achieve high-quality texture generation and enhanced realism. Real-ESRGAN further enhances its adaptability to real low-quality images and has been widely used in fields such as photo restoration and video enhancement. However, in the context of capsule endoscopy, existing deep learning super-resolution methods still suffer from the following key deficiencies:

[0019] Lack of medical perception capabilities: Existing models are primarily trained on natural images, making it difficult to perceive the structural importance of medical images. They may overlook small lesions or mistakenly enhance meaningless areas.

[0020] Poor data generalization: Existing model training data lacks representative contaminated and occluded images of intestinal scenes, resulting in unstable results in practical applications;

[0021] Structural information reconstruction is insufficient: most methods optimize the overall perceptual quality or visual style, but do not focus on the anatomical structure integrity and diagnostic elements in medical images;

[0022] Model parameters are complex: the ESRGAN and Real-ESRGAN models have large model parameter quantities and high computing resource occupancy, and are not suitable for deployment in embedded scenarios such as capsule devices and edge computing modules;

[0023] The loss function design is insufficient: conventional loss functions (such as MSE, VGG perceptual loss, and adversarial loss) cannot effectively guide the model to focus on the fine-grained structure of medical images, which may lead to degradation of the diagnostic value of the generated images.

[0024] In summary, neither traditional interpolation nor image reconstruction methods nor current mainstream deep learning super-resolution models can achieve super-resolution reconstruction with clarity, structural fidelity, pollution adaptability, and deployment feasibility in the special medical scenario of capsule endoscopy images. Therefore, there is an urgent need for a super-resolution method that is customized and optimized for the characteristics of intestinal endoscopy images, which can ensure diagnostic value while having efficient deployability. SUMMARY

[0025] The present application provides a deep learning-based capsule endoscopy image super-resolution reconstruction method, which has the following advantages: introducing a data enhancement and label mechanism designed for scenes such as pollution and occlusion; optimizing the ESRGAN model with a lightweight structure and attention mechanism; and proposing a loss function design for medical structures to improve lesion visualization.

[0026] To solve the above technical problems, the technical solution adopted by the present application is as follows:

[0027] A deep learning-based capsule endoscopy image super-resolution reconstruction method, comprising the following steps:

[0028] Step 1: Design a data enhancement and label mechanism for scenes such as pollution and occlusion:

[0029] Specifically, to simulate common pollution problems (such as foam, mucus, liquid reflection, and occlusion) in real endoscopy scenarios, the present application first collects more than 10,000 capsule endoscopy images as training samples q, then manually (using the image labeling tool LabelMe, by professional or labeling personnel browsing capsule endoscopy images frame by frame) or semi-automatically (using image brightness, contrast, color saturation, and texture change detection to extract candidate regions, and then manually selecting or fine-tuning) extracts pollution region image blocks from the training samples q to form a structured pollution texture library.

[0030] The contaminated area image block is taken as a random data enhancement method to simulate contamination of the training image with a probability of 20-30%. The contaminated area image block in the contamination texture library is used to contaminate 20-30% of the capsule endoscopy images in the training sample q. The method for contaminating any one capsule endoscopy image is as follows: 2-3 contaminated area image blocks are randomly sampled from the contamination texture library, the contamination degree is determined according to the contamination proportion, and after random scaling, rotation and cropping, they are superimposed on the randomly selected capsule endoscopy image to simulate real imaging and are added to the training sample q to obtain the data set q1, so as to improve the diversity. After this processing, the model will be exposed to training samples with different contamination degrees and modes, effectively improving its fault tolerance and reconstruction ability for contaminated images.

[0031] In addition to contamination superposition, image occlusion synthesis technology is also introduced, that is, 5-10% of the images in the data set q1 are subjected to occlusion processing. The method for occluding any one image is as follows: a part of the image is randomly cropped and occluded from itself or another image to construct local occlusion (occluding blood vessels, mucosal edges, etc.), local blur (using Gaussian blur to simulate shooting motion blur) and / or non-structural occlusion (using color patches or shape patches to simulate external occlusion), which are added to the data set q1 to obtain the data set q2. In this way, the inference ability of the model under the condition of structural loss and local invisibility is enhanced.

[0032] At the same time, binary masks of the contaminated areas and the occluded areas are generated synchronously for subsequent loss function weighting.

[0033] In summary, this step extracts foam, water stains, light spots and other contaminated areas from the actually collected capsule endoscopy images through manual or semi-automatic methods to construct a contamination texture library. In the training data, contaminated image blocks are randomly selected, scaled, rotated and superimposed on the original image to generate contaminated images, and image occlusion is introduced at the same time, and the corresponding binary mask is generated.

[0034] Step two, optimize the ESRGAN model by combining lightweight structure and attention mechanism: design the following network structure: specifically, to adapt to the multiple requirements of the capsule endoscope device on inference speed, model size and image quality, the present application combines a lightweight module replacement strategy and an attention mechanism guided enhancement strategy on the basis of the classic ESRGAN (Enhanced Super-Resolution Generative Adversarial Network) model to construct an efficient super-resolution model architecture suitable for endoscopic images; although the RRDB (Residual-in-Residual Dense Block) in the traditional ESRGAN has excellent effect, the standard RRDB stacking leads to long inference time, which is not suitable for deployment in embedded systems or portable medical terminals, and cannot dynamically focus on key structures in medical images (such as edges, blood vessels, and lesions), and is prone to learn false textures such as foam and occlusion. The present application proposes the following double structure optimization scheme to solve the above problems: on the basis of preserving the perceptual modeling ability of ESRGAN, a lightweight module is used to replace part of the RRDB structure to reduce the amount of calculation, and the specific replacement is shown in the following table:

[0035]

[0036] Each residual block in ESRGAN contains multiple standard 3x3 convolutions and ReLU activation functions. This structure has strong universality, but the calculation and parameter amount of each layer of convolution is large (especially when the number of channels is large), the present application uses Depthwise Separable Convolution (Depthwise Separable Convolution) to replace the standard convolution, and replaces the activation function with GELU (Gaussian Error Linear Unit) The parameter amount is reduced by about 80%: from O(C 2 ) to O(C);

[0037] Replace the original dense connection structure with a local residual connection plus channel compression strategy, use 1x1 convolution for channel compression to reduce subsequent calculation overhead, which can control the growth of channel number, reduce memory burden, improve forward and backward propagation efficiency, preserve necessary feature channels, and improve reconstruction quality;

[0038] Finally, the Mobile Bottleneck Block in MobileNetV2 is introduced and embedded into the ESRGAN backbone structure to form a mixed super-resolution backbone network with RRDB, which reduces redundant calculations while maintaining expressiveness, which can significantly reduce model parameters and FLOPs. The result obtained in step two is a lightweight super-resolution model structure optimized by pollution perception guidance, denoted as m.

[0039] Step three, structure-aware label generation and loss function design for medical structures: Based on the model structure m constructed in step two, further design the loss function and label mechanism. Specifically, since the intestinal image contains a large amount of structural information regions (such as mucosal texture, folds, blood vessel network, lesion edge, etc.), these regions are particularly important for diagnosis. The present application introduces a structure-aware mechanism to improve the reconstruction capability of structural information regions: a lightweight image segmentation network (such as UNet, DeepLabV3) is used to pre-label the structural information regions on the original image, and the resulting structure-aware mask map is used for weighted control in the subsequent training stage;

[0040] Step four, model training: using the data set q2 obtained in step one and the model structure m obtained in step two, the model is trained by supervised learning, and finally the super-resolution model x for image inference is obtained;

[0041] In the training process, the structure-aware mask map obtained in step three is used to weight the loss function for different regions. For regions with clear structures, the weight coefficient is increased, so that the model pays more attention to the restoration quality of these regions; for contaminated regions or uncertain regions, low weight or mask is given to these regions to prevent the model from learning false textures from contaminated regions and avoid training noise; this mechanism ensures that the model prioritizes processing medical-related structural information when reconstructing, effectively avoiding the problem of "contaminated artifact enhancement", and improving the visualization quality of the lesion area;

[0042] Specifically, the structure-aware mask map obtained in step three is used to assign weights to each pixel in the image: structural information regions (such as blood vessels, folds, and lesions) are given higher weights (such as w = 2) to enhance the restoration accuracy of these regions; contaminated regions or uncertain regions are given lower (w = 0.5) or zero weight to avoid overfitting to contaminated artifacts. Loss function example (weighted L1):

[0043]

[0044] Where: M(x, y) is the binary mask of the contaminated and occluded regions generated in step one (M(x, y) = 0 or M(x, y) = 1), w is the amplification weight of the structural region; I SR , I HR is the super-resolution image and the ground truth image;

[0045] At the same time, the present application also adds an edge-aware function during model training, which calculates the image gradient difference or edge map difference in the structural information region (such as the ulcer edge) to enhance the model's ability to recover the edge shape. The specific formula is as follows:

[0046]

[0047] wherein, is the value of the super-resolution image generated by the model at pixel point i, j; I i,j is the value of the original high-resolution image at pixel point i, j; M i,j is the weight of the structure mask at position i, j; ||·||1 is the L1 loss, i.e., the absolute value difference; λ is the loss scaling factor of the non-structure area, used to reduce the influence of the pollution or invalid area;

[0048] The final training loss is a combination of multiple structure-aware losses:

[0049] L total = λ1L weight-L1 + λ2L perceptual + λ3L edge

[0050] wherein, λ1, λ2, λ3 respectively control the weight of each loss term, which can be adjusted according to the experiment.

[0051] In addition, in step four, in the training stage, multiple loss functions are introduced to jointly optimize the performance of the generator to further improve the image quality: first, pixel loss (Pixel Loss) is used to ensure that the generated image is as close as possible to the original high-resolution image at the pixel level; second, multi-scale perceptual loss (Multi-Scale Perceptual Loss) is used to extract image features at different scales using the VGG network, so that the generated image is consistent with the real image at the semantic level, significantly improving the subjective perception quality;

[0052] In the above step four, in the training stage, to further improve the realism of the image, an adversarial loss (Adversarial Loss) is also used in the training, i.e., using the adversarial mechanism of the discriminator and the generator to push the generated result to approach the real image distribution; finally, through the cooperative training of multiple loss functions, the super-resolution model x can be obtained, which can infer any input low-resolution endoscopic image to generate a high-quality, high-structure restoration super-resolution image, providing a strong image basis for subsequent lesion identification or three-dimensional reconstruction tasks.

[0053] In the above step four, in the super-resolution model x, a structure preservation loss (Structural Consistency Loss) is added, for example, through MS-SSIM or gradient preservation mechanism, to strengthen the clarity of image edges and textures, which is particularly suitable for the restoration of small details in medical images. At the same time, a feature similarity loss (Feature Similarity Loss) is introduced to finely adjust the high-frequency details of the generated image to reduce the blurring phenomenon.

[0054] After the data preparation and model construction in steps one to three, step four focuses on the training process of the model and the generation of super-resolution images. Specifically, using the low-resolution-high-resolution image pair dataset q2 constructed in step one, and the improved network structure m designed in steps two and three, the model is trained through supervised learning, and finally a super-resolution model x that can be used for image inference is obtained.

[0055] In step one, the pollution texture library includes transparent or white foam groups, blurred occlusion areas formed by viscous liquid, reflective or light spot (brightness saturation) areas, foreign object occlusion or background black area (such as the pattern formed when the camera is blocked).

[0056] In step one, during pollution processing, any one capsule endoscopy image is processed at most once. That is, there is no need to repeatedly process the pollution-processed image.

[0057] In step one, during occlusion processing, any one image is processed at most once. That is, there is no need to repeatedly process the occlusion-processed image.

[0058] In step one, during superposition, transparent superposition is used for foam and liquid areas to simulate the occlusion layer in real imaging; hard occlusion (such as texture mapping mode) is used for occlusion objects to simulate the case where the lens is blocked by foreign objects.

[0059] In step three, the structural information area mainly includes: mucosal folds and boundaries, visible blood vessels (red texture), suspected lesions (edge blur or color difference area), and high-texture detail areas.

[0060] The present application proposes an image super-resolution reconstruction method fusing medical structure perception to solve the problems of low resolution, serious pollution and occlusion, and difficult identification of structural details in capsule endoscopy images in actual clinical environment. The method effectively improves the attention ability of the model to key medical information by constructing a pollution simulation data enhancement mechanism and a structure mask label system; combines a lightweight improved super-resolution network architecture with an attention mechanism optimization to significantly reduce the computational resource demand while ensuring the reconstruction quality; designs a structure weighted loss function and an edge perception loss strategy to significantly enhance the visualization performance of the lesion area.

[0061] The present application not only has innovation in algorithm structure and training mechanism, but also fully considers the image complexity and resource constraints in clinical application, providing a feasible technical path for intelligent enhancement of capsule endoscopy images, and has good application prospect and popularization value.

[0062] The technologies not mentioned in the present application refer to the prior art.

[0063] The application is based on a deep learning based capsule endoscopy image super-resolution reconstruction method, constructs a pollution simulation data enhancement mechanism and a structure mask label system, effectively improves the attention ability of the model to key information, combines a lightweight improved and attention mechanism optimized super-resolution network architecture, greatly reduces the calculation resource demand while ensuring the reconstruction quality, and significantly enhances the visualization performance of the lesion area. BRIEF DESCRIPTION OF DRAWINGS

[0064] Figure 1 A flowchart of the deep learning based capsule endoscopy image super-resolution reconstruction method of the application;

[0065] Figure 2 A model training flowchart of the application;

[0066] Figure 3 An improved RRDB structure diagram of the application;

[0067] Figure 4 Image one before and after processing by the super-resolution model x of the application (the left side is a colon part image in the body taken by a capsule endoscope, and the right side is an image improved by using the super-resolution model x of the application);

[0068] Figure 5 Image two before and after processing by the super-resolution model x of the application (the left side is a colon part image in the body taken by a capsule endoscope, and the right side is an image improved by using the super-resolution model x of the application);

[0069] Figure 6 Image three before and after processing by the super-resolution model x of the application (the left side is a colon part image in the body taken by a capsule endoscope, and the right side is an image improved by using the super-resolution model x of the application); DETAILED DESCRIPTION

[0070] In order to better understand the application, the content of the application will be further illustrated below in combination with examples, but the content of the application is not limited to the following examples.

[0071] Example 1

[0072] As shown in the following: Figure 1 A deep learning based capsule endoscopy image super-resolution reconstruction method includes the following steps:

[0073] Step one, design data enhancement and label mechanism for scenes such as pollution and occlusion:

[0074] Specifically, to simulate the common pollution problems (such as foam, mucus, liquid reflection, occlusion, etc.) in real endoscopic scenes, the application first actually collects 11000 capsule endoscopy images as training samples q, and then extracts pollution region image blocks from the training samples q by manual (using the image labeling tool LabelMe, by professional or labeling personnel frame by frame browsing capsule endoscopy images) or semi-automatic (using image brightness, contrast, color saturation, texture change detection "abnormal area", automatically extracting candidate areas, and then manually selecting or fine-tuning) to form a structured pollution texture library;

[0075] The pollution region image blocks are used as a random data enhancement method to simulate pollution on the training images with a probability of 30%. The pollution region image blocks in the pollution texture library are used to process 30% of the capsule endoscopy images in the training samples q. The method for processing any one capsule endoscopy image is as follows: three pollution region image blocks are randomly sampled from the pollution texture library, the pollution degree is determined according to the pollution proportion, and after random scaling, rotation and cropping, they are superimposed on the randomly extracted capsule endoscopy image to simulate real imaging and are added to the training samples q to obtain a data set q1 to improve diversity. After this processing, the model will be exposed to training samples with different pollution degrees and modes, effectively improving its fault tolerance and reconstruction ability for pollution images.

[0076] In addition to pollution superposition, image occlusion synthesis technology is also introduced, that is, 5% of the images in the data set q1 are processed for occlusion. The method for processing any one image is as follows: a part of the region is randomly cropped from the randomly extracted image in the data set q1 and is occluded itself or other images, and is constructed: local occlusion (occluding blood vessels, mucosal edge and other regions), local blur (using Gaussian blur to simulate shooting motion blur) and / or non-structural occlusion (using color patch or shape patch to simulate external occlusion), added to the data set q1 to obtain a data set q2. In this way, the inference ability of the model in the case of structural loss and local invisibility is enhanced.

[0077] At the same time, binary masks of pollution regions and occlusion regions are generated synchronously for subsequent loss function weighting.

[0078] In summary, this step extracts foam, water stains, light spots and other pollution regions from actually collected capsule endoscopy images by manual or semi-automatic methods to construct a pollution texture library. In the training data, pollution blocks are randomly selected, scaled, rotated and superimposed on the original image to generate polluted images, and image occlusion is introduced at the same time, and the corresponding binary mask is generated.

[0079] Step two, optimize the ESRGAN model by combining lightweight structure and attention mechanism: design the following network structure: specifically, to adapt to the multiple requirements of the capsule endoscope device on inference speed, model size and image quality, the present application combines a lightweight module replacement strategy and an attention mechanism guided enhancement strategy on the basis of the classic ESRGAN (Enhanced Super-Resolution Generative Adversarial Network) model to construct an efficient super-resolution model architecture suitable for endoscopic images; although the RRDB (Residual-in-Residual Dense Block) in the traditional ESRGAN has excellent effect, the standard RRDB stacking leads to long inference time, which is not suitable for deployment in embedded systems or portable medical terminals, and cannot dynamically focus on key structures in medical images (such as edges, blood vessels, and lesions), and is prone to learn false textures such as foam and occlusion. The present application proposes the following double structure optimization scheme to solve the above problems: as shown in Figure 3 , on the basis of retaining the perceptual modeling ability of ESRGAN, a lightweight module is used to replace part of the RRDB structure to reduce the amount of calculation, and the specific replacement is shown in the following table:

[0080]

[0081] Each residual block in ESRGAN contains multiple standard 3x3 convolutions and ReLU activation functions. This structure has strong versatility, but the calculation and parameter amount of each layer of convolution is large (especially when the number of channels is large), the present application uses Depthwise Separable Convolution (depthwise separable convolution) to replace the standard convolution, and replaces the activation function with GELU (Gaussian Error Linear Unit) The parameter amount is reduced by about 80%: from O(C 2 ) to O(C);

[0082] Replace the original dense connection structure with a local residual connection plus channel compression strategy, use 1x1 convolution for channel compression to reduce subsequent calculation overhead, which can control the growth of channel number, reduce memory burden, improve forward and backward propagation efficiency, preserve necessary feature channels, and improve reconstruction quality;

[0083] Finally, the Mobile Bottleneck Block in MobileNetV2 is introduced and embedded into the ESRGAN backbone structure to form a mixed super-resolution backbone network with RRDB, which reduces redundant calculations while maintaining expressiveness, which can significantly reduce model parameters and FLOPs. The result obtained in step two is a lightweight super-resolution model structure optimized by pollution perception guidance, denoted as m.

[0084] Step three, structure-aware label generation and medical structure-oriented loss function: based on the model structure m constructed in step two, further design the loss function and label mechanism, specifically, since the intestinal image contains a large amount of structural information area (such as mucosal texture, folds, blood vessel network, lesion edge, etc.), these areas are particularly important for diagnosis. The present application introduces a structure-aware mechanism to improve the reconstruction ability of structural information areas: a lightweight image segmentation network (such as UNet, DeepLabV3) is used to pre-label the structural information area on the original image, and the structure-aware mask map obtained is used for subsequent training stage weighting control;

[0085] Step four, model training: using the data set q2 obtained in step one, and the model structure m obtained in step two, the model is trained by supervised learning, and the model training process is as shown in Figure 2 The final super-resolution model x for image inference is obtained;

[0086] In the training process, the structure-aware mask map obtained in step three is used for loss function weighting processing of different regions, and the weight coefficient is increased for the structural clear area, so that the model pays more attention to the restoration quality of these areas; for the pollution area or uncertain area, give low weight or even mask these areas, prevent the pseudo-texture of the pollution area from being learned by the model, and avoid training noise; this mechanism ensures that the model prioritizes processing medical related structural information when reconstructing, effectively avoids the problem of "pollution artifact enhancement", and improves the visualization quality of the lesion area;

[0087] Specifically, the structure-aware mask map obtained in step three is used to assign weights to each pixel in the image: structural information areas (such as blood vessels, wrinkles, lesions) are given higher weights (w = 2) to enhance the restoration accuracy of these areas; pollution area or uncertain area: give lower (w = 0.5) or zero weight, avoid overfitting pollution artifacts. Loss function example (weighted L1):

[0088]

[0089] Where: M(x, y) is the binary mask of the pollution area and the occluded area generated in step one (M(x, y) = 0 or M(x, y) = 1), w is the amplification weight of the structural area; I SR , I HR is the super-resolution image and the true value image;

[0090] At the same time, the present application also adds an edge-aware function during model training, which calculates the image gradient difference or edge map difference in the structural information area (such as ulcer edge) to enhance the model's ability to recover the edge shape, and the specific formula is as follows:

[0091]

[0092] wherein, is the value of the super-resolution image generated by the model at pixel point i, j; I i,j is the value of the original high-resolution image at pixel point i, j; M i,j is the weight of the structure mask at position i, j; ||·||1 is the L1 loss, i.e., the absolute value difference; λ is the loss scaling factor of the non-structure area, used to reduce the influence of the pollution or invalid area;

[0093] The final training loss is a combination of multiple structure-aware losses:

[0094] L total = λ1L weight-L1 + λ2L perceptual + λ3L edge

[0095] wherein, λ1, λ2, λ3 respectively control the weight of each loss term, which can be adjusted according to the experiment.

[0096] In addition, in step four, in the training stage, multiple loss functions are introduced to jointly optimize the performance of the generator to further improve the image quality: first, pixel loss (Pixel Loss) is used to ensure that the generated image is as close as possible to the original high-resolution image at the pixel level; second, multi-scale perceptual loss (Multi-Scale Perceptual Loss) is used to extract image features at different scales using the VGG network, so that the generated image is consistent with the real image at the semantic level, significantly improving the subjective perception quality;

[0097] In the above step four, in the training stage, to further improve the realism of the image, an adversarial loss (Adversarial Loss) is also used in the training, i.e., using the adversarial mechanism of the discriminator and the generator to push the generated result to approach the real image distribution; finally, through the cooperative training of multiple loss functions, the super-resolution model x can be obtained, which can infer any input low-resolution endoscopic image to generate a high-quality, high-structure restoration super-resolution image, providing a strong image basis for subsequent lesion identification or three-dimensional reconstruction tasks.

[0098] In the above step four, in the super-resolution model x, a structure consistency loss (Structural Consistency Loss) is added, such as through MS-SSIM or gradient preservation mechanism, to strengthen the clarity of image edges and textures, which is particularly suitable for the restoration of small details in medical images. At the same time, a feature similarity loss (Feature Similarity Loss) is introduced to finely adjust the high-frequency details of the generated image to reduce the blurring phenomenon.

[0099] Step four, after the data preparation and model construction in steps one to three, step four focuses on the training process of the model and the generation of super-resolution images. Specifically, using the low-resolution-high-resolution image pair dataset q2 constructed in step one, and the improved network structure m designed in steps two and three, the model is trained through supervised learning, and finally a super-resolution model x that can be used for image inference is obtained.

[0100] In step one, the pollution texture library contains transparent or white foam groups, blurred occlusion areas formed by viscous liquid, reflective or light spot (brightness saturation) areas, foreign object occlusion or background black area (such as the pattern formed when the camera is blocked). During pollution processing, each capsule endoscopy image is processed at most once. That is, there is no need to repeat the pollution processing of the processed image. During occlusion processing, each image is processed at most once. That is, there is no need to repeat the occlusion processing of the processed image. During superposition, transparent superposition is used for foam and liquid areas to simulate the occlusion layer in real imaging; hard occlusion (such as texture mapping mode) is used for occlusion objects to simulate the case where the lens is blocked by foreign objects.

[0101] In step three, the structural information area mainly includes: mucosal folds and boundaries, visible blood vessels (red texture), suspected lesions (edge blur or color difference area), and high-texture detail areas.

[0102] The super-resolution model x obtained above is used to enhance the low-resolution image data (resolution 480x480) taken by the capsule endoscope to obtain the processed high-definition image data. For example, we use PSNR and SSIM as objective evaluation indicators for image enhancement and super-resolution reconstruction. PSNR can intuitively reflect the numerical difference between the reconstructed image and the original image, which is a traditional evaluation standard in the field of image processing. The higher the PSNR, the smaller the image distortion and the better the image quality. Generally, a PSNR value greater than 30 dB indicates that the human eye can hardly perceive the difference. SSIM simulates the sensitivity of the human eye to image structure and texture, and can better reflect the performance of image in terms of detail fidelity and perceptual quality; SSIM ∈ [0, 1], the closer to 1, the more similar the image structure; the two are complementary and can comprehensively evaluate the algorithm performance from the aspects of numerical accuracy and structural perception. Figure 4-6 In the figure, the left side is the image of the colon part in the body taken by the capsule endoscope, and the right side is the image after super-resolution enhancement using the model trained by the above method, with a resolution of 1920x1920 from the original 480x480. It can be seen that the small dirt in the left image is greatly eliminated after image super-resolution. And according to the PNSR and SSIM index display Figure 4 In the figure, PSNR: 43.82 dB, SSIM: 0.9863; Figure 5In the middle, PNSR: 39.88dB, SSIM: 0.9774; Figure 6 In the middle, PNSR: 42.00dB, SSIM: 0.9770), it can be seen that the image quality of the super-resolution of the application is greatly improved, and the enhanced image is also closer to the original image in structure detail and texture fidelity.

[0103] The application aims at the problems of low resolution, serious pollution and shielding, and difficult to identify structure details of capsule endoscopy images in actual clinical environment, and proposes an image super-resolution reconstruction method fusing medical structure perception. The method effectively improves the attention ability of the model to key medical information by constructing a pollution simulation data enhancement mechanism and a structure mask label system; combines a lightweight improved and attention mechanism optimized super-resolution network architecture, greatly reduces the demand for computing resources while ensuring the reconstruction quality; designs a structure weighted loss function and an edge perception loss strategy, which significantly enhances the visualization performance of the lesion area.

Claims

1. A method for super-resolution reconstruction of capsule endoscopy images based on deep learning, characterized by: The steps include: Step 1: Design data enhancement and labeling mechanisms for pollution occlusion: Collect more than 10,000 capsule endoscopy images as training samples q, and then extract contaminated area image blocks from the training samples q to form a structured contamination texture library; 20-30% of the capsule endoscopy images in the training sample q are contaminated using contaminated area image blocks in the contamination texture library. The method for contaminating any capsule endoscopy image is as follows: 2-3 contaminated area image blocks are randomly sampled from the contamination texture library, the degree of contamination is determined according to the contamination ratio, and the contamination is randomly scaled, rotated, and cropped. The contamination is then superimposed on a randomly sampled capsule endoscopy image to simulate real imaging and added to the training sample q to obtain the dataset q1 to improve diversity. Perform occlusion processing on 5-10% of the images in dataset q1. The method for occlusion processing for any image is as follows: randomly crop a part of the image randomly extracted from dataset q1 and occlude itself or other images, construct: local occlusion, local blur and / or unstructured occlusion, and add it to dataset q1 to obtain dataset q2; Generate binary masks of polluted and occluded areas simultaneously for subsequent weighting of loss functions; Step 2: Optimize the ESRGAN model by combining lightweight structure and attention mechanism: To meet the multiple requirements of capsule endoscopy equipment for inference speed, model size, and image quality, we built an efficient super-resolution model architecture for endoscopic images by combining a lightweight module replacement strategy with an attention-guided enhancement strategy based on the classic ESRGAN model. While retaining the perceptual modeling capabilities of ESRGAN, we replaced some RRDB structures with lightweight modules to reduce computational complexity. Specifically, the replacements are: Use Depthwise Separable Convolution to replace the standard convolution, and replace the activation function with GELU to reduce the number of parameters by about 80%: from O(C 2 ) is reduced to O(C); The original dense connection structure is replaced by a local residual connection plus channel compression strategy, and 1×1 convolution is used for channel compression to reduce subsequent computational overhead, thereby controlling the growth of the number of channels, alleviating memory burden, improving forward and backward propagation efficiency, retaining necessary feature channels, and improving reconstruction quality. The Mobile Bottleneck Block in MobileNetV2 is introduced and embedded in the ESRGAN backbone structure and stacked alternately with RRDB to form a structurally hybrid super-resolution backbone network. This reduces redundant calculations while maintaining expressiveness, resulting in a lightweight super-resolution model structure optimized through pollution-awareness guidance, denoted as m. Step 3: Structure-aware label generation and medical structure-oriented loss function: Based on the model structure m constructed in step 2, we further designed the loss function and labeling mechanism, and introduced a structure-aware mechanism to improve the reconstruction ability of structural information areas: we used a lightweight image segmentation network to pre-label the structural information areas on the original image, and obtained a structure-aware mask map; Step 4: Model training: Using the dataset q2 obtained in step 1 and the model structure m obtained in step 2, the model is trained through supervised learning to obtain a super-resolution model x that can be used for image reasoning. During the training process, the structure-aware mask obtained in step 3 is used to perform weighted loss function processing on different regions. The weight coefficient is increased for regions with clear structures, so that the model pays more attention to the restoration quality of these regions. For polluted or uncertain regions, low weights are given or even masked out to prevent the pseudo-texture in polluted areas from being learned by the model and avoid training noise. Specifically, the structure-aware mask obtained in step 3 is used to assign weights to each pixel in the image: regions with structural information are given higher weights to enhance the restoration accuracy of these regions. Contaminated or uncertain areas: assign low or zero weights to avoid overfitting contamination artifacts, and the loss function is weighted L1: Where: M(x,y) is the binary mask of the contaminated area and the occluded area generated in step 1 (M(x,y)=0 or M(x,y)=1), w is the magnification weight of the structure area; I SR , I HR is the super-resolution graph and the true value graph; At the same time, an edge perception function is added during model training to calculate the image gradient difference or edge map difference in the structural information area to enhance the model's ability to recover edge shapes. The specific formula is as follows: in, The value of the super-resolution image generated by the model at pixel i, j; I i,j is the value of the original high-resolution image at pixel i, j; M i,j is the weight of the structure mask at position i, j; ||·||1 is the L1 loss, i.e., the absolute value difference; λ is the loss scaling factor of the non-structured area, which is used to reduce the impact of contamination or invalid areas; The final training loss is a combination of multiple structure-aware losses: L total =λ1L weight-L1 +λ2L perceptual +λ3L edge Among them, λ1, λ2, and λ3 control the weights of each loss term respectively and can be tuned according to experiments.

2. The method for super-resolution reconstruction of capsule endoscopy images based on deep learning according to claim 1, characterized in that: In step 4, during the training phase, multiple loss functions are introduced to jointly optimize the performance of the generator to further improve image quality: first, pixel loss is used to ensure that the generated image is as close as possible to the original high-resolution image at the pixel level; second, combined with multi-scale perceptual loss, the VGG network is used to extract image features at different scales, so that the generated image is consistent with the real image at the semantic level.

3. The method for super-resolution reconstruction of capsule endoscopy images based on deep learning according to claim 2, characterized in that: In step 4, during the training phase, adversarial loss is also used in training, that is, the adversarial mechanism between the discriminator and the generator is utilized to push the generated results closer to the real image distribution; finally, through the collaborative training of multiple loss functions, the obtained super-resolution model x can infer any input low-resolution endoscopic image and generate high-quality, high-structural restoration super-resolution images, providing a strong image foundation for subsequent tasks such as lesion recognition or three-dimensional reconstruction.

4. The method for super-resolution reconstruction of capsule endoscopy images based on deep learning according to claim 3, characterized in that: In step 4, a structure-preserving loss is added to the super-resolution model x to enhance the clarity of image edges and textures through the MS-SSIM or gradient-preserving mechanism. At the same time, a feature similarity loss is introduced to fine-tune the high-frequency details of the generated image to reduce blurring.

5. The method for super-resolution reconstruction of capsule endoscopy images based on deep learning according to any one of claims 1 to 4, characterized in that: In step 1, the pollution texture library contains transparent or white foam balls, fuzzy occlusion areas formed by viscous liquids, reflective or light spot areas, foreign body occlusions or background black areas.

6. The method for super-resolution reconstruction of capsule endoscopy images based on deep learning according to any one of claims 1 to 4, characterized in that: In step 1, during contamination processing, any capsule endoscopy image can be processed at most once.

7. The method for super-resolution reconstruction of capsule endoscopy images based on deep learning according to any one of claims 1 to 4, characterized in that: In step 1, during occlusion processing, any image is processed at most once.

8. The method for super-resolution reconstruction of capsule endoscopy images based on deep learning according to any one of claims 1 to 4, characterized in that: In step 1, when superimposing, transparent superposition is used for the foam and liquid areas to simulate the occlusion layer in real imaging; hard occlusion is used for the occlusion to simulate the situation where the lens is blocked by foreign objects.

9. The method for super-resolution reconstruction of capsule endoscopy images based on deep learning according to any one of claims 1 to 4, characterized in that: In step three, the structural information areas include: mucosal folds and boundaries, visible blood vessels, suspected lesions, and high texture detail areas.

Citation Information

Cited By

  • Automatic artifact detection and restoration method and system based on endoscope image

    CN121392066A

  • Ground penetrating radar image sharpening method based on super-resolution reconstruction

    CN121639516A

  • Unblocking method, device and equipment for security check image and storage medium

    CN122243802A