Enteroscope image enhancement method based on improved generative adversarial network
Through improved generative adversarial network and dual-branch attention mechanism, the multi-scale features of colonoscopic images are extracted, which solves the problem of insufficient image quality in the existing technology under low light and high noise conditions, and achieves high-quality colonoscopic image enhancement and improves the accuracy of clinical diagnosis.
Patent Information
- Application Number
- CN202510204043.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-02-24
AI Technical Summary
Existing colonoscopic image enhancement methods are difficult to effectively process image quality under low light and high noise conditions, especially color distortion and noise amplification will affect image quality and clinical diagnostic effects.
Using an improved generative adversarial network, combined with ResUNet's encoder-decoder structure and dual-branch attention mechanism, multi-scale features of colonoscopic images are extracted, and high-quality colonoscopic enhanced images are generated through feature fusion and feedback mechanisms.
It effectively improves the quality of colonoscopy images under low light and high noise conditions, reduces image blur and detail loss, improves color fidelity and noise control, and improves the accuracy of clinical diagnosis.
Smart Images

Figure CN119941561A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image data processing, and in particular relates to a colonoscopy image enhancement method based on an improved generative adversarial network. Background Art
[0002] Colonoscopy generally refers to colonoscopy. Colonoscopy is currently the most direct and accurate method for diagnosing colorectal diseases. Colonoscopy is widely used in clinical practice for the screening of colorectal diseases and the prevention of early colorectal cancer. Colonoscopic images play a vital role in the diagnosis and treatment of diseases. Although colonoscopy technology has made significant progress in the medical field, colonoscopic image capture is often affected by multiple factors, such as intestinal movement, insufficient lighting, etc., resulting in poor image quality, especially in low light and high noise environments.
[0003] Traditional colonoscopy image enhancement methods improve the visual effect of images by improving the brightness and contrast of images. However, these traditional methods often cause over-enhancement or distortion of images under low-light and high-noise conditions, which in turn affects the image quality. In addition, traditional methods are usually unable to effectively deal with color distortion problems in images, which may cause unnatural color changes in the image during the enhancement process. Therefore, although traditional methods can improve the brightness and contrast of colonoscopy images to a certain extent, these methods are insufficient in low light, noise and color fidelity, making the image enhancement effect unable to meet the high standards of clinical applications.
[0004] The colonoscopic image enhancement method based on deep learning automatically extracts useful features from the image by learning a large amount of colonoscopic image data, and enhances the image using a deep learning model. However, the colonoscopic image enhancement method based on deep learning still has certain limitations in the enhancement effect, and often has problems such as image blur and loss of details. Secondly, the deep learning-based method has insufficient control over the color and noise of the image, which may cause color distortion or noise amplification, thereby affecting the image quality and clinical diagnosis effect. Therefore, in view of the limitations of traditional methods and existing deep learning methods in the field of colonoscopic image enhancement, a colonoscopic image enhancement method combining generative adversarial networks and attention mechanisms is proposed. Summary of the invention
[0005] The present invention provides a colonoscopy image enhancement method based on an improved generative adversarial network, aiming to propose a colonoscopy image enhancement model, wherein the generator is based on the encoder-decoder structure of ResUNet, combined with a dual-branch attention mechanism, to extract multi-scale features of colonoscopy images, and enhances the features of key areas of the image by fusing features of different scales to generate enhanced colonoscopy images; wherein the discriminator is composed of multiple layers of convolution, combined with a specific loss function to evaluate the enhancement effect of the generated image, and guides the generator to generate high-quality colonoscopy enhanced images through a feedback mechanism.
[0006] The present invention aims to propose a colonoscopy image enhancement model and provide a colonoscopy image enhancement method based on an improved generative adversarial network, comprising the following steps.
[0007] S1. Obtain a colonoscopy image dataset, preprocess the colonoscopy image dataset, including format conversion, size adjustment, normalization and denoising operations, and divide the preprocessed colonoscopy image dataset into a training set and a test set.
[0008] S2. Construct a generator through an improved ResUNet network. The improved ResUNet network is based on the encoder-decoder structure of ResUNet, combined with a dual-branch attention mechanism, to extract multi-scale features of colonoscopy images, design a learnable parameter ρ to adaptively combine features of different scales, enhance features of key areas of colonoscopy images, and generate enhanced colonoscopy images through a decoder.
[0009] S3. Construct a discriminator, which consists of multiple layers of convolution, combines a specific loss function to evaluate the enhancement effect of the generated image, and guides the generator to generate high-quality colonoscopy enhanced images through a feedback mechanism.
[0010] S4. Construct a colonoscopy image enhancement model, which consists of an input, a generator, a discriminator and an output. It combines a dual-branch attention mechanism on the basis of a generative adversarial network to achieve the enhancement of colonoscopy images.
[0011] S5. Obtain colonoscopy enhanced images, input the preprocessed colonoscopy images into the colonoscopy image enhancement model, and through the adversarial game between the generator and the discriminator, the colonoscopy image enhancement model continuously optimizes the image enhancement ability of the generator and the discrimination ability of the discriminator, and finally generates colonoscopy enhanced images.
[0012] Preferably, the construction method of the generator in S2 is: The encoder extracts the features of colonoscopy images layer by layer through convolution operations. The dual-branch attention mechanism consists of a channel attention branch and a spatial attention branch. The channel attention branch extracts important features of colonoscopy images through weighted pooling and mixed pooling. The multi-layer perceptron is used to optimize the pooled features, generate channel attention weights, and multiply the channel attention weights with the input features element by element to obtain the enhanced channel features of the colonoscopy images. The spatial attention branch obtains the local and global features of the colonoscopy images through average pooling and maximum pooling, concatenates the pooled features in the channel dimension, generates spatial attention weights through convolution, and multiplies the spatial attention weights with the input features element by element to obtain the enhanced spatial features of the colonoscopy images. A learnable parameter ρ is designed to adaptively combine the channel and spatial features of the colonoscopy images after enhancement, ρ = Sigmoid(W″·ReLU(W′·[μ(F),σ(F)]+b′)+b″), F∈R H×W×C is the input feature map, H, W and C are the height, width and number of channels of the input feature map, W′ and W″ are weight matrices, b′ and b″ are bias parameters, μ(F) is the channel mean of the input feature F, and the global mean of each channel is calculated as σ(F) is the channel standard deviation of the input feature F, and the global standard deviation of each channel is calculated as The channel mean and channel standard deviation of the input feature F are concatenated to obtain the statistical feature vector S, S = [μ(F), σ(F)], and the weight matrix W′ is used to reduce the dimension of the statistical feature vector S to obtain the reduced feature h, h = ReLU(W′·S+b′), and the weight matrix W′ is used to map the reduced feature h to the scalar space to obtain the parameter ρ raw , and use the Sigmoid function to set the parameter ρ raw The value of is limited to the range of [0, 1], and finally a learnable parameter ρ is obtained. The enhanced colonoscopy features output by the dual-branch attention mechanism are input into the decoder, and the resolution of the feature map is gradually restored through convolution and upsampling operations, and finally an enhanced colonoscopy image is generated.
[0013] Preferably, the generator construction method in S2, the generator combines the dual-branch attention mechanism and the adaptive feature fusion technology, can accurately extract and enhance the key feature areas in the colonoscopy image, while reducing the interference of irrelevant noise, and realizes the adaptive optimization of feature fusion by designing the learnable parameter ρ, so that the generator can dynamically adjust the output features according to the distribution of the input data.
[0014] Preferably, in S2, the method for generating the enhanced colonoscopy image is: S21. Input colonoscopy image F in To the encoder, F in ∈R H×W×C, H, W and C are the height, width and number of channels of the colonoscopy image respectively. The encoder consists of three blocks, each of which consists of three convolutions and residual connections. The blocks are connected by maximum pooling. The input F in In the first block of the encoder, three 3×3 convolutional layers are used, each with 64 convolution kernels. Each convolution is followed by a ReLU activation function, where the input F in The number of channels is adjusted to 64 through a 1×1 convolutional layer, and a residual connection is made with the output of the second convolutional layer. The output is used as the input of the third layer, and finally the output feature F1 of the first block of the encoder is obtained, F1∈R H×W×64 , feature F1 is downsampled by the maximum pooling operation to obtain feature F pool1 , Enter F pool1 In the second block of the encoder, the structure of the second block of the encoder is the same as that of the first block of the encoder, except that the number of convolution kernels in each layer is 128, and the output feature F2 of the second block of the encoder is obtained. Feature F2 is downsampled by the maximum pooling operation to obtain feature F pool2 , Enter F pool2 In the third block of the encoder, the structure of the third block of the encoder is the same as that of the first block of the encoder, except that the number of convolution kernels in each layer is 256, and the output feature F3 of the third block of the encoder is obtained. S22, input feature F3 to the dual-branch attention module. The dual-branch attention module consists of a channel attention branch and a spatial attention branch. For the channel attention branch, input feature F3 and obtain feature F through weighted pooling operation. W , w(i, j) is the weight coefficient, that is, the weight assigned to each pixel in the weighted pooling process. The feature F is obtained through the mixed pooling operation. M , F M =α·WeightedPooling(F3)+(1-α)·MaxPooling(F3), α is a trade-off parameter that controls the ratio between weighted pooling and maximum pooling, WeightedPooling is a weighted pooling operation, and MaxPooling is a maximum pooling operation. W and feature F M Input into the multi-layer perceptron for feature optimization, F W =W1(ReLU(W0F W )), F M =W1(ReLU(W0F W), W0 and W1 are the weights in the multilayer perceptron, ReLU is the activation function, and the optimized feature F is adaptively combined using the learnable parameter ρ W and F M Get feature F com ,F com =ρ·F W +(1-ρ)·F M , for feature F com Use the Sigmoid activation function to get the channel attention weight, multiply the channel attention weight by the input feature F3 element by element to get the enhanced feature F C ; For the spatial attention branch, input feature F3 and obtain feature F through average pooling operation A , The feature F is obtained through the maximum pooling operation MP , The feature F is transformed into A and feature F MP Concatenate to get feature F con , the feature F after convolutional layer processing con Use the Sigmoid activation function to get the spatial attention weight, and multiply the spatial attention weight by the input feature F3 element by element to get the enhanced feature F S ; Feature F output by fusion channel attention branch C And the feature F output by the spatial attention branch S , and obtain the output feature F of the dual-branch attention module att , S23, input feature F att In the decoder, the structure of the decoder is similar to that of the encoder and consists of three blocks, each of which consists of three convolutions and residual connections. The blocks are connected by transposed convolution upsampling. The features output by each block are concatenated with the features output by the corresponding block in the encoder. The number of convolution kernels in each layer of the first block is 256, the number of convolution kernels in each layer of the second block is 128, and the number of convolution kernels in each layer of the third block is 64. Finally, through the 1×1 convolution layer, the enhanced colonoscopy image F generated by the generator is obtained. out .
[0015] Preferably, in S2, the method for generating enhanced colonoscopy images adopts a dual-branch fusion mechanism of channel attention and spatial attention, which enhances the accuracy of feature extraction and expression through dynamic weight allocation, improves the ability of the colonoscopy image enhancement model to capture key area features, and enables the colonoscopy image enhancement model to effectively distinguish important areas from irrelevant information.
[0016] Preferably, in S3, the method for generating a high-quality colonoscopy enhanced image is: Input colonoscopy original image F in And the enhanced colonoscopy image F generated by the generator out , F in ∈R H×W×C , F out ∈R H ×W×C , using the sliding window method to in and F out Divide the image into p×p blocks, and transform F in and F out The input is sent to five layers of convolution, where the size of each convolution kernel is p×p, and the number of convolution kernels is 64, 128, 256, 512 and 1 respectively. The first four layers of convolution are followed by a Leaky ReLU activation function, and the last layer of convolution is followed by a Sigmoid activation function. Finally, the score of the image block is output, and the weighted sum of the scores of all image blocks is obtained to obtain the score of the entire colonoscopy image. Assume that the original colonoscopy image F in For x, the colonoscopy enhanced image F generated by the generator out G(x) is G(x), and the binary cross entropy loss function is used to minimize the error of the discriminator. For each input sample, the adversarial loss of the discriminator is L adversarial , D(x, G(x))) is the evaluation score of the discriminator on the enhanced colonoscopy image; the perceptual loss is introduced to calculate the feature space difference between the original colonoscopy image x and the enhanced colonoscopy image G(x), and the perceptual loss of the discriminator is L perceptual , is the feature extracted by multi-layer convolution operation, ||·||2 is the L2 norm, which is used to measure the difference in feature space; the colonoscopy image enhancement model is trained in a self-supervised manner, and the discriminator evaluates the similarity of colonoscopy images before and after enhancement through the structural similarity index. The self-supervised loss of the discriminator is L SSIM , L SSIM =1-SSIM(x, G(x)), SSIM(x, G(x)) calculates the structural similarity between the original colonoscopy image x and the generated colonoscopy enhanced image G(x); the final discriminator loss is the weighted sum of adversarial loss, perceptual loss and self-supervisory loss L D , L D =λ1L adversarial +λ2L perceptual +λ3L SSIM , λ1, λ2 and λ3 are the weights of the loss terms, which are used to control the contribution of different loss terms to the total loss; the discriminator is based on the loss function L DCalculate the gradient update parameters and repeat the process until the loss of the discriminator converges and the generator is able to generate high-quality colonoscopy enhanced images, and finally obtain high-quality colonoscopy enhanced images.
[0017] Preferably, in S3, the method for generating high-quality colonoscopy enhanced images, the discriminator combines sliding window evaluation and adversarial, perceptual, and structural similarity losses to achieve high-quality enhancement of colonoscopy images; the generator can effectively extract key features, improve the clarity and contrast of colonoscopy enhanced images, suppress noise and artifacts, and the discriminator finely evaluates the quality of colonoscopy enhanced images, guiding the generator to generate more realistic colonoscopy enhanced images.
[0018] Preferably, the method for constructing the colonoscopy image enhancement model in S4 is: The preprocessed colonoscopy image training set is input into the generator to generate enhanced colonoscopy images. The generated colonoscopy enhanced images are input into the discriminator together with the real colonoscopy images. The discriminator evaluates the enhancement effect of the generated colonoscopy enhanced images. The generator and the discriminator are continuously optimized through adversarial training. During the training process, the generator and the discriminator continuously adjust their respective weights and parameters, and finally reach a balance point by optimizing the loss function. After the training, the generator can generate high-quality colonoscopy enhanced images based on the input colonoscopy images.
[0019] Preferably, the method for constructing the colonoscopic image enhancement model in S4 continuously improves the clarity, contrast and realism of the generated colonoscopic enhanced images through the collaborative optimization of the generator and the discriminator, so that the generator generates high-quality colonoscopic enhanced images. The colonoscopic image enhancement model can not only effectively suppress noise and artifacts, but also optimize the details of the colonoscopic images, making the enhanced colonoscopic images more suitable for medical diagnosis and analysis, thereby providing doctors with more intuitive and reliable auxiliary diagnosis support and improving the efficiency and accuracy of doctors' diagnosis.
[0020] Compared with the prior art, the present invention has the following technical effects: The technical solution provided by the present invention proposes a colonoscopy image enhancement model, which generates enhanced colonoscopy images in combination with an improved generative adversarial network and an attention mechanism. The generator in the colonoscopy image enhancement model is based on the encoder-decoder structure of ResUNet combined with a dual-branch attention mechanism to extract multi-scale features of colonoscopy images. The features of colonoscopy images can be effectively enhanced by fusing features of different scales, wherein the dual-branch attention module enables the generator to adaptively focus on key areas in colonoscopy images, especially lesion areas; the discriminator in the colonoscopy image enhancement model extracts features through multi-layer convolution operations, evaluates the quality of the generated image in combination with a specific loss function, and introduces a feedback mechanism to guide the generator to generate high-quality colonoscopy enhanced images. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 It is a flow chart of colonoscopy image data processing provided by the present invention.
[0022] Figure 2 It is a structural diagram of the generator provided by the present invention.
[0023] Figure 3 This is a structural diagram of the generator dual-branch attention module provided by the present invention.
[0024] Figure 4 This is a rendering of the colonoscopy image before enhancement provided by the present invention.
[0025] Figure 5 This is a diagram showing the enhanced effect of the colonoscopy image provided by the present invention. DETAILED DESCRIPTION
[0026] The present invention aims to propose a colonoscopy image enhancement method based on an improved generative adversarial network, and proposes a colonoscopy image enhancement model, wherein the generator is based on the encoder-decoder structure of ResUNet, combined with a dual-branch attention mechanism, to extract multi-scale features of colonoscopy images, and enhance the features of key areas of the image by fusing features of different scales to generate enhanced colonoscopy images; wherein the discriminator is composed of multiple layers of convolution, combined with a specific loss function to evaluate the enhancement effect of the generated image, and guides the generator to generate high-quality colonoscopy enhanced images through a feedback mechanism.
[0027] See also Figure 1 As shown, a colonoscopy image enhancement method based on an improved generative adversarial network in an embodiment of the present application.
[0028] S1. Obtain a colonoscopy image dataset, preprocess the colonoscopy image dataset, including format conversion, size adjustment, normalization and denoising operations, and divide the preprocessed colonoscopy image dataset into a training set and a test set.
[0029] Furthermore, in step S1, for obtaining the colonoscopy image dataset, the colonoscopy image dataset in DICOM format is obtained through the electronic colonoscopy system, the Pydicom library and Pillow library in Python are used to convert the colonoscopy images in DICOM format into colonoscopy images in PNG format, the bilinear interpolation method is used to resize the colonoscopy images to a size of 224×224, the mean variance normalization method is used to normalize the colonoscopy images, the bilateral filtering method is used to remove the noise contained in the colonoscopy images, and the preprocessed colonoscopy image dataset is divided into a training set and a test set in a ratio of 8:2.
[0030] S2. Construct a generator through an improved ResUNet network. The improved ResUNet network is based on the encoder-decoder structure of ResUNet, combined with a dual-branch attention mechanism, to extract multi-scale features of colonoscopy images, design a learnable parameter ρ to adaptively combine features of different scales, enhance features of key areas of colonoscopy images, and generate enhanced colonoscopy images through a decoder.
[0031] Furthermore, the construction method of the generator in S2 is: The structure of the generator is as follows Figure 1 As shown in FIG. 1 , the encoder extracts the features of the colonoscopy image layer by layer through convolution operations. The dual-branch attention mechanism consists of a channel attention branch and a spatial attention branch. The channel attention branch extracts the important features of the colonoscopy image through weighted pooling and mixed pooling. The multi-layer perceptron is used to optimize the pooled features, generate channel attention weights, and multiply the channel attention weights with the input features element by element to obtain the enhanced channel features of the colonoscopy image. The spatial attention branch obtains the local and global features of the colonoscopy image through average pooling and maximum pooling, concatenates the pooled features in the channel dimension, generates spatial attention weights through convolution, and multiplies the spatial attention weights with the input features element by element to obtain the enhanced spatial features of the colonoscopy image. A learnable parameter ρ is designed to adaptively combine the channel and spatial features of the colonoscopy image after enhancement, ρ = Sigmoid(W″·ReLU(W′·[μ(F),σ(F)]+b′)+b″), F∈R H×W×C is the input feature map, H, W and C are the height, width and number of channels of the input feature map, W′ and W″ are weight matrices, b′ and b″ are bias parameters, μ(F) is the channel mean of the input feature F, and the global mean of each channel is calculated as σ(F) is the channel standard deviation of the input feature F, and the global standard deviation of each channel is calculated as The channel mean and channel standard deviation of the input feature F are concatenated to obtain the statistical feature vector S, S = [μ(F), σ(F)], and the weight matrix W′ is used to reduce the dimension of the statistical feature vector S to obtain the reduced feature h, h = ReLU(W′·S+b′), and the weight matrix W′ is used to map the reduced feature h to the scalar space to obtain the parameter ρ raw , and use the Sigmoid function to set the parameter ρ raw The value of is limited to the range of [0, 1], and finally a learnable parameter ρ is obtained. The enhanced colonoscopy features output by the dual-branch attention mechanism are input into the decoder, and the resolution of the feature map is gradually restored through convolution and upsampling operations, and finally an enhanced colonoscopy image is generated.
[0032] Furthermore, in S2, the method for generating the enhanced colonoscopy image is: S21. Input colonoscopy image F in To the encoder, F in ∈R 224×224×3 , 224, 224 and 3 are the height, width and number of channels of the colonoscopy image respectively. The encoder consists of three blocks, each of which consists of three convolutions and residual connections. The blocks are connected by maximum pooling. The input F in In the first block of the encoder, three 3×3 convolutional layers are used, each with 64 convolution kernels. Each convolution is followed by a ReLU activation function, where the input F in The number of channels is adjusted to 64 through a 1×1 convolutional layer, and a residual connection is made with the output of the second convolutional layer. The output is used as the input of the third layer, and finally the output feature F1 of the first block of the encoder is obtained, F1∈R 224×224×64 , feature F1 is downsampled by the maximum pooling operation to obtain feature F pool1 , F pool1 ∈R 112×112×64 , enter F pool1 In the second block of the encoder, the structure of the second block of the encoder is the same as that of the first block of the encoder, except that the number of convolution kernels in each layer is 128, and the output feature F2 of the second block of the encoder is obtained, F2∈R 112×112×128 , feature F2 is downsampled by the maximum pooling operation to obtain feature F pool2 , F pool2 ∈R 56×56×128 , enter F pool2 In the third block of the encoder, the structure of the third block of the encoder is the same as that of the first block of the encoder, except that the number of convolution kernels in each layer is 256, and the output feature F3 of the third block of the encoder is obtained, F3∈R 56×56×256 ; S22, input feature F3 to the dual-branch attention module, whose structure is as follows Figure 3 As shown in Figure 1, the dual-branch attention module consists of a channel attention branch and a spatial attention branch. For the channel attention branch, the feature F3 is input and the feature FW is obtained through weighted pooling operation. w(i, j) is the weight coefficient, that is, the weight assigned to each pixel in the weighted pooling process. The feature F is obtained through the mixed pooling operation. M , F M =α·WeightedPooling(F3)+(1-α)·MaxPooling(F3), α is a trade-off parameter that controls the ratio between weighted pooling and maximum pooling, WeightedPooling is a weighted pooling operation, and MaxPooling is a maximum pooling operation. W and feature F M Input into the multi-layer perceptron for feature optimization, F W=W1(ReLU(W0F W )), F M =W1(ReLU(W0F W ), W0 and W1 are the weights in the multilayer perceptron, ReLU is the activation function, and the optimized feature F is adaptively combined using the learnable parameter ρ W and F M Get feature F com , F com =ρ·F W +(1-ρ)·F M , for feature F com Use the Sigmoid activation function to get the channel attention weight, multiply the channel attention weight by the input feature F3 element by element to get the enhanced feature F C ; For the spatial attention branch, input feature F3 and obtain feature F through average pooling operation A , F A ∈R 56×56×1 , the feature F is obtained through the maximum pooling operation MP , F MP ∈R 56×56×1 , the feature F is transformed into A and feature F MP Concatenate to get feature F con , the feature F after convolutional layer processing con Use the Sigmoid activation function to get the spatial attention weight, and multiply the spatial attention weight by the input feature F3 element by element to get the enhanced feature F S ; Feature F output by fusion channel attention branch C And the feature F output by the spatial attention branch S , and obtain the output feature F of the dual-branch attention module att , F att ∈R 56×56×256 ; S23, input feature F att In the decoder, the decoder contains three blocks, each of which consists of three convolutions and residual connections. The blocks are connected by transposed convolution upsampling. The features output by each block are concatenated with the features output by the corresponding block in the encoder. The number of convolution kernels in each layer of the first block is 256, the number of convolution kernels in each layer of the second block is 128, and the number of convolution kernels in each layer of the third block is 64. Finally, through the 1×1 convolution layer, the enhanced colonoscopy image F generated by the generator is obtained. out .
[0033] S3. Construct a discriminator, which consists of multiple layers of convolution, combines a specific loss function to evaluate the enhancement effect of the generated image, and guides the generator to generate high-quality colonoscopy enhanced images through a feedback mechanism.
[0034] Furthermore, in S3, the method for generating a high-quality colonoscopy enhanced image is: Input colonoscopy original image F in And the enhanced colonoscopy image F generated by the generator out , F in ∈R 224×224×3 , F out ∈R 224×224×3 , using the sliding window method to transform the image F in and F out Divide the image into 4×4 blocks and transform the image F in and F out The input is sent to five layers of convolution, where the size of each convolution kernel is 4×4, and the number of convolution kernels is 64, 128, 256, 512 and 1 respectively. The first four convolution layers are followed by a Leaky ReLU activation function, and the last convolution layer is followed by a Sigmoid activation function. Finally, the score of the image block is output, and the weighted sum of the scores of all image blocks is obtained to obtain the score of the entire colonoscopy image. Assume that the original colonoscopy image F in For x, the colonoscopy enhanced image F generated by the generator out g(x) is used to minimize the error of the discriminator using the binary cross entropy loss function. For each input sample, the adversarial loss of the discriminator is L adversarial , D(x, G(x))) is the evaluation score of the discriminator on the enhanced colonoscopy image; the perceptual loss is introduced to calculate the feature space difference between the original colonoscopy image x and the enhanced colonoscopy image G(x), and the perceptual loss of the discriminator is L perceptual , is the feature extracted by multi-layer convolution operation, ||·||2 is the L2 norm, which is used to measure the difference in feature space; the colonoscopy image enhancement model is trained in a self-supervised manner, and the discriminator evaluates the similarity of colonoscopy images before and after enhancement through the structural similarity index. The self-supervised loss of the discriminator is L SSIM , L SSIM =1-SSIM(x, G(x)), SSIM(x, G(x)) calculates the structural similarity between the original colonoscopy image x and the generated colonoscopy enhanced image G (x); the final discriminator loss is the weighted sum of adversarial loss, perceptual loss and self-supervisory loss L D , L D =λ1L adversarial +λ2L perceptual +λ3L SSIM, λ1, λ2 and λ3 are the weights of the loss terms, which are used to control the contribution of different loss terms to the total loss; the discriminator is based on the loss function L D Calculate the gradient update parameters and repeat the process until the loss of the discriminator converges and the generator is able to generate high-quality colonoscopy enhanced images, and finally obtain colonoscopy enhanced images.
[0035] S4. Construct a colonoscopy image enhancement model, which consists of an input, a generator, a discriminator and an output. It combines a dual-branch attention mechanism on the basis of a generative adversarial network to achieve the enhancement of colonoscopy images.
[0036] Furthermore, the method for constructing the colonoscopy image enhancement model in S4 is: The preprocessed colonoscopy image training set is input into the generator to generate enhanced colonoscopy images. The generated colonoscopy enhanced images are input into the discriminator together with the real colonoscopy images. The discriminator evaluates the enhancement effect of the generated colonoscopy enhanced images. The generator and the discriminator are continuously optimized through adversarial training. During the training process, the generator and the discriminator continuously adjust their respective weights and parameters, and finally reach a balance point by optimizing the loss function. After the training, the generator can generate high-quality colonoscopy enhanced images based on the input colonoscopy images.
[0037] Furthermore, in step S4, the colonoscopy image enhancement model was implemented based on the Python language using the Pytorch framework, the optimizer used Adam, the generator in the improved generative adversarial network used a binary cross entropy loss function, and the discriminator used the weighted sum of adversarial loss, perceptual loss, and feature consistency loss as the total loss of the discriminator. The learning rate was set to 0.0002, the batch size was 64, the training rounds were 50, and indicators such as peak signal-to-noise ratio and structural similarity index were used to evaluate the performance of the colonoscopy image enhancement model.
[0038] S5. Obtain colonoscopy enhanced images, input the preprocessed colonoscopy images into the colonoscopy image enhancement model, and through the adversarial game between the generator and the discriminator, the colonoscopy image enhancement model continuously optimizes the image enhancement ability of the generator and the discrimination ability of the discriminator, and finally generates colonoscopy enhanced images.
[0039] Further, in step S5, for obtaining colonoscopy enhanced images, the colonoscopy images preprocessed in step S1 are input into the colonoscopy image enhancement model, such as Figure 4 As shown, Figure 4 The effect of colonoscopy images before enhancement is shown. The enhanced colonoscopy images are generated by the generator, the enhancement effect of the generated images is evaluated by the discriminator, and the generator is guided to generate high-quality colonoscopy enhanced images through the feedback mechanism. Figure 5 As shown, Figure 5The image shows the enhanced effect of the colonoscopy image after being processed by the colonoscopy image enhancement model.
[0040] The above are only preferred embodiments of the present invention. It should be pointed out that a person skilled in the art can make several modifications and improvements without departing from the inventive concept of the present invention, and these all fall within the protection scope of the present invention.
Claims
1. A colonoscopy image enhancement method based on an improved generative adversarial network, characterized in that: The following steps are involved: S1, obtaining a colonoscopy image dataset, preprocessing the colonoscopy image dataset, including format conversion, size adjustment, normalization and denoising operations, and dividing the preprocessed colonoscopy image dataset into a training set and a test set; S2. Construct a generator by using an improved ResUNet network. The improved ResUNet network is based on the encoder-decoder structure of ResUNet, combined with a dual-branch attention mechanism, to extract multi-scale features of colonoscopy images, design a learnable parameter ρ to adaptively combine features of different scales, enhance features of key areas of colonoscopy images, and generate enhanced colonoscopy images through a decoder; S3. Construct a discriminator, which consists of multiple layers of convolution, evaluates the enhancement effect of the generated image in combination with a specific loss function, and guides the generator to generate high-quality colonoscopy enhanced images through a feedback mechanism; S4. Construct a colonoscopy image enhancement model, which consists of an input, a generator, a discriminator, and an output. The colonoscopy image enhancement model is combined with a dual-branch attention mechanism on the basis of a generative adversarial network to enhance the colonoscopy image. S5. Obtain colonoscopy enhanced images, input the preprocessed colonoscopy images into the colonoscopy image enhancement model, and through the adversarial game between the generator and the discriminator, the colonoscopy image enhancement model continuously optimizes the image enhancement ability of the generator and the discrimination ability of the discriminator, and finally generates colonoscopy enhanced images.
2. According to claim 1, a colonoscopy image enhancement method based on an improved generative adversarial network is characterized in that: The construction method of the generator in S2 is: The encoder extracts the features of colonoscopy images layer by layer through convolution operations. The dual-branch attention mechanism consists of a channel attention branch and a spatial attention branch. The channel attention branch extracts important features of colonoscopy images through weighted pooling and mixed pooling. The multi-layer perceptron is used to optimize the pooled features, generate channel attention weights, and multiply the channel attention weights with the input features element by element to obtain the enhanced channel features of the colonoscopy images. The spatial attention branch obtains the local and global features of the colonoscopy images through average pooling and maximum pooling, concatenates the pooled features in the channel dimension, generates spatial attention weights through convolution, and multiplies the spatial attention weights with the input features element by element to obtain the enhanced spatial features of the colonoscopy images. A learnable parameter ρ is designed to adaptively combine the channel and spatial features of the colonoscopy images after enhancement, ρ = Sigmoid(W″·ReLU(W′·[μ(F),σ(F)]+b′)+b″), F∈R H×W×C is the input feature map, H, W and C are the height, width and number of channels of the input feature map, W′ and W″ are weight matrices, b′ and b″ are bias parameters, μ(F) is the channel mean of the input feature F, and the global mean of each channel is calculated as σ(F) is the channel standard deviation of the input feature F, and the global standard deviation of each channel is calculated as The channel mean and channel standard deviation of the input feature F are concatenated to obtain the statistical feature vector S, S = [μ(F), σ(F)], and the weight matrix W′ is used to reduce the dimension of the statistical feature vector S to obtain the reduced feature h, h = ReLU(W′·S+b′), and the weight matrix W′ is used to map the reduced feature h to the scalar space to obtain the parameter ρ raw , and use the Sigmoid function to set the parameter ρ raw The value of is limited to the range of [0,1], and finally a learnable parameter ρ is obtained. The enhanced colonoscopy features output by the dual-branch attention mechanism are input into the decoder, and the resolution of the feature map is gradually restored through convolution and upsampling operations, and finally an enhanced colonoscopy image is generated.
3. The colonoscopy image enhancement method based on improved generative adversarial network according to claim 2, characterized in that: In S2, the method for generating the enhanced colonoscopy image is: S21. Input colonoscopy image F in To the encoder, F in ∈R H×W×C , H, W and C are the height, width and number of channels of the colonoscopy image respectively. The encoder consists of three blocks, each of which consists of three convolutions and residual connections. The blocks are connected by maximum pooling. The input F in In the first block of the encoder, three 3×3 convolutional layers are used, each with 64 convolution kernels. Each convolution is followed by a ReLU activation function, where the input F in The number of channels is adjusted to 64 through a 1×1 convolutional layer, and a residual connection is made with the output of the second convolutional layer. The output is used as the input of the third layer, and finally the output feature F1 of the first block of the encoder is obtained, F1∈R H×W×64 , feature F1 is downsampled by the maximum pooling operation to obtain feature F pool1 , Enter F pool1 In the second block of the encoder, the structure of the second block of the encoder is the same as that of the first block of the encoder, except that the number of convolution kernels in each layer is 128, and the output feature F2 of the second block of the encoder is obtained. Feature F2 is downsampled by the maximum pooling operation to obtain feature F pool2 , Enter F pool2 In the third block of the encoder, the structure of the third block of the encoder is the same as that of the first block of the encoder, except that the number of convolution kernels in each layer is 256, and the output feature F3 of the third block of the encoder is obtained. S22, input feature F3 to the dual-branch attention module. The dual-branch attention module consists of a channel attention branch and a spatial attention branch. For the channel attention branch, input feature F3 and obtain feature F through weighted pooling operation. W , w(i, j) is the weight coefficient, that is, the weight assigned to each pixel in the weighted pooling process. The feature F is obtained through the mixed pooling operation. M , F M =α·WeightedPooling(F3)+(1-α)·MaxPooling(F3), α is a trade-off parameter that controls the ratio between weighted pooling and maximum pooling, WeightedPooling is a weighted pooling operation, and MaxPooling is a maximum pooling operation. W and feature F M Input into the multi-layer perceptron for feature optimization, F W =W1(ReLU(W0F W )), F M =W1(ReLU(W0F W ), W0 and W1 are the weights in the multilayer perceptron, ReLU is the activation function, and the optimized feature F is adaptively combined using the learnable parameter ρ W and F M Get feature F com , F com =ρ·F W +(1-ρ)·F M , for feature F com Use the Sigmoid activation function to get the channel attention weight, multiply the channel attention weight by the input feature F3 element by element to get the enhanced feature F C ; For the spatial attention branch, input feature F3 and obtain feature F through average pooling operation A , The feature F is obtained through the maximum pooling operation MP , The feature F is transformed into A and feature F MP Concatenate to get feature F con , the feature F after convolutional layer processing con Use the Sigmoid activation function to get the spatial attention weight, and multiply the spatial attention weight by the input feature F3 element by element to get the enhanced feature F S ; Feature F output by fusion channel attention branch C And the feature F output by the spatial attention branch S , and obtain the output feature F of the dual-branch attention module att , S23. Input feature F att In the decoder, the decoder contains three blocks, each of which consists of three convolutions and residual connections. The blocks are connected by transposed convolution upsampling. The features output by each block are concatenated with the features output by the corresponding block in the encoder. The number of convolution kernels in each layer of the first block is 256, the number of convolution kernels in each layer of the second block is 128, and the number of convolution kernels in each layer of the third block is 64. Finally, through the 1×1 convolution layer, the enhanced colonoscopy image F generated by the generator is obtained. out .
4. The colonoscopy image enhancement method based on improved generative adversarial network according to claim 3, characterized in that: In S3, the method for generating a high-quality colonoscopy enhanced image is: Input colonoscopy original image F in And the enhanced colonoscopy image F generated by the generator out , F in ∈R H×W×C , F out ∈R H×W×C , using the sliding window method to in and F out Divide the image into p×p blocks, and transform F in and F out The input is sent to five layers of convolution, where the size of each convolution kernel is p×p, and the number of convolution kernels is 64, 128, 256, 512 and 1 respectively. The first four layers of convolution are followed by a Leaky ReLU activation function, and the last layer of convolution is followed by a Sigmoid activation function. Finally, the score of the image block is output, and the weighted sum of the scores of all image blocks is obtained to obtain the score of the entire colonoscopy image. Assume that the original colonoscopy image F in For x, the colonoscopy enhanced image F generated by the generator out G(x) is G(x), and the binary cross entropy loss function is used to minimize the error of the discriminator. For each input sample, the adversarial loss of the discriminator is L adversarial , D(x, G(x))) is the evaluation score of the discriminator on the enhanced colonoscopy image; the perceptual loss is introduced to calculate the feature space difference between the original colonoscopy image x and the enhanced colonoscopy image G(x), and the perceptual loss of the discriminator is L perceptual , is the feature extracted by multi-layer convolution operation, ||·||2 is the L2 norm, which is used to measure the difference in feature space; the colonoscopy image enhancement model is trained in a self-supervised manner, and the discriminator evaluates the similarity of colonoscopy images before and after enhancement through the structural similarity index. The self-supervised loss of the discriminator is L SSIM , L SSIM =1-SSIM(x, G(x)), SSIM(x, G(x)) calculates the structural similarity between the original colonoscopy image x and the generated colonoscopy enhanced image G(x); the final discriminator loss is the weighted sum of adversarial loss, perceptual loss and self-supervisory loss L D , L D =λ1L adversarial +λ2L perceptual +λ3L SSIM , λ1, λ2 and λ3 are the weights of the loss terms, which are used to control the contribution of different loss terms to the total loss; The discriminator is based on the loss function L D Calculate the gradient update parameters and repeat the process until the loss of the discriminator converges and the generator is able to generate high-quality colonoscopy enhanced images, and finally obtain high-quality colonoscopy enhanced images.
5. The colonoscopy image enhancement method based on improved generative adversarial network according to claim 4, characterized in that: The construction method of the colonoscopy image enhancement model in S4 is: The preprocessed colonoscopy image training set is input into the generator to generate enhanced colonoscopy images. The generated colonoscopy enhanced images and real colonoscopy images are input into the discriminator together. The discriminator evaluates the enhancement effect of the generated colonoscopy enhanced images. The generator and discriminator are continuously optimized through adversarial training. During the training process, the generator and discriminator continuously adjust their respective weights and parameters, and finally reach a balance point by optimizing the loss function. After the training, the generator can generate high-quality colonoscopy enhanced images based on the input colonoscopy images.
Citation Information
Patent Citations
Image enhancement method based on residual self-attention and generative adversarial network
CN112561838A
Colorectal cancer T staging method and system based on tumor area CT image
CN115100165A
Colorectal cancer liver metastasis MRI image enhancement method based on generative adversarial network
CN115908333A
Medical image segmentation method and system applying multi-attention mechanism
CN115984296A
Intestinal endoscope image enhancement method based on image fusion
CN116188340A
Cited By
Oral CT image quality enhancement method and device based on generative adversarial network
CN121639495A
Colorectoscope polyp intelligent classification and detection method based on data enhancement and improved YOLO
CN121707930A