A method based on a fuzzy self-guided structure protection generative adversarial network

By adopting a fuzzy self-guided structural protection generation adversarial network (FS-GAN) based on fuzzy self-guided structural protection module and light distribution correction module, combined with fuzzy discriminator and cyclic consistency loss, the shortcomings of traditional Chinese medicine image enhancement are solved, and the medical image enhancement effect is achieved with high quality and good fidelity.

CN115796264BActive Publication Date: 2025-06-10GUANGZHOU UNIVERSITY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211455273.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-21
Publication Date
2025-06-10
Estimated Expiration
2042-11-21

AI Technical Summary

Technical Problem

Existing medical image enhancement methods have problems with high-quality regions overenhancement and noise amplification, and it is difficult to effectively process non-paired medical images.

Method used

The fuzzy self-guided structure protection generation adversarial network (FS-GAN) is adopted, and the neural fiber structure information is captured and the illumination distribution correction module (IDCM) is used to capture and correct the illumination distribution through the self-guided structure protection module (SSRM) and the illumination distribution correction module (IDCM). The fuzzy discriminator is used to distinguish input images from enhanced images in the real domain and the fuzzy domain, and the loss function is optimized using cyclic consistency loss and texture fidelity loss.

Benefits of technology

It improves the quality of medical images, avoids excessive enhancement and noise amplification of high-quality areas, and is suitable for non-paired medical image enhancement, significantly improving the fidelity of the image content information and texture details.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115796264B_ABST
    Figure CN115796264B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of high-quality medical images, and discloses a method based on a fuzzy self-guided structure-preserving generative adversarial network, comprising the following steps: exploring and retaining the structural information of an input image based on a self-guided structure-preserving module; correcting the illuminance distribution of an enhanced image based on an illumination distribution correction module; distinguishing the input image and the enhanced image in the real domain and the fuzzy domain based on a fuzzy discriminator; and using a cyclic consistency loss and a texture fidelity loss to represent a loss function. The method based on the fuzzy self-guided structure-preserving generative adversarial network of the present invention uses a fuzzy discriminator to distinguish the differences between the input image and the enhanced image in the real domain and the fuzzy domain, thereby improving the recognition ability and optimizing the enhancement effect. The present invention uses a cyclic consistency loss and a texture fidelity loss, enabling the network to be applicable to the case of unpaired training, so that the content information and texture details of the image will not migrate and be lost during the enhancement process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of high-quality medical images, and specifically to a method based on a fuzzy self-guided structure-preserving generative adversarial network. Background Art

[0002] Medical images contain a large amount of information related to biological or anatomical tissues and are an important basis for clinical diagnosis and treatment. However, medical images collected in actual scenarios are usually of low quality. In recent years, medical image enhancement has attracted much attention due to its high practical value and wide application scenarios. The purpose of medical image enhancement is to unify image brightness, improve image structure, restore texture details of the image, and more clearly display specific pathological content expressed by the image. Machine learning and computer vision technologies have become research hotspots for improving the quality of medical images.

[0003] In existing image enhancement methods, traditional techniques for improving image quality include: histogram equalization, adaptive filters using wavelet transforms, gray correction, etc. Traditional image enhancement methods have some obvious defects, such as over-enhancement of high-quality regions and amplification of noise. Taking histogram equalization as an example, the cumulative distribution function is used as the mapping function. This principle is simple and fast, and the image contrast can be intuitively improved, but the detail preservation rate is low, and noise will inevitably be enhanced. In addition, the method based on histogram equalization ignores the different requirements of different parts of the image for enhancement and has limitations in medical image processing. In recent years, deep learning technologies have greatly improved the performance of image enhancement methods. Among them, the generative adversarial network (GAN), as a classic unsupervised network, has been widely applied to image enhancement tasks. GANs for image enhancement usually use paired images for training. However, paired medical images are difficult to obtain, and the enhancement effect of the model is not ideal. Therefore, the present invention proposes a fuzzy self-guided structure-preserving generative adversarial network system. Summary of the Invention

[0004] (1) Technical Problems to be Solved

[0005] Aiming at the deficiencies of the prior art, the present invention provides a method based on a fuzzy self-guided structure-preserving generative adversarial network. By a uniquely designed fuzzy discriminator to distinguish the differences between the input image and the enhanced image in the real domain and the fuzzy domain, a self-guided structure-preserving module (SSRM) and an illumination distribution correction module (IDCM) are adopted to capture the structural information of nerve fibers and correct the illumination distribution of the image to solve the above problems.

[0006] (2) Technical Solutions

[0007] To achieve the above object, the present invention provides the following technical solutions:

[0008] A method based on a fuzzy self-guided structure protection generative adversarial network, comprising the following steps:

[0009] The first step: Use the self-guided structure protection module in the fuzzy self-guided structure protection generative adversarial network system to explore and retain the structural information of the input image;

[0010] The second step: Correct the illuminance distribution of the enhanced image based on the illumination distribution correction module;

[0011] The third step: Distinguish the input image and the enhanced image in the fuzzy domain based on the fuzzy discriminator;

[0012] The fourth step: Use the cycle consistency loss and the texture fidelity loss to represent the loss function.

[0013] Preferably, the fuzzy self-guided structure protection generative adversarial network system includes two generators and a fuzzy discriminator respectively connected to the two generators. The generator is composed of an image enhancement module, a self-guided structure protection module, and an illumination distribution correction module. The self-guided structure protection module is responsible for exploring the nerve fiber structure in the medical image and adding the feature map to the backbone network proportionally during decoding. The illumination distribution correction module corrects the illumination distribution of the output image of the backbone network by using the prior attention map of the low-quality input image, and two generators with the same architecture are used. The output image of one generator will be input into the other generator to achieve end-to-end self-learning.

[0014] Preferably, the image enhancement module uses U-Net as the backbone network. U-Net has eight encoding layers and eight decoding layers. The encoding layer is composed of a residual convolution module and a max pooling layer, and the decoding layer is composed of an upsampling module and a residual convolution module. The residual convolution modules of the encoding layer and the decoding layer are the same, including two stacked convolutional layers + batch normalization layer + LeakyReLU, where the input and output are connected in the form of residuals. The upsampling module is a serial structure of a bilinear interpolation upsampling layer + 3×3 convolutional layer + batch normalization layer + LeakyReLU + dropout layer.

[0015] Preferably, the image structure information extracted by the self-guided structure protection module will be serially connected with the encoding features of U-Net and the decoding features of the previous layer, and injected into the decoding path of the image enhancement module.

[0016] Preferably, the self-guided structure protection module is used to explore and retain the structural information of the low-quality input image and guide the image enhancement module and the illumination distribution correction module to reconstruct the medical image. The self-guided structure protection module is composed of two encoder-decoder networks with the same structure, and the encoder is initialized by the pre-trained Res-Net-34.

[0017] The self-guided structure protection module compresses the original input and the IEM output into the deep semantic space, and then restricts the distance between them to make the latent feature representations of the two stages close. The encoded features of the two networks are respectively denoted as E. 1 and E. 2 , and the self-guided loss function can be expressed as:

[0018]

[0019] where LQ and HQ respectively represent the low-quality image set and the high-quality image set.

[0020] Preferably, the specific steps of the fuzzy operation in the fuzzy discriminator are as follows:

[0021] (1) Two input images m 1 and m 2 respectively pass through a 1×1 convolutional layer to obtain feature maps of the same size and dimension;

[0022] (2) The two feature maps are non-linearly activated by the sigmoid function;

[0023] (3) The two activated feature maps perform a fuzzy AND operation, that is, take the smaller value of the same dimension, and finally obtain the fuzzy feature map f.

[0024] The above steps can be expressed as:

[0025] f = Fuzzy And[σ(c(m 1 )), σ(c(m 2 ))];

[0026] where c and σ respectively represent the convolutional layer and the sigmoid function;

[0027] Let G LQ→HQ and G HQ→LQ respectively represent the high-quality image generator and the low-quality image generator, FD HQ and FD LQ respectively represent the high-quality fuzzy discriminator and the low-quality fuzzy discriminator. The loss functions of the fuzzy discriminator and the generator can be expressed as: L adv1 = E x∈LQ [logFD LQ (x)] + E y∈HQ [log(1 - FD LQ (G HQ→LQ (y)))] + E y∈HQ [logFD HQ (y)] + E x∈LQ [log(1 - FD HQ (G LQ→HQ (x)))];

[0028] The specific structure of the image patch discriminator includes 3 4×4 convolutional layers with a stride of 2 and 2 4×3 convolutional layers with a stride of 1. It distinguishes between real and fake based on image patches, and the image patch size is set to 50×50. Similarly, the adversarial loss between the image patch discriminator and the generator can be expressed as:

[0029] L adv2 =E x∈LQ [logD LQ (x)] + E y∈HQ [log(1 - D LQ (G HQ→LQ (y)))] + E y∈HQ [logD HQ (y)] + E x∈LQ [log(1 - D HQ (G LQ→HQ (x)))];

[0030] where D HQ and D LQ represent the high-quality discriminator and the low-quality discriminator respectively.

[0031] Finally, the adversarial loss of FS-GAN is expressed as:

[0032] L adv = L adv1 + L adv2 ;

[0033] Preferably, the loss function includes the cycle consistency loss L cyc and the identity mapping loss L idt . The main purpose of the cycle consistency loss is to achieve the mutual conversion between the low-quality domain and the high-quality domain. The low-quality image x ∈ LQ is input into the high-quality generator G LQ→HQ to obtain image enhancement, and then the enhanced image is input into the low-quality generator G HQ→LQ to restore the image as much as possible. Therefore, x ≈ G HQ→LQ (G LQ→HQ (x)), and for backward cycle consistency, y ≈ G LQ→HQ (G HQ→LQ (y)), where y ∈ HQ represents the high-quality image. The cycle consistency loss can be expressed as:

[0034] L cyc = E x∈LQ [||G HQ→LQ (G LQ→HQ (x)) - x|| 1 + E y∈HQ [||G LQ→HQ (G HQ→LQ||(y)) - y|| 1 ;

[0035] The identity mapping loss can be expressed as:

[0036] L idt = E x∈LQ [||G HQ→LQ (x) - x|| 1 + E y∈HQ [||G LQ→HQ (y) - y|| 1 ;

[0037] The texture fidelity loss L TF , divides the image into multiple image patches for fine processing and comparison, and is specifically expressed as:

[0038]

[0039] where y i and x i respectively represent the i-th local image patch of the high-quality image y and the low-quality image x, G HQ→LQ (y) i and G LQ→HQ (x) i respectively represent the i-th local image patch of the two generated images, and represent the covariance matrix between the i-th image patch of the original image and the i-th image patch of the corresponding generated image, and respectively represent the standard deviation matrices corresponding to y i , G HQ→LQ (y) i , x i and G LQ→HQ (x) i c is a small constant used to avoid numerical instability, p is the number of image patches, and it should be noted that SSRM can protect the global structure information, and the texture fidelity loss L TF can enhance the texture recovery in the local area;

[0040] Combined with the fuzzy discriminator, the loss function of FS-GAN can be expressed as:

[0041] L FS-GAN = L adv + αL cyc + βL idt + γL TF + ηL enc ;

[0042] α, β, γ, and η are the weights of the cyclic consistency loss, identity mapping loss, texture fidelity loss, and self-guidance loss, respectively.

[0043] (III) Beneficial Effects

[0044] Compared with the prior art, the method of the fuzzy self-guided structure protection generative adversarial network (FS-GAN) provided by the present invention has the following beneficial effects:

[0045] 1. The method of the fuzzy self-guided structure protection generative adversarial network develops a fuzzy discriminator to distinguish the differences between the input image and the enhanced image in the real domain and the fuzzy domain, thereby improving the recognition ability and optimizing the enhancement effect. Using the cyclic consistency loss and texture fidelity loss, FS-GAN is applicable to the case of unpaired training, so that the content information and texture details of the image will not migrate and be lost during the enhancement process.

[0046] 2. The method of the fuzzy self-guided structure protection generative adversarial network embeds a self-guided structure protection module (SSRM) in the generator, significantly improving the network's feature learning and perception ability and avoiding the situation where the nerve fiber structure is homogenized into the background. At the same time, an illumination distribution correction module (IDCM) is designed, which can correct the illumination distribution of the enhanced image using the prior attention map.

[0047] 3. The method of the fuzzy self-guided structure protection generative adversarial network compares FS-GAN with advanced methods through a large number of comparative experiments, including visual observation, evaluation metrics, and downstream task performance. The experimental results show that FS-GAN has the best enhancement performance and can adapt to unpaired medical image enhancement. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 It is a schematic diagram of the module composition structure and process of the adversarial network method in the embodiment of the present invention;

[0049] Figure 2 It is a schematic diagram of the structure and process of the fuzzy discriminator of the adversarial network method in the embodiment of the present invention;

[0050] Figure 3 It is a schematic diagram for comparing the texture structures of the enhanced images of different methods of the adversarial network method in the embodiment of the present invention;

[0051] Figure 4 It is a schematic diagram for comparing the illumination distributions of the enhanced images of different methods of the adversarial network method in the embodiment of the present invention;

[0052] Figure 5 It is a schematic diagram for comparing the segmentation results of the enhanced images of different methods of the adversarial network method in the embodiment of the present invention;

[0053] Figure 6 Schematic diagram for the effectiveness analysis of SSRM and IDCM of the adversarial network method in the embodiments of the present invention;

[0054] Figure 7 Schematic diagram for the effectiveness analysis of the fuzzy discriminator of the adversarial network method in the embodiments of the present invention. Detailed implementation manners

[0055] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0056] Embodiment

[0057] The method for a fuzzy self-guided structure protection generative adversarial network provided by the embodiments of the present invention provides high-quality medical images as an important basis for doctors for clinical diagnosis and treatment, and includes the following steps:

[0058] The first step: Explore and retain the structural information of the input image based on the self-guided structure protection module;

[0059] The second step: Correct the illuminance distribution of the enhanced image based on the illumination distribution correction module;

[0060] The third step: Distinguish the input image and the enhanced image in the fuzzy domain based on the fuzzy discriminator;

[0061] The fourth step: Use the cycle consistency loss and the texture fidelity loss to represent the loss function.

[0062] Specifically, please refer to Figure 1-7 , in order to obtain high-quality medical images with clear texture structures and balanced contrast, the method for a fuzzy self-guided structure protection generative adversarial network provided by the present invention constructs a fuzzy self-guided structure protection generative adversarial network system (FS-GAN). The architecture of FS-GAN is as Figure 1 shown. It embeds SSRM and IDCM to better capture the fiber structure details and the uniform illumination distribution of the image. The self-guidance mechanism and cycle consistency enable it to be used for unpaired training. In addition, FS-GAN combines fuzzy theory to distinguish low-quality input images and high-quality enhanced images in the real domain and the fuzzy domain respectively, thereby improving the quality of the enhanced images. The important components of FS-GAN are described in detail below.

[0063] Figure 1Schematic diagram of the generator structure of the method based on fuzzy self-guided structure protection generative adversarial network provided by the embodiments of the present invention. The generator consists of three modules, namely the Image Enhancement Module (IEM), the Self-guided Structure Protection Module (SSRM), and the Illumination Distribution Correction Module (IDCM). The SSRM is responsible for exploring the nerve fiber structure in medical images and adding the feature maps to the backbone network proportionally during decoding. The IDCM corrects the illumination distribution of the output image of the backbone network by using the prior attention map of the low-quality input image. As shown above for the cycle consistency, two generators with the same architecture are used. The output image of one generator will be input into the other generator to achieve end-to-end self-learning.

[0064] Image Enhancement Module

[0065] U-Net can fully extract multi-layer features of images and performs excellently in the fields of medical image segmentation and reconstruction. Therefore, the Image Enhancement Module (IEM) of the present invention uses U-Net as the backbone network. Specifically, U-Net has eight encoding layers and eight decoding layers. The encoding layers consist of residual convolution modules and max pooling layers, and the decoding layers consist of upsampling modules and residual convolution modules. The specific structure is shown in Table 1. As can be seen from Table 1, the residual convolution modules of the encoding layer and the decoding layer are the same, including two stacked convolution layers + batch normalization layer + LeakyReLU. Among them, the input and output are connected in the form of residuals. The upsampling module is a serial structure of a bilinear interpolation upsampling layer + 3×3 convolution layer + batch normalization layer + LeakyReLU + dropout layer.

[0066] Structure.

[0067] The image structure information extracted by the Self-guided Structure Protection Module (SSRM) will be serially connected with the encoding features of U-Net and the decoding features of the previous layer, and injected into the decoding path of the IEM to guide the IEM to better learn and restore texture details.

[0068] Table 1: Specific structure of the backbone network: It includes 8 encoding layers and 8 decoding layers, where ic represents the input dimension and oc represents the output dimension.

[0069]

[0070] Self-guided Structure Protection Module

[0071] When using GAN to enhance images, it is usually desired that the model generates images with different backgrounds, while neglecting the protection of the global structure and local details of the input images. However, as the basis for doctors' diagnosis, the specific structural details expressed by medical images should be very important. The loss of some nerve fiber structures will lead to diagnostic errors and distorted pathological descriptions. Some current single-channel methods do not pay attention to distinguishing the background and foreground during the enhancement process, which results in some nerve fibers being homogenized into the background.

[0072] To solve this problem, the present invention designs a self-guided structure protection module (SSRM). It can be used to explore and retain the structural information of low-quality input images and guide the IEM and illumination distribution correction module (IDCM) to reconstruct medical images. Figure 1 As can be seen, the SSRM consists of two encoder-decoder networks with the same structure. The encoder is initialized by the pre-trained Res-Net-34. The information extracted by each decoder operator of the first encoder-decoder network is injected into the U-Net through skip connections to guide the IEM to restore texture details. The other encoder-decoder network is responsible for correcting the illumination distribution of the IEM output. Inspired by multi-stage refinement work, the present invention designs the two encoder-decoder networks for sequential refinement and mutual guidance. First, the SSRM compresses the original input and the IEM output into the deep semantic space, and then restricts the distance between them to make the latent feature representations of the two stages close. The encoded features of the two networks are respectively denoted as E· 1 and E· 2 , and the self-guided loss function can be expressed as:

[0073]

[0074] where LQ and HQ represent the low-quality image set and the high-quality image set respectively. The self-guided design can well solve the unpaired tasks, so that the details of the input images will not be mis-transferred and the original structural information will be retained. In addition, the front and rear modules learn and guide each other on an end-to-end basis, which is beneficial to image reconstruction.

[0075] Illumination distribution correction module

[0076] For low-quality medical images with unbalanced illumination distribution, the dark areas can be enhanced to visually match other areas with normal brightness. Inspired by regularization research, the present invention normalizes the illumination channel of the original input image to [0,1] to obtain I, and then takes 1-I as the attention map A, providing prior knowledge for IDCM to correct the illumination distribution. The input of IDCM is the residual between the IEM output and A, passing through four 3×3 convolutional blocks. It is worth mentioning that although this helps improve the uniformity of illumination, the contrast will decrease during this process, and some important texture details are easily lost. The SSRM provided by the present invention can suppress IDCM to ensure the integrity of the nerve fiber structure.

[0077] Fuzzy discriminator

[0078] The working principle of GAN is to optimize the generator and discriminator in confrontation. The better the performance of the discriminator, the higher-quality images the generator can produce. To further improve the discrimination ability of the model, the present invention designs a fuzzy discriminator. The structure of the fuzzy discriminator is as Figure 2 shown. Its working principle is that after using four convolutional blocks to extract features, the input image is projected into the fuzzy domain through fuzzy operations, thereby reducing the difference between the enhanced image and the real high-quality image and further increasing the difficulty of discrimination.

[0079] The specific steps of the fuzzy operation are as follows:

[0080] (1) Two input images m 1 and m 2 respectively pass through a 1×1 convolutional layer to obtain feature maps of the same size and dimensions;

[0081] (2) The two feature maps are non-linearly activated by the sigmoid function;

[0082] (3) The two activated feature maps are subjected to a fuzzy AND operation, that is, taking the smaller value of the same dimension, and finally obtaining the fuzzy feature map f.

[0083] The above steps can be expressed as:

[0084] f = Fuzzy And[σ(c(m 1 )), σ(c(m 2 ))] (2)

[0085] where c and σ represent the convolutional layer and the sigmoid function respectively.

[0086] Let G LQ→HQ and G HQ→LQ respectively represent the high-quality image generator and the low-quality image generator, FD HQ and FD LQrespectively represent the high-quality and low-quality blur discriminators. The loss functions of the blur discriminator and the generator can be expressed as:

[0087] L adv1 = E x∈LQ [logFD LQ (x)] + E y∈ H Q [log(1 - FD LQ (G HQ→LQ (y)))] +

[0088] E y∈HQ [logFD HQ (y)] + E x∈LQ [log(1 - FD HQ (G LQ→HQ (x)))] (3)

[0089] The specific structure of the patch discriminator includes 3 4×4 convolutional layers with a stride of 2 and 2 4×3 convolutional layers with a stride of 1. It distinguishes between real and fake based on patches. The patch size is set to 50×50. Similarly, the adversarial loss between the patch discriminator and the generator can be expressed as:

[0090] L adv2 = E x∈LQ [logD LQ (x)] + E y∈HQ [log(1 - D LQ (G HQ→LQ (y)))] +

[0091] E y∈HQ [logD HQ (y)] + E x∈LQ [log(1 - D HQ (G LQ→HQ (x)))] (4)

[0092] where D HQ and D LQ respectively represent the high-quality discriminator and the low-quality discriminator.

[0093] Finally, the adversarial loss of FS-GAN is expressed as:

[0094] L adv = L adv1 + L adv2 (5)

[0095] The loss function

[0096] As a bidirectional GAN framework, FS-GAN is suitable for unpaired training and includes two basic loss terms: the cycle-consistency loss L cycand the identity mapping loss L idt . The main purpose of the cycle consistency loss is to achieve the mutual conversion between the low-quality domain and the high-quality domain. For the previous enhancement task, the low-quality image x ∈ LQ is input into the high-quality generator G LQ→HQ to obtain image enhancement, and then the enhanced image is input into the low-quality generator G HQ→LQ to restore the image as much as possible. Therefore, x ≈ G HQ→LQ (G LQ→HQ (x)). Similarly, for backward cycle consistency, it is preferable that y ≈ G LQ→HQ (G HQ→LQ (y)), where y ∈ HQ represents the high-quality image. Therefore, the cycle consistency loss can be expressed as:

[0097] L cyc = E x∈LQ [||G HQ→LQ (G LQ→HQ (x)) - x|| 1 + E y∈HQ [||G LQ→HQ (G HQ→LQ (y)) - y|| 1 (6)

[0098] The identity mapping loss is to ensure the specificity of the generator, that is, when an image in the same domain is input into the generator, identity mapping should be performed. The identity mapping loss can be expressed as:

[0099] L idt = E x∈LQ [||G HQ→LQ (x) - x|| 1 + E y∈HQ [||G LQ→HQ (y) - y|| 1 (7)

[0100] When correcting the illumination distribution of medical images, some foreground structures that are important for medical diagnosis may be homogenized into the background, which is not the original intention of image enhancement. Although the traditional structural similarity loss comprehensively considers brightness, contrast, and structure, it only compares images from a global perspective, which is relatively rough, especially for medical images. The enhancement difficulty and requirements for each region of low-quality medical images are different. Some background regions do not require key attention, while some regions do. The present invention adopts the texture fidelity loss L TF , divides the image into multiple image blocks for fine processing and comparison, and is specifically expressed as:

[0101]

[0102] where y iand x i respectively represent the i-th local image patch of the high-quality image y and the low-quality image x. G HQ→LQ (y) i and G LQ→HQ (x) i respectively represent the i-th local image patch of the two generated images. and represent the covariance matrix between the i-th image patch of the original image and the i-th image patch of the corresponding generated image. and respectively represent the standard deviation matrices corresponding to y i , G HQ→LQ (y) i , x i and G LQ→HQ (x) i . c is a small constant used to avoid numerical instability. p is the number of blocks into which the image is segmented. It should be noted that SSRM can protect the global structural information, and the texture fidelity loss L TF can enhance the texture recovery in the local region.

[0103] Finally, the loss function of FS-GAN can be expressed as:

[0104] L FS-GAN = L adv + αL cyc + βL idt + γL TF + ηL enc (9)

[0105] where α, β, γ, and η are the weights of the cyclic consistency loss, identity mapping loss, texture fidelity loss, and self-guidance loss, respectively.

[0106] Experimental Example

[0107] Experimental Setup

[0108] Experimental Dataset

[0109] The experiment was conducted on the publicly available corneal confocal microscopy (CCM) dataset CORN-2. CORN-2 contains 688 confocal microscopy images with a size of 384×384. Two professionals divided these images into high-quality images and low-quality images for training, 340 and 288 respectively. In addition, there are 60 low-quality images available for testing. Among them, the low-quality confocal images are characterized by low contrast, speckle noise, and non-uniform illumination.

[0110] Comparison Algorithms

[0111] To demonstrate the superior performance of FS-GAN, several SOTA methods were selected for comparative experiments, including two traditional methods, CLAHE and DCP, and five deep learning methods, NST, MSG Net, EnlightenGAN, CycleGAN, and StillGAN. These methods are suitable for unpaired training, and their relevant parameter settings refer to the original papers and the released codes. For fair comparison, these five deep learning methods use the same data preprocessing process as FS-GAN.

[0112] Image Quality Assessment

[0113] Qualitative Analysis

[0114] In this section, we use visual observation to qualitatively analyze the quality of the generated images and evaluate the integrity of the texture structure and the rationality of the illumination distribution.

[0115] From Figure 3 it can be seen that the texture details of the original low-quality images are relatively blurred, which is not conducive to medical diagnosis. The improvement of the texture structure of corneal confocal images by the two traditional methods, CLAHE and DCP, is very limited, while the six deep learning-based methods perform better. Among all the images, the images generated by the method of the present invention not only have a clear overall structure, but also can restore the most complete texture details, which is the best among the eight methods.

[0116] From Figure 4 it can be seen that the illumination distribution of the original low-quality images is very uneven, with some areas overexposed and some areas very dark, covering many important pathological information. Although DCP and EnlightinGAN, which rely on prior information, adjust the illumination distribution to a certain extent, it is difficult for them to achieve good enhancement effects in some overexposed areas. Generally speaking, the illumination distributions of the images generated by StillGAN and FS-GAN (the method of the present invention) are closest to those of high-quality images, more uniform, and more suitable for clinical diagnosis. If the local area is magnified, the method of the present invention can not only adjust the illumination distribution, but also restore some texture details in the process. From the above analysis, it can be seen that the method of the present invention has the strongest ability to enhance texture details, and the images generated by it are most consistent with human vision.

[0117] Table 2: Comparison of the image quality enhancement of different methods in 5 evaluation metrics. Among these image quality evaluation metrics, Entropy and AVG are positive metrics, while Brisque, NIQE, and PIQE are negative metrics.

[0118] Entropy↑ AvG↑ Brisque↓ NIQE↓ PIQE↓ Original 4.743±0.095 5.137±1.302 0.498±0.165 28.939±5.097 9.498±2.248 CLAHE 6.964±0.983 6.935±0.168 0.492±0.011 25.622±3.869 10.655±2.427 DCP 5.565±0.041 7.327±0.217 0.532±0.207 22.601±6.002 11.494±2.479 NST 5.897±0.537 6.547±0.084 0.492±0.002 27.687±3.045 22.044±3.575 MSG-Net 6.583±0.096 6.544±0.043 0.493±0.011 31.863±5.235 2.289±0.227 CycleGAN 6.479±0.213 6.511±0.214 0.508±0.027 30.104±3.563 2.658±0.324 EnlightenGAN 6.229±1.078 6.695±0.338 0.487±0.116 26.290±4.690 7.098±1.974 StillGAN 6.546±0.399 6.587±0.125 0.492±0.001 31.649±4.246 1.861±0.165 The method of the present invention 6.785±0.138 7.332±0.024 0.484±0.005 28.107±6.073 1.774±0.235

[0119] Quantitative Analysis

[0120] To quantitatively analyze the image quality, the present invention uses five metrics to evaluate the image quality, including Entropy, AVG, Brisuqe, NIQE, and PIQE.

[0121] Table 2 shows the evaluation metric values of each enhanced image. Among all the methods, the AvG, Brisk, and PIQE of the method of the present invention are the best, being 0.005, 0.008, and 0.087 higher than the second-ranked method respectively, which indicates that the images generated by the method of the present invention have the highest quality.

[0122] Evaluation of Application Effect

[0123] Qualitative Analysis

[0124] The application value of medical images is also a way to measure their quality. Therefore, the present invention uses CS-Net to segment each enhanced image and evaluates its enhancement effect through the segmentation performance of each method. From Figure 5 It can be seen that due to factors such as uneven illumination and blurred structure, the segmentation effect of low-quality images is very poor, which is reflected in the difficulty of identifying and continuing nerve fibers. The nerve fiber structure segmented from the images generated by FS-GAN (the method of the present invention) has strong continuity and is closest to the real situation. The method of the present invention can solve the problem of mis-segmentation in some areas and greatly improve the subsequent application value of medical images.

[0125] Quantitative Analysis

[0126] The segmentation evaluation metric values of each enhanced image are shown in Table 3. Among them, the six segmentation metric values of the method of the present invention are significantly higher than those of other methods, and the segmentation effect is the best. It can be considered that the method of the present invention has the highest quality and application value for enhanced images.

[0127] Table 3: Comparison of segmentation metrics of enhanced images by different methods. These segmentation metrics are all positive, that is, the larger the value, the better the segmentation effect.

[0128]

[0129] Ablation Study

[0130] To verify the effectiveness of the components of FS-GAN (the method of the present invention), the present invention conducts two ablation experiments.

[0131] First, the present invention uses IEM as the backbone network and embeds SSRM and IDCM into the network respectively to verify the necessity and effectiveness of these two modules for image enhancement. Figure 6The enhanced images of each ablation scheme are shown. It can be seen that after embedding the SSRM, the nerve fiber structure of the generated images is more complete. The IDCM significantly corrects the illumination distribution, and the images generated by it are more suitable for clinical diagnosis. Therefore, it can be said that the SSRM and IDCM have played their respective roles.

[0132] Similarly, the effect of the fuzzy discriminator is also verified. From Figure 7 it can be seen that after embedding the fuzzy discriminator, the FS-GAN (the method of the present invention) can suppress noise in some areas, as shown by the dashed boxes in the first and second columns. The visual effect of the enhanced image of the FS-GAN is the best, indicating that the fuzzy discriminator can effectively enhance the discriminative ability of the model.

[0133] To generate high-quality medical images, an embodiment of the present invention proposes a fuzzy self-guided structure-preserving generative adversarial network system (FS-GAN), which can be applied to the training of unpaired data. In particular, the present invention develops a fuzzy discriminator to distinguish real images and generated images in the fuzzy domain, which can improve the enhancement performance of the model. In addition, a self-guided structure-preserving module (SSRM) and an illumination distribution correction module (IDCM) are designed to capture the structural information of nerve fibers in a self-guided manner and correct the illumination distribution of the images to improve the visual effect. The comparative experimental results show that the FS-GAN can significantly improve the quality of medical images and perform well in downstream application tasks, and can provide strong support for doctors' clinical diagnosis and treatment.

[0134] In the figure:

[0135] Figure 1 : The structure of the method in this paper. The generator consists of three modules, namely an image enhancement module (IEM), a self-guided structure-preserving module (SSRM), and an illumination distribution correction module (IDCM). The SSRM is responsible for exploring the nerve fiber structure in medical images and adding the feature maps to the backbone network proportionally during decoding. The IDCM corrects the illumination distribution of the output image of the backbone network by using the prior attention map of the low-quality input image. The cycle consistency is as shown above, and two generators with the same architecture are used. The output image of one generator will be input into the other generator to achieve end-to-end self-learning.

[0136] Figure 2 : The structure of the fuzzy discriminator. Different from the traditional discriminator, the fuzzy discriminator embeds a fuzzy operation unit to fuse the input image and the corresponding depth features, that is, projects the input image into the fuzzy domain for true / false discrimination.

[0137] Figure 3: Comparison of different methods for enhancing image texture structure. The first, third, and fifth rows are the complete images, and the second, fourth, and sixth rows are the enlarged images of the regions framed in red in the first, third, and fifth rows respectively. The images from left to right are the original image, CLAHE, DCP, NST, MSG Net, CycleGAN, Enlighten GAN, StillGAN, and FS-GAN (the method of the present invention).

[0138] Figure 4 : Comparison of the illumination distribution of images enhanced by different methods. The first, third, and fifth rows are the complete images, and the second, fourth, and sixth rows are the enlarged images of the regions framed in red in the first, third, and fifth rows respectively. The images from left to right are the original image, CLAHE, DCP, NST, MSG Net, CycleGAN, Enlighten GAN, StillGAN, and FS-GAN (the method of the present invention).

[0139] Figure 5 : Comparison of the image segmentation results enhanced by different methods. The first, third, and fifth rows are the enhanced images, and the second, fourth, and sixth rows are the segmentation images corresponding to the first, third, and fifth rows respectively. The images from left to right are the original image, CLAHE, DCP, NST, MSGNet, CycleGAN, EnlightenGAN, StillGAN, FS-GAN (the method of the present invention), and GroundTruth.

[0140] Figure 6 : Effectiveness analysis of SSRM and IDCM. The first and third rows are the complete images, and the second and fourth rows are the enlarged images of the regions framed in red in the first and third rows respectively. The images from left to right are the original image, Backbone (main network), With SSRM, With IDCM, and FS-GAN (the method of the present invention).

[0141] Figure 7 : Effectiveness analysis of the fuzzy discriminator. The first and second columns are the noisy images, and the fourth column is the enlarged local image patch of the red frame in the third column. The images from top to bottom are FS-GAN (the method of the present invention), -FuzzyD (the network without the fuzzy discriminator), and the original image.

[0142] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for a structure - protected generative adversarial network based on fuzzy self - guidance, characterized in that, it includes the following steps: The first step: Use the self - guided structure - protection module in the structure - protected generative adversarial network system based on fuzzy self - guidance to explore and retain the structural information of the input image; The second step: Correct the illuminance distribution of the enhanced image based on the illumination distribution correction module; The third step: Distinguish the input image and the enhanced image in the fuzzy domain based on the fuzzy discriminator; The fourth step: Use the cycle - consistency loss and the texture fidelity loss to represent the loss function; The self - guided structure - protection module is used to explore and retain the structural information of low - quality input images and guide the image enhancement module and the illumination distribution correction module to reconstruct medical images. The self - guided structure - protection module consists of two encoder - decoder networks with the same structure. The encoder is initialized by a pre - trained Res - Net - 34; The self-guided structure protection module compresses the original input and the output of the image enhancement module into the deep semantic space, and then restricts the distance between them to make the latent feature representations of the two stages close. The encoded features of the two networks are respectively denoted as E· 1 and E· 2 , and the self-guided loss function is expressed as: where LQ and HQ respectively represent the low - quality image set and the high - quality image set; The specific steps of the fuzzy operation in the fuzzy discriminator are as follows: (1) Two input images m 1 and m 2 respectively pass through a 1×1 convolutional layer to obtain feature maps of the same size and dimensions; (2) The two feature maps are non - linearly activated by the sigmoid function; (3) The two activated feature maps perform a fuzzy sum operation, that is, take the smaller value of the same dimension, and finally obtain the fuzzy feature map f.

2. A method for a structure - protected generative adversarial network based on fuzzy self - guidance according to claim 1, characterized in that: The specific steps of the fuzzy operation in the fuzzy discriminator are expressed as: f = FuzzyAnd[σ(c(m 1 )), σ(c(m 2 ))]; where c and σ respectively represent the convolutional layer and the sigmoid function; Let G LQ→HQ and G HQ→LQ represent a high-quality image generator and a low-quality image generator respectively, and FD HQ and FD LQ represent a high-quality blur discriminator and a low-quality blur discriminator respectively. The loss functions of the blur discriminator and the generator are expressed as: L adv1 = E x∈LQ [logFD LQ (x)] + E y∈HQ [log(1 - FD LQ (G HQ→LQ (y)))] + E y∈HQ [logFD HQ (y)]+E x∈LQ [log(1-FD HQ (G LQ→HQ (x)))]; The specific structure of the image patch discriminator includes three 4×4 convolutional layers with a stride of 2 and two 4×3 convolutional layers with a stride of 1. It distinguishes between true and false based on image patches, and the size of the image patches is set to 50×50. Similarly, the adversarial loss between the image patch discriminator and the generator is expressed as: L adv2 = E x∈LQ [logD LQ (x)] + E y∈HQ [log(1 - D LQ (G HQ→LQ (y)))] + E y∈HQ [logD HQ (y)] + E x∈LQ [log(1 - D HQ (G LQ→HQ (x)))]; Among which D HQ and D LQ represent a high-quality discriminator and a low-quality discriminator respectively; Finally, the adversarial loss of FS - GAN is expressed as: L adv = L adv1 + L adv2 .

3. A method for a structure - protected generative adversarial network based on fuzzy self - guidance according to claim 1, characterized in that: The structure - protected generative adversarial network system based on fuzzy self - guidance includes two generators and a fuzzy discriminator respectively connected to the two generators. The generator consists of an image enhancement module, a self - guided structure - protection module, and an illumination distribution correction module. The self - guided structure - protection module is responsible for exploring the nerve fiber structure in the medical image and adding the feature map to the backbone network proportionally during decoding. The illumination distribution correction module corrects the illumination distribution of the output image of the backbone network by using the prior attention map of the low - quality input image; Two generators with the same architecture are used, and the output image of one generator will be input into the other generator to achieve end - to - end self - learning.

4. A method for a structure - protected generative adversarial network based on fuzzy self - guidance according to claim 3, characterized in that: The image enhancement module uses U - Net as the backbone network. U - Net has eight encoding layers and eight decoding layers. The encoding layer consists of a residual convolution module and a max - pooling layer. The decoding layer consists of an up - sampling module and a residual convolution model. The residual convolution modules of the encoding layer and the decoding layer are the same, including two stacked convolutional layers + batch normalization layer + LeakyReLU, where the input and output are connected in the form of a residual; The up - sampling module is a serial structure of a bilinear interpolation up - sampling layer + 3×3 convolutional layer + batch normalization layer + LeakyReLU + dropout layer.

5. A method for a structure-protected generative adversarial network based on fuzzy self-guidance according to claim 3, characterized in that: The image structure information extracted by the self-guided structure protection module will be serially connected with the encoded features of the U-Net and the decoded features of the previous layer level by level, and injected into the decoding path of the image enhancement module.

6. A method for a structure-protected generative adversarial network based on fuzzy self-guidance according to claim 1, characterized in that: The loss function includes a cycle-consistency loss \(L\) cyc and an identity mapping loss \(L\) idt . The main purpose of the cycle-consistency loss is to achieve the mutual conversion between the low-quality domain and the high-quality domain. The low-quality image \(x\in LQ\) is input into the high-quality generator \(G\) LQ→HQ to obtain image enhancement, and then the enhanced image is input into the low-quality generator \(G\) HQ→LQ to recover the image as much as possible. Therefore, \(x\approx G\) HQ→LQ (G LQ→HQ (x)). For backward cycle-consistency, \(y\approx G\) LQ→HQ (G HQ→LQ (y)), where \(y\in HQ\) represents the high-quality image. The cycle-consistency loss is expressed as: L cyc = E x∈LQ [||G HQ→LQ (G LQ→HQ (x)) - x|| 1 + E y∈HQ [||G LQ→HQ (G HQ→LQ (y)) - y|| 1 ; The identity mapping loss is expressed as: L idt = E x∈LQ [||G HQ→LQ (x) - x|| 1 + E y∈HQ [||G LQ→HQ (y) - y|| 1 ; Texture fidelity loss L TF , the image is divided into multiple image patches for fine processing and comparison, specifically expressed as: where y i and x i represent the i-th local image patch of the high-quality image y and the low-quality image x respectively, G HQ→LQ (y) i and G LQ→HQ (x) i represent the i-th local image patch of the two generated images respectively, and represent the covariance matrix between the i-th image patch of the original image and the i-th image patch of the corresponding generated image, and represent the standard deviation matrices corresponding to y i , G HQ→LQ (y) i , x i and G LQ→HQ (x) i respectively, c represents a constant used to avoid numerical instability, p is the number of image blocks into which the image is segmented, SSRM protects the global structure information, and the texture fidelity loss L TF enhances the texture recovery in the local regions; The loss function of FS-GAN is expressed as: L FS-GAN = L adv + αL cyc + βL idt + γL TF + ηL enc ; α, β, γ, and η are the weights of the cycle consistency loss, identity mapping loss, texture fidelity loss, and self-guidance loss, respectively.

Citation Information

Patent Citations

  • Medical image enhancement method

    CN113763288A