Retinal Image Synthesis Method Based on Lesion-Attention Conditional Generative Adversarial Network
By generating an adversarial network based on lesion attention conditions, combined with technical means such as retinal vascular mask map and multi-output discriminator, the problem of blurred lesions and single samples of retinal synthetic images is solved, and the lesion details and diversity of high-resolution images are enhanced and the disease recognition effect is improved.
Patent Information
- Application Number
- CN202111243429.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-25
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-10-25
AI Technical Summary
In the prior art, the lesions of retinal synthetic images are blurred and the samples are single, resulting in poor results in the model on high-resolution images, especially in the case of small data volumes.
A generative adversarial network is adopted based on lesion attention conditions. By obtaining the retinal vascular mask map and inputting it into the trained generative adversarial network, combining the weight-sharing multi-output discriminator, random forest classifier, reverse activation network and attention module, the lesion feature fusion is enhanced, and the activation feature matching loss and perceived loss are used to improve image quality.
The lesion details of the synthetic images are enhanced, the disease recognition effect of high-resolution images is improved, and the diversity of images and the generalization performance of the model is enhanced.
Smart Images

Figure CN114155190B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to a method for synthesizing retinal images based on a lesion attention conditional generative adversarial network. Background Art
[0002] Color fundus photography is currently the most economical and non-invasive imaging method for detecting retinal diseases. Its wide availability makes it an ideal choice for evaluating various ophthalmic diseases. Currently, conventional fundus cameras have been widely used to detect retinal diseases, although there are some limitations. For example, conventional fundus images only cover the area of the central 30 to 60 degrees of the retina. In contrast, the range of ultra-wide field (UWF) retinal images based on the Optos camera is 200 degrees, covering 80% of the retina. It allows more clinically relevant lesions from the peripheral retina to be detected, which is important for lesions that vary in the peripheral retina, such as retinal degeneration, detachment, hemorrhage, exudation, etc. Deep learning has been successfully applied to the screening of conventional fundus images and achieved good results in the detection of various retinal diseases. In recent years, deep learning has also been applied to UWF fundus images. However, there are still some challenges in automatic diagnosis based on UWF images. First, some lesion information appears very small compared to the global image, making the lesion area very inconspicuous. Second, limited samples are prone to overfitting, causing the model to degrade in performance in new datasets. To alleviate the above problems, many researchers have developed many useful solutions based on the generative adversarial network (GAN). Using GAN to synthesize training samples to make up for the deficiencies of the original dataset can effectively enhance image samples. However, typical GANs have only achieved good results on low-resolution images and still have defects for high-resolution images.
[0003] Therefore, the prior art still needs to be improved and developed. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a method for synthesizing retinal images based on a lesion attention conditional generative adversarial network in view of the above-mentioned defects of the prior art, aiming to solve the problems of blurred lesion details in the synthesized retinal images and insufficient single synthesized retinal image samples in the prior art.
[0005] The technical solution adopted by the present invention to solve the problem is as follows:
[0006] In a first aspect, an embodiment of the present invention provides a method for synthesizing retinal images based on a lesion attention conditional generative adversarial network, wherein the method includes:
[0007] Obtain a retinal blood vessel mask image;
[0008] Input the retinal vessel mask image into a trained lesion-attention conditional generative adversarial network to obtain a synthetic retinal image.
[0009] In one implementation, the lesion-attention conditional generative adversarial network includes a generator, a weight-sharing multi-output discriminator, a random forest classifier, a reverse activation network, and an attention module. The weight-sharing multi-output discriminator outputs a disease discrimination result and a disease classification result. The random forest classifier is used to identify the lesion features of the disease classification result output by the weight-sharing multi-output discriminator. The reverse activation network is used to activate the lesion features. The attention module is used to fuse the lesion features into the decoder of the generator.
[0010] In one implementation, the random forest classifier is used to identify the lesion features of the disease classification result output by the weight-sharing multi-output discriminator specifically as follows:
[0011] The random forest classifier is used to obtain the lesion features by statistically analyzing the classification feature frequencies in the disease classification result output by the weight-sharing multi-output discriminator.
[0012] In one implementation, the training process of the lesion-attention conditional generative adversarial network includes:
[0013] Obtain a random Gaussian noise vector;
[0014] Obtain training samples, where the training samples include real retinal images, retinal vessel masks, and disease classification labels, and the retinal vessel masks are obtained by transforming the real retinal images.
[0015] Input the random Gaussian noise vector, the retinal vessel mask, and the real retinal image into a first network model, and output a predicted disease classification result corresponding to the real retinal image through the first network model.
[0016] According to the disease classification label and the predicted disease classification result, obtain a loss function;
[0017] Based on the loss function, train the first network model to obtain a lesion-attention conditional generative adversarial network.
[0018] In one implementation, the retinal vessel mask is obtained by transforming the real retinal image specifically as follows:
[0019] Perform retinal vessel segmentation on the real retinal image to obtain a retinal vessel segmentation image;
[0020] Filter the retinal vessel segmentation image to obtain a retinal vessel mask.
[0021] In one implementation, the loss function is obtained by adding an adversarial loss function, a classification loss function, an activation feature matching loss function, and a perceptual loss.
[0022] In one implementation, the activation feature matching loss function is used to calculate the lesion features of the reverse activation network.
[0023] In one implementation, the classification loss function is used to learn the lesion features of the real retinal image.
[0024] In a second aspect, an embodiment of the present invention further provides a retinal image synthesis device based on a lesion attention conditional generative adversarial network, where the device includes:
[0025] A retinal vessel mask map acquisition module, configured to acquire a retinal vessel mask map;
[0026] A retinal synthetic image acquisition module, configured to input the retinal vessel mask map into a trained lesion attention conditional generative adversarial network to obtain a retinal synthetic image.
[0027] In a third aspect, an embodiment of the present invention further provides an intelligent terminal, including a memory, and one or more programs, where one or more programs are stored in the memory and are configured to be executed by one or more processors, and the one or more programs include methods for executing the retinal image synthesis method based on a lesion attention conditional generative adversarial network as described in any one of the above.
[0028] In a fourth aspect, an embodiment of the present invention further provides a non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the retinal image synthesis method based on a lesion attention conditional generative adversarial network as described in any one of the above.
[0029] Advantages of the present invention: In the embodiments of the present invention, first, a retinal vessel mask map is acquired; finally, the retinal vessel mask map is input into a trained lesion attention conditional generative adversarial network to obtain a retinal synthetic image; it can be seen that the retinal synthetic image obtained by the lesion attention conditional generative adversarial network in the embodiments of the present invention can enhance the lesion details of the synthetic image, improve the diversity of the synthetic image, and enhance the disease recognition effect of high-resolution images. Description of the Drawings
[0030] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0031] Figure 1 Schematic flow chart of the retinal image synthesis method based on the lesion attention conditional generative adversarial network provided by the embodiment of the present invention.
[0032] Figure 2 Overall technical flow chart of an embodiment of the retinal image synthesis method based on the lesion attention conditional generative adversarial network provided by the embodiment of the present invention.
[0033] Figure 3 Example diagram of an implementation manner of the details of the synthesized DR image provided by the embodiment of the present invention.
[0034] Figure 4 Example diagram of an implementation manner of the details of the synthesized RP image provided by the embodiment of the present invention.
[0035] Figure 5 Experimental result diagram of an implementation manner of using synthetic data to increase the DR training set provided by the embodiment of the present invention.
[0036] Figure 6 Experimental result diagram of an implementation manner of using synthetic data to increase the RP training set provided by the embodiment of the present invention.
[0037] Figure 7 Principle block diagram of the retinal image synthesis device based on the lesion attention conditional generative adversarial network provided by the embodiment of the present invention.
[0038] Figure 8 Internal structure principle block diagram of the intelligent terminal provided by the embodiment of the present invention. Detailed implementation manners
[0039] The present invention discloses a retinal image synthesis method based on a lesion attention conditional generative adversarial network. To make the purpose, technical solutions and effects of the present invention clearer and more definite, the following further details the present invention with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0040] Those skilled in the art can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present invention means the presence of the stated features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes any and all combinations of one or more of the associated listed items.
[0041] Those skilled in the art can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention pertains. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.
[0042] In the prior art, in order to improve the effect of GAN in synthetic images, the basic GAN framework is enhanced with side information in traditional retinal image synthesis methods. One strategy is to provide class labels to the generator and discriminator to generate class-conditional samples. Adding an auxiliary classifier in the GAN framework can make the synthetic image have higher global consistency and obtain higher classification performance. However, this method requires a large amount of unlabeled data to train the GAN, is not applicable to the UWF dataset with less data, and currently only achieves good results on low-resolution images and has poor effects on high-resolution images. Another strategy is to synthesize realistic retinal images conditioned on the retinal vessel mask. However, the current method cannot effectively retain the details of small lesions, and the synthetic image may be blurred in many key details. These defects will cause the model to be unable to learn effective lesion features.
[0043] To solve the problems of the prior art, this embodiment provides a retinal image synthesis method based on a lesion attention conditional generative adversarial network. The retinal synthetic image obtained through the lesion attention conditional generative adversarial network can enhance the lesion details of the synthetic image, improve the diversity of the synthetic image, and enhance the disease recognition effect of high-resolution images. Specifically, during implementation, first obtain a retinal vessel mask image; then input the retinal vessel mask image into a trained lesion attention conditional generative adversarial network to obtain a retinal synthetic image.
[0044] Exemplary Method
[0045] This embodiment provides a method for synthesizing retinal images based on a lesion attention conditional generative adversarial network, which can be applied to intelligent terminals for image processing. Specifically, as Figure 1 shown, the method includes:
[0046] Step S100: Obtain a retinal blood vessel mask image;
[0047] In the embodiment of the present invention, the retinal image is an ultra-widefield retinal image (UWF). Because the ultra-widefield retinal image can detect peripheral retinal lesions simultaneously and obtain more complete retinal information, it has become an important tool for screening retinal diseases. Automatic detection of retinal diseases based on deep learning plays an important role in clinical practice. However, training a model with strong generalization ability requires a large amount of training data. Therefore, the present invention first obtains a retinal blood vessel mask image from a segmentation network that has been trained on a retinal blood vessel segmentation dataset, preparing for subsequent generation of diverse retinal synthetic images.
[0048] After obtaining the retinal blood vessel mask image, the following steps as Figure 1 shown can be executed: Step S200: Input the retinal blood vessel mask image into a trained lesion attention conditional generative adversarial network to obtain a retinal synthetic image;
[0049] In practice, the effects of typical data augmentation methods are limited and cannot generate diverse data. Therefore, the obtained retinal blood vessel mask image is input into a trained lesion attention conditional generative adversarial network. Since this network is based on lesion attention conditions, it can obtain more lesion information of retinal images, and the synthesized images have better effects at high resolutions.
[0050] In one implementation, as Figure 2 shown, the lesion attention conditional generative adversarial network includes a generator, a multi-output discriminator with weight sharing, a random forest classifier, a reverse activation network, and an attention module. Among them, the multi-output discriminator with weight sharing outputs a disease discrimination result and a disease classification result; the random forest classifier is used to identify the lesion features of the disease classification result output by the multi-output discriminator with weight sharing; the reverse activation network is used to activate the lesion features; the attention module is used to fuse the lesion features into the decoder of the generator. In this embodiment,
[0051] The generator adopts a common encoder - decoder structure to synthesize 512×512 images. First, the encoder encodes the retinal vessel mask, and then introduces a noise code before decoding. The noise code is random noise with a normal distribution. Finally, the decoder gradually generates the synthetic retinal image. To preserve the vessel structure during the synthesis of the retinal image, skip connections are used in the encoding path and the decoding path. Thus, the entire encoder and decoder perform 6 times of downsampling and upsampling respectively. To improve the model performance, a series of residual modules are used in the generator. Each module consists of two BatchNorm - ReLU - Convolution layers. Each convolutional layer has a 3x3 convolutional kernel, and residual connections are used to reduce overfitting. Among them, the downsampling module sets stride = 2 in the first convolutional layer, and the upsampling module uses nearest - neighbor interpolation to increase the feature size. The upsampling residual module is as shown in Figure 2 (b). During the upsampling process, the features of the discriminator's reverse activation network are fused through an attention module. Finally, the generator outputs the synthetic retinal image through a 3x3 convolutional layer and a tanh function, and the pixel values are kept between [-1, 1]. The discriminator does not use residual connections. The discriminator also performs 6 times of downsampling, which is carried out by using max - pooling layers. And the discriminator's module is composed of two groups of Convolution - LeakyReLU layers without using batch normalization layers because research shows that using batch normalization layers in the discriminator network of GAN will reduce the diversity of the generated images. The discriminator outputs both discriminative results and classification results. The output results and classification results share the feature extraction network but construct their own output layers respectively. In addition, an auxiliary classifier C is constructed in the discriminator to learn the lesion features of the image through the classification loss function. The present invention proposes a lesion feature attention mechanism that uses these lesion features to enhance the generator to generate real - looking retinal images. Specifically, first, the disease classification is learned through a random forest classifier. According to the research of Gu et al., the lesion features are obtained by statistically counting the classification feature frequencies in the disease classification results output by the multi - output discriminator with shared weights. That is to say, the key features, namely the lesion features, can be identified by calculating the frequency of each contributing classification feature. Then, these lesion features are activated through the reverse activation network. Finally, these lesion features are fused through the attention module to provide information for the generator. In practice, only a small number of features extracted by the discriminator are important for the prediction of lesions. After identifying these key features through the random forest, they can be output to the activation network to obtain the activation projection of the key features. As shown in Figure 2As shown, the activation network is the reverse operation of the discriminator. For each convolutional layer in the discriminator, there is a corresponding transposed convolutional layer, where the stride and the convolutional kernel size are the same. The transposed convolutional layers share the same weight parameters, except that the convolutional kernels are vertically and horizontally flipped. For each max pooling layer, there is an unpooling layer that performs a partial reverse operation, where the max elements are located through skip connections and the non-max elements are padded with zeros. For a LeakyReLU function, there is a corresponding LeakyReLU function in the activation network, which is used to ensure that the output activation values of each layer are positive. Figure 3 and Figure 4 Some examples of the activation of lesion features are shown in Figure 2 . As can be seen from the figure, although the classification training only uses image-level labels for training, the reverse activation map of the key features accurately locates the position of the lesion, indicating that the activation network contains the lesion information of the retina. The attention module uses the lesion features obtained by the activation network to provide lesion information for the generator, so that the generated retinal synthetic image can generate better lesion details. As shown in Figure 2 (c), at the l-th layer, the feature representation of the generator module is F l , and the feature representation of the activation network is A l . After the two features are normalized respectively, they are multiplied, and the ReLU activation function and 1×1 convolution are used to generate the attention features. Finally, the attention features are fused into the generator. This attention mechanism can be expressed as:
[0052] F attn = Conv(ReLU(Norm(F l )) × Norm(A l ))),
[0053]
[0054] where A l is obtained by reverse-activating the key features and contains the lesion information. The role of using the Norm function to normalize the input features is to ensure that the feature learning remains standardized and controllable. Based on this attention module, each generation module will incorporate the lesion features, which can enhance the lesion details of the generated retinal synthetic image.
[0055] In one implementation, the training process of the lesion attention conditional generative adversarial network includes the following steps: obtaining a random Gaussian noise vector; obtaining training samples, where the training samples include real retinal images, retinal vascular masks, and disease classification labels, and the retinal vascular masks are obtained by transforming the real retinal images; inputting the random Gaussian noise vector, the retinal vascular masks, and the real retinal images into a first network model, and outputting a predicted disease classification result corresponding to the real retinal image through the first network model; obtaining a loss function according to the disease classification labels and the predicted disease classification results; and training the first network model based on the loss function to obtain a lesion attention conditional generative adversarial network.
[0056] Specifically, during actual training, the training samples are the classification labels corresponding to the labeled real retinal images, forming a real retinal image-label pair, that is, one sample contains (real retinal image, disease classification label, retinal vascular mask). Among them, the real retinal images are obtained by shooting with hospital equipment, the disease classification labels are marked by expert doctors, and the retinal vascular masks are obtained by a segmentation network trained on a retinal vascular segmentation dataset and are obtained by transforming the real retinal images. Specifically, the real retinal images are subjected to retinal vascular segmentation to obtain retinal vascular segmentation images; then the retinal vascular segmentation images are subjected to stripe filtering such as rotation, retaining small bright spots in the images and removing large-area bright spots, and a retinal vascular mask is obtained through several filtering combinations. In this embodiment, for a retinal image training set where is an RGB color real retinal image with a width of W and a height of H, v i ∈ {0, 1} W×H is the retinal vascular mask corresponding to x i , and y i is the diagnostic classification label corresponding to the x i image. Inputting the random Gaussian noise vector Z and the retinal vascular mask V into the generator in the first network model, where Z is connected to a class label vector and outputs a synthesized image, which is formally expressed as The random noise Z is used to introduce appearance diversity, and the generator G learns a multi-valued mapping. After the generator generates a synthesized image, it and the real retinal image are input into the multi-output discriminator D with weight sharing in the first network model through several affine transformations, and the loss function is calculated by the discriminator D. The goal of the discriminator D is to distinguish the synthesized retinal image from the real retinal image x is expressed as D(X, v) → d ∈ [0, 1]. When X is the real retinal image x, d should tend to 1, and vice versa when X is the synthesized retinal image When d should tend to 0. For such a minimax two-player game setting, the generative adversarial network (GANs) method is adopted, and the following optimization problem characterizing the interaction between G and D is considered:
[0057]
[0058] The random forest in the first network model learns disease classification using the features of the last layer extracted by the discriminator D, and identifies lesion features by calculating the frequency of each contribution to classification. The activation network is the reverse operation of the discriminator, and its weights are imported by the discriminator. Inputting the lesion features identified by the random forest into the activation network of the first network model can obtain the lesion features of the corresponding layer, which are fused into the generator through the attention module during the decoding process of the generator in the first network model.
[0059] Furthermore, combining the above formula with the global L1 loss can provide more consistent results, ensuring that the synthesized retinal images do not deviate significantly from the real retinal images. The global L1 loss is as follows:
[0060]
[0061] Furthermore, the performance of the GAN can be enhanced by providing side information. For example, the auxiliary classifier GAN (AC-GAN) constructs a conditional model, which improves the performance by extending the objective function of the GAN through an auxiliary classifier. Therefore, in addition to the discriminator D, the classification labels of the dataset can also be utilized simultaneously, and an auxiliary classifier C is constructed. The classification loss function is as follows:
[0062]
[0063] The training of the GAN can significantly increase the amount of data. Data augmentation and regularization methods have been applied to GAN training to improve the performance of the GAN. According to the research of Tran et al., data augmentation can improve the learning of the original data distribution and further enhance the quality of the generated images by adding transformed data. For the discriminator, a multi-output discriminator based on weight sharing is designed to utilize the transformed data and improve the generator's learning of the original data. As Figure 2 (a) shown on the right, in addition to inputting the generated retinal synthetic images and real retinal images into the discriminator, a series of transformation functions T k , k = {1,..., K}, are constructed to perform K - 1 transformations on the retinal synthetic images and real retinal images respectively, and then input them into the discriminator. The synthetic retinal images and real retinal images undergo the same transformation. Therefore, the distribution of the original data will not be changed when using data augmentation. Here, T1 is represented as the operation without transformation, then The adversarial loss function in will become:
[0064]
[0065] Similarly, the classifier is also extended to K, and each performs the classification of the transformed image. Therefore, the classification loss function becomes:
[0066]
[0067] To further improve the quality of the synthetic image, the present invention proposes an activation feature matching loss based on a reverse activation network. The feature matching loss minimizes the statistical difference between the real retinal image and the retinal synthetic image at multiple scales, obtaining better image distribution information. Due to changes in color, texture, illumination, etc., fundus images have appearance diversity. Using lesion features at different scales for learning can constrain the lesion distribution information of the synthetic image. The activation feature matching loss is calculated using the features of the reverse activation network. It focuses on the lesion features and ignores the background. Let p represent the p-th layer of the activation network, and the activation feature matching loss function is expressed as:
[0068]
[0069] Data augmentation regularization improves feature learning, especially crucial for the diversity learning of generated images. Since the activation feature matching loss focuses on the lesion features, the background features are not greatly constrained. Therefore, randomness is introduced into the generator to enhance diversity. In addition, to ensure the physiological details of the retinal image, a perceptual loss is applied. The perceptual loss is based on the pre-trained VGG-19. Let q represent a certain layer of the VGG19 network, and the difference between the real image and the synthetic image is calculated at this layer. The perceptual loss function is expressed as:
[0070]
[0071] Finally, the loss function is obtained by adding the adversarial loss function, the classification loss function, the activation feature matching loss function, and the perceptual loss.
[0072] Experimental results of the present invention
[0073] 1. Dataset
[0074] To verify the effectiveness of the model, three datasets with diagnostic labels were used for diabetic retinopathy (DR) and retinitis pigmentosa (RP). The first dataset was collected from Shenzhen Eye Hospital, obtained through an Optos camera, including 398 DR lesion images, 473 RP lesion images, and 948 normal retina images. The original images were collected from different patients registered in the hospital from 2016 to 2019. Three experts participated in labeling these images, and the image labels were retained only when at least two ophthalmologists agreed on the disease labels. In the experiment, normal retina images were randomly divided for DR and RP studies, with 70% used for training and 30% for testing. The images used for DR and RP experiments were abbreviated as DR-1 and RP-1, respectively.
[0075] The second dataset is the publicly available dataset DeepDR provided by ISBI2020. This dataset contains 2000 conventional fundus images and 256 ultra-wide field images. However, in the experiment, only 150 training set ultra-wide field images with publicly available labels were used, among which there were 106 lesions and 44 normals. This image set was only used to test the method proposed in the present invention to verify the generalization performance of the method. The dataset was abbreviated as DR-2.
[0076] The third dataset is the publicly available dataset Masumoto from a Japanese hospital, Tsukazaki, which contains 150 RP lesion images and 223 normal retina images. The original images were collected between 2011 and 2017. This image set was also only used to test the method proposed in the present invention to verify the generalization performance of the method, and was abbreviated as RP-2.
[0077] 2. Effect of synthetic images
[0078] In Figure 3 and Figure 4 show some synthetic retina images and lesion details trained using DR and RP data. It can be seen from the figure that the method proposed in the present invention can synthesize retina images and related lesion details with better effects.
[0079] 3. Data augmentation for lesion classification using synthetic images
[0080] To evaluate whether the synthesized data can enhance the training of the classification model, three different classical classification models, namely VGG-16, ResNet-50, and Inception-V3, were trained as benchmarks in the experiment. And three experimental settings were set up. The first setting was to train the classification model using the real training set and test it on the synthesized images. The second setting was to train using the real training set and test in the real test set. The third setting was to use a mixture of real and synthesized images as the training set and test in the real test set. In the three experimental settings, the number of samples in the synthesized image set was the same as that in the formal image training set. The experimental results using DR data and RP data are shown in Table 1 and Table 2 respectively.
[0081] As can be seen from these tables, the images synthesized by the method proposed in the present invention can enhance the training of the model. For the results of the first experimental setting, the model trained on real images can achieve the best accuracy on the synthesized images. The synthesized images were obtained from the real image training set. Therefore, the distribution of the realistic synthesized images is similar to that of the training set. From experimental settings 2 and 3, it can be seen that using synthesized images improves the performance of the model, especially the performance of the model on DR-2 and RP-2, which indicates that the synthesized images can improve the feature diversity of the images, thereby improving the generalization performance of the model.
[0082] Table 1 shows the experimental results using DR data to evaluate the enhancement performance of the synthesized images. Through three comparative experimental settings, it is shown that the synthesized images can improve the accuracy and generalization performance of the model. Real in the table represents the real image training set, and Fake represents the synthesized images.
[0083]
[0084] Table 2 shows the experimental results using RP data.
[0085]
[0086] To verify the degree to which the synthesized images enhance the model training, the training set was multiplied with the synthesized images. The experimental results are as Figure 5 and Figure 6 shown. "1x" in the figure represents using only the training set of real retinal images, "2x" represents the mixed training set with the data doubled using the synthesized images, and so on. As can be seen from the figure, as more retinal synthesized images are added to the augmented mixed training set, the baseline classification performance is improved, and the best performance is achieved after being augmented to about 7 - 8 times.
[0087] Innovative points of the present invention:
[0088] (1) A GAN for UWF image synthesis based on conditional input and lesion feature attention is proposed. The proposed lesion feature attention mechanism can enhance the lesion details of the synthesized images.
[0089] (2) The feature matching loss is improved to enhance the performance of GAN training. This loss function focuses on lesion features, which can enhance lesion details and improve the diversity of the synthesized images at the same time.
[0090] (3) A multi-discriminator output with weight sharing is designed to utilize the model performance enhanced by affine transformation.
[0091] In summary, a lesion feature attention mechanism is proposed in the present invention to enhance the lesion details of the synthesized retinal images in the conditional GAN, improve the feature matching loss to enhance the diversity of the synthesized retinal images, and the designed multi-discriminator with weight sharing can utilize affine transformation to enhance model training.
[0092] Exemplary Device
[0093] As Figure 7 shown, an embodiment of the present invention provides a retinal image synthesis device based on a lesion attention conditional generative adversarial network. The device includes a retinal vascular mask map acquisition module 301 and a retinal synthesized image acquisition module 302, where:
[0094] The retinal vascular mask map acquisition module 301 is used to acquire a retinal vascular mask map;
[0095] The retinal synthesized image acquisition module 302 is used to input the retinal vascular mask map into a trained lesion attention conditional generative adversarial network to obtain a retinal synthesized image.
[0096] Based on the above embodiment, the present invention also provides an intelligent terminal, and its principle block diagram can be as Figure 8 shown. The intelligent terminal includes a processor, a memory, a network interface, a display screen, and a temperature sensor connected through a system bus. Among them, the processor of the intelligent terminal is used to provide computing and control capabilities. The memory of the intelligent terminal includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the intelligent terminal is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it realizes a retinal image synthesis method based on a lesion attention conditional generative adversarial network. The display screen of the intelligent terminal can be a liquid crystal display screen or an electronic ink display screen, and the temperature sensor of the intelligent terminal is pre-set inside the intelligent terminal to detect the operating temperature of the internal device.
[0097] Those skilled in the art can understand that Figure 8 the schematic diagram in
[0098] In one embodiment, an intelligent terminal is provided, including a memory, and one or more programs, where one or more programs are stored in the memory and are configured to be executed by one or more processors. The one or more programs include instructions for performing the following operations:
[0099] Obtain a retinal vessel mask image;
[0100] Input the retinal vessel mask image into a trained lesion attention conditional generative adversarial network to obtain a retinal synthetic image.
[0101] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to the memory, storage, database, or other media used in the embodiments provided by the present invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0102] In summary, the present invention discloses a method for synthesizing retinal images based on a lesion attention conditional generative adversarial network, and the method includes: obtaining a retinal blood vessel mask image; inputting the retinal blood vessel mask image into a trained lesion attention conditional generative adversarial network to obtain a retinal synthetic image. The retinal synthetic image obtained by the present invention through the lesion attention conditional generative adversarial network can enhance the lesion details of the synthetic image, improve the diversity of the synthetic image, and enhance the disease recognition effect of high-resolution images.
[0103] Based on the above embodiments, the present invention discloses a method for synthesizing retinal images based on a lesion attention conditional generative adversarial network. It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present invention.
Claims
1. A method for synthesizing retinal images based on a lesion attention conditional generative adversarial network, characterized in that The method includes: Obtaining a retinal vessel mask image; Inputting the retinal vessel mask image into a trained lesion attention conditional generative adversarial network to obtain a retinal synthetic image; The lesion attention conditional generative adversarial network includes a generator, a multi-output discriminator with shared weights, a random forest classifier, a reverse activation network, and an attention module; Among them, the generator has an encoder-decoder structure and is used to synthesize images; Moreover, the multi-output discriminator with shared weights outputs a disease discrimination result and a disease classification result; The random forest classifier is used to obtain lesion features by statistically analyzing the classification feature frequencies in the disease classification results output by the multi-output discriminator with shared weights; The reverse activation network is used to activate the lesion features for localizing the lesion positions; The attention module fuses the lesion features into the decoder of the generator.
2. The method for synthesizing retinal images based on a lesion attention conditional generative adversarial network according to claim 1, wherein The training process of the lesion attention conditional generative adversarial network includes: Obtaining a random Gaussian noise vector; Obtaining training samples, where the training samples include real retinal images, retinal vessel masks, and disease classification labels, and the retinal vessel masks are obtained by transforming the real retinal images; Inputting the random Gaussian noise vector, the retinal vessel mask, and the real retinal image into a first network model, and outputting a predicted disease classification result corresponding to the real retinal image through the first network model; Obtaining a loss function according to the disease classification label and the predicted disease classification result; Training the first network model based on the loss function to obtain a lesion attention conditional generative adversarial network.
3. The method for synthesizing retinal images based on a lesion attention conditional generative adversarial network according to claim 2, wherein The retinal vessel mask obtained by transforming the real retinal image is specifically: Performing retinal vessel segmentation on the real retinal image to obtain a retinal vessel segmentation image; Filtering the retinal vessel segmentation image to obtain a retinal vessel mask.
4. The method for synthesizing a retinal image based on a lesion attention conditional generative adversarial network according to claim 2, wherein The loss function is obtained by adding an adversarial loss function, a classification loss function, an activation feature matching loss function, and a perceptual loss.
5. The method for synthesizing a retinal image based on a lesion attention conditional generative adversarial network according to claim 4, wherein The activation feature matching loss function is used to calculate the lesion features of the reverse activation network.
6. The method for synthesizing retinal images based on a lesion attention conditional generative adversarial network according to claim 4, wherein The classification loss function is used to learn the lesion features of the real retinal image.
7. An intelligent terminal, characterized in that, It includes a memory and one or more programs, where one or more programs are stored in the memory and are configured to be executed by one or more processors. The one or more programs include those for executing the method according to any one of claims 1-6.
8. A non-transitory computer-readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device can execute the method according to any one of claims 1-6.
Citation Information
Patent Citations
Image synthesis method and system combining adversarial auto-encoder and generative adversarial network
CN111402179A
Medical image interpretation method and apparatus, computer device and storage medium
WO2020215557A1